The DecoderMatthias Bastian
Qwen3.8-Omni-Flash Matches Gemini Flash on Multimodal Benchmarks at a Fraction of the Cost
Original title:Qwen3.8-Omni-Flash undercuts Google's Gemini Flash pricing while matching its multimodal benchmarks
Models83
Qwen has introduced Qwen3.8-Omni-Flash, its first multimodal model tailored for autonomous agent tasks. Capable of simultaneous audio-video processing and native tool use, it automates workflows like vlog editing, clip translation, and long-form video summarization. Benchmark results place its audio-visual capabilities in close competition with Gemini 3.8 Flash while undercuting its API pricing, underscoring the rapid decline in multimodal inference expenses.
Why it's worth reading
It represents a notable cost reduction for audio-visual agent workflows, bringing continuous multimodal perception much closer to widespread commercial viability.
Tags
QwenMultimodalAI AgentsGemini FlashAPI PricingAudio VideoOpen Source