Alibaba's Qwen team released Qwen3.8-Omni-Flash on September 19, 2026, a multimodal model that processes audio and video together and operates as an AI agent capable of using tools independently. The model matches Google's Gemini 3.8 Flash on key benchmarks while undercutting its pricing by roughly 80 percent on input and nearly 88 percent on output tokens.
Qwen3.8-Omni-Flash is the company's first model purpose-built for agentic workflows. It draws conclusions from combined audio-video inputs and can edit vlogs, translate short videos, or summarize movies without step-by-step human direction. The context window holds one million tokens.
Pricing that resets expectations
API costs sit at $0.15 per million input tokens and $0.47 per million output tokens. Qwen estimates audio input at under $0.01 per hour. Processing 720p video with audio at one frame per second runs about $0.20, excluding response costs. By contrast, Gemini 3.8 Flash charges $0.75 for input and $3.75 for output per million tokens at its introductory rate. Those prices double on January 1, 2027.
The gap matters most for teams running high-volume multimodal workloads. A customer support operation processing thousands of video inquiries or a marketing team analyzing user-generated clips will see the difference compound quickly. Qwen said the model performs "on par with Gemini Flash 3.8 in multimodal benchmarks."
Tools and real-time interaction
The model is available through Qwen Studio, Qwen Cloud, and the API. The open-source Qwen-MM-Plugins extend its reach into practical workflows - video editing, speaker recognition, PDF video notes, and reusable task chains. These plugins work with agents like Claude Code, Gemini CLI, and Qwen Code.
Qwen-Live Harness adds real-time interaction through a camera and microphone. That opens use cases in live customer support, event coverage, and on-the-fly content production where waiting for batch processing isn't an option.
Why this matters for creatives, marketers, and support teams
For teams producing video content, handling customer inquiries, or running events, the price difference is the immediate headline. A marketing department that processes hundreds of hours of footage monthly can shift costs from thousands of dollars to hundreds. Creatives editing vlogs or short-form video gain agent-driven tooling that handles repetitive cuts and translations without manual intervention. Customer support operations can experiment with real-time video analysis at a fraction of the cost of incumbent providers. The model's tool-use capability means fewer handoffs between AI output and the software that acts on it - the agent edits, translates, or summarizes directly.
Your membership also unlocks: