Qwen has launched Qwen3.8-Omni-Flash, its next-generation native omnimodal model designed to handle text, images, audio, and video while also planning tasks, using tools, and completing multi-step work.
The model supports a 1-million-token context window and is now available through the Qianwen AI Platform. Qwen is targeting workflows including video editing, music video creation, film commentary, audio-visual summarization, and real-time conversations.
Major Upgrade Over Qwen3.5-Omni-Plus
Qwen says the new model improves its average score across 29 evaluations by more than 25% compared with Qwen3.5-Omni-Plus.
The company also claims audio input costs have fallen by more than 98%, while audio-visual input costs are down by over 93% under its pricing methodology.
Qwen says Qwen3.8-Omni-Flash delivers audio-visual performance close to Gemini 3.8 Flash while exceeding it in overall audio performance. These figures come from Qwen’s own testing.
| Benchmark / Capability | Qwen3.8-Omni-Flash |
|---|---|
| WildClawBench-MM | 71.0 |
| UniClawBench | 69.6 |
| DailyOmni | 85.1 |
| StreamingBench | 80.8 |
| SWE-bench Pro | 63.3 |
| LiveCodeBench v6 | 92.6 |
| CoWorkBench | 75.3 |
Cases and more details: pic.twitter.com/Mg6dBqWgGT
— Qwen (@Alibaba_Qwen) September 18, 2026
AI That Decides What to Watch
For long videos, Qwen3.8-Omni-Flash does not necessarily process every frame. Its agent can start with a question, decide which parts of the recording matter, and progressively gather evidence.
On OmniVideoBench, this approach increased accuracy from 63.4 to 67.8 while reducing token use by around 45.7%.
The model can also analyze specific elements such as characters, camera shots, lighting, and sound, based on what the user asks it to examine.
Meetings Can Turn Into Actions
Qwen3.8-Omni-Flash supports up to one hour of audio-visual meeting input.
It can identify speakers, transcribe discussions, generate minutes, extract action items, and analyze project risks. When connected to tools, it can also send emails, organize tasks, or begin coding based on meeting requirements.
Video Editing and Content Creation
Qwen is also using the model for end-to-end media production.
It can analyze music before planning and creating music videos, while its video translation workflow can handle transcription, translation, voice cloning, dubbing, audio mixing, and final review.
For full-length films, Qwen says an agent can handle plot extraction, script planning, voiceover, music, editing, rendering, and final quality checks.
Real-Time Version Also Available
Qwen has also introduced Qwen3.8-Omni-Flash-Realtime, which processes live audio and video while responding and using tools.
The company says it can combine visual information with spatial sound to determine where sounds are coming from and assist with localization and navigation.
The model supports speech recognition in 74 languages, including Urdu and Punjabi, while speech generation covers 29 languages.
Qwen has also expanded Qwen-MM-Plugins and open-sourced Qwen-Live Harness for multimodal agent workflows, long-term memory, task delegation, and real-time interaction.
The post Qwen3.8-Omni-Flash Launches With 1M Context and Advanced Audio-Visual Agents appeared first on ProPakistani.
