QwQ-Max-Preview from Qwen is a powerful new LLM excelling in reasoning, math, coding, and agent tasks. Features a "thinking mode" for complex problems. Open-source coming soon!
Qwen's omni-modal model built around agentic capabilities
Qwen3.8-Omni-Flash understands text, image, audio, and video with a 1M-token context, then plans, calls tools, and completes real work: editing video, making music videos, dubbing/translating shows, summarizing meetings into action items, even coding off what it hears and sees.
No reviews yetBe the first to leave a review for QWQ-Max
Hunter
📌
Qwen3.8-Omni-Flash is Alibaba's next-gen native omnimodal model, built to shift omni AI from just understanding audio/video to actually planning, calling tools, and finishing the work.
It takes text, image, audio, and video input with a 1M-token context window. Across 29 evaluations it improves over 25% on average vs the previous Qwen3.5-Omni-Plus, while audio input pricing drops over 98% and audio-visual pricing drops over 93%.
What it can actually do:
Turn a song into a full music video (Music2MV) - timed lyrics, scene and character design based on rhythm and mood
Translate a whole short drama in one instruction - speaker-aware transcription, translation, voice cloning/dubbing, remixing, QC
Watch a 2-3 hour film and produce a full commentary video - plot extraction, script, voiceover, music, editing, render
Turn a multi-speaker meeting recording into minutes, action items, and even trigger emails or coding tasks
Do agentic long-video search - locate the relevant few minutes in hours of footage instead of processing everything, cutting token use ~46% while improving accuracy
Run real-time audio-visual conversation via a separate Realtime variant, including spatial sound localization ("go see what's making that noise")
It also ships with two open-source pieces: Qwen-MM-Plugins (plugs omni capabilities into agent harnesses like Claude Code, Codex, Qwen Code) and Qwen-Live Harness (runtime for real-time voice/video agents, npm install).
Qwen3.8-Omni-Flash is Alibaba's next-gen native omnimodal model, built to shift omni AI from just understanding audio/video to actually planning, calling tools, and finishing the work.
It takes text, image, audio, and video input with a 1M-token context window. Across 29 evaluations it improves over 25% on average vs the previous Qwen3.5-Omni-Plus, while audio input pricing drops over 98% and audio-visual pricing drops over 93%.
What it can actually do:
Turn a song into a full music video (Music2MV) - timed lyrics, scene and character design based on rhythm and mood
Translate a whole short drama in one instruction - speaker-aware transcription, translation, voice cloning/dubbing, remixing, QC
Watch a 2-3 hour film and produce a full commentary video - plot extraction, script, voiceover, music, editing, render
Turn a multi-speaker meeting recording into minutes, action items, and even trigger emails or coding tasks
Do agentic long-video search - locate the relevant few minutes in hours of footage instead of processing everything, cutting token use ~46% while improving accuracy
Run real-time audio-visual conversation via a separate Realtime variant, including spatial sound localization ("go see what's making that noise")
It also ships with two open-source pieces: Qwen-MM-Plugins (plugs omni capabilities into agent harnesses like Claude Code, Codex, Qwen Code) and Qwen-Live Harness (runtime for real-time voice/video agents, npm install).
Try it: Qwen Studio · API · Qwen-MM-Plugins