Qwen3.5-Omni - A native omni model for voice, video, and tools

Qwen3.5-Omni is Qwen"s new native omni model for text, images, audio, and video, with stronger multilingual speech, realtime voice interaction, web search, function calling, voice cloning, and long-context audio/video understanding.

Add a comment

Replies

Best

Hi everyone!

Qwen3.5-Omni is the latest native omni model from the Qwen family. It handles text, images, audio, and video in one system, pushes hard on multilingual speech, and adds a lot of the interaction stuff that actually matters in practice: semantic interruption, realtime voice control, WebSearch, Function Calling, and voice cloning. The audio/video captioning and "audio-visual vibe coding" angle is especially wild.

It is not open-sourced yet. Right now, the way to try it is through the Hugging Face or demos, or through the .

Would love to see this land in the soon!

If there were no security risks, I’d actually like to use Chinese LLMs more. Their prices are low and the performance is pretty good. But that security aspect always worries me a bit.