Qwen3.8-Flash-Next is a 125B multimodal MoE with only 6B active parameters and a new architecture built around QSA, Gated Residual, N-gram embeddings, and Muon. Its open weights give an early look at the architecture Qwen is building toward Qwen4.
Reviewers describe Qwen3 as a practical, strong all-round model, especially for coding, prototyping, long-running agent tasks, and cases where other AI tools fall short. Users praise its fast iteration, solid response quality, and, in one comparison, better speed and lower cost than GPT-4o, though another user reports latency bottlenecks for real-time chat. Founder feedback is narrower but consistent: makers of Knowlify, Zesty by DoorDash, and Juno cite creativity, agentic search, and formatting-heavy workflows.
Qwen has done this once before. Qwen3-Next gave everyone an early look at the architecture that later showed up in Qwen3.5.
Qwen3.8-Flash-Next is doing the same for Qwen4.
It’s a 125B model with only 6B active parameters, and a lot of the new design is about getting more capability without dragging compute up with it. QSA makes long-context retrieval cheaper, Gated Residual gives information more paths through the network, and the new N-gram memory adds capacity with very little per-token compute.
Qwen keeps putting these architecture previews out as real models people can actually run, well before the next generation arrives.
for qwen latest model: qwen.3.8 - Agentic performance is great. I use for database agent, gemma, mistral, gemini and qwen. qwen latest model my favorite for the long running task(60-100 steps).
What needs improvement
fast performance (1)
Generation Speed and Latency, while highly intelligent, qwen3.8 can experience latency bottlenecks in raw tokens-per-second throughput compared to highly optimized smaller models or specialized "thinking" samplers, which can be a friction point for real-time chat applications.
vs Alternatives
some benchmark says: qwen3.8 is them same claude-opus-4.6, but when start to work qwen, I feel it(I used before opus and sonnet. not opus-5
I’ve been using Qwen for building a simple code and website generator, and it works really well for fast iterations. Great for prototyping and lightweight generation.
What needs improvement
I need more on the history pages, a section when we can re-edit the input/process/output with easy UX. Basically, better handling of edge cases without extra prompting
vs Alternatives
I choose Qwen because it’s fast, lightweight, and great for turning ideas into simple, working code or websites. It was also the first web-based tool I explored for code generation, which made it easy to start prototyping right away.
Great launch! Qwen has been incredibly useful, especially when I reach a point where other AI services can no longer technically deliver what I need. I’m also excited to see it matching the “big players” in benchmark results. 2026 is shaping up to be very interesting.
Flowtica Scribe
Hi everyone!
Qwen has done this once before. Qwen3-Next gave everyone an early look at the architecture that later showed up in Qwen3.5.
Qwen3.8-Flash-Next is doing the same for Qwen4.
It’s a 125B model with only 6B active parameters, and a lot of the new design is about getting more capability without dragging compute up with it. QSA makes long-context retrieval cheaper, Gated Residual gives information more paths through the network, and the new N-gram memory adds capacity with very little per-token compute.
Qwen keeps putting these architecture previews out as real models people can actually run, well before the next generation arrives.
And the weights are already up :) (with a license worth reading)