Kimi K3 - The world's first open 3T-class model
by•
Kimi K3 is a 2.8T-parameter open model featuring native vision capabilities, a 1-million-token context window, and Moonshot AI's Kimi Delta Attention and Attention Residuals architectures. Built as the world's first open 3T-class model, it delivers frontier-level performance in long-horizon coding, compiler development, digital creation, and scientific reasoning, outperforming previous open models in scaling efficiency and agentic capabilities.


Replies
Mom Clock
Hi everyone! 👋
While most of the industry is focused on scaling compute, Moonshot is focused on scaling intelligence.
Instead of just scaling up model parameters (which they did anyway—hitting a massive 2.8T parameters!), Kimi K3 introduces a 2.5x improvement in scaling efficiency using their custom Kimi Delta Attention and Attention Residuals architectures.
It is the world’s first open 3T-class model, and its long-horizon agentic workflows are wild:
🤯 1M Context + Native Vision: Built for massive data ingestion, from video and screens to complex systems.
💻 Autonomous Engineering: It built its own GPU compiler (MiniTriton) and optimized complex GPU kernels competitively with the strongest proprietary models.
🧠 Chip Design & Astrophysics: In a single 48-hour run, it autonomously designed and verified its own microchip. It also bridged astrophysics literature with executable code to reproduce complex stellar relations.
It's impressive to see a 2.8T model with this level of long-horizon reasoning being open-sourced.
How do you see open-source weights of this scale shifting the balance with proprietary AI?
please let there be a smaller distill for the rest of us.
scaling efficiency claim is the interesting part honestly, the chip design/compiler stuff feels more like a flex until someone outside moonshot replicates it.
great, is there a free version to test it first ?
Tried to put K3 through a real test before commenting instead of just reading the benchmarks. Signed in, picked K3 Max from the model list, and asked it a nested Navigator Hero animation question straight from my Flutter client work. Two attempts, both came back with Task paused due to system peak. Honestly that says more about launch day demand than about the model, but I did not get my answer yet.
Two things worth flagging for the team. The composer still defaults to K2.6 Fast for signed-in users, so a lot of people arriving from this page and typing straight into the box are probably testing the old model without realizing it. And when K3 pauses under load it would be great to see queue position or an ETA instead of a bare retry link.
The open weights angle is the genuinely exciting part for me as an agency dev. A 1M context window plus native vision at open 3T class scale changes what small teams can even consider self-hosting. Congrats on shipping, upvoted, and I will retry the Flutter question once the servers cool down.
Said I would retry once the servers cooled down, so here is the retry.
First, credit where it is due. Signed in, the composer now defaults to K3 Max instead of K2.6. That was the main thing I flagged and it is fixed, which is a fast turnaround on launch week feedback.
The rest did not go as well. My paused task from yesterday still showed Task paused due to system peak with a Continue Task button. Clicking it did not resume anything. It dropped me into a fresh empty chat, the model reverted to K2.6, and the original thread was gone. So the pause is not really recoverable, it just looks like it is.
One more thing worth knowing. Signed out, the model picker is locked to K2.6 with no K3 option at all. Anyone landing here from Product Hunt and typing into the box without an account is still testing the previous model and has no way to tell.
Still chasing my answer on that Flutter question, will report back when I get it. The open weights work is the part I am most interested in either way.
Coming back to correct myself on two things.
I said the composer defaults to K3 Max now for signed in users. That was true on the account I was using, but I tried a fresh one and it came up on K2.6 with K3 sitting there unselected. So it depends on the account, not fixed everywhere.
I also said Continue Task dumps you into a new chat and loses the thread. Retested and it stayed put, so that was almost certainly my session expiring rather than the button. My mistake.
The thing I actually learned is more useful. Free account, K3 Max picked on purpose, same Flutter question. Task paused due to system peak, agent credits refunded. Hit Continue Task and got the honest version: too many people are chatting with Kimi right now, subscribe to enter a dedicated priority queue.
That reads very differently from a launch day spike. Three tries across two days and two accounts, and I have still not got K3 to answer me once. The page says Free, so people come in expecting to try the flagship and instead meet a queue they cannot leave without paying. I would just say it up front. Something like K3 is priority queued for subscribers at peak times. People take a limit fine when they can see it coming.
Still want to try the model properly. Still do not have my Flutter answer.
Tycoon AI
I'm really impressed by its front-end coding capabilities.
Hope it gets open-sourced soon!
Uploaded a dense methodology section from a paper and it gave me a clean summary in seconds, way better than skimming for ten minutes. Going to keep using it for my lit reviews.
Finally tried Kimi for breaking down a dense research paper and it actually pulled out the methodology cleanly in like 10 seconds, saved me a real headache honestly.
honestly the paper interpretation part was pretty solid for me, kind of speeds up the whole reading process when you just need the gist.
That's impressive scale. How does it perform on long-horizon coding tasks versus other open models?