SWE-2 is Cognition's new coding model, post-trained from Kimi K3 with RL that optimizes for cost and capability at the same time. It hits 50.0% on FrontierCode 1.1 Main, within a point of Fable 5.1 at 64% less, and lands within a few points of GPT-6 Astra at a quarter of the cost. Compared to SWE-1.7 it takes 58% fewer turns and costs 81% less while scoring higher. Available now in Devin Desktop and CLI.
SWE-2 is Cognition's new coding model, built to sit near the frontier without frontier pricing.
The problem: coding agents get smarter by getting more expensive, and often by burning tokens. Their own SWE-1.7 had this issue, users said it over-explored simple tasks.
What they did: post-trained Kimi K3 with RL that puts the actual dollar cost of a run into the training reward, and trains all three effort levels (medium, high, max) in one go instead of stitching separate models together. So the whole cost-performance curve moves up, not just the top end.
The numbers:
50.0% on FrontierCode 1.1 Main, within a point of Fable 5.1 at 64% less cost
Within a few points of GPT-6 Astra at a quarter of the price
92.8% on Terminal-Bench 2.1, top of their comparison table
vs SWE-1.7: 58% fewer turns, 81% cheaper, higher score. First code edit after 18 steps instead of 48
Also worth noting: better end-to-end test writing, and it re-derives conclusions when you push back instead of just agreeing with you.
Good fit if you're running coding agents at volume and per-task cost is an actual line item. Available now in Devin Desktop and CLI, rolling out to Web and Fusion.
SWE-2 is Cognition's new coding model, built to sit near the frontier without frontier pricing.
The problem: coding agents get smarter by getting more expensive, and often by burning tokens. Their own SWE-1.7 had this issue, users said it over-explored simple tasks.
What they did: post-trained Kimi K3 with RL that puts the actual dollar cost of a run into the training reward, and trains all three effort levels (medium, high, max) in one go instead of stitching separate models together. So the whole cost-performance curve moves up, not just the top end.
The numbers:
50.0% on FrontierCode 1.1 Main, within a point of Fable 5.1 at 64% less cost
Within a few points of GPT-6 Astra at a quarter of the price
92.8% on Terminal-Bench 2.1, top of their comparison table
vs SWE-1.7: 58% fewer turns, 81% cheaper, higher score. First code edit after 18 steps instead of 48
Also worth noting: better end-to-end test writing, and it re-derives conclusions when you push back instead of just agreeing with you.
Good fit if you're running coding agents at volume and per-task cost is an actual line item. Available now in Devin Desktop and CLI, rolling out to Web and Fusion.
Try it: the SWE-2 writeup
P.S. I hunt the latest and greatest launches in tech, SaaS and AI, follow to be notified → @rohanrecommends