SWE-2 Hits 50.0% on FrontierCode 1.1 Main as Cost Drops 64%
Updated
Updated · cognition.com · Sep 10
SWE-2 Hits 50.0% on FrontierCode 1.1 Main as Cost Drops 64%
3 articles · Updated · cognition.com · Sep 10
Summary
Cognition introduced SWE-2 as its new flagship coding model, saying it scores 50.0% on FrontierCode 1.1 Main—within one point of Fable 5.1—while costing 64% less.
A 2.8T-parameter Kimi K3 base and a new single-run RL method drove the gain, training multiple reasoning-effort levels together with linear cost penalties tuned to the model’s Pareto frontier.
Against prior models, SWE-2 posted 73.0% on DeepSWE 1.1 and 92.8% on Terminal-Bench 2.1, beating SWE-1.7 and Grok 4.6 on both score and cost while nearing GPT-6 Astra at roughly a quarter of the price.
On practical efficiency, SWE-2 medium matched or exceeded SWE-1.7 while using 58% fewer turns and 81% lower average cost on FrontierCode 1.1 Main, making its first real edit after a median 18 steps versus 48.
Starting today, SWE-2 is available in Devin Desktop and CLI, with rollout to Devin Web and Fusion, as Cognition argues stronger coding judgment can improve both reliability and cost efficiency.