Claude Opus 5 Outscores GPT-5.6 Sol on Cognition's Code-Mergeability Benchmark
Cognition's FrontierCode 1.1 leaderboard has Claude Opus 5 beating OpenAI's GPT-5.6 Sol on whether a maintainer would actually merge the generated patch, 53.4% to 47.5%, while costing 32% less per rollout.
Cognition's FrontierCode 1.1 leaderboard, which grades whether a human maintainer would actually merge a model's proposed patch rather than just whether it executes, put Claude Opus 5 ahead of GPT-5.6 Sol in results published this week.
- Claude Opus 5: 53.4% mergeability score at $4.30 per rollout, run at medium effort
- GPT-5.6 Sol: 47.5% mergeability score at $6.30 per rollout, run at max effort
- Opus 5's edge: 5.9 points higher mergeability at 32% lower cost per rollout
The two ran at different effort tiers, so it isn't a matched-compute comparison of raw capability, only of the configurations developers are likely to actually pick. It's the second independent benchmark this week to land in Opus 5's favor since Anthropic launched it at roughly half of Fable 5's price. Compare both yourself in our coding arena.