← All posts
Analysis·5 min read

Kimi K3 Beat Fable 5 in One Category. It Lost the Other 34.

Kimi K3 topped a single blind-voted leaderboard category and the coverage read it as China closing the gap on Anthropic. Moonshot's own 35-benchmark battery, GDPval scores, and hallucination rate tell a narrower story.


On July 16, Moonshot AI shipped Kimi K3, a 2.8 trillion parameter open-weight model, and within 48 hours the coverage had gone from "impressive open model" to "China has leveled up" to a semiconductor stock selloff. The receipt behind that story is one category, one leaderboard, one win rate. The other 34 benchmarks Moonshot ran against Fable 5 and GPT-5.6 Sol tell a different story, and it's worth separating what got measured from what got written.

What actually shipped

K3 is a mixture-of-experts model with 896 experts, of which only 16 fire on any given token, about 1.8% of the network active per step. That's how Moonshot gets a 2.8T-parameter model down to workable per-token compute. It ships with two new architecture pieces, Kimi Delta Attention (a hybrid linear-attention mechanism) and Attention Residuals (a replacement for standard residual connections meant to keep scaling gains consistent), plus a 1 million token context window and native vision. Full weights aren't out yet: Moonshot says they land July 27, so most of what's circulating right now is Moonshot's own numbers plus whatever independent labs could test against the hosted API in the first three days.

The category it actually won

The number driving headlines is a single leaderboard: Arena's Frontend Code category, where K3 posted a 76% pairwise win rate against Fable 5's 63% in blind voting, enough to take the top spot. That's a real result in a real category, and blind pairwise voting on code output is close to how we run the coding board here. But it's one slice of frontend work, not a verdict on the model.

  • Moonshot's own self-reported battery: K3 won roughly 7 of 35 benchmarks tested; Fable 5 won more individual tests than any other model in the same set.
  • GDPval v2 (Elo, real-world tasks across 44 occupations): K3 scored 1,668, ahead of GPT-5.5 and Opus 4.8, behind Fable 5's 1,760.
  • Hallucination rate rose from 39% to 51% versus K3's predecessor, even as raw accuracy improved.
  • Pricing: $3/M input and $15/M output tokens, averaging about $0.94 per task, close to GPT-5.6 Sol's $1.04 and cheaper than Opus 4.8's $1.80, but far above DeepSeek V4 Pro's roughly $0.04.

Put those together and the honest summary is: K3 is a genuinely strong, genuinely large open-weight model that leads in one blind-voted category and trails both closed frontier models on the broader batteries, while costing meaningfully more than the open-weight competition it's supposed to be undercutting. That's a solid release. It is not a coronation.

How one leaderboard becomes "China leveled up"

The market didn't read it as a solid release. Semiconductor stocks slid the same week, and outlets ran with framing like "China's AI Has Leveled Up" and warnings that the US lead was being threatened. Some of that conflated with an unrelated story, Google reportedly delaying Gemini 3.5 Pro over missed internal targets, which hit Alphabet's stock in the same news cycle and made the whole week read as a single "US AI is struggling, China is surging" narrative when the two stories don't actually share a cause.

That connection has been mostly severed now - Simon Willison, on why his long-running pelican-riding-a-bicycle test no longer tracks overall model quality across labs.

Willison's point generalizes past pelicans. A single benchmark, even a well-designed blind arena category, stops being a reliable stand-in for "which model is better" the moment every lab starts optimizing toward it. Frontend Code Arena is worth watching. It is not a substitute for testing the model on your own task.

What to actually check before you switch

If you build web UI and want a frontend-specific signal, K3's win rate there is a legitimate reason to try it. If your work looks like the other 34 categories Moonshot tested, or like GDPval-style mixed real-world tasks, the receipts say Fable 5 and GPT-5.6 Sol are still ahead, and K3 costs more than most of the open-weight field for that edge. We already covered Kimi's pricing shift when the K2 family started closing the gap with closed models; K3 continues that trend rather than reversing it.

Full weights land July 27. Once they're up, run your own frontend prompt through the coding arena blind, against Fable 5 and GPT-5.6 Sol, and see if the 76-to-63 split holds on your task instead of Moonshot's. That's the same reason we keep saying output tokens aren't quality: a headline number from someone else's test run tells you less than one blind vote on your own prompt.

Chinese open-weight releases are landing fast enough now that a new flagship shows up most weeks. If you want that beat covered in Polish, nowosci.ai tracks the same releases as they land. We'll keep testing them the way we test everything else here: live, blind, one prompt at a time.

Don’t take the post’s word for it

The arena runs every model’s real output live. Pick a challenge, go blind, and cast a vote that counts in the public tally.

Open the arena