Kimi K3 Prices Like Claude Sonnet Now. The Cheap Open-Weight Era Just Ended.
Moonshot's 2.8-trillion-parameter Kimi K3 landed at $3/$15 per million tokens, matching Claude Sonnet, and ships with exactly one reasoning gear: max. Two signals worth reading before you swap models to save money.
On July 16, Moonshot AI shipped Kimi K3, a 2.8-trillion-parameter open-weight model, and priced it at $3 per million input tokens and $15 per million output tokens. That is not a discount. It is what Anthropic charges for Claude Sonnet, dollar for dollar. If your mental model of Chinese open-weight labs is "frontier-adjacent capability at a fifth of the price," K3 is the release that breaks it.
From loss leader to price-matcher
Kimi K2.6, Moonshot's previous flagship, charged $0.95 per million input tokens and $4 per million output tokens. K3 is roughly 3x the input cost and nearly 4x the output cost of its own predecessor. It still undercuts Claude Opus by about 40% on both ends, but it no longer undercuts Sonnet at all. It sits exactly on top of it.
- K2.6 pricing: $0.95 input / $4 output per million tokens
- K3 pricing: $3 input / $15 output per million tokens (cache-hit input drops to $0.30)
- K3 vs Claude Sonnet: exact parity on both input and output
- K3 vs Claude Opus: roughly 40% cheaper on both ends
- Artificial Analysis Elo on long-horizon knowledge work: 1547, described as +732 points over K2.6, trailing only Claude Fable 5
- Context window: 1M tokens, flat pricing across the full window
The framing that dominated the last two years of open-weight coverage, ours included when we compared GLM-5.2 against DeepSeek V4, was that open weights meant cheap by default and closed weights meant a capability tax. K3 is evidence that the relationship was never structural. It was a market-entry strategy. Once a lab believes its model is genuinely close to frontier, it prices like it.
One reasoning gear, stuck on max
The more interesting number isn't the sticker price, it's the ratio underneath it. K3 currently ships with a single reasoning effort setting, and testing by outside reviewers found it burning 13,241 reasoning tokens to produce 3,417 tokens of visible response, a nearly 4-to-1 ratio of thinking to answer. Every one of those reasoning tokens bills as output at $15 per million. There is no low-effort mode to fall back to when a prompt doesn't need it.
It only has one reasoning effort right now, 'max,' and it shows. — Simon Willison, on Kimi K3's token consumption
We wrote about why effort settings change the shape of a bill, not just the shape of an answer, when we looked at whether thinking effort actually matters. K3 is a live case of what happens when a lab hasn't built that dial yet: every query, however trivial, pays the max-effort reasoning tax whether the task needed it or not.
What this means before you switch
None of this makes K3 a bad model. Early benchmark placement has it just below Claude Fable 5 and GPT-5.6 Sol, which is a real result for an open-weight release, and it reportedly leads a third-party frontend coding benchmark outright. But "open weight" stopped being a reliable shorthand for "cheap" the moment Moonshot priced it at Sonnet parity, and the missing effort dial means the effective cost per task can run higher than the headline number suggests until Moonshot ships a lighter mode.
Check the output price, not just the input price, since reasoning tokens bill as output. Check whether there's an effort or thinking-budget control you can turn down. And run the same prompt blind against whatever you're currently using before trusting a benchmark Elo number to predict your bill.
We already run Kimi's earlier generation head to head against Claude in our coding arena, and we'll add K3 to the rotation once it's stable enough to test the way we test everything else: same one-shot prompt, live output, blind vote, no cherry-picked screenshots. Go compare current models yourself rather than taking either Moonshot's Elo claim or our word for it.
Readers following the same story from the European side, where Moonshot's pricing move landed as evidence that China's compute-constrained labs are chasing margin as hard as capability, can track it in Polish at nowosci.ai.
Don’t take the post’s word for it
The arena runs every model’s real output live. Pick a challenge, go blind, and cast a vote that counts in the public tally.
Open the arena