Everyone is talking about Kimi K3—and for good reason.
Moonshot AI’s new model is not definitively the best in the world. Moonshot itself says K3 still trails Claude Fable 5 and GPT‑5.6 Sol overall, while independent testing tells a similar story.
But K3 is close enough to force a serious comparison. It tops Arena’s frontend-coding leaderboard, ranks first on Artificial Analysis’s AutomationBench, and places second only to Fable 5 on a long-horizon knowledge-work evaluation. These are useful tasks, not just tests of obscure knowledge.
The economics are just as interesting. K3 costs $3 per million input tokens and $15 per million output tokens. GPT‑5.6 Sol costs $5 and $30, while Fable 5 costs $10 and $50.
K3 is expensive compared with other Chinese models, but still much cheaper per token than the American frontier. And prices should fall as inference improves and more providers compete to host it.
One caveat: cheaper tokens do not always translate into equally large savings. Artificial Analysis estimates K3 at $0.94 per task, versus $1.04 for GPT‑5.6 Sol, because models use different numbers of tokens to complete the same work.
The other big moment comes on 27 July, when Moonshot says it will release K3’s full weights. A 2.8-trillion-parameter model will not be cheap or easy to run, but open weights allow providers to compete, companies to customise it and researchers to inspect it.
None of this means businesses should abandon OpenAI or Anthropic. Their models remain stronger overall, with more mature tooling and a better user experience.
But Kimi K3 changes the burden of proof. The question is no longer whether an open-weight Chinese model can approach the US frontier. It is whether the remaining gap justifies the premium.
Increasingly, the answer may be no.
Sources: Moonshot AI · Artificial Analysis · OpenAI · Anthropic
Leave a Comment