K3's published results are strong across reasoning, coding, agentic, and search benchmarks. Below is what has actually been reported, with sourcing, followed by the caveats that matter — because a benchmark table is evidence for running your own test, not a substitute for it.
Published scores
| Benchmark | What it measures | Kimi K3 |
|---|---|---|
| GPQA Diamond | Graduate-level science reasoning | 93.5 |
| Terminal-Bench 2.1 | Multi-step terminal and tool use | 88.3 |
| BrowseComp | Multi-hop web research | 91.2 |
| DeepSearchQA (F1) | Search-based question answering | 95.0 |
| DeepSWE | Software engineering tasks | 67.5 |
| SWE Marathon | Long-horizon software work | 42.0 |
| Program Bench | Program synthesis | 77.8 |
| MMMU-Pro | Multimodal reasoning | 81.6 / 83.4 |
| Frontend Code Arena | Blind human preference on generated UI | 1st, 1,679 points |
Figures are as reported on Moonshot's model card, except the Frontend Code Arena result, which is third-party and decided by blind developer voting.
The result that matters most
The Frontend Code Arena placement is the strongest single claim K3 has, for a specific reason: it is blind human preference, not an automated scorer. Developers compared generated interfaces without knowing which model produced them, and K3 came first at 1,679 points — ahead of Claude Fable 5, a model priced at $10/$50 per million tokens against K3's $3.00/$15.00.
Blind preference is harder to game than a benchmark with a public test set. It is also narrow: it measures generated frontend code, not correctness, accessibility, maintainability, or anything outside that domain.
Where K3 sits overall
On third-party aggregate rankings — the Artificial Analysis Intelligence Index — K3 places in the top five of all frontier models, behind Claude Fable 5 and GPT-5.6 Sol, with a reported Intelligence score of 57.1 and a Coding score of 76.2.
It is the first open-weight model to reach the 3-trillion-parameter class, and the strongest open-weight entry on that index at the time of release.
How to read all of this
Four caveats, in order of importance:
- Vendor-published scores are self-reported. Most of the table above comes from Moonshot's own model card. That is normal practice across the industry and it still means the evaluation conditions were chosen by the vendor.
- Index scores are not comparable across versions or across time. Aggregate indices get revised. A number from one release cannot be set against a number from another, which is why we do not compare K3's index score against Kimi K2's.
- No published head-to-head exists against Claude Opus 5 or GPT-5.6 Sol. Anyone claiming a definitive ranking between those models is interpolating.
- Your workload is not in this table. Twenty representative tasks from your real queue will tell you more than every row above combined.
More on Kimi K3
Start with the complete Kimi K3 guide for the overview, or go deeper:
Ready to go deeper?
Read the full Kimi K3 guideFrequently Asked Questions
What are Kimi K3's benchmark scores?
The headline published figures: GPQA Diamond 93.5, Terminal-Bench 2.1 88.3, BrowseComp 91.2, DeepSearchQA 95.0 F1, DeepSWE 67.5, SWE Marathon 42.0, Program Bench 77.8, and MMMU-Pro 81.6/83.4. It also placed first on the Frontend Code Arena at 1,679 points. Most are vendor-reported; the arena result is third-party.
Is Kimi K3 the best open-weight model?
At release it was the strongest open-weight entry on the Artificial Analysis Intelligence Index and the first open-weight model in the 3-trillion-parameter class. "Best" depends on your task and on how much you weight the licence terms, which are more restrictive than MIT.
How does Kimi K3 rank against closed models?
On the Artificial Analysis Intelligence Index it places in the top five overall, behind Claude Fable 5 and GPT-5.6 Sol. On the Frontend Code Arena specifically it placed first, ahead of Fable 5. There is no published head-to-head against Claude Opus 5.


