Kimi K3 vs Fable 5, GPT-5.6, Grok: 2026 Benchmarks
July 22, 2026

Kimi K3 is Moonshot AI's 2.8-trillion-parameter model, launched July 16, 2026, and billed as the first "open 3T-class" model. On the independent Artificial Analysis Intelligence Index it scores 57 — third overall, behind only Claude Fable 5 and GPT-5.6 Sol, and ahead of Grok 4.5 and Gemini.
TL;DR: Moonshot AI launched Kimi K3 on July 16, 2026, a 2.8T-parameter, natively multimodal model with a 1M-token context window.1 On the independent Artificial Analysis Intelligence Index, K3 scores 57 and ranks #3, comparable to Claude Opus 4.8 and GPT-5.5 but behind Claude Fable 5 (59.9) and GPT-5.6 Sol (58.9); it sits ahead of Grok 4.5 (54) and of Google's newest model, the fast-tier Gemini 3.6 Flash (50).234
Moonshot is refreshingly blunt about this: it says K3 "still trails the most powerful proprietary models, Claude Fable 5 and GPT 5.6 Sol."1 The real story is that an open-weight model — with weights promised by July 27 — now lands inside the closed frontier's top three, tops Arena's Frontend Code arena over Fable 5, and costs less per task than GPT-5.6 Sol or Opus 4.8.25
The catches are just as real: it isn't open yet, its price is a 3–4x jump over Kimi K2.6, it launched with only one (expensive) "max" thinking mode, and Moonshot's own notes flag a user-experience gap versus Fable 5 and a hallucination-rate regression.12
What You'll Learn
- What Kimi K3 is, and why "open 3T-class" is a milestone
- Where K3 ranks against Claude Fable 5, GPT-5.6 Sol, Grok 4.5, and Gemini on a neutral benchmark
- How K3 performs on agentic and coding tasks specifically
- What K3 costs to run — and why cost-per-task, not price-per-token, is the number that matters
- The catches Moonshot itself discloses
- What K3's arrival means for the open-vs-closed AI frontier
Why This Matters Now
Kimi K3 didn't land in a quiet week. Artificial Analysis counted four frontier launches in eight days — Grok 4.5, GPT-5.6, Meta's Muse Spark 1.1, and Kimi K3 — after which six labs now field a model scoring above 50 on its Intelligence Index, up from two in early June.6
The headline from that surge is that the price of near-frontier intelligence has collapsed.6 K3 is the sharpest example: an open-weight challenger scoring within three points of the best closed model on the market.
For anyone building agents, the practical question isn't "who has the single smartest model." It's "how close is the cheap, open, self-hostable option to the frontier — and what do I give up by using it." K3 is the strongest answer that question has had yet.
Google's entry is the outlier. It did ship on July 21 — Gemini 3.6 Flash — but that's a fast-tier model, scoring 50 on the same index, and Google still has no refreshed Pro flagship: Gemini 3.5 Pro was announced at I/O in May and has slipped repeatedly, unshipped as of publication.4 Google's best-scoring shipping model now sits seven points below an open-weight challenger.
What Kimi K3 Is
Kimi K3 is a 2.8-trillion-parameter Mixture-of-Experts model that activates 16 of 896 experts per token.1 Moonshot calls it "the world's first open 3T-class model" — rounding 2.8T up — and says it takes the open-size crown from DeepSeek's 1.6T V4 Pro.15
It's built on two new architectural pieces, Kimi Delta Attention and Attention Residuals, which Moonshot credits for roughly 2.5x better scaling efficiency than Kimi K2.1 It has native multimodal input (text and images, text-only output) and a 1-million-token context window.12
Crucially, this is an agentic model, not a chatbot. Moonshot positions it for "long-horizon coding, knowledge work, and reasoning" — sustaining long engineering sessions, navigating large repositories, and orchestrating terminal tools.1 It ships alongside Kimi Code (a terminal/IDE coding agent) and Kimi Work.1
The showpieces are agentic: in one 48-hour autonomous run K3 designed a small chip using open-source EDA tools, and in another it built "MiniTriton," a compact Triton-like GPU compiler, from scratch.1 These are Moonshot's own demos, so treat them as vendor claims — but they signal where K3 is aimed. This is the same open-agentic lineage as Kimi K2.6's agent-swarm coding push.
The Independent Scorecard
Vendor benchmark tables are hard to compare because each lab runs its own harness. The cleanest apples-to-apples read comes from Artificial Analysis, which runs every model through the same methodology.
Here is where the five models you'd actually weigh land on its Intelligence Index, mid-July 2026:
| Model | AA Intelligence Index | Open weights? |
|---|---|---|
| Claude Fable 5 | 59.9 | No |
| GPT-5.6 Sol | 58.9 | No |
| Kimi K3 | 57 | Pending (by Jul 27) |
| Claude Opus 4.8 | 55.7 | No |
| Grok 4.5 | 54 | No |
| Gemini 3.6 Flash (fast tier) | 50 | No |
Table: Artificial Analysis Intelligence Index (v4.1). Kimi K3, Grok 4.5, and Gemini 3.6 Flash per Artificial Analysis; Fable 5, GPT-5.6 Sol, and Opus 4.8 per the same index as reported on OpenAI's GPT-5.6 page. Gemini 3.6 Flash (July 21) is Google's newest model but a fast-tier one — Google has not shipped a refreshed Pro flagship.234
Chart: Artificial Analysis Intelligence Index (v4.1), higher is better. Source: Artificial Analysis; Fable 5, GPT-5.6 Sol, and Opus 4.8 values per the same index as reported on OpenAI's GPT-5.6 launch page (July 2026).
The takeaway: K3 is a genuine top-three model on a neutral test, clears Grok 4.5 by three points, and sits seven points above Google's best shipping model. Among open-weight models it has no rival — GLM-5.2 scores 51 and DeepSeek V4 Pro 44 — though at 2.8T parameters it is far larger than either.2
Agentic and Coding Performance
Because the Intelligence Index blends nine evaluations, the agentic-specific numbers are where K3's coding pitch lives — and Artificial Analysis breaks those out.
On GDPval-AA v2, an agentic real-world task evaluation, K3 reaches an Elo of 1668 — up from Kimi K2.6's 1190 — passing GLM-5.2 (1514), GPT-5.5 (1494), and Claude Opus 4.8 (1600), while still trailing Claude Fable 5 (1760).2
Chart: GDPval-AA v2 agentic real-world task Elo ratings, higher is better. Source: Artificial Analysis, "Kimi K3 achieves #3 in the Artificial Analysis Intelligence Index" (July 17, 2026).
On AA-Briefcase, a private long-horizon knowledge-work evaluation, K3 posts an Elo of 1547 — a 732-point jump over K2.6 — placing second behind only Fable 5.2 And on AutomationBench-AA, Artificial Analysis's version of Zapier's agentic-workflow test, K3 scores 53% and takes the #1 spot.2
Its most eye-catching win is on Arena's Frontend Code arena, where K3 became the leading model, surpassing even Claude Fable 5.5 For a model that Moonshot admits trails Fable 5 overall, topping a head-to-head coding arena is a real result.
Moonshot's own harness-specific coding numbers (SWE Marathon, Terminal-Bench 2.1, FrontierSWE and others) are strong but reported under a mix of KimiCode, Claude Code, and Codex harnesses, so they aren't directly comparable across vendors — read them as Moonshot-reported rather than independently confirmed.1
The Cost Story
K3's pricing is where the "cheap frontier" narrative gets complicated. The first-party API is $3 per million input tokens and $15 per million output, with cached input discounted 90% to $0.30.12
That is a 3–4x jump over Kimi K2.6's $0.95/$4 — its output price alone rose from $4 to $15 — making K3 the most expensive model a Chinese lab has released, priced at the level of Anthropic's Claude Sonnet series.5 The days of Moonshot undercutting everyone on sticker price are over.
But sticker price isn't the bill. As covered in our breakdown of why agent cost-per-task beats price-per-token, what matters for agents is tokens consumed per finished task. On that measure K3 looks better:
| Model | Input ($/M) | Output ($/M) | AA cost per task |
|---|---|---|---|
| Kimi K3 | $3 | $15 | $0.94 |
| GPT-5.6 Sol | $5 | $30 | $1.04 |
| Claude Opus 4.8 | $5 | $25 | $1.80 |
| Grok 4.5 | $2 | $6 | — |
| Claude Fable 5 | $10 | $50 | — |
Table: First-party list prices; "AA cost per task" is Artificial Analysis's average cost to complete one Intelligence Index task, reported for a subset of models.1237
Chart: Average cost (USD) per Artificial Analysis Intelligence Index task, lower is better. Source: Artificial Analysis, "Kimi K3 achieves #3 in the Artificial Analysis Intelligence Index" (July 17, 2026).
The Catches
The launch framing is stronger than the shipping reality, and Moonshot is unusually candid about it.
It isn't open yet. At launch K3 is API-only; weights are promised "by July 27, 2026."1 The "open" headline is a commitment, not today's status.
One thinking mode, and it's the expensive one. K3 launched with only "max" reasoning effort, with lower-effort modes to follow.1 Simon Willison's simple SVG test burned 13,241 reasoning tokens on a single prompt — the max-only default makes casual use pricey and slow.5
Moonshot flags a UX gap. Its own notes concede "a noticeable gap in user experience compared with Claude Fable 5 and GPT 5.6 Sol," plus two agentic quirks: K3 can be unstable if a harness drops its thinking history mid-session, and it can be "excessively proactive," making decisions on your behalf on ambiguous tasks unless you constrain it explicitly in the system prompt or AGENTS.md.1
Accuracy up, but hallucinations up too. Artificial Analysis found K3's factual accuracy improved over K2.6 (33% to 46%), yet its hallucination rate regressed, rising from 39% to 51%.2 For agentic use, that's a meaningful reliability caveat.
What It Means
Kimi K3 doesn't dethrone the closed frontier — Moonshot says as much.1 What it does is compress the gap to about three points on a neutral index, while undercutting the leaders on cost per task and promising open weights within days.2
For teams building agents, that changes the calculus. The self-hostable, open option is no longer a generation behind; it's a rounding error behind on capability and ahead on cost, with the trade-offs (a UX gap, a higher hallucination rate, huge 2.8T hardware needs) now the deciding factors rather than raw intelligence.
It also reframes the competitive map. This is the same dynamic driving the whole July surge and the managed-runtime land grab from AWS, Google, and Alibaba: near-frontier intelligence is becoming a commodity, and the moat is moving to price, tooling, and reliability. If K3's weights ship on schedule July 27, the best open model in the world will be one you can run yourself.
The Bottom Line
Kimi K3 is the moment the open frontier stopped being a step behind. A 2.8T open-weight model now ranks third on a neutral intelligence test, beats the best closed model in a coding arena, and does it for less per task — while its maker openly admits it isn't quite Fable-5 or GPT-5.6-Sol class on polish and reliability.12
Watch July 27. If the weights ship as promised, the calculus for every team choosing between a closed API and a self-hosted agent just got a lot more interesting.
References
Footnotes
-
Moonshot AI, "Kimi K3: Open Frontier Intelligence" (official blog). 2.8T parameters, Kimi Delta Attention + Attention Residuals, 16-of-896 MoE, native multimodal input, 1M context, pricing ($3 input / $15 output / $0.30 cached), weights "by July 27, 2026," "trails the most powerful proprietary models, Claude Fable 5 and GPT 5.6 Sol," coding case studies, and limitations (thinking-history sensitivity, excessive proactiveness, UX gap). https://www.kimi.com/blog/kimi-k3 ↩ ↩2 ↩3 ↩4 ↩5 ↩6 ↩7 ↩8 ↩9 ↩10 ↩11 ↩12 ↩13 ↩14 ↩15 ↩16 ↩17 ↩18 ↩19 ↩20 ↩21
-
Artificial Analysis, "Kimi K3 achieves #3 in the Artificial Analysis Intelligence Index, comparable to Opus 4.8 and GPT-5.5," July 17, 2026. Index score 57 (#3); GDPval-AA v2 Elo 1668; AA-Briefcase Elo 1547; AutomationBench-AA 53% (#1); cost per task $0.94 vs GPT-5.6 Sol $1.04 and Opus 4.8 $1.80; 21% fewer output tokens than K2.6; accuracy 33%→46%; hallucination rate 39%→51%; open peers GLM-5.2 (51) and DeepSeek V4 Pro (44). https://artificialanalysis.ai/articles/kimi-k3-achieves-3-in-the-artificial-analysis-intelligence-index-comparable-to-opus-4-8-and-gpt-5-5 ↩ ↩2 ↩3 ↩4 ↩5 ↩6 ↩7 ↩8 ↩9 ↩10 ↩11 ↩12 ↩13 ↩14 ↩15 ↩16 ↩17 ↩18 ↩19 ↩20
-
OpenAI, "GPT-5.6," July 9, 2026 — Artificial Analysis Intelligence Index v4.1 values used for Claude Fable 5 (59.9), GPT-5.6 Sol (58.9), Claude Opus 4.8 (55.7), GPT-5.5 (54.8), and Gemini 3.1 Pro Preview (46.5), plus GPT-5.6 pricing. https://openai.com/index/gpt-5-6/ ↩ ↩2 ↩3
-
Google's newest model as of publication is Gemini 3.6 Flash, released July 21, 2026 — a fast-tier model scoring 50 on the Artificial Analysis Intelligence Index (input/output $1.50/$7.50 per 1M tokens). Its predecessor Pro model, Gemini 3.1 Pro Preview, scores ≈46 on the current index; Google's next Pro flagship, Gemini 3.5 Pro, was announced at Google I/O (May 20, 2026) but slipped repeatedly and was not shipped as of publication, and there is no Gemini 3.6 Pro. https://artificialanalysis.ai/models/gemini-3-6-flash ↩ ↩2 ↩3 ↩4 ↩5
-
Simon Willison, "Kimi K3, and what we can still learn from the pelican benchmark," July 16, 2026 — launch details, #1 on Arena's Frontend Code arena over Claude Fable 5, pricing context (4x jump over Kimi K2.6's $0.95/$4; most expensive Chinese-lab model), and max-only thinking effort at launch. https://simonwillison.net/2026/Jul/16/kimi-k3/ ↩ ↩2 ↩3 ↩4 ↩5 ↩6 ↩7 ↩8
-
Artificial Analysis, "Four frontier launches in eight days: six labs now field a model above 50 on the Artificial Analysis Intelligence Index," July 17, 2026. Grok 4.5, GPT-5.6, Muse Spark 1.1, and Kimi K3 launched within eight days; six labs above 50 (up from two in early June); the price of near-frontier intelligence has collapsed. https://artificialanalysis.ai/articles/four-frontier-launches-in-eight-days-six-labs-now-field-a-model-above-50-on-the-artificial-analysis-intelligence-index ↩ ↩2
-
Competitor list pricing from official sources: xAI Grok 4.5 ($2 input / $6 output, https://x.ai/news/grok-4-5); Anthropic Claude Fable 5 ($10 / $50) and Claude Opus 4.8 ($5 / $25), https://platform.claude.com/docs/en/about-claude/pricing. ↩

