Moonshot AI launched Kimi K3 on July 16, 2026 and committed the full weights to HuggingFace eleven days later, on July 27. At 2.8 trillion parameters it is the first open model in the 3-trillion class — a scale that used to exist only inside OpenAI, Anthropic, and Google. The open-weight drop hit #1 on HuggingFace's trending board within an hour, and Moonshot's valuation was already being repriced at $50 billion (pre-IPO G round) by the time the weights landed.
The pitch is straightforward: open weights at frontier-adjacent quality, with native vision and a 1M-token context, priced at $3 input / $15 output per million tokens — roughly a third of Claude Fable 5's $10/$50. The honest caveat, which Moonshot itself puts in the model card, is that K3 still trails Fable 5 and GPT-5.6 Sol on the broad agentic indices. What it wins — and what builders should actually care about — is a specific bundle: open weights, 1M context, native multimodality, and a price point that makes overnight batch work economically reasonable.
What's New

2.8T total, 104B active — the largest open model ever. K3 is a sparse Mixture-of-Experts with 896 routed experts, 16 activated per token, plus 2 shared experts. That is 2.8 trillion parameters on disk and roughly 104 billion active per forward pass — so it trains like a giant and serves like a mid-size model. The full weights download is about 1.56 TB, which means "open weights" here mostly means provider choice and price competition, not self-hosting in your garage.
Native vision, not a bolt-on. A single MoonViT-V2 encoder (401M parameters) feeds images and video into the same backbone that handles text — no post-hoc alignment stage. K3 can iterate code against live screenshots, edit video by watching its own cuts, and build playable games from a one-paragraph concept. This is the difference between "multimodal input" and "vision in the loop."
1M context, served at $3/M. K3 supports 1,048,576 tokens. Moonshot's serving stack (Mooncake disaggregated inference) reports a >90% cache-hit rate on coding workloads, which is why the cache-hit input price is $0.30/M — ten times cheaper than a cache miss. Flat across the full 1M window, no long-context surcharge.
Quantization-aware from SFT. Weights are natively MXFP4 and activations MXFP8 — the model was trained that way, so there is no train-to-inference mismatch. Moonshot also contributed a KDA-aware prefix-cache implementation to vLLM, which matters because KDA breaks conventional prefix caching.
Architecture

K3's architecture is where the 2.5× scaling-efficiency claim over K2 actually comes from. Five pieces do the work:
| Component | Detail |
|---|---|
| Total / active params | 2.8T / 104B |
| Layers | 93 (69 KDA + 24 Gated MLA) |
| Experts | 896 routed, 16 active, +2 shared |
| Attention | KDA (Kimi Delta Attention) + Gated MLA + Attention Residuals |
| MoE routing | Stable LatentMoE, Quantile Balancing |
| Activations | SiTU-GLU (Sigmoid Tanh Unit) |
| Vision encoder | MoonViT-V2 (401M) |
| Context | 1,048,576 tokens |
| Quantization | MXFP4 weights / MXFP8 activations (QAT from SFT) |
KDA + Gated MLA is the long-context trick. Kimi Delta Attention is a hybrid linear-attention layer; Gated MLA is the Multi-head Latent Attention used on the dense-attention layers. Interleaved across 93 layers, they give K3 the KV-cache efficiency that makes 1M context practical at 2.8T scale. Conventional prefix caching does not work on KDA, hence the vLLM contribution.
Attention Residuals improve depth scaling. Instead of accumulating residual stream uniformly, AttnRes selectively retrieves representations across depth — a mechanism Moonshot credits for stable training past the trillion-parameter regime.
Stable LatentMoE fixes routing at 896 experts. With 16-of-896 routing, expert imbalance can tank throughput. Quantile Balancing derives expert allocation directly from router-score quantiles (no heuristic updates), and a fully balanced expert-parallel training method with static shapes removes host synchronization from the critical path. Per-Head Muon optimizes attention heads independently.
SiTU-GLU activation. Sigmoid Tanh Unit replaces the usual GELU/SwiGLU with tighter activation control — the short version is that it trains more stably at this scale.
Benchmarks

Moonshot's card runs everything at max reasoning effort, temperature 1.0, and uses its own Kimi Code harness for K3 while comparing against Claude Code / Codex harnesses. Treat agentic scores as "model + scaffold."
| Benchmark | Kimi K3 | Claude Fable 5 | GPT-5.6 Sol | Claude Opus 4.8 |
|---|---|---|---|---|
| AA Intelligence Index v4.1.1 (max) | 60 | 62 | 61 | — |
| LMArena Frontend Code Arena (Elo) | 1,679 (#1) | 1,631 | 1,618 | 1,562 |
| Terminal-Bench 2.1 | 88.3 | 88.0 | 88.8 | 84.6 |
| DeepSWE 1.1 | 67.5 | 70.0 | 73.0 | 59.0 |
| FrontierSWE | 81.2 | 86.6 | 71.3 | 66.7 |
| SWE-Marathon | 42.0 | 35.0 | 39.0 | 40.0 |
| ProgramBench | 77.8 | 76.8 | 77.6 | 71.9 |
| GPQA Diamond | 93.5 | 92.6 | 94.1 | 91.0 |
| BrowseComp | 91.2 | 88.0 | 90.4 | — |
| HLE-Full (no tools) | 43.5 | 53.3 | 44.5 | 49.8 |
| MMLU | 89.2% | — | — | — |
Read the table in two halves. K3 wins frontend coding outright — 1,679 Elo is #1 on LMArena's Frontend Code Arena, six of seven domains, and it beats Fable 5 on SWE-Marathon, ProgramBench, BrowseComp, and MCPMark. K3 loses hard agentic and hardest reasoning: Fable 5 leads HLE-Full, FrontierSWE, OSWorld, GDPval-AA; Sol leads GPQA, DeepSWE, and the Vals Index v2 (57.8% vs Sol's 63.7%, Opus 5's 67.2%).
Pricing tells the other half. K3 is $3/$15 with $0.30 cache; Fable 5 is $10/$50. Per Artificial Analysis's 7:2:1 blended mix, K3 runs at $2.31 per million tokens versus Opus 5's $3.85. Independent per-ticket testing found K3 matching Opus 4.8 quality at ~25% of the cost — but taking ~44 minutes per ticket, roughly twice Opus 4.8 and four times the GPT-5.6 family.
What It Means for Developers

Frontend and UI generation — default to K3. It is the human-preference #1 on the Frontend Code Arena and costs a third of Fable 5. If your product builds UIs from prompts, landing pages, dashboards, or design-to-code, this is the headline result of July 2026.
Asynchronous batch coding — K3's sweet spot. Overnight refactors, bulk code review, migration passes, repo-wide lint fixes. Opus-4.8-class output at roughly a quarter of the token cost. Nobody is waiting on the 37 tokens/second throughput or the 44-minute ticket time, so the latency penalty disappears.
Interactive pairing — look elsewhere. At 37 tok/s and 3-second time-to-first-token, K3 is not the model you sit and watch. GPT-5.6 Sol (72 tok/s) or Opus 5 (56 tok/s) wins the pairing UX.
Cache engineering is worth real money. The official API reports >90% cache-hit rates on coding workloads. At $0.30 vs $3.00, that is a 10× swing on the input side. If your agent loop reuses a system prompt, tool schema, or retrieved context, architect for prefix caching — and remember KDA needs the vLLM implementation Moonshot contributed.
Watch the reasoning-token budget. Independent testers have hit requests where K3 burned 4,093 of 4,096 completion tokens on reasoning and returned nothing. Set generous output budgets; do not cap at the model's default.
Limitations to design around. Moonshot's own card flags three: (1) K3 is sensitive to thinking-history truncation — use a harness verified for it (Kimi Code) and don't switch mid-session; (2) it is over-eager to improvise on ambiguous requests — constrain it in the system prompt or AGENTS.md; (3) hallucination rate measured at 51% on AA-Omniscience — keep verification in the loop.
Self-hosting is theoretical. 2.8T parameters, 1.56 TB download, Moonshot's own Kimi K3 License (not MIT/Apache), recommended on 64+ accelerator supernodes. For nearly every team, "open weights" means thirteen API providers competing on price, not your own rack.
Bottom Line
Kimi K3 is the strongest open-weight model of July 2026 and the first 3T-class release anyone can actually call via API. It is not a frontier model — on the two cross-vendor indices that weight broad economic and agentic work, it sits behind Claude Opus 5, Fable 5, and GPT-5.6 Sol. But it is #1 on human-judged frontend code, it matches Opus 4.8 on quality at ~25% of the token cost, and it does it with open weights, native vision, and 1M context. Route it to frontend generation and asynchronous batch work; keep Claude and GPT tiers for hard interactive agentic work.
Resources
- Kimi K3 Official Blog — release post, architecture, full benchmark table
- Kimi API Platform —
kimi-k3model, $0.30/$3.00/$15.00 per M tokens - Kimi K3 Technical Report (arXiv) — architecture, training, evaluations
- HuggingFace — moonshotai/Kimi-K3 — open weights (July 27)
- Together.ai K3 Quickstart — third-party serving at ~$2.60/$13.00
Kimi K3 is available via API at platform.kimi.ai under the model ID kimi-k3, with open weights released July 27, 2026 under the Kimi K3 License.