Moonshot AI released Kimi K2.7 Code on June 12, 2026 — and the name tells you exactly what it is. This isn't a new base model. It's a coding-first, agent-tuned fork of the 1T/32B K2.7 MoE base, the same architecture family that shipped as K2.6 two months earlier. Moonshot took a proven general-purpose base and ran coding-specialized post-training on top, then pushed the result as the default model in Kimi Code.
That's a meaningful strategic shift. Up through K2.6, Moonshot's identity was "the general model that also codes well." K2.7 Code says the quiet part out loud: long-horizon, agentic software engineering is the use case worth optimizing for as a separate product. It's the same move OpenAI made with Codex and Anthropic made with Claude Code-specific tuning — take a strong general base, spend the post-training budget on code, tool use, and long sessions, and sell the result to developers who live in terminals.
What's New

Coding-first post-training on a fixed 1T base. K2.7 Code keeps the K2.7 architecture — 1T total parameters, 32B active per token, 384 routed experts with 8 selected plus 1 shared, MLA attention, 400M MoonViT vision encoder, 160K vocab. What changed is the post-training objective: long-horizon code generation, code review, debugging across multi-file repositories. Moonshot explicitly recommends the balanced K2.6 for writing, analysis, and general chat — K2.7 Code is the coding specialist.
262K context, always on. The context window is 262,144 tokens (~256K) — enough for a whole repository or a long multi-session debugging thread in one pass. And there's no off-switch for thinking: K2.7 Code always runs in thinking mode. If you request a non-thinking response in Kimi Code, the request silently gets served by K2.6 instead. That's a deliberate quality guarantee — no degraded fast path.
~30% fewer thinking tokens than K2.6. This is the efficiency win that matters for cost and latency. Reasoning models tend to over-think — burning thousands of tokens on questions that don't need them. K2.7 Code cuts average thinking-token usage by about 30% while scoring higher on every coding benchmark than K2.6. Better answers, fewer tokens spent reaching them.
Open weights, Modified MIT. The model weights are on HuggingFace under a Modified MIT license — commercial-friendly with the standard caveat to read the license file before production use. Cloudflare Workers AI picked it up on day one (@cf/moonshotai/kimi-k2.7-code), and it plugs into the usual coding harnesses: Kimi Code, Claude Code, Cline, RooCode.
Benchmarks

Moonshot compares K2.7 Code against its own K2.6, plus GPT-5.5 (evaluated in Codex at xhigh) and Claude Opus 4.8 (evaluated in Claude Code at xhigh), all with thinking on and a 262,144-token context.
Coding benchmarks:
| Benchmark | Kimi K2.6 | K2.7 Code | GPT-5.5 | Claude Opus 4.8 |
|---|---|---|---|---|
| Kimi Code Bench v2 | 50.9 | 62.0 (+21.8%) | 69.0 | 67.4 |
| Program Bench | 48.3 | 53.6 (+11.0%) | 69.1 | 63.8 |
| MLS Bench Lite | 26.7 | 35.1 (+31.5%) | 35.5 | 42.8 |
Agent benchmarks:
| Benchmark | Kimi K2.6 | K2.7 Code | GPT-5.5 | Claude Opus 4.8 |
|---|---|---|---|---|
| Kimi Claw 24/7 Bench | 42.9 | 46.9 | 52.8 | 50.4 |
| MCP Atlas | 69.4 | 76.0 | 79.4 | 81.3 |
| MCP Mark Verified | 72.8 | 81.1 | 92.9 | 76.4 |
The honest read: K2.7 Code is a clear step up from K2.6 across the board, but it is not beating GPT-5.5 or Claude Opus 4.8 on raw coding score. Kimi Code Bench v2 goes 50.9 → 62.0, a real +21.8% jump, but GPT-5.5 is still at 69.0 and Opus 4.8 at 67.4. The one place it beats Opus 4.8 outright is MCP Mark Verified (81.1 vs 76.4) — the tool-calling/agent benchmark where it also gains ~10 points over K2.6.
So the pitch isn't "we're the best coder." It's "we're the best coding specialist in the open-weight, cost-competitive tier, and we're genuinely strong at the tool-use layer where agentic coding actually breaks." The +30% thinking-token saving makes every one of those benchmark scores cheaper to reproduce in production.
Comparison: Coding-First Fork vs Dedicated Coding Bases

It helps to place K2.7 Code against the two names developers always bring up — DeepSeek Coder and Qwen Coder — because they optimize the problem differently.
DeepSeek Coder and Qwen Coder are from-scratch coding base models. They're trained on code-heavy corpora from the ground up: the pre-training objective, the data mix, the tokenizer weighting all tilt toward code. That gives them deep, specialized code knowledge — strong on obscure languages, tricky algorithms, low-level systems code. The tradeoff is that their general capabilities are narrower; you use them when the job is almost all code.
K2.7 Code is a coding-specialized post-trained variant of a general 1T MoE base. The base already absorbed broad web, math, tool-use, and multilingual data before the coding fine-tune. That means K2.7 Code ships with the whole general-purpose vocabulary still underneath it — it can reason about a business requirement, read a design doc, plan a refactor across languages, and then write the code. Its edge isn't "knows more Python than DeepSeek Coder"; it's "can take a messy real-world ticket and drive it to completion through tools over a long session."
That's why its strongest benchmark is MCP Mark Verified (tool calling) and why it's positioned as an agent model, not a code-completion model. If your task is a single-file LeetCode-style generation, a dedicated coding base may still win. If your task is "here's a repo, fix this bug across three files, run the tests, update the docs," the general-base-plus-coding-post-training recipe is what wins that kind of end-to-end work.
The other differentiator is cost. K2.7 Code's API starts at ¥6.50/M input and ¥27.00/M output, with cached input at ¥1.30/M. For agent loops that send the same repo context on every turn, the cache discount (80% off repeated input) is the number that matters — a 200K repo prompt re-sent across 50 turns costs ¥1.30/M for every turn after the first.
What It Means for Developers

You're choosing the K2.7 tier, then picking a specialization. K2.6 is the balanced general model; K2.7 Code is the coding specialist. Moonshot's own guidance is explicit: writing, analysis, and conversation → K2.6; long-horizon coding and agentic engineering → K2.7 Code. The practical workflow is to keep both keys and route by task, exactly the way teams route between a reasoning model and a fast model.
The always-on thinking rule is a feature, not a restriction. You can't turn thinking off on K2.7 Code, and that's intentional — it prevents the lazy, low-quality fast path. The trade is that output always costs thinking tokens, which is why the 30% thinking-token reduction is so important: it keeps the always-thinking mode affordable. Budget output tokens accordingly, and use the cache for repeated context.
Two price tiers, pick by latency tolerance. Standard kimi-k2.7-code runs at the ¥6.50/¥27 rate. The kimi-k2.7-code-highspeed variant doubles both input and output prices (¥13.00/¥54.00) but outputs at ~180 tokens/second — up to 260 tokens/second on short context. For interactive terminal use where you watch generation stream, the highspeed tier is worth the 2×; for batch pipelines, standard is the default.
It drops into your existing harness. K2.7 Code isn't a walled garden. Cloudflare Workers AI shipped it on launch day, and it works through Claude Code, Cline, and RooCode. If you already have a coding-agent setup, swapping the model ID is a config change, not a rearchitecture. The open weights mean you can also self-host if your compliance needs require it.
Where it isn't the pick. For pure single-shot code generation or obscure-language depth, DeepSeek Coder / Qwen Coder may still hold an edge. For top-of-the-leaderboard coding score, GPT-5.5 and Opus 4.8 still beat it on Kimi Code Bench v2. K2.7 Code's sweet spot is agentic, multi-file, tool-driven work where the 262K context, the MCP tool-use strength, and the cost structure compound.
Bottom Line
Kimi K2.7 Code is the clearest signal yet that Moonshot is splitting its general model into productized specialists. It keeps the proven 1T/32B base, adds coding-first post-training, cuts thinking tokens 30%, and lands at a price tier (¥6.50/¥27 per M tokens, 80% off cached input) that makes long agent loops economically sane. It doesn't topple GPT-5.5 or Opus 4.8 on raw coding benchmarks, but it beats Opus 4.8 on MCP Mark Verified and wins on the cost-and-context combination that real developer workflows run on. For teams already in the Kimi ecosystem, it's the obvious upgrade for agentic coding; for everyone else, it's the open-weight coding specialist to route your multi-file, tool-driven work to.
Resources
- Kimi K2.7 Code Official Page — benchmarks, architecture, and entry points
- Kimi K2.7 Code Pricing — API token rates and subscription tiers
- Kimi API Platform — developer API access
- Cloudflare Workers AI Changelog — day-one hosting on
@cf/moonshotai/kimi-k2.7-code - HuggingFace Model Card — open weights under Modified MIT
Kimi K2.7 Code is available via API at platform.kimi.com. Model ID kimi-k2.7-code and kimi-k2.7-code-highspeed, 262,144-token context, Modified MIT weights on HuggingFace.