A year ago, a 1M-token context window was a slide-deck claim — impressive in a blog post, uneconomical to actually use. In mid-2026 it became the Chinese floor: DeepSeek V4 ships 1M natively at $0.14/M, Qwen3.8-Flash reaches 1M via YaRN at $0.16/M, and Kimi K2-series carries 256K context with agent-swarm tooling built around it. The price of "read the whole book" stopped being a rounding error and became a budget line.
But the surface similarity masks a real architectural fork. Three teams arrived at long context through three different theories of what the problem actually is. DeepSeek treats attention arithmetic as the enemy and compresses structurally. Qwen treats positional extrapolation as the enemy and scales it cheaply with a host-memory lookup trick. Kimi treats context as an orchestration problem and splits long work across parallel sub-agents. If you are picking a long-context provider for RAG, codebase analysis, or document pipelines, the route matters as much as the window size. This comparison lays the three bets side by side.
What's New

1M is the new 128K. DeepSeek V4-Pro and V4-Flash both ship 1M-token windows. Qwen3.8-Flash-Next ships 262K native and extends to 1M with YaRN. The implication is not just "bigger input box" — it is that full-repository code analysis and whole-contract legal review are now single-call workflows rather than chunked pipelines.
The architectures diverged where they used to converge. A year ago every Chinese long-context story was "RoPE scaling + context extension." Now DeepSeek replaced MLA entirely with CSA+HCA; Qwen pairs Gated DeltaNet with sparse attention and a 51B n-gram lookup table; Kimi keeps a standard MoE but orchestrates long context across up to 300 parallel sub-agents in Agent Swarm mode.
Cost follows the architecture. DeepSeek V4-Flash: $0.14/M input, 1M context, 13B active. Qwen3.8-Flash: $0.16/M input, 1M via YaRN, 6B active, with up to 8.6× the prefill throughput of its predecessor at 1M. Kimi K2: 256K context at 32B active, but its Agent Swarm pattern processes far longer effective tasks by decomposing them.
DeepSeek: CSA+HCA Compression, Layer by Layer

DeepSeek V4-Pro stacks 61 transformer layers, and the long-context design is visible in the layer assignment. Layers 0–1 run Heavily Compressed Attention (HCA). From layer 2 to layer 60, CSA and HCA alternate. This is not arbitrary — it is a deliberate division of labor written into the graph.
CSA (Compressed Sparse Attention) compresses KV by 4:1 along the sequence axis, then a Lightning Indexer selects top-1024 compressed entries per query (top-512 on V4-Flash) using FP4 math. This is the "precise look" path — it knows which compressed blocks matter and only attends to those. HCA (Heavily Compressed Attention) compresses harder, at 128:1, then runs dense attention over the tiny result. This is the "global glance" path — cheap enough to run everywhere, dense enough to catch long-range structure that sparse selection might drop. A 128-token sliding-window branch keeps the most recent tokens raw, so numbers, names, code identifiers, and function arguments stay exact.
The arithmetic result is the headline: at 1M context, V4-Pro uses 27% of V3.2's single-token inference FLOPs and 10% of its KV cache — a 73% FLOPs reduction and 90% KV cache reduction. V4-Flash, with fewer active parameters, pushes further to 10% and 7%. Against a vanilla BF16 GQA8 baseline, the V4 KV cache at 1M context lands around 2% of what uncompressed dense attention would require.
mHC (Manifold-Constrained Hyper-Connections) is the stability glue. Long context at this depth stresses residual paths, so V4 replaces plain residual connections with a 4-stream hyper-connection whose mixing matrix is projected onto the Birkhoff polytope — the set of doubly stochastic matrices — via 20 Sinkhorn-Knopp iterations. The spectral norm stays bounded at 1, so signals cannot blow up across 61 layers. Fused kernels and selective recompute keep mHC's wall-time overhead to just 6.7% of the overlapped pipeline stage. Without it, the compression-only architecture would train unstably.
The DeepSeek bet is: long context is a compression problem. Pay for exactness only where you need it (sliding window + sparse top-k), pay for global awareness cheaply (128:1 dense), and make the residual stream mathematically safe across the deep stack.
Qwen and Kimi: Scaling vs Orchestration

Qwen3.8-Flash-Next takes the opposite architectural bet: keep the transformer mostly intact, and extend context through positional scaling plus a deterministic lookup table.
YaRN (Yet another RoPE extensioN) is the context-extension primitive. The model natively trains and serves 262,144 tokens; applying YaRN at inference stretches it to 1,000,000 without retraining. This is the "stretch the position encoding" school — lower engineering risk, but with the usual long-context recall penalty that YaRN mitigates rather than eliminates.
The GDN + QSA hybrid does the real compression work. Every block of four layers uses three Gated DeltaNet layers (which fold history into a fixed-size recurrent state) and one Qwen Sparse Attention layer. This is Qwen's analogue of DeepSeek's CSA+HCA — a recurrent-compression path mixed with a sparse-attention path — but with a recurrent-state flavor rather than sequence-level block compression.
The N-gram Embedding is the headline trick. Qwen adds 51B parameters as a 20M-entry bigram/trigram lookup table. Because the lookup address is known deterministically from local tokens, these 51B parameters live in host RAM, not GPU VRAM, and are prefetched asynchronously. Effective model capacity jumps without touching the per-token GPU matmul budget — the active parameter count stays at 6B. At 1M tokens this architecture delivers up to 8.6× the prefill throughput of Qwen3.7-Plus, with 7.6× faster prompt processing and 4.9× faster generation in benchmark setups.
Kimi K2 takes the third route. Its 1T-total / 32B-active MoE keeps a native 256K context window rather than chasing 1M. Instead, Moonshot bets that very long tasks are better decomposed than swallowed. Agent Swarm mode — introduced with K2.5 and extended to 300 parallel sub-agents executing up to 4,000 coordinated steps in K2.6 — breaks a long-horizon task into parallel domain-specialized subtasks, each running in its own smaller context. The effective "context" is no longer a single string but a swarm of workers with shared memory. Kimi's bet: 1M is a crutch; the right answer for long work is orchestration, not one giant prompt.
Engineering Comparison: Cache, Speed, Recall, Cost

| Dimension | DeepSeek V4 (Pro/Flash) | Qwen3.8-Flash-Next | Kimi K2 Series |
|---|---|---|---|
| Native context | 1M | 262K (→1M YaRN) | 256K |
| Active params | 49B / 13B | 6B (+51B host n-gram) | 32B |
| Attention design | CSA 4:1 + HCA 128:1 + 128 SWA | 3× GDN + 1× QSA per block | Standard MoE + Agent Swarm |
| KV cache @ 1M | ~10% / 7% of V3.2 (~2% of GQA8) | Recurrent-state compression | Standard KV at 256K; swarms split work |
| Throughput @ 1M | Flash: 10% FLOPs of V3.2 | 8.6× prefill vs Qwen3.7-Plus | N/A (task decomposed) |
| Input price | $0.435 / $0.14 per MTok | $0.16 per MTok | K2 API tier |
| Long-task pattern | Single 1M call | Single 1M call (YaRN) | 100–300 parallel sub-agents |
KV cache is where the routes diverge hardest. DeepSeek's structural compression makes a 1M request's cache roughly an order of magnitude smaller than a dense 1M model — that is why $0.14/M with 1M is even possible. Qwen's GDN recurrent state is smaller per token than full KV but relies on host-memory n-gram offload for capacity; expect lower GPU VRAM per concurrent request than a vanilla transformer, but not as aggressive as CSA+HCA. Kimi's 256K-native design means per-request memory is bounded, but very long work only fits by fanning out.
Throughput favors whoever activates fewer parameters. Qwen3.8-Flash at 6B active and 8.6× prefill throughput is the speed champion for 1M prompts. DeepSeek V4-Flash at 13B active is the quality/speed middle. Kimi at 32B active is the slowest per call but the only one that parallelizes long tasks across agents.
Long-document recall is the fine print. DeepSeek reports 83.5 MMR on MRCR 1M (V4-Pro Max), second only to Claude Opus 4.6 at the far end of the window. Qwen's YaRN extension works well up to 1M but community reports flag brittleness on very long agent chains. Kimi's 256K window avoids the extreme-end recall penalty entirely — it never pretends to read a million tokens in one pass.
Cost per million tokens at 1M effective context therefore depends on the task shape, not just the sticker price. A 500-document RAG corpus that fits one DeepSeek or Qwen call is cheaper than 10 Kimi swarm calls; a 4,000-step engineering task that requires tool use and web browsing is exactly where Kimi's swarm model earns its keep.
What It Means for Developers
Full-repository code analysis is a single call now. A 200K-token codebase fits one V4-Flash or Qwen3.8-Flash prompt. You no longer need chunk-and-embed-and-rerank pipelines for monorepos under ~750K tokens. Point the agent at the repo, let it read, accept that exact identifier recall at the very start of the file is slightly degraded by compression, and rely on the sliding-window branch for the file it is currently editing.
Long-document RAG is economically viable. Processing a 1,000-page contract through Qwen3.8-Flash at $0.16/M costs cents. The old economics — where RAG existed because the context window was too expensive — are inverted. Use RAG for retrieval quality, not for cost. Reserve the giant single call for extraction and summarization where global context matters more than precise quote recall.
Pick the route, not just the number. If your workload is retrieval-heavy and latency-sensitive, Qwen's 6B active + 8.6× prefill wins. If you need the best reported 1M recall and can tolerate 13B active, DeepSeek V4-Pro. If your "long context" is actually a multi-hour agentic task with tools and browsing, Kimi's Agent Swarm decomposition beats any single-context approach regardless of window size.
Test recall at your actual length. All three vendors self-report strong long-context numbers. Run your own needle-in-a-haystack at your document length before committing — YaRN extension in particular degrades non-linearly near the top of the range, and compressed-attention models can drop rare tokens buried in the middle of a 1M prompt.
Bottom Line
1M context is no longer a Chinese model's differentiator — it is the entry ticket. What differentiates them now is the bet underneath: DeepSeek compresses attention into submission (73% fewer FLOPs, 90% less KV cache at 1M), Qwen scales positions cheaply and hides 51B parameters in host RAM (8.6× prefill, 6B active), and Kimi refuses the single-prompt premise entirely and runs long work across hundreds of parallel agents. For builders, the practical question is no longer "which model has the biggest window" but "which route matches how my long tasks actually look."
Resources
- DeepSeek-V4 Technical Report (arXiv:2606.19348) — CSA+HCA layer layout, mHC, MRCR 1M recall, FLOPs/KV cache comparisons
- Qwen3.8-Flash-Next Blog — YaRN 262K→1M, GDN+QSA hybrid, 51B N-gram host offload, 8.6× prefill throughput
- Kimi K2 Official Blog — 1T/32B MoE, 256K context, Agent Swarm agentic design
- CAICT Inference Optimization Report (March 2026) — industry context on long-context serving economics
DeepSeek-V4 is available via API at https://platform.deepseek.com.