If you have priced a Chinese frontier model against GPT or Claude this year, the sticker price looks almost too good to be true. DeepSeek V4-Flash serves 1M-token context at $0.14 per million input tokens. Qwen3.8-Flash does coding-grade work at $0.16/M. Tencent's Hunyuan runs a 1.8B model on a phone from a 600 MB download. None of this is charity — it is the result of a decade of inference engineering being compressed into a single release cycle.

The China Academy of Information and Communications Technology (CAICT) put it plainly in its March 2026 report on inference optimization: the industry has moved from "can it run" to "can it run cheaply at scale." That shift is happening on three simultaneous axes — attention arithmetic, numeric precision, and the hardware itself. This article maps the technical stack that makes Chinese models cheap, with the actual numbers attached.

What's New

What's New

The headline trend is that inference optimization stopped being a serving-layer problem and became an architecture problem. Three releases in 2026 capture the state of play:

DeepSeek V4 (April 2026) replaced dense MLA attention with a CSA+HCA hybrid, pushing single-token inference FLOPs at 1M context down to 27% of V3.2 (V4-Pro) and 10% (V4-Flash). KV cache storage dropped to 10% and 7% respectively. Routing expert weights ship in MXFP4 (FP4), with a lossless FP4→FP8 dequantization path.

Tencent Hunyuan AngelSlim (February–April 2026) shipped HY-1.8B-2Bit, a 1.8B model compressed to an effective ~0.3B parameter footprint at ~600 MB on disk — smaller than many mobile games. The follow-up Hy-MT2 reaches 1.25-bit extreme quantization at 440 MB, running 1.5× faster than the 4-bit build on an Apple A15.

China Mobile Cloud heterogeneous inference (September 2026) took the hardware route. A rack pairing Tianshu Zhixin domestic GPUs with LingXi neuromorphic chips — splitting attention onto GPU and feed-forward networks onto neuromorphic silicon — doubles inference output and energy efficiency on DeepSeek V4-Flash while cutting operating cost by more than 40%. No NVIDIA silicon in the stack.

The throughline: every layer of the inference stack — attention math, weight format, scheduling, and now silicon partitioning — is being optimized in lockstep. The rest of this article drills into each.

KV Cache Optimization: Compress Before You Sparse

KV Cache Optimization

The KV cache is where long-context inference lives or dies. Every past token's key and value tensors must sit in HBM, and at 1M tokens that footprint becomes gigabytes per concurrent request. Chinese labs attacked this with structural compression — not just sparse selection, but actually shrinking the cache itself.

DeepSeek CSA+HCA hybrid attention is the reference design. V4-Pro stacks 61 transformer layers: layers 0–1 run Heavily Compressed Attention (HCA), and layers 2–60 alternate Compressed Sparse Attention (CSA) and HCA. CSA compresses KV by 4:1 (four tokens → one entry) then runs a Lightning Indexer that selects top-1024 compressed entries per query using FP4 math. HCA compresses harder, at 128:1, then runs dense attention over the tiny compressed result — cheap global awareness that backfills what sparse selection misses. A 128-token sliding-window branch keeps the most recent tokens raw so numbers, code, and function arguments stay exact.

The payoff, per the DeepSeek V4 technical report, is concrete: at 1M context, V4-Pro uses 27% of V3.2's single-token FLOPs and 10% of its KV cache; V4-Flash uses 10% and 7%. Against a BF16 GQA8 baseline, the V4 KV cache shrinks to roughly 2% of what vanilla attention would require. That is the difference between 1M context being a demo and being a default.

CED (Causal Encoder-Decoder) takes it further. DeepSeek V4.1-Flash (September 2026) splits its 40-layer backbone into a 20-layer encoder and a 20-layer decoder. The decoder does not compute its own global KV at all — it projects it directly from the encoder's final hidden state. Result: global KV cache lands at ~890 bytes per token, a 4× reduction over V4-Flash, and persistent SSD storage for cached prefixes drops to 1/8. Prefill activates only 8B parameters; decoding activates 16B — asymmetrical compute that is cheap for input-heavy agent workloads.

On the serving side, the CAICT report documents the standard playbook now universal in Chinese stacks: PagedAttention-style non-contiguous KV cache allocation that lifts concurrent requests by 3× or more on the same hardware, and RadixAttention-style prefix caching that reuses shared system prompts and conversation history across requests. DeepSeek's own on-disk prefix reuse goes a step further, storing compressed CSA/HCA KV to disk so long agent prefixes are recomputed once, not per request.

Quantization: FP4, 2-Bit, and Lossless Dequant

Quantization

If KV compression shrinks runtime memory, weight quantization shrinks the model itself. Chinese labs pushed quantization from post-training afterthought to something baked into training.

DeepSeek MXFP4 QAT is the most aggressive server-side example. In the V4 checkpoint, 96% of the model — the routed experts that dominate GPU memory — ships in MXFP4 format; the remaining 4% (shared expert, attention, norms) uses FP8 or BF16. The clever part is the dequantization contract: FP4 (E2M1) to FP8 (E4M3) dequantization is provably lossless, because FP8 has two extra exponent bits. As long as the scale-ratio within a 128×128 FP8 block stays under threshold, the fine-grained FP4 sub-block scaling is fully absorbed by FP8's dynamic range. That means the entire quantization-aware training pipeline can reuse the FP8 training framework — gradients flow to FP8 master weights, RL rollout and inference run true FP4 weights, with no train/deploy mismatch. The Lightning Indexer's QK path runs in FP4 directly; index scores quantized to BF16 give a 2× faster top-k selector while holding 99.7% KV recall.

Tencent Hunyuan 2-bit (AngelSlim) is the edge-side counterpart. HY-1.8B-2Bit takes a 1.8B model and, using Tencent's Stretched Elastic Quantization, compresses it to an effective ~0.3B parameter equivalent — roughly one-sixth the compute — at ~600 MB on disk. The translation-specific Hy-MT1.5-1.8B-2bit drops from 3.3 GB FP16 to 574 MB. Hy-MT2's 1.25-bit extreme variant reaches 440 MB and runs 1.5× faster than its own 4-bit predecessor on Apple A15 silicon, with the 2-bit/1.25-bit builds targeting mid-range and budget phones respectively. AngelSlim bundles this with FP8/INT8 PTQ, speculative decoding, and distillation into one pipeline.

INT8 multi-chip adaptation is the workhorse layer underneath. The CAICT report and Tencent's own distributed inference writeups document INT8 KV cache quantization with GPU kernels hitting ~1,694× over CPU baselines, plus INT8 weight serving ported across NVIDIA, Huawei Ascend, and domestic GPU targets. The point is not peak precision — it is that a single quantized artifact runs across multiple silicon vendors without per-chip retraining.

System-Level Optimization: MoE Sparsity, Pipelining, and Heterogeneous Silicon

System-Level Optimization

Architecture and quantization only pay off if the serving stack can actually feed the model fast enough. The Chinese playbook combines ultra-sparse MoE, communication-compute overlap, and — newest of all — deliberate silicon partitioning.

Ultra-sparse MoE cuts the arithmetic per token. Qwen3-Next is the extreme: 80B total parameters but only ~3B activated per token — a 3.7% activation ratio — across 512 routed experts with 1 shared and 10 activated. Alibaba reports up to 90% compute cost reduction versus dense equivalents at that capability tier. Qwen3.8-Flash-Next pushes the idea further with a 125B backbone, 6B active, plus 51B of N-gram embedding parameters that live in host RAM and are deterministically looked up — free capacity that never touches GPU matmuls.

Communication and pipeline scheduling is where the MoE efficiency actually survives to the user. DeepSeek's MegaMoE fuses expert-parallelism dispatch, GEMM, activation, and combine into one kernel split across waves, so wave-1 GEMMs overlap with wave-2 dispatches. Measured speedup is 1.50–1.73× on general inference and up to 1.96× on short-batch RL rollouts. The CAICT report echoes this pattern: double-batch overlap, deferred token scheduling that hides CPU overhead, and prefix-aware batching all compound on top of the raw model savings.

Heterogeneous inference is the newest lever. China Mobile Cloud's September 2026 system uses an Attention-FeedForward (AF) split: Tianshu Zhixin domestic GPUs handle prefill and attention layers, while LingXi neuromorphic chips handle the feed-forward network layers — each hardware type running the layer shape it is best at. Validated on DeepSeek V4-Flash, the setup more than doubles inference output and energy efficiency versus a pure-GPU cluster of equivalent investment, and cuts business operating cost by more than 40%. It ships with 15 invention patents and is in small-batch trial production in China. This is the logical endpoint of the trend: stop pretending one chip type should run every layer, and partition the network to the hardware.

What It Means for Developers

If you are building on top of these models, the practical takeaways are not "Chinese models are cheap" — they are where the cheapness actually comes from and what it lets you do.

1M context is an economic artifact, not a marketing number. The 90% KV cache reduction means a 200K-document RAG call costs roughly what a 32K call cost two years ago. Stop over-chunking. Retrieval-augmented pipelines that used to need sliding-window RAG or hierarchical summarization can now dump 200–500 chunks into one prompt and let CSA/HCA recall handle it.

Watch the active-parameter number, not the total. Qwen3-Next at 80B/3B and Qwen3.8-Flash at 125B/6B deliver flagship-adjacent quality at small-model serving cost. When you route traffic, route on active parameters and price — not on total parameter count, which is training capacity, not inference tax.

Self-hosting economics just changed on two axes. MXFP4 artifacts (V4-Flash unsloth builds run ~103 GB at 3-bit, ~162 GB at 8-bit) put a 284B model on a well-equipped workstation. At the edge, Hunyuan 2-bit puts a translation-capable model in a 600 MB mobile app with no network dependency.

Heterogeneous silicon is a bet to watch. The China Mobile Cloud result (2× throughput, 40% cost cut, no NVIDIA parts) suggests that inference cost is about to be decoupled from the high-end GPU procurement cycle. For teams outside the US export-control bubble, this matters more than any single model release.

Bottom Line

Chinese inference cheapness is not a single discount — it is a stack of structural wins: CSA+HCA and CED cut KV cache by 90%+; MXFP4 with lossless dequantization halves weight memory versus FP8; ultra-sparse MoE activates 3–6B of 80–284B; and GPU-plus-neuromorphic partitioning now squeezes another 2× out of the metal. Any one of these would be a story. Together they are why $0.14/M and 1M context became the floor rather than the ceiling.

Resources


DeepSeek-V4 is available via API at https://platform.deepseek.com.