On September 17, Yicai reported a number that should stop anyone still treating the US–China AI race as a footnote: Chinese large models processed 61.17 trillion API tokens in the week of September 7–13, against 21.76 trillion for US models. That is 20 consecutive weeks of Chinese volume leadership on OpenRouter, the largest model-aggregation platform with over 5 million developers. Global weekly inference hit 127 trillion tokens that week, up 10.43% week-over-week.
Here is the tension. The same week, Artificial Analysis published Intelligence Index v4.3. The top five are still Western: Claude Fable 5.1 and GPT-6 Astra tied at 53, Claude Opus 5 at 51, Claude Fable 5 at 50, Meta's Muse Spark 1.3 at 48. The highest-ranked Chinese model — Zhipu's GLM-5.3 (max) — sits at 45, in 18th place. Kimi K3 (max) is 23rd at 44. DeepSeek V4-Pro (high) is 70th at 30. China owns 71.6% of OpenRouter's top-10 inference volume and still does not have a model in the intelligence index top 10.
That gap — volume vs. intelligence — is the whole story of Chinese AI in September 2026.
What's New: 20 Weeks of Volume Leadership

The OpenRouter weekly data, as reported by Sina Finance, quantifies the shift:
- Global weekly inference: 127 trillion tokens (+10.43% WoW)
- China weekly: 61.17 trillion tokens (+7.85% WoW)
- US weekly: 21.76 trillion tokens (+31.56% WoW)
- Consecutive weeks China > US: 20
The US grew faster that week (+31.56% vs. +7.85%), but off a much smaller base. The gap is now nearly 3×. China's lead is not a one-week spike — it is a structural pattern that has held since late April.
The top-5 model list for the week is where the China story shows up most clearly:
| Rank | Model | Weekly tokens | WoW |
|---|---|---|---|
| #1 | (Western model — unnamed in report) | — | — |
| #2 | Tencent Hunyuan Hy4 preview | 16.8T | +15% |
| #3 | Zhipu GLM 5.3 Flash | 11.9T | — |
| #4 | DeepSeek V4 Flash 0731 | 11.6T | — |
| #5 | Xiaomi MiMo-V2.5 | 7.77T | +230% |
| #6 | DeepSeek V4.1 Flash (launched Sept 10) | 4.94T | new |
Four of the top five are Chinese. Seven of the top 10 are Chinese. Combined Chinese share in the top 10: approximately 71.6%.
The launch cadence in September alone underscores how fast this stack moves: Qwen3.8-Max-0902 on September 2, DeepSeek V4.1 Flash on September 10, and Kimi K2.8 Preview on September 11. DeepSeek's V4.1 Flash reached #6 in its first three days. The Chinese labs are not shipping one model and waiting — they are running overlapping model lines.
The Data: Volume Dominance, Intelligence Gap

The Artificial Analysis Intelligence Index v4.3 (September 7, 2026) composites 10 benchmarks — GDPval-AA, Terminal-Bench, GPQA Diamond, SciCode, τ²-Bench, and others — into a single intelligence score. The leaderboard tells a different story than OpenRouter volume:
| Rank | Model | Provider | Intelligence Index |
|---|---|---|---|
| T-1 | Claude Fable 5.1 (max, fallback) | Anthropic | 53 |
| T-1 | GPT-6 Astra (max) | OpenAI | 53 |
| 3 | Claude Opus 5 (max) | Anthropic | 51 |
| 4 | Claude Fable 5 (fallback) | Anthropic | 50 |
| 5 | Muse Spark 1.3 (max) | Meta | 48 |
| 18 | GLM-5.3 (max) | Zhipu | 45 |
| 23 | Kimi K3 (max) | Moonshot | 44 |
| 27 | GLM-5.3-Flash | Zhipu | 42 |
| 70 | DeepSeek V4-Pro (high) | DeepSeek | 30 |
The gap from #1 (53) to the top Chinese model (45) is 8 index points. To #23 (Kimi K3 at 44) is 9 points. To DeepSeek V4-Pro at rank 70, the gap is 23 points. On the hardest reasoning tasks — Humanity's Last Exam, FrontierSWE, the long-horizon agent evals — the gap is wider still.
But price tells the other half of the story. The Chinese LLM pricing comparison data for September 2026:
| Model | Input $/M | Output $/M | Blend vs Claude Opus 5 |
|---|---|---|---|
| DeepSeek V4 Flash | ~$0.14 | ~$0.28 | ~50× cheaper |
| DeepSeek V4 Pro | ~$1.98 | ~$3.96 | ~8× cheaper |
| Qwen3.8-Max-0902 | $2.00 | $6.00 | ~4× cheaper |
| GLM-5.2 | ~$0.85 | ~$2.65 | ~8× cheaper |
| Kimi K2.5 | $0.59 | $3.00 | ~10× cheaper |
| Claude Opus 5 | $15.00 | $75.00 | baseline |
| Claude Sonnet 5 | $2.00 | $10.00 | ~3× cheaper |
The math that drives OpenRouter volume is straightforward: Chinese models are 7–10× cheaper at the flagship tier, and 30–50× cheaper at the speed tier. When developers are choosing what to route production traffic to, the price gap is large enough that a 5–10 point intelligence deficit on the hardest tasks does not matter for the bulk of inference — summarization, extraction, chat, RAG, simple coding.
Analysis: Why Volume Outruns Intelligence

Three structural forces explain why Chinese model volume has outpaced intelligence parity.
1. Price compression is the moat, not the bug. Chinese labs operate on domestic infrastructure with lower GPU amortization and higher volume density. DeepSeek V4 Flash at $0.14/$0.28 is not a promotional price — it is a structural cost position. When a model is 10× cheaper than the Western frontier, developers route everything that does not strictly need the frontier to the cheaper option. This is why OpenRouter volume skews Chinese even though the intelligence index skews Western: most inference is not frontier-level reasoning.
2. The September release cadence is a matrix, not a model. Tencent runs Hy4 preview and Hy3 in parallel. DeepSeek runs V4 Flash 0731, V4 Flash 0423, and V4 Pro 0423 simultaneously. Zhipu operates GLM 5.3 Flash and Ox Alpha side by side. This is not one lab shipping one flagship — it is a portfolio of models tuned for different price/performance points, all competing for the same OpenRouter slots. The US labs (Anthropic, OpenAI, Google) each ship one or two frontier models and price them at the top of the market. The Chinese approach is breadth: cover every price tier with a model that is "good enough."
3. The intelligence gap is narrowing, but not on the hardest tasks. Qwen3.8-Max-0902's CodeArena WebDev #1 finish (1,691 Elo, 3 ahead of Claude Opus 5 Max) shows that on applied, agentic, web-development coding, Chinese models are now at parity. But on TerminalBench 3.0 (29.0 vs. Claude's 42.7), FrontierSWE, and Humanity's Last Exam, the double-digit gaps remain. The pattern is consistent: Chinese models close the gap on tasks with concrete, buildable outputs; they trail on open-ended, multi-step reasoning where the evaluation is hardest to game.
The Yicai report frames this as China "leading global AI calls." That is true on volume. But the intelligence index data suggests the more accurate framing is: China has won the inference infrastructure war at the mid and low price tiers, and is now pressing on the frontier. The top five intelligence positions are still Western. The question for 2027 is whether the Chinese labs can close that 8-point gap at the top while keeping the price advantage — or whether the volume lead becomes a permanent structural advantage that funds the intelligence catch-up.
What It Means for Developers

If you are building on the US frontier stack (Claude / GPT): The intelligence index still says these are the best models for the hardest reasoning tasks. Claude Fable 5.1 and GPT-6 Astra tied at 53 are not going to be replaced by a Chinese model on SWE-Pro, HLE, or FrontierSWE this quarter. But the price gap means routing even 30–40% of your traffic to a Chinese model for the easier tasks is a direct 4–10× cost saving on those calls.
If you are price-sensitive: The Chinese flagship tier at $2–$6/M output is now the default for production coding agents, RAG pipelines, and document processing. Qwen3.8-Max-0902, DeepSeek V4 Pro, and Kimi K3 are all within a few points of Claude Opus 5 on applied coding benchmarks at a quarter of the price. The 20-week volume streak is not a coincidence — it is developers voting with their API keys.
If you need the absolute frontier: Keep a US model in your routing stack for the hardest 10–15% of tasks. The intelligence index gap (53 vs. 45) is real on reasoning evals. Build a router that sends hard reasoning to Claude Fable 5.1 or GPT-6 Astra and everything else to Qwen / DeepSeek / GLM.
If you are self-hosting: The open-weight advantage tilts further Chinese. Qwen3.8-2.4T-A95B, DeepSeek V4, and GLM weights are all available for on-prem deployment at levels of capability that did not exist a year ago. The US frontier models (Claude, GPT) remain API-only.
One structural watch-item: The US weekly growth rate (+31.56% that week) is faster than China's (+7.85%). The gap is widening on volume for now, but the US growth rate suggests catch-up in usage — not necessarily in capability. Watch whether Western labs respond with price cuts at the mid tier.
Bottom Line
Twenty weeks of volume leadership is not a blip. Chinese models now process 61 trillion tokens a week on OpenRouter against the US's 21 trillion, at prices 7–10× lower than the Western frontier. But the intelligence index still has Claude Fable 5.1 and GPT-6 Astra at the top, with the best Chinese model (GLM-5.3) 8 points behind in 18th place. China won the inference volume war by competing on price and shipping breadth; the intelligence war is still being fought at the frontier. For developers, the operational conclusion is a router: Chinese models for volume and cost, Western models for the hardest reasoning tasks.
Resources
- OpenRouter — live model usage and pricing data
- Artificial Analysis Intelligence Index v4.3 — composite benchmark leaderboard
- Sina Finance: China 20-week report — weekly token volume breakdown
- Chinese LLM Pricing Guide 2026 — price comparison across providers
- DeepSeek V4.1 Flash release — September 10 architecture update
Chinese first-tier models are available via API at OpenRouter: Qwen3.8-Max ($2/$6), DeepSeek V4 Pro ($1.98/$3.96), Kimi K3 (~$2/$6), GLM-5.3 (~$0.85/$2.65). Western frontier: Claude Fable 5.1 and GPT-6 Astra lead the Intelligence Index at 53, priced at the top of the market.