Meituan released LongCat-2.0 on June 30, 2026, and the headline isn't the parameter count. A 1.6T MoE model with 1M context is table stakes in 2026 — DeepSeek-V4-Pro shipped the same 1.6T/49B/1M recipe back in April. What makes LongCat-2.0 the release that developers actually argued about on X and Reddit is the stack it ran on: from random initialization through pre-training to production inference, not a single NVIDIA GPU touched the model.
That sentence matters because "domestic compute" has been a story with a lot of asterisks. Previous Chinese frontier models used domestic chips for inference, or for post-training on top of NVIDIA-trained weights. LongCat-2.0 is the first trillion-parameter model to do the whole loop — pre-training from scratch on tens of trillions of tokens, then serving inference — on domestic silicon. As one developer put it to 36Kr: earlier attempts were furnishing a house already built; LongCat-2.0 poured the foundation, framed the walls, and moved in.
What's New

The model itself is an evolution of Meituan's LongCat-Flash architecture, not a from-scratch reinvention.
1.6T total, ~48B activated, dynamically. LongCat-2.0 is a Mixture-of-Experts model with 1.6 trillion total parameters and an average activation of roughly 48B per token. Crucially, activation isn't a fixed number — it ranges from 33B to 56B depending on token difficulty. The router spends compute on hard tokens and skips it on easy ones, tuned by a PID controller that converges to the target allocation after about 20B tokens.
1M native context, agentic coding as the design center. LongCat-2.0 natively supports a 1M-token context window — enough for a million-character single pass. But the architecture was tuned around a specific target: real-world agentic coding. Every design choice, from the sparse attention to the post-training data mix, was oriented toward code agents that work across multi-file repositories over long sessions.
Anonymous traffic as the real benchmark. Before the reveal, LongCat-2.0 preview ran on OpenRouter under the codename "Owl Alpha" for over a month. By the end of June it had climbed into the global top three for total token call volume, hit #1 on the Hermes harness, and ranked #2 worldwide on Claude Code agent traffic — behind only Claude Opus 4.8 itself. That's not a leaderboard score; that's developers picking it with their wallets while not knowing whose model it was.
Open-sourced days later. On July 6, Meituan open-sourced the model weights, the inference engine, and core technical documentation. The same day, Huawei Ascend, Moore Threads, and MetaX all announced inference adaptations — the multi-vendor domestic inference story landed the same week.
Architecture

LongCat-2.0 inherits two signature ideas from LongCat-Flash, then adds new attention and embedding work.
| Component | Details |
|---|---|
| Architecture | MoE (Mixture of Experts) |
| Total parameters | 1.6T |
| Activated per token | ~48B average (33B–56B dynamic) |
| Sparse attention | LongCat Sparse Attention (LSA) |
| Context window | 1M tokens (native) |
| Embedding | N-gram Embedding |
| Optimizer | Muon |
| Parallelism | 6D parallel training |
| Post-training | MOPD multi-objective |
Zero-computation experts (ScMoE). This is the genuinely clever part. Meituan puts "identity-mapping" experts into the pool that do no computation at all — they pass their input straight through. The router decides per token how many real experts and how many identity experts to use. That's how activation became a 33B–56B range instead of a fixed number, and it's why the model can spend more on hard reasoning tokens and less on trivial ones.
MoE sparsity at ~97% — the asymptote. Meituan disclosed that LongCat-2.0's MoE sparsity is roughly 97% (excluding N-gram Embedding). At that level, adding another 135B of expert parameters yields negligible gains. For context, DeepSeek-V3 ran ~94% and V4-Pro ~97% — the field has converged on the same sparsity wall. The implication for the whole industry: just throwing more experts at the pool has stopped paying off. LongCat-2.0's real innovation is efficiency at a fixed budget, not scale.
LongCat Sparse Attention (LSA) and N-gram Embedding. LSA is Meituan's evolution of DeepSeek's Sparse Attention family, extended with a 3-step Multi-Token Prediction module for speculative decoding. N-gram Embedding improves token-level representation. The combined effect is what makes the 1M context window commercially viable — the architecture's center of gravity shifted from "bigger expert pool" to "squeeze more out of a fixed pool."
Domestic Compute

This is the part that rewrote the narrative, so let's be precise about what's confirmed versus what's community reconstruction.
What Meituan officially states: LongCat-2.0 was trained and inferred entirely on domestic AI ASIC compute; peak training scale exceeded 50,000 domestic AI cards, making it the largest training task ever run on domestic silicon; pre-training consumed over 30T tokens; daily hardware failure rates dropped 70%+; training MFU improved 1.5×; MFU stayed above 30%; key operator efficiency improved 14%. The training ran end-to-end without rollbacks or unrecoverable loss spikes.
What media and community reconstruct: The "50,000 cards" figure comes from tech press reports (ITHome and similar), with estimates ranging from 50,000 to 60,000 cards. The chip is widely inferred by the community to be Huawei Ascend 910C — a chiplet package combining two 910B dies, roughly SMIC N+2 (7nm-class), with ~64GB HBM per die and 200Gbps RDMA — but neither Meituan nor Huawei has confirmed the specific model in the LongCat context. Treat the 910C attribution as a strong community inference, not an official spec.
Why from-scratch pre-training is the bar that matters. Early June 2026 saw a related project that continued pre-training DeepSeek-V4-Pro weights on Ascend for 1,500+ steps. That proved domestic chips could fine-tune a trillion-parameter model. Meituan did something harder: it took a random-initialized 1.6T network and pushed it to convergence across tens of trillions of tokens. Post-training walks a converged loss landscape; from-scratch pre-training has to survive every NCCL-style timeout, every silent data corruption, every loss spike across a 50,000-card cluster.
The cost claim. Meituan states that LongCat-2.0's training and inference cost came in below other trillion-parameter models globally, thanks to compute optimization and the domestic stack. On the inference side, longcat.ai prices at ¥9.9 per 50M tokens and ¥399 per 1B tokens — a price tier that helps explain why free-tier OpenRouter traffic flooded to "Owl Alpha."
Ecosystem validation. At the July 6 open-source drop, Ascend, Moore Threads, and MetaX all shipped inference adaptations on the same day. That multi-vector adaptation speed is the actual signal that the domestic inference ecosystem matured — not one vendor proving it works, but three shipping on launch day.
Benchmarks and Real-World Signal
The honest read on LongCat-2.0's benchmark position is that it's strong but not SOTA, and Meituan isn't pretending otherwise.
Technical community assessments place LongCat-2.0's agent core near Claude Opus 4.6, behind the newer Opus 4.8. Its coding edge is slightly above Zhipu GLM-5.1 but trails the late-June GLM-5.2. On X and Reddit, side-by-side tests against GPT-5.5 consistently favored GPT-5.5 on prose quality, while developers praised LongCat's coding and agent feel. Pure copywriting and creative writing are not its strengths.
So why does it matter? Because the benchmark leaderboard was never the point. LongCat-2.0 simultaneously holds two cards no other Chinese open model has: it's a trillion-parameter open model that runs entirely on domestic chips, and it proved itself on real paid agent traffic before anyone knew its name. When an open model is both cheap and actually runnable on domestic hardware, token traffic votes differently than MMLU does.
Meituan's underlying asset is operational. A decade of food delivery, in-store services, and rider scheduling generated a mountain of real-world agent data — dispatch, tool use, multi-step workflows — that generic benchmark data can't replicate. That's the training moat behind the agentic coding lead, not the architecture alone.
What It Means for Developers

The "domestic chips can only infer" claim is dead. For the past two years, the soft consensus was that domestic silicon could run inference but couldn't train frontier models. LongCat-2.0, trained from scratch to convergence on 50,000 domestic cards with no rollbacks, is the industrial-scale counterexample. The bottleneck was never raw FLOPS — it was the system software (CANN/HCCL-level operator libraries, interconnect stack, failure recovery), and Meituan's 1.5× MFU and 70% failure-rate reduction numbers show where that engineering got spent.
MoE scale has hit an asymptote — build for efficiency. The 97% sparsity disclosure is the developer takeaway most worth internalizing. Adding another 135B experts doesn't move the needle. If you're planning your own MoE roadmap, the frontier has moved from "make the pool bigger" to "extract more from a fixed pool" — sparse attention (LSA/DSA family), dynamic activation budgeting, and KV-cache compression are where the next gains live.
Domestic-inference deployment is now multi-vendor. The same-day adaptations by Ascend, Moore Threads, and MetaX mean that if you're building for a regulated or supply-constrained region, you no longer have a single-vendor lock-in story. Weights are on HuggingFace-style repos, inference code is open, and vLLM/llama.cpp community ports followed within weeks.
Pricing as a strategic weapon. ¥9.9 per 50M tokens is the price point that won the OpenRouter traffic war. For developers shipping high-volume agent pipelines, LongCat-2.0's cost-per-token, combined with the 1M context, makes it a serious default for code-agent workloads where you don't need the very top of the reasoning leaderboard.
Where it isn't the right call: long-form prose, creative writing, and anything where GPT-5.5 or Opus 4.8's prose quality is the deciding factor. Use it where agentic coding throughput and cost matter.
Bottom Line
LongCat-2.0 isn't the most capable model on the leaderboard, and Meituan isn't claiming it is. It's the model that proved the hard thing: a 1.6T, 1M-context, agentic-coding-grade LLM can be trained from scratch and served in production without a single NVIDIA chip. That's the line between "domestic chips for inference" and "a sovereign training loop," and it was crossed on June 30, 2026. Combined with open weights, same-day multi-vendor inference adaptation, and traffic that already voted for it anonymously, LongCat-2.0 is the reference point for the domestic-compute path — not the best model, but the proof that the alternative stack is real.
Resources
- Meituan Official Release — June 30 announcement: training fully on domestic chips
- Meituan Open-Source Post — July 6 weights, inference engine, core docs
- Xinhua Report — domestic-compute training coverage
- Zhongguancun Science City via NCSI — 50,000-card scale and 1.6T/48B specs
- 36Kr Analysis — architecture deep-dive, "Owl Alpha" traffic signal, 97% sparsity
LongCat-2.0 is available via API at longcat.ai. Model weights and inference engine open-sourced July 6, 2026, with adaptations across Ascend, Moore Threads, and MetaX.