Huawei took the stage at HDC 2026 on June 12 and dropped something the community had been waiting years for: openPangu 2.0, the first open-source release in the Pangu lineage. Announced by Yu Chengdong himself, it ships as two MoE variants — a 505B Pro and a 92B Flash — both trained and inferred entirely on Huawei's own Ascend NPUs, both with a 512K-token context window.

The significance isn't just another open model. Pangu has existed since 2020 as a closed, enterprise-commercial model. Opening it signals a strategic shift: Huawei is no longer treating the model layer as a proprietary moat. It's treating it as the bait that pulls developers onto the Ascend and HarmonyOS stack. When the company that builds the chips, the interconnect, the training framework (MindSpore/CANN), the model, and the OS all at once finally opens the model, the whole vertical stack gets a reference implementation.

What's New

What's New

Two versions, one ecosystem. openPangu 2.0 arrives as a deliberate size pair covering two deployment points. The Pro variant is 505B total parameters with 18B activated per token — a knowledge-capacity model with inference cost in the 18B-dense bracket. The Flash variant is 92B total with just 6B active — closer to a 7B-dense compute footprint. Both share the same 512K context window, so long-context capability isn't a Pro-only luxury.

Pangu's first open-source release. Since 2020, Pangu has been a closed, commercial model served through Huawei Cloud to enterprise and government customers. openPangu 2.0 is the first time Pangu weights, base inference code, and a tech report have been offered openly. That's the strategic pivot: closed commercial Pangu continues for enterprise contracts; openPangu is the developer-facing entry point.

Ascend-native by design, not by port. Every other frontier model in this generation was trained on NVIDIA GPUs. openPangu 2.0 was trained on Ascend NPUs from day one — weights, kernels, and routing tuned to the CANN/MindSpore stack rather than CUDA. The model is the reference implementation that proves the Ascend training path works at frontier scale, not an NVIDIA model retrofitted to run on Ascend.

Phased open-source rollout. Huawei announced that starting June 30, seven core components would be open-sourced in sequence — model architecture, weights, training and inference tooling — rather than a single big-bang drop. The Flash variant leads the rollout on GitCode, with Pro weights and the full technical report following.

Architecture

Architecture

Both variants use the same MoE recipe: sparse experts, a fixed small active budget, and a long context window. The engineering bet is the same one DeepSeek and Mixtral made — pack in the knowledge capacity of a huge model, pay inference cost of a small one.

Spec openPangu 2.0 Pro openPangu 2.0 Flash
Total parameters 505B 92B
Activated per token 18B 6B
Architecture MoE (sparse routing) MoE (sparse routing)
Context window 512K tokens 512K tokens
Training hardware Ascend NPU Ascend NPU
Inference target Ascend-optimized Ascend-optimized

What 18B and 6B active mean in practice. At 18B active, Pro sits in a compute bracket similar to running an 18B dense model per forward pass, but with access to a 505B expert pool — the standard MoE quality/cost trade. Flash at 6B active is lightweight enough to target much cheaper inference, making it the volume-tier model for high-throughput applications where latency and cost beat peak accuracy.

512K context across both tiers. Both models ship with the same 512K window. That's enough to process whole codebases, long legal filings, or multi-hour conversation histories in a single pass — and Huawei is not making context a Pro-only upsell. For long-context workloads, Flash carries the same window as Pro.

The hardware stack behind it. The Ascend 910B is Huawei's current flagship training chip, manufactured on a mature 7nm-class process via SMIC. Under US sanctions since 2020, Huawei built the full stack around it: the chip, the interconnect, MindSpore/CANN, cluster management, and optimization tooling. Training 505B parameters required running thousands of these accelerators in parallel with high-bandwidth interconnect — and the model is the proof the stack held up.

Ecosystem

Ecosystem

The release only makes sense as the centerpiece of a broader stack play, and Huawei was explicit about it on stage.

HarmonyOS's native AI backbone. openPangu 2.0 is being integrated directly into HarmonyOS as the system-level intelligence layer. If you develop for HarmonyOS phones, tablets, IoT, or cars, this is the native model your apps will call into. The vertical integration here is unique in the industry: chips (Ascend) → training infrastructure (CANN/MindSpore) → model (openPangu) → OS (HarmonyOS) → devices. Apple has a version of this for on-device models, but nowhere near this parameter scale.

The Ascend software ecosystem's missing piece. Ascend had hardware and a cloud service, but it lacked the open-weight "reference model" that developers actually study, fork, and port. openPangu 2.0 is that reference. When an open model is trained end-to-end on your chips, every operator kernel, every communication primitive, every optimization in the training stack gets public validation. The model is the best possible documentation for the Ascend toolchain.

License: permissive, royalty-free. openPangu 2.0 ships under the Huawei openPangu License — permissive for commercial use, royalty-free (no per-token fees or revenue sharing), and non-exclusive (it doesn't lock you out of other models). That's roughly MIT-like in spirit, with the standard caveat to read the actual license text for export-compliance and attribution clauses in your jurisdiction.

How to use it today. Two paths: the easiest is Huawei Cloud ModelArts, where hosted inference is available through the AI Gallery API with no hardware to manage. The sovereign/custom path is downloading the weights and self-hosting — Pro needs a multi-Ascend setup (think 4+ high-end NPUs), while Flash's 6B active footprint makes it the more accessible entry point for local experimentation.

Honest benchmark caveat. At announcement, Huawei shared limited public benchmark numbers. Independent community evaluations take time to materialize. The fair expectation: Pro competes in the 70B-dense quality class given its 18B active budget and large expert pool; Flash trades peak accuracy for speed; both lead on long context. Treat the quality claims as pending independent verification, not as launch-day certainty.

What It Means for Developers

What It Means for Developers

If you're on NVIDIA APIs today, nothing changes tomorrow. openPangu 2.0 is unlikely to beat the models you're already using on most tasks. Don't rip out your stack. But its existence matters even if you never run it: it means NVIDIA is no longer the only path to a frontier-scale trained model, and that competition benefits the whole market.

The sovereign / supply-chain option is real now. If you operate under US chip export restrictions, or you build for customers with data-residency and Xinchuang (信创) compliance requirements, there is now a frontier-tier model that was trained on non-US silicon from scratch. Before openPangu, your compliant options were weaker models or post-training on foreign weights. This is the first end-to-end sovereign training option at this scale.

Flash is the practical developer entry point. At 6B active with 512K context and a permissive license, openPangu-2.0-Flash is the variant to actually kick the tires on. It's small enough to explore on accessible hardware, long enough to handle whole repositories in one pass, and permissively licensed enough to build commercial products on. Pro is for organizations with Ascend clusters already.

Watch the Ascend portability story. Weights exported from the MindSpore/CANN stack are format-agnostic in principle. Community efforts to run openPangu inference on standard GPU stacks (via format conversion) are expected, with Flash the most likely candidate. If that ecosystem matures, openPangu becomes usable even outside Ascend hardware — at which point the license and 512K context, not the silicon, become the reasons to choose it.

End-cloud协同 is the long game. The real bet is that HarmonyOS developers build against openPangu on-device and in cloud in the same SDK. If you're building HarmonyOS-native AI features, openPangu 2.0 is going to be the default model you're wired into whether you choose it or not — so learning its API shape now is not optional for that platform.

Bottom Line

openPangu 2.0 is a strategy release more than a benchmark release. It's Pangu's first open-source move ever, it ships two MoE models (505B/18B and 92B/6B) with 512K context across both tiers, and — most importantly — it was trained end-to-end on Ascend NPUs with no NVIDIA hardware. For developers already on the NVIDIA stack, it's not an immediate switch. For sovereign-AI, Xinchuang-compliant, and HarmonyOS-native builds, it's the first frontier-scale option whose entire stack you can actually own end to end. The open-sourcing phasing (Flash first, Pro and tech report to follow) gives the community a practical on-ramp starting June 30.

Resources


openPangu 2.0 is available via Huawei Cloud ModelArts API at huaweicloud.com. Weights and inference code released under the Huawei openPangu License, phased from June 30, 2026.