Huawei opened its 11th annual Huawei Connect (HC2026) in Shanghai on September 17, 2026, and the headline number was not a new chip — it was a deployment count. Rotating chairman Wang Tao announced that Ascend SuperPods are now deployed in 1,000-plus units across 370-plus customers. After years of chip, server, interconnect, and software validation, Huawei's next-generation AI compute stack has crossed from engineering trials into commercial deployment.

The same keynote surfaced the next product a quarter earlier than planned. The Ascend 960 silicon has doubled in performance and is ready for Q1 2027, three quarters ahead of the original roadmap. Pair it with the Atlas 960 SuperPod — the industry's first Near-Packaged Optics (NPO) SuperPod — and the message is clear: as model parameters move from hundreds of billions to trillions, the 8-GPU server is dead, and the tightly-coupled SuperPod is the default. The SLAI T-Rex technical report, which full-parameter post-trained DeepSeek-V4 on Ascend NPU SuperPOD hardware, is independent proof that the stack already runs frontier models.

What's New

What's New

1,000+ SuperPods, 370+ customers: Wang Tao put the Ascend SuperPod fleet in concrete commercial terms. The prior-gen Ascend 910C SuperPod alone is deployed in more than 1,000 units, and the Ascend 950 SuperPod has entered volume commercial use. This is the number that matters for builders: it is no longer a pilot — it is installed base.

Ascend 960 surfaced early: The next-gen Ascend 960 chip has doubled performance and is ahead of schedule by three quarters. Ascend 960 DT (development/test) is ready Q1 2027; Ascend 960 PR (production) is ready Q3 2027, one quarter early. The original roadmap had it landing Q4 2027.

Atlas 960 SuperPod — world's first NPO SuperPod: Built on Huawei's own UnifiedBus ("灵衢") interconnect plus the new Hi-ONE optical engine, Atlas 960 is the industry's first SuperPod to ship with NPO. The liquid-cooled Atlas 960 SuperPod lands Q3 2027; the air-cooled Atlas 860 SuperPod lands Q2 2027.

One-gen-per-year cadence locked in: Ascend 970 is planned for 2028 and Ascend 980 for 2029. Huawei frames this as "Tao's Law" — continued doubling of compute scale, HBM bandwidth/capacity, and interconnect bandwidth each year.

Architecture

Architecture

A SuperPod fuses tens to thousands of chips over a dedicated high-speed bus into one tightly-coupled compute system, so a whole rack — or several racks — behaves as a single supercomputer at the software layer. As models scale toward trillion and ten-trillion parameter training, a standard 8-card server simply cannot carry the inter-chip traffic; the SuperPod is how every serious compute vendor in 2026 has to play.

The Atlas 960 SuperPod pushes that one step further by moving optics closer to the silicon:

Spec Atlas 960 SuperPod
NPU interconnect 4,096 NPUs in one SuperPod
Precision compute 8 EFLOPS FP8 / 16 EFLOPS FP4
Interconnect fabric UnifiedBus ("灵衢") + Hi-ONE NPO optical engine
Topology Orthogonal architecture, full liquid cooling
Memory model Unified memory addressing over UnifiedBus
RTT latency As low as 2 microseconds
Optical swap 5,500 Hi-ONE replaces ~48,000 × 800G optical modules
Power saving >550 kW lower than the 800G-module design
Availability 99.8%, MTBF doubled

NPO, not CPO. NPO (Near-Packaged Optics) packages the optical engine independently and places it next to the chip, shortening the electrical run before the signal converts to light. CPO (Co-Packaged Optics) goes further — optics and switch silicon on the same substrate — for the highest integration, but a failed optical engine cannot be swapped. NPO keeps the optical engine pluggable and serviceable while still cutting signal loss and power. Hi-ONE is the world's first mass-produced NPO product, at 7.2T total capacity and the only NPO with a built-in light source.

Unified memory addressing across 4,096 chips is the other architectural bet. Tens of thousands of chips communicate with near-linear bandwidth and latency, so MFU (model FLOPs utilization) — the number that actually decides whether a trillion-parameter train runs on schedule — climbs instead of collapsing under communication overhead.

Deployment

Deployment

The product timeline, as of the September 17 keynote:

Product Readiness
Ascend 910C SuperPod Deployed: 1,000+ units, 370+ customers
Ascend 950 SuperPod Volume commercial use now
Ascend 960 DT Q1 2027 (3 quarters early)
Atlas 860 air-cooled SuperPod Q2 2027
Ascend 960 PR / Atlas 960 liquid-cooled SuperPod Q3 2027 (1 quarter early)
Ascend 970 2028
Ascend 980 2029

The workload claim is not just marketing. The SLAI T-Rex technical report (73 pages, submitted July 2026) presents end-to-end full-parameter post-training of the DeepSeek-V4 model family on an Ascend NPU SuperPOD. It built a hierarchical optimization stack spanning model parallelism, computation-communication overlap, and low-level kernels, reaching 34.22% Model FLOPs Utilization — a 2.93× improvement over the open-source baseline recipe while keeping training stable. On top of that infrastructure, the team ran a CPT + SFT workflow for Operations Research tasks and produced a domain-specialized DeepSeek-V4-Flash variant that hit 71.81% zero-shot Pass@1, beating GPT-5.4-Mini by 3.98 points and the base V4-Flash by 11.27 points.

That combination matters: the hardware can already post-train a trillion-parameter MoE family, and a research group can actually get useful MFU out of it rather than watching communication eat the budget.

Ecosystem

Ecosystem

Huawei is deliberately presenting Ascend as a system, not a chip. To build SuperPods and SuperPod clusters, it has developed 11 serial chips spanning compute, interconnect, and storage: Ascend and Kunpeng carry core compute, with dedicated interconnect and storage chips around them. The strategy is seven pillars, but the three that matter for builders are hardware monetization, system-architecture competition via "SuperPod + cluster," and an open ecosystem that lets mainstream models train natively on Ascend.

Two cross-cutting signals reinforce that the ecosystem is real:

What It Means for Developers

Ascend is now a deployment option, not a footnote. 1,000+ installed SuperPods and 370+ customers mean the software stack (CANN, MindSpore, the parallel frameworks used in SLAI T-Rex) is battle-tested at volume. If you are choosing compute for trillion-parameter training or high-concurrency inference, Ascend SuperPods are no longer a speculative bet — they are a commercial product line with a known roadmap.

Plan around the NPO inflection in 2027. The Atlas 960's 4,096-NPU / 8 EFLOPS FP8 design and 2µs RTT are aimed squarely at trillion-parameter training and high-concurrency inference. Teams signing multi-year capacity deals in 2027 should compare Atlas 860 (air-cooled, Q2 2027) against Atlas 960 (liquid-cooled, Q3 2027) on power budget and rack density — the >550 kW saving and 99.8% availability are datacenter-capEx arguments, not chip-spec arguments.

MFU is the metric to watch. The SLAI T-Rex result (34.22% MFU, 2.93× over open-source) is the number that tells you whether the SuperPod's near-linear bandwidth actually translates to wall-clock training time. Benchmark your own parallelism recipe against that baseline before committing.

Bottom Line

HC2026 is the moment Huawei Ascend SuperPod crossed from validation to commercial product: 1,000+ deployed, 370+ customers, and the Ascend 960 / Atlas 960 NPO platform surfaced three quarters early. The bet is structural — unified memory across thousands of chips, optics moved to the package, and a one-generation-per-year cadence. The SLAI T-Rex report shows the stack already post-trains DeepSeek-V4-class models at useful MFU. For teams planning 2027+ compute capacity, Ascend is now a first-choice rather than a backup plan.

Resources


Huawei Ascend SuperPod and Atlas 960 product details are available at huawei.com; the underlying frontier-model workload is trained and served via platforms such as DeepSeek's API.