Huawei has introduced the Atlas 960E SuperPoD, an AI system built around near-packaged optics and designed to connect thousands of accelerators as one tightly coupled machine. In its official HUAWEI CONNECT 2026 announcement, the company says the liquid-cooled system can scale to 4,096 NPUs and deliver 8 exaflops of FP8 performance.
The launch targets a problem that becomes harder as models grow: moving data between processors, memory and storage without spending too much power or time on the interconnect. Huawei says the Atlas 960E uses its UnifiedBus architecture and Hi-ONE optical engines to keep those transfers inside a single system design. The announcement is part of a larger September 17 keynote about infrastructure for models approaching 10 trillion parameters.
Optics move closer to the accelerator
Huawei’s headline change is the use of near-packaged optics, or NPO, for a SuperPoD. The company says Hi-ONE provides 7.2 terabits per second of transmission capacity per engine and includes an integrated light source. That is a vendor claim about the interconnect component, not an independent benchmark of an entire AI workload.
A single Atlas 960E contains 5,500 Hi-ONE units, according to Huawei. The company says that replaces the need for 48,000 conventional 800G optical modules to connect the NPUs, cutting power consumption by more than 550 kilowatts. Huawei also claims the system doubles fault-free operating time and reaches 99.8% availability. Those figures describe Huawei’s stated design targets and should not be read as a tested result from an outside lab.
The system uses Huawei’s Ascend 960 family and unified memory addressing through UnifiedBus. Huawei lists 8 exaflops at FP8 and 16 exaflops at FP4 for a single SuperPoD, alongside up to 1 petabyte of HBM capacity. The company has also said the Ascend 960DT is due in the first quarter of 2027, with the 960PR following in the third quarter.
That emphasis on the connective tissue of an AI system matters because accelerator speed alone does not determine useful throughput. Our memory-bandwidth explainer covers why moving weights and intermediate data can become the limiting step, while our TPU guide explains how specialized accelerators differ from general-purpose GPUs.
A bigger system, with a careful caveat
Huawei says multiple Atlas 960E systems can be combined with its UnifiedBus network or RoCE. Its stated architecture reaches 512,000 NPUs with a two-tier, four-plane Clos design and can scale to one million NPUs with a multi-rail topology. Those are architecture claims, not evidence that a million-NPU installation is already operating in production.
The company also used HUAWEI CONNECT to present a 4,096-node TaiShan 950 SuperPoD, OceanStor M900 context-memory storage and an open-source direction for its CANN software stack. Together, the announcements show Huawei treating compute, optical networking and cache storage as one infrastructure product rather than as isolated components.
For buyers, the immediate question is less whether a large number on a slide beats a rival chip and more whether the whole system can be supplied, programmed and operated at scale. Huawei’s Atlas 960E is a clear statement that the interconnect is now part of that competition. Availability, independent testing and real customer deployments will determine how much of the claimed advantage reaches working AI clusters.