Huawei Technologies unveiled a new computing architecture on September 17, 2026, that it says is designed to let as many as one million processors function together as a single computer. The announcement came at Huawei Connect 2026 in Shanghai, where the company also detailed new hardware built on the design and laid out an accelerated roadmap for its Ascend AI chips.

Huawei calls the new design the Peerium Computing Architecture. In its own announcement, the company said the architecture "breaks through the Turing paradigm" through a technique it calls Nested BSP, short for Nested Bulk Synchronous Parallel processing, extends the traditional von Neumann single-machine model, and moves away from the master-slave computing structure that has dominated large systems for decades.

It is worth being precise about what this claim means. Huawei is not describing a single physical chip with a million processing cores. It is describing a distributed system architecture intended to coordinate very large numbers of separate processors, memory units and storage devices so they behave, from a software and workload perspective, more like one large machine than a collection of separate computers linked by a network.

How UnifiedBus ties the system together

The interconnect technology at the center of Peerium is called UnifiedBus, known in Chinese as Lingqu. According to Huawei, UnifiedBus is built on a single open protocol and functions as a high-speed bus connecting CPUs, NPUs, memory, solid-state drives, network interface cards and switches. The company describes this as peer-to-peer interconnection, meaning components communicate directly with one another across compute, storage and networking layers rather than passing everything through a central controller.

Huawei says Peerium reaches million-processor scale through three combined elements: nested parallelism, unified memory addressing and this peer-to-peer interconnect. Nested parallelism organizes work in layers, so that large computing jobs are broken into parallel phases within parallel phases, rather than being scheduled as one flat set of tasks. Unified memory addressing allows processors across the system to reference a shared memory space, reducing the need to copy data back and forth between separate pools of memory.

Huawei's rationale for this approach is that raw chip speed alone is no longer enough to meet the computing demands of modern AI systems. As AI models and AI agents grow larger and more complex, the bottleneck increasingly shifts from individual processor performance to how efficiently thousands or millions of processors can share data and stay synchronized. Building a faster single chip does little if the surrounding system cannot move information between chips quickly enough to keep them all working productively.

Atlas 950 and the coming Atlas 960

Huawei said the Atlas 950 SuperPoD and the larger SuperPoD-based SuperClusters built from it are the first generation of products designed around the Peerium architecture. According to the company, an Atlas 950 SuperCluster containing 256,000 computing cards is already being deployed. Huawei has not published independent, third-party benchmark data confirming performance at that full scale, and this deployment figure should be understood as a company-reported number rather than an externally audited result.

Huawei also confirmed that its next-generation system, the Atlas 960, is currently under testing rather than in commercial deployment. At Huawei Connect 2026, the company introduced a specific version of this system, the Atlas 960E SuperPoD, which it described as the industry's first to use near-packaged optics. Huawei said each Atlas 960E pod holds 4,096 NPUs, delivers 8 exaflops of FP8 compute, and carries up to a petabyte of high-bandwidth memory. The optical technology behind it, called Hi-ONE, is Huawei's own near-packaged optical engine with a built-in light source, which the company said moves data at 7.2 terabits per second per unit.

An accelerated chip roadmap

Huawei first laid out a multi-year Ascend chip roadmap at its September 2025 Huawei Connect event, which included the Ascend 950PR and 950DT chips arriving in 2026, followed by the Ascend 960 in 2027 and the Ascend 970 in 2028. At this year's event, Huawei updated that timeline. The company said its Ascend 960DT chip is now scheduled for the first quarter of 2027, three quarters earlier than the date given a year earlier, with an inference-focused Ascend 960PR variant following in the third quarter of 2027. Huawei also extended its roadmap further out than before, adding an Ascend 970 chip for 2028 and an Ascend 980 for 2029. The company attributed this faster cadence to what it calls its Tau Scaling Law, a framework it uses internally to plan chip generations.

Why system-level scale matters for China

Huawei's emphasis on connecting enormous numbers of processors, rather than competing purely on the speed of any single chip, reflects the constraints the company faces under continued US restrictions on advanced semiconductor manufacturing technology available to Chinese firms. Those restrictions limit Huawei's access to the most advanced foundry processes used by rivals, making it harder to match competitors on a chip-for-chip basis. By instead focusing on how efficiently large numbers of chips can be linked, coordinated and kept fed with data, Huawei is pursuing a strategy where system architecture, interconnects and software coordination compensate for a gap in individual processor manufacturing technology.

This approach also fits into China's broader push to build domestic AI computing infrastructure that does not depend on foreign suppliers such as Nvidia. Huawei has said it has developed its own high-bandwidth memory technology for use in its chips, reducing reliance on external memory suppliers, a step that is separate from but related to its system-level interconnect work.

How this compares with Nvidia's approach

Nvidia's dominant position in AI computing rests heavily on its NVLink interconnect technology and a tightly integrated software stack that keeps large clusters of GPUs operating efficiently together. Huawei has published its own comparative figures against Nvidia's upcoming NVL144 system, claiming substantially greater scale, computing power, memory capacity and interconnect bandwidth for its Atlas 950 SuperPoD. These figures come directly from Huawei and have not been independently verified through neutral, third-party testing, so they should be read as the company's own projections rather than confirmed, audited benchmarks.

Analysts covering the announcement have noted that publishing an ambitious roadmap and comparative numbers is not the same as demonstrating sustained real-world performance across large model training and inference workloads. Nvidia's advantage has historically come as much from software maturity and ecosystem support as from raw hardware specifications, an area where Huawei's newer platform has less of a track record to point to.

What remains unproven

Huawei's claim that Peerium can scale to one million processors describes the architecture's intended design goal, not a benchmark result Huawei has published from a system actually running at that scale. The confirmed, deployed figure the company has disclosed is the 256,000-card Atlas 950 SuperCluster, and even that number rests on Huawei's own disclosure rather than independent verification. The Atlas 960 system, along with the accelerated Ascend 960DT and 960PR chips, remain in testing or pre-launch stages rather than commercial availability. As with any newly announced computing architecture, its real value will depend on how it performs once deployed at scale for actual AI training and inference workloads, evidence that is not yet available.

Further reading and useful links

Reader questions

Frequently asked questions

Does Huawei's "one million processors" claim mean it built a single chip with a million cores?

No. Peerium is a distributed system architecture designed to coordinate very large numbers of separate processors so they work together efficiently, not a single physical processor with that many cores.

Has Huawei demonstrated a system actually running at the one-million-processor scale?

Not according to available information. The confirmed, deployed system Huawei has disclosed is a 256,000-card Atlas 950 SuperCluster. The million-processor figure describes the architecture's intended scaling capability.

What is UnifiedBus and why does it matter?

UnifiedBus is Huawei's high-speed interconnect protocol that links CPUs, NPUs, memory, storage and networking hardware through peer-to-peer communication. It is the technology Huawei says makes coordinating large numbers of processors possible.

How does this relate to US restrictions on advanced chips for China?

US export controls limit Huawei's access to the most advanced chip manufacturing processes. By focusing on system-level interconnects and architecture rather than competing purely on individual chip speed, Huawei aims to build competitive AI computing capacity using domestically available technology.


Corrections and updates

Nexuswild welcomes factual corrections. Email [email protected] with evidence and the article URL.