Huawei details AI accelerator roadmap, pulls in next-generation Ascend NPUs by several quarters — FP4 performance of the Ascend 960PR doubles expectations

Wait 5 sec.

Huawei has updated its AI hardware roadmap by adding new accelerators and supporting processors and pulling in next-generation Ascend 960 accelerators at its annual Huawei Connect event. Specifically, the company accelerated its Ascend 960 roadmap, disclosed Ascend 970 and 980 specifications, introduced its Peerium architecture based on the UnifiedBus, and expanded its vertically integrated AI infrastructure portfolio.Huawei is currently in the middle of transitioning from its SIMD architectures that it has used for almost a decade with its Ascend accelerators (or neural processing units, how the company prefers to call them) to its all-new SIMD+SIMT architectures that bring together vector-based processing and thread-level parallelism to improve hardware utilization and performance across a variety of AI workloads (SIMD for data parallel operations and SIMT for branch-heavy workloads). Image is for illustrative purposes only. (Image credit: Huawei)The first Ascend NPUs to adopt Huawei's new architecture are Ascend 950PR for prefill and recommendation, as well as Ascend 950DT for decoding and training. Huawei said at the event that its Ascend 950 platform is gaining traction as the Atlas 950 SuperPoD systems are already in large-scale commercial use, though it did not elaborate. The company said tests of its training-oriented Ascend 950DT have produced 'good results' and expects numerous Chinese AI developers to begin training models on 950DT-based systems next year. Meanwhile, Huawei acknowledged that its production capacity remains insufficient to satisfy domestic demand.Indeed, in September 2025, Huawei announced the maximum Atlas 950 SuperPoD configuration as 2,048 Kungpeng 950 CPUs, 8,192 Ascend 950DT NPUs, 160 cabinets (128 compute + 32 communications), 8 FP8 EFLOPS, 16 FP4 EFLOPS, and 16 PB/s of aggregate interconnect bandwidth. However, in July 2026 Huawei publicly showed a real Atlas 950 SuperPoD implementation with 256 CPUs as well as 1,024 accelerator cards, which is well below the maximum configuration. While the company still describes the architecture as scaling up to 8,192 NPUs, it is not listed on its website, so we can only wonder which systems are now in large-scale commercial use. For now, the adoption of the Atlas 950 SuperPod does not seem to be proceeding rapidly, perhaps because of insufficient supply, or maybe because of the all-new architecture that requires major redesign of software. In any case, the Atlas 950 SuperPod will in many ways be a pipecleaner for the company to clear the road for more capable Ascend 960-series accelerators and their successors. Speaking of the Ascend 960, this family will start with the Ascend 960DT in Q1 2027, when it is set to be formally available, three quarters earlier than previously planned. (Image credit: Huawei)The Ascend 960DT accelerator is expected to deliver 2 FP8 PFLOPS and 4 FP4 PFLOPS, carries 288 GB of presumably HiZQ memory with 9.6 TB/s bandwidth, and features a 2.2-TB/s interconnect. The Ascend 960PR NPU follows in Q3 2027, one quarter earlier than originally planned, with 2 FP8 PFLOPS for training, but 8 FP4 PFLOPS for inference (2X higher than Huawei announced last year). The unit carries 192 GB of memory providing 2.4 TB/s of bandwidth and retains the 2.2-TB/s interconnect. For comparison: Nvidia's VR200 GPU due in Q4 2026 can deliver 35 NVFP4 PFLOPS for training and 50 NVFP4 PFLOPS for inference while carrying 288 GB of HBM4 memory."We are evolving our Ascend chip series on a one-generation-a-year cycle," said David Wang, the Deputy Chairman of the Board and Rotating Chairman at Huawei, in his keynote. "In 2028 and 2029, we will roll out the Ascend 970 and 980 chips, respectively. Thanks to the Tau (τ) Scaling Law, not only will their compute specifications continue to double, but you can also expect to see huge improvements across the board in terms of memory bandwidth, memory capacity, interconnect bandwidth, and more."Huawei Ascend roadmapNPUTargeted ReleaseArchitectureFP8 PerformanceFP4 PerfMemoryMemory BandwidthInterconnect BandwidthSupported Formats Ascend 910C2025 Q1SIMD––128 GB3.2 TB/s784 GB/sFP32, HF32, FP16, BF16, INT8 Ascend 950PR2026 Q1SIMD + SIMT1 PFLOPS2 PFLOPS128 GB of HiBL 1.01.6 TB/s2.0 TB/sFP32, HF32, FP16, BF16, FP8, MXFP8, HiF8, MXFP4 Ascend 950DT2026 Q4SIMD + SIMT1 PFLOPS2 PFLOPS144 GB of HiZQ 2.04.0 TB/s2.0 TB/sFP32, HF32, FP16, BF16, FP8, MXFP8, HiF8, MXFP4 Ascend 960DT2027 Q1SIMD + SIMT2 PFLOPS4 PFLOPS288 GB9.6 TB/s2.2 TB/sFP32, HF32, FP16, BF16, FP8, MXFP8, HiF8, MXFP4, HiF4 Ascend 960PR2027 Q3SIMD + SIMT2 PFLOPS8 PFLOPS192 GB2.4 TB/s2.2 TB/sFP32, HF32, FP16, BF16, FP8, MXFP8, HiF8, MXFP4, HiF4Ascend 9702028SIMD + SIMT3.6 PFLOPS14 PFLOPS288 GB14.4 TB/s4.4 TB/sFP32, HF32, FP16, BF16, FP8, MXFP8, HiF8, MXFP4, HiF4Ascend 9802029SIMD + SIMT7.2 PFLOPS*28 PFLOPS*384 GB38.4 TB/s*8 TB/s*FP32, HF32, FP16, BF16, FP8, MXFP8, HiF8, MXFP4, HiF4*Starting with the Ascend 960-series and onwards, Huawei plans to maintain a one-generation-per-year cadence for its AI accelerators. Pulling in the Ascend 960DT by several quarters is, without any doubt, a remarkable achievement. However, what is even more extraordinary is that Huawei has managed to increase FP4 performance of the Ascend 960PR by two times compared to original expectations, which likely means that the company has substantially reworked the processor's low-precision compute capabilities rather than merely adjusted its memory subsystem or clock speeds. In fact, four-fold higher FP4 performance compared to FP8 is set to be a distinctive feature of Ascend 970 and 980.The Ascend 970 is due in 2028 with 3.6 FP8 PFLOPS, 14 FP4 PFLOPS, 288 GB of memory providing 14.4 TB/s, and 4.4 TB/s of interconnect bandwidth. Ascend 980 follows in 2029 with 7.2 FP8 PFLOPS and 28 FP4 PFLOPS, along with 384 GB of memory reaching 38.4 TB/s and an 8-TB/s interconnect. Huawei marks the Ascend 980 figures as preliminary.