Huawei's impressive next-generation Ascend 900-series AI accelerators will be offered only in China, not internationally, as the company struggles to meet domestic demand amid capacity constraints, the company announced this week. While the upcoming Ascend 960-series neural processing units (NPUs) could rival some of AMD's and Nvidia's existing AI GPUs, demand for these units outside of China was not guaranteed anyway.Go deeper with TH Premium: AI shortages(Image credit: Nvidia)AI data centers are swallowing the world's memory and storage supplyDemand for data center CPUs has surged, and AI agents are responsibleChip scarcity assaults auto industry amid the worsening Nexperia and DRAM crisisThe custom AI ASIC state of play"Since we do not have enough capacity to even satisfy the demand in China, we do not have a plan to expand into the international market in a fully-fledged way," said Eric Xu, rotating chairman of Huawei, on the sidelines of the company's Huawei Connect conference, Reuters reports. He added that Huawei supplies limited volumes to 'some countries where demand is particularly strong,' though he did not elaborate.Huawei this week unveiled its latest AI accelerator roadmap, revealing major training and inference performance gains for its next-generation Ascend 960, 970, and 980 NPUs over the existing Ascend 910C and Ascend 950-series. The Ascend 960DT and 960PR are set to increase their FP8 training performance to 2 PFLOPS and their FP4 inference performance to 4 PFLOPS and 8 PFLOPS, respectively, in 2027. Meanwhile, their successors, Ascend 970 and Ascend 980, are projected to increase their FP4 performance to 14 PFLOPS and 28 PFLOPS, respectively, in the coming years.Huawei Ascend vs Nvidia AI GPUsNPUFP8 PerformanceFP4 PerfMemoryMemory BandwidthInterconnect BandwidthTargeted ReleaseNvidia H2004 PFLOPS-141 GB HBM3E4.8 TB/s900 GB/s2023 Q4Nvidia B30010 PFLOPS15/20 S/D PFLOPS279 GB HBM3E8 TB/s1.8 TB/s2025 Q4Ascend 950PR1 PFLOPS2 PFLOPS128 GB of HiBL 1.01.6 TB/s2 TB/s2026 Q1Ascend 950DT1 PFLOPS2 PFLOPS144 GB of HiZQ 2.04.0 TB/s2 TB/s2026 Q4Nvidia R20017.5 PFLOPS35/50 T/I PFLOPS 288 GB HBM419.2 TB/s3 TB/s2026 Q4Ascend 960DT2 PFLOPS4 PFLOPS288 GB9.6 TB/s2.2 TB/s2027 Q1Ascend 960PR2 PFLOPS8 PFLOPS192 GB2.4 TB/s2.2 TB/s2027 Q3Ascend 9703.6 PFLOPS14 PFLOPS288 GB14.4 TB/s4.4 TB/s2028Ascend 9807.2 PFLOPS*28 PFLOPS*384 GB38.4 TB/s*8 TB/s2029*Preliminary dataS/D - Sparse and DenseT/I - Training and InferenceBut while the upcoming Ascend NPUs will be considerably faster than their predecessors, particularly for inference, they will remain well behind Nvidia's previous- and current-generation accelerators, at least in raw compute performance. Huawei's 2027 Ascend 960DT is projected to deliver 2 FP8 TFLOPS for training, compared with Nvidia's 4 FP8 TFLOPS for the H200, released in 2023. The Ascend 960PR is expected to offer 8 FP4 PFLOPS for training, which is far behind Nvidia's B300, which delivers 15–20 NVFP4 PFLOPS. Even the Ascend 980, targeted for 2029, is projected to reach 7.2 FP8 PFLOPS and 28 FP4 PFLOPS, well below Nvidia's R200, which is on track to deliver 17.5 FP8 PFLOPS and 35/50 FP4 PFLOPS this year.Such a massive performance difference with leading AI hardware will reinforce Huawei's reliance on massive system-level scaling rather than chip-for-chip performance to compete with Nvidia. But massive system-level scaling comes with massive power consumption, which will make Huawei's next-generation Atlas SuperPoDs and SuperClusters considerably less competitive in markets that can access hardware from AMD or Nvidia.Huawei is in an interesting paradoxical situation. On the one hand, its integration efforts like near-package optics (NPO) clearly free up capacity on 'older' nodes that can be used for other components of AI platforms. But on the other hand, SMIC's inability to ramp production on 7nm and 6nm-class nodes limits Huawei's ability to supply its AI hardware anyway, which is why it can barely meet demand.Then again, while Huawei's Atlas SuperPoDs with up to 15,488 Ascend 960 NPUs can deliver up to 30 FP8 EFLOPS and 120 FP4 EFLOPS performance by far exceeding the capabilities of Nvidia's NVL72 clusters with a 72-GPU scale-up world size, their performance-per-watt is poised to be dramatically lower compared to Nvidia's architectures, which means that demand for such hardware outside of China will be limited at best. That said, a global AI hardware push doesn't make much sense for Huawei right now. What perhaps does make sense is offering cloud access to its hardware to various academic and research customers to popularize its CANN software stack.