Huawei unveils Ascend 950PR Atlas 350 with 2.9x Nvidia H20 performance as China scales AI inference

Huawei unveils Ascend 950PR Atlas 350 with 2.9x Nvidia H20 performance as China scales AI inference

Huawei Technologies used its China Partner Conference 2026 to push its Ascend computing stack into broader commercial deployment, unveiling the Ascend 950PR-powered Atlas 350 accelerator card and positioning it as a higher-throughput alternative to Nvidia’s H20 for inference workloads.

At the event, Huawei said the Atlas 350 delivers 2.87 times the single-card compute of the Nvidia H20 and is the only domestically available product in China that supports FP4 low-precision inference, a feature aimed at cutting latency and raising throughput in high-concurrency recommendation systems.

The launch immediately broadened Huawei’s server supply chain: seven core partners, including Kunlun, Hua Kun Zhenyu, ShenZhou KunTai, Changjiang Computing, Powerleader, SoftStone Huafang and Bx, announced complete server products based on Atlas 350, marking Ascend 950-generation inference compute entering the commercial phase, according to the Shanghai Securities News.

Partners Roll Out Server SKUs to Speed Commercial Adoption

SoftStone Information Technology subsidiary SoftStone Huafang introduced a 6U dual-socket AI server product, “Super A860 A5,” designed to support up to eight Atlas 350 accelerator cards and based on a new model of the Kunpeng 920 processor, according to statements made at the conference.

iFlytek said its next-generation Spark large model will be adapted to the Ascend 910/950 compute base, signaling that Chinese model developers are aligning product roadmaps more tightly with domestic accelerators to standardize training and inference deployment.

Huawei Highlights FP4 Inference and Memory Bandwidth as Differentiators

Huawei’s Ascend computing president Zhang Dixuan said Atlas 350 pairs FP4 support with higher memory capacity and bandwidth to improve multimodal generation and inference efficiency. The card’s HBM capacity reaches 112GB, which Huawei said is 1.16 times that of the H20, while multimodal generation speed can rise by 60%.

On-site specifications displayed at the booth put Atlas 350’s FP4 compute at 1.56P and memory bandwidth at 1.4TB/s. Huawei also disclosed a 600W power draw, which it said is about 1.5 times the H20, underscoring a tradeoff that data-center buyers will need to evaluate against throughput gains and rack-level power budgets.

Huawei staff at the exhibition said supporting FP16, FP8 and FP4 allows servers integrating Atlas 350 to run larger models with lower inference latency. In measured tests for internet recommendation scenarios, the card showed lower latency and faster response, making it suitable for short-video, e-commerce and advertising recommendation workloads. For multimodal inference such as text-to-image and text-to-video, Huawei said performance is comparable to Nvidia’s L20.

Huawei Expands “Three-Tier” Compute Strategy to Map Chips to Model Sizes

Huawei framed Atlas 350 as part of a broader Ascend roadmap that targets three core scenarios across model scales. For trillion-parameter models, Huawei promoted its Ascend 384 “supernode,” citing “ultra-large bandwidth, ultra-low latency and unified memory addressing” to enable linear scaling of effective compute and support training and inference, and said it has already landed in multiple industries.

For hundred-billion-parameter models, Huawei is positioning “out-of-the-box” single-server systems as a balance between fast deployment and controllable cost. For ten-billion-parameter models, it is offering more compute tiers and higher-integration modules and cards, paired with wider OS compatibility and scenario SDKs to enable partners to build diversified products.

Industry Solutions and “All-in-One” Boxes Aim for Repeatable Deployment

Huawei said it jointly released 2026 Ascend AI scenario solutions with 20 industry partners, covering use cases including office assistance, AI training labs, electronic medical records, intelligent customer service and government office workflows, emphasizing lightweight deployment and scalability.

Huawei Vice President Ma Haixu said demand for all-in-one machines has risen again, and that in the past month-plus more than a dozen partners launched Ascend-based OpenClaw all-in-one products. Huawei said it has built more than 400 industry all-in-one models with partners, serving more than 2,700 customers and accounting for over 80% of China’s domestic all-in-one market.

Related Coverage:

Huawei Aggressively Prices Band 11 Series to Corner China’s Entry-Level Wearable Market

Huawei’s QiJing GT7 Debuts With Three First-to-Market QianKun Features, Signaling Deeper GAC Tie-Up

Subscribe to ChinaBiz Insider

Don’t miss out on the latest issues. Sign up now to get access to the library of members-only issues.
[email protected]
Subscribe