Moore Threads Unveils New Architecture and Roadmap to Scale AI Clusters to 100,000 Chips
Moore Threads Technology unveiled an aggressive technical roadmap and a next-generation GPU architecture at its inaugural developer conference in Beijing on Saturday, signaling a significant leap in China's domestic computing capabilities. The announcement marks an intensified effort to build an autonomous accelerated computing ecosystem capable of supporting large-scale artificial intelligence model training amid ongoing global semiconductor trade restrictions.
Company founder and CEO Zhang Jianzhong introduced "Flower Harbor," a new full-function GPU architecture designed to increase computing density by 50% and energy efficiency by 10 times compared to previous iterations. Based on this architecture, Moore Threads is preparing to launch two key chips: "Huashan" for combined AI training and inference, and "Lushan" for high-performance graphics rendering. The company asserts that the new specifications for memory bandwidth and interconnect speed will exceed current industry benchmarks, positioning the hardware to compete with global tier-one products.
The Beijing-based chipmaker also demonstrated the operational status of its "Kuae" computing cluster, which currently links 10,000 GPUs, and outlined concrete plans to scale this infrastructure to support clusters of 100,000 cards. This expansion is critical for training trillion-parameter models, a necessity for competing in the generative AI sector. In a demonstration of real-world application, the company showcased performance data for the DeepSeek R1 large model, claiming its S5000 card achieved inference throughput speeds that set new records for domestic GPUs.
The event, held in late 2025, underscores the rapid maturation of China’s local supply chain for advanced computing. By releasing a full-stack update covering hardware, software, and cloud infrastructure, Moore Threads is attempting to reassure investors and developers that domestic platforms can offer a viable alternative to restricted foreign technology for both high-end model training and edge computing applications.
New Architecture and High-Performance Silicon
The centerpiece of the roadmap is the "Flower Harbor" architecture, which supports end-to-end full-precision computing ranging from FP4 to FP64. The architecture introduces "MTLink" high-speed interconnect technology, supporting chip-to-chip speeds of 1,314 GB/s, a critical feature for building the massive clusters required for modern AI workloads.
Based on this architecture, the upcoming "Huashan" chip focuses on AI training and ultra-large-scale intelligent computing. It integrates a new generation of asynchronous programming and full-precision tensor calculation units. Moore Threads claims the chip’s floating-point computing power, memory bandwidth, and interconnect speeds outperform the industry standard "HXXX" products and approach the specifications of "BXXX" series products in certain configurations.
Simultaneously, the company announced the "Lushan" chip for the graphics market. Compared to the previous MTT S80, the Lushan chip reportedly delivers a 15-fold increase in AAA gaming performance, a 64-fold boost in AI computing performance, and a quadrupling of video memory capacity. It also features a new hardware ray-tracing engine, aiming to close the gap in high-fidelity rendering.
Scaling Infrastructure and Inference Breakthroughs
Addressing the demands of data centers, Moore Threads confirmed the completion of its "Kuae" 10,000-card intelligent computing cluster. The system boasts a floating-point operation capability of 10 ExaFLOPS (EFLOPS) and supports the training of trillion-parameter models. The company reported a training operation linearization efficiency of 95% and an effective training time ratio exceeding 90%, metrics that are vital for the economic viability of large model training.
On the inference front, the company highlighted the performance of its S5000 card running the DeepSeek R1 671B model. Through joint optimization with SiliconFlow, the single-card throughput for "Prefill" exceeded 4,000 tokens per second, while "Decode" speeds surpassed 1,000 tokens per second. The company stated these figures represent a new ceiling for domestic GPU inference performance, significantly lowering the cost barrier for deploying large models.
Looking ahead, Zhang announced plans for the MTT C256 super-node architecture, designed to facilitate the construction of 100,000-card clusters. The design utilizes a high-density hardware architecture to minimize bandwidth loss and latency, directly targeting the needs of next-generation hyperscale data centers.
Software Ecosystem and Edge Computing
To support its hardware, Moore Threads released MUSA 5.0, a significant upgrade to its full-stack software platform. The update focuses on compatibility, supporting both CUDA C and the native MUSA C programming languages to lower the migration threshold for developers. The company also announced plans to open-source core components, including computation and communication libraries, to foster a broader developer community.
Extending beyond the data center, the company introduced the "Changjiang" intelligent SoC and the MTT AIBOOK, a laptop-style device aimed at developers. The "Changjiang" chip integrates a high-performance CPU, GPU, and NPU, offering 50 TOPS of heterogenous AI computing power. The AIBOOK is positioned as an "out-of-the-box" AI development platform capable of running models up to 30 billion parameters locally, bridging the gap between heavy cloud training and edge deployment.
This comprehensive product suite reflects a strategy to capture value across the entire computing spectrum, from cloud infrastructure to end-user devices, thereby solidifying the MUSA ecosystem as a foundational layer for China's digital economy.