China Telecom AI Technology Co., Ltd., with significant backing from Huawei, has officially launched its next-generation large language model, Xing4.0-29B-A4B. Announced on September 17th, this advanced model is designed to excel in complex task comprehension, planning, tool utilization, and autonomous execution. It boasts support for long context windows and is poised for application across a variety of domains, including software development, data analysis, office automation, and everyday services.
A Leap Forward in AI Capabilities
The Xing4.0-29B-A4B model features a total of 29 billion parameters, with only 4 billion actively engaged, making it a highly efficient yet powerful AI. Notably, it marks the first large-scale model in China to be trained on Huawei’s Ascend Atlas 900 A3 SuperPoD liquid-cooled supercomputing cluster and utilizing the MindSpore deep learning framework. This powerful combination has been optimized for demanding engineering tasks.
Leveraging a new architectural design, Xing4.0-29B-A4B significantly enhances its ability to handle long-range tasks and multi-step operations. The model can autonomously devise task strategies, integrate necessary tools, and complete intricate objectives based on user prompts.
Collaborative Development and Technical Prowess
Huawei highlighted the collaborative effort between its teams and China Telecom’s AI division, emphasizing a systemic approach to overcoming challenges in advanced AI development. This collaboration spanned data management, model architecture, training engineering, and reinforcement learning.
Data Foundation
A robust data ecosystem was built, encompassing trillions of tokens for both pre-training and post-training phases. This involved processing petabytes of multi-source heterogeneous data during pre-training, incorporating vast amounts of agent behavioral data and structured knowledge graphs. For post-training, a high-concurrency sandbox container environment was established to generate end-to-end long-range task synthetic data, supported by an automated verification mechanism to ensure data accuracy.
Architectural Innovation
The model’s architecture is specifically engineered for long-range agent tasks. Key innovations include MLA multi-head latent attention for KV cache compression, MT multi-token prediction to improve generation coherence, mHC manifold-constrained hyper-connectivity to address training numerical instability, and Muon and QK-Clip optimizers for training stability. The training process employed a phased strategy, progressively extending sequence lengths from 4K to 32K, 128K, and ultimately 256K, systematically building its long-range capabilities.
Engineering Optimization
Significant engineering optimizations were implemented on the Ascend supercomputing cluster, resulting in a remarkable 96% increase in training performance. This was achieved through multi-level hardware and software co-optimization. The MindSpore framework was fully adapted for mHC residual connections and DSA long sequence mechanisms. For fine-grained MoE (64 routes / Top4 activation), optimizations like expert load balancing, Grouped GEMM, communication-computation overlap, and DVM graph-computation fusion boosted throughput by 32%. Selective recomputation at layer and operator levels provided an additional 19% improvement. Custom Ascend C mHC fused operators addressed Sinkhorn fragmentation, enhancing computational efficiency by 30% and further boosting network throughput by 25%.
Multi-Expert Reinforcement Learning
The Slime reinforcement learning framework was adapted to the Ascend ecosystem, enabling multi-expert reinforcement learning for agent, coding, and reasoning tasks. The MOPD technique was employed for efficient fusion of multi-expert capabilities, significantly enhancing performance in complex reasoning and agent-based tasks.
Ascend Ecosystem Strength
Huawei’s Ascend platform continues to demonstrate its robust capabilities, having supported the native training of over 40 leading industry models. It stands as China’s sole provider supporting commercial-scale pre-training and the deployment of State-of-the-Art (SOTA) large models.
Open Access and Further Information
The Xing4.0-29B-A4B model and its associated resources are available through various platforms:
- GitHub: https://github.com/XingChen-AGI/Xing4.0-29B-A4B
- HuggingFace: https://huggingface.co/XingChen-AGI/Xing4.0-29B-A4B
- ModelScope: https://modelscope.cn/models/XingChen-AGI/Xing4.0-29B-A4B
- Gitee: https://gitee.com/xingchen-agi/xing4.0-29b-a4b
- Modelers: https://modelers.cn/models/XingChen-AGI/Xing4.0-29B-A4B









