Li, Dongfang, Luo, Xiaodong, Sun, Ruoyu, Chen, Xuhui, Qiu, Linyuan, Meng, Jian, Lu, Zhengxuan, Wang, Yiting, Xie, Yucheng, Guo, Tao, Fang, Tianxiang, Li, Jing, Chen, Sihang, Hong, Shihao, Liu, Chang, Dai, Weihua, Zeng, Zirong, Zhu, Ziwei, Wang, Zhuohan, Yue, Zhengjun, Vasilyev, Igor, Liu, Min, Sun, Weijian, Chen, Xin, Gao, Yingmeng, Zhou, Jinhua, Chen, Taolue, Wu, Chenwei, Zhang, Dong, Jin, Wenlong, Xiang, Jinmin, Maria, Barkova, Anton, Ushakov, Jin, Xianfei, Ding, Tian, Lin, Zhihang, Chen, Qian, Yang, Linxin, Yang, Mingzhe, Zhang, Bingwei, Yang, Hongzhang, Zhang, Fangxue, Qin, Shijun, Yu, Jie, Hu, Cuihua, Vasiliy, Tolstykh, Ivan, Nosov, Amir, Abdullin, Zhou, Zhichen, Zhang, Xin, Ning, Zhixiong, Zhao, Xutong, Huang, Junjie, Liu, Jiajun, Kong, Weiyan, Zhang, Zheng, Luo, Wenhan, Hu, Lin, Guo, Yangbo, Zeng, Li, Zeng, Shihao, Hu, Baotian, Zhang, Min, Li, Haizhou, Luo, Zhiquan
Abstract
Full-parameter post-training of trillion-parameter-scale MoE models introduces substantial system-level challenges for large-scale distributed training, including severe memory pressure, non-overlapped communication overhead, and inefficient kernel execution. While most large-scale LLM training systems are built around GPU-based clusters, this report presents an end-to-end optimization practice on the Ascend NPU SuperPOD. Using the DeepSeek-V4 model family as the target workload, we develop a hierarchical optimization framework spanning model-level parallelism, computation-communication orchestration, and low-level kernel execution. The resulting system achieves 34.22% Model FLOPs Utilization (MFU) with a 2.93x improvement over the open-source baseline recipe while maintaining training stability. Building on this optimized infrastructure, we further establish a CPT and SFT workflow for complex Operations Research (OR) tasks. We refer to the integrated framework as SLAI T-Rex. Using DeepSeek-V4-Flash, we develop OR-oriented CPT and SFT data pipelines that combine collected domain resources with solver-verified synthetic optimization documents. The resulting dataset contains 10K high-quality SFT samples spanning four task categories and three problem representations. The specialized model achieves the highest average zero-shot Pass@1 score among the evaluated models, reaching 71.81% and outperforming GPT-5.4-Mini and the base DeepSeek-V4-Flash model by 3.98 and 11.27 percentage points, respectively. Overall, this work demonstrates a full-stack pathway from efficient trillion-parameter model post-training on Ascend infra to domain-specialized Flash models for solver-grounded mathematical modeling, advancing frontier-model systems for complex reasoning.
Chinese Translation
对万亿参数规模的MoE模型进行全参数后训练引入了大规模分布式训练的重大系统级挑战,包括严重的内存压力、非重叠的通信开销和低效的内核执行。尽管大多数大规模LLM训练系统是基于GPU集群构建的,但本报告展示了在Ascend NPU SuperPOD上的端到端优化实践。以DeepSeek-V4模型系列作为目标工作负载,我们开发了一个层次化优化框架,涵盖模型级并行性、计算-通信协调和低级内核执行。最终系统实现了34.22%的模型FLOPs利用率(MFU),相较于开源基线方案提高了2.93倍,同时保持了训练的稳定性。在此优化基础设施上,我们进一步建立了复杂运筹学(OR)任务的CPT和SFT工作流程。我们将这一集成框架称为SLAI T-Rex。利用DeepSeek-V4-Flash,我们开发了面向运筹学的CPT和SFT数据管道,将收集的领域资源与求解器验证的合成优化文档相结合。最终数据集包含10K个高质量的SFT样本,涵盖四个任务类别和三种问题表示。该专用模型在评估的模型中实现了最高的平均零-shot Pass@1得分,达到71.81%,分别比GPT-5.4-Mini和基础DeepSeek-V4-Flash模型高出3.98和11.27个百分点。总体而言,本研究展示了从在Ascend基础设施上高效的万亿参数模型后训练到针对求解器基础的数学建模的领域专用Flash模型的全栈路径,推动了复杂推理的前沿模型系统的发展。