This report describes an end-to-end system for full-parameter post-training of trillion-parameter-scale DeepSeek-V4 models on an Ascend NPU SuperPOD. The work optimizes model parallelism, computation-communication orchestration, and kernel execution, reaching 34.22% Model FLOPs Utilization, reported as a 2.93x improvement over an open-source baseline recipe. The authors also build CPT and SFT pipelines for operations-research tasks using DeepSeek-V4-Flash, including solver-verified synthetic documents. A 10K-sample dataset covers four task categories and three problem representations. The specialized model reports a zero-shot average Pass@1 of 71.81%, exceeding GPT-5.4-Mini by 3.98 points and the base DeepSeek-V4-Flash by 11.27 points.
No heat snapshots are available in the last 24 hours.