This paper studies zero reinforcement learning with verifiable rewards at trillion-parameter scale, a regime largely unexplored because of compute constraints. It introduces a training pipeline using clipped importance sampling, training-inference ratio correction, and mixed-precision control. The authors report that Ring-2.5-1T-Zero reaches competitive results across seven mathematical benchmarks. They also describe a two-stage training pattern, discovery followed by sharpening, and claim emergent behaviors such as self-verification, structured formatting, parallel reasoning, anthropomorphism, and context anxiety. A three-dimensional framework evaluates chain-of-thought quality through comprehensibility, reproducibility, and efficiency.
No heat snapshots are available in the last 24 hours.