The paper attributes instability in reinforcement learning for large language models partly to training-inference discrepancy caused by separate training and inference engines and by lower-precision inference quantization. It proposes Adaptive Control Reinforcement Learning (ACRL), which adaptively keeps this discrepancy within a reasonable range. According to the abstract, experiments with an FP8 inference engine show that ACRL stabilizes RL training, matches the accuracy of a BF16 baseline, and outperforms importance-sampling fixes. The provided abstract does not report model names, task-level metrics, or the exact control mechanism.
No heat snapshots are available in the last 24 hours.