Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

ACRL: Adaptive Control of Training-Inference Discrepancy for Stable Reinforcement Learning

First seen · 7/27/2026, 03:05 PMLatest activity · 7/27/2026, 03:05 PM

The paper attributes instability in reinforcement learning for large language models partly to training-inference discrepancy caused by separate training and inference engines and by lower-precision inference quantization. It proposes Adaptive Control Reinforcement Learning (ACRL), which adaptively keeps this discrepancy within a reasonable range. According to the abstract, experiments with an FP8 inference engine show that ACRL stabilizes RL training, matches the accuracy of a BF16 baseline, and outperforms importance-sampling fixes. The provided abstract does not report model names, task-level metrics, or the exact control mechanism.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorarXiv7/27, 03:05 PMnot independentRepresentative
    ACRL: Adaptive Control of Training-Inference Discrepancy for Stable Reinforcement Learning