Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

Understanding Reasoning from Pretraining to Post-Training

First seen · 7/20/2026, 12:00 PMLatest activity · 7/20/2026, 12:00 PM

This paper studies the interface between pretraining and reinforcement-learning post-training using chess as a controlled testbed. Models ranging from 5M to 1B parameters are pretrained on human chess games, supervised-finetuned on synthetic reasoning traces, and trained with verifiable-reward RL on chess puzzles. The authors report that pretraining loss predicts post-RL performance at a given RL compute budget, while the slope of RL reward curves improves approximately linearly with pretraining tokens. RL reinforces already-preferred correct moves on easy puzzles but can uncover correct moves nearly absent from the SFT policy on hard puzzles. A 1B math-domain model shows a similar pattern.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorarXiv7/18, 12:31 AMnot independent
    Understanding Reasoning from Pretraining to Post-Training
  2. AggregatorHuggingFace Daily Papers7/20, 12:00 PMnot independentRepresentative
    Understanding Reasoning from Pretraining to Post-Training