Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

Beyond Rephrasing: Book-Level Organization Improves Synthetic Textbook Data for Mid-Training

First seen · 7/30/2026, 08:15 PMLatest activity · 7/30/2026, 08:15 PM

This paper studies whether synthetic textbook data benefits from book-level organization, beyond content quality or local rewriting. Its pipeline retrieves material from a pre-training corpus, clusters it by topic, plans hierarchical tables of contents, and assembles source-grounded sections into complete books. The Full setting produces 686K textbooks containing 32B tokens across more than 15,000 disciplines. In controlled mid-training comparisons, packaging identical sections as coherent books yields a +1.02 mean downstream gain over the Split setting, while RandomConcat and document-level Rephrase perform worse. Results are also reported on Llama3-8B.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorarXiv7/30, 08:15 PMnot independentRepresentative
    Beyond Rephrasing: Book-Level Organization Improves Synthetic Textbook Data for Mid-Training