This paper studies whether synthetic textbook data benefits from book-level organization, beyond content quality or local rewriting. Its pipeline retrieves material from a pre-training corpus, clusters it by topic, plans hierarchical tables of contents, and assembles source-grounded sections into complete books. The Full setting produces 686K textbooks containing 32B tokens across more than 15,000 disciplines. In controlled mid-training comparisons, packaging identical sections as coherent books yields a +1.02 mean downstream gain over the Split setting, while RandomConcat and document-level Rephrase perform worse. Results are also reported on Llama3-8B.
No heat snapshots are available in the last 24 hours.