This paper proposes Logic-PPT, an initialization stage that trains language models on formal derivations before natural-language pretraining. At a reported 100B-token evaluation scale, the authors say Logic-PPT reaches 80% accuracy on linguistic tasks using 36B fewer tokens than standard initialization and exceeds alternative symbolic pre-pretraining baselines. They further associate the intervention with lower-rank, spectrally concentrated representations and report that pruning to about 33% sparsity preserves dense-baseline performance. These claims are potentially significant, but the supplied arXiv record is dated August 4, 2026 and was not independently verified here.
No heat snapshots are available in the last 24 hours.