This paper studies why language models struggle to perform multi-hop reasoning within a single forward pass. In two-hop tasks, the authors report that a looped Transformer can often decode the correct bridge entity after its first loop, but its hidden state is poorly aligned with the corresponding token embedding. DiscoLoop addresses this representational bottleneck by carrying both a discrete embedding channel and a continuous hidden-state channel through recurrence. The abstract reports near-perfect accuracy with fewer training steps on symbolic and synthetic-language tasks, plus lower training loss and stronger benchmark performance than looped-Transformer baselines in real-world pretraining.
No heat snapshots are available in the last 24 hours.