This paper proposes that Transformers use one forward computation stream for two different roles: predicting the next token and storing state useful for future predictions. It introduces a two-stream Transformer that separates these functions and evaluates it through pretraining at multiple scales. According to the abstract, the design improves validation loss and data/compute efficiency, while outperforming standard Transformers by an average of 2–3 percentage points on downstream tasks. The authors also report analyses addressing confounders and identifying differences in the resulting gradients.
No heat snapshots are available in the last 24 hours.