This paper conducts controlled pretraining experiments on attention-only decoder Transformers, called Simple Attention Networks (SANs), against standard Transformers while matching parameter count, training FLOPs, or depth. Across 6M to 87M parameters, 2 to 48 layers, and up to 105B tokens, deleting feed-forward layers directly causes a substantial loss. However, reallocating the freed budget to attention depth nearly closes the gap: the matched-parameter difference is 0.006 nats, or 0.27% of loss. The remaining deficit is concentrated in parametric recall rather than context-grounded answering. QK-normalization is reported as essential for training deep attention-only stacks.
No heat snapshots are available in the last 24 hours.