Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

A Controlled Study of Attention-Only Transformers

First seen · 7/20/2026, 11:29 PMLatest activity · 7/20/2026, 11:29 PM

This paper conducts controlled pretraining experiments on attention-only decoder Transformers, called Simple Attention Networks (SANs), against standard Transformers while matching parameter count, training FLOPs, or depth. Across 6M to 87M parameters, 2 to 48 layers, and up to 105B tokens, deleting feed-forward layers directly causes a substantial loss. However, reallocating the freed budget to attention depth nearly closes the gap: the matched-parameter difference is 0.006 nats, or 0.27% of loss. The remaining deficit is concentrated in parametric recall rather than context-grounded answering. QK-normalization is reported as essential for training deep attention-only stacks.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorarXiv7/20, 11:29 PMnot independentRepresentative
    A Controlled Study of Attention-Only Transformers