Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

Linear Attention Architectures: Mechanisms, Trade-offs, and Cross-Layer Routing

First seen · 7/10/2026, 12:00 PMLatest activity · 7/10/2026, 12:00 PM

This paper presents a common recurrent-memory notation for comparing softmax attention with DeltaNet, Gated DeltaNet, Kimi Delta Attention, and Gated DeltaNet-2. Its main experiments use 350M-parameter models trained on 15B tokens, with additional optimizer, learning-rate, hybrid-stack, sequence-length, scaling, and downstream evaluations. In the reported sweep, Kimi Delta Attention with Muon achieves the lowest final validation loss, while a pure Gated DeltaNet stack with AdamW delivers the highest normalized training throughput. Hybrid stacks generally trade throughput for lower loss. The proposed Cross-Layer Value Routing improves matched DeltaNet and Gated DeltaNet runs modestly, but no empirical inference-speed benchmark is provided.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorHuggingFace Daily Papers7/10, 12:00 PMnot independentRepresentative
    Linear Attention Architectures: Mechanisms, Trade-offs, and Cross-Layer Routing