This paper presents a common recurrent-memory notation for comparing softmax attention with DeltaNet, Gated DeltaNet, Kimi Delta Attention, and Gated DeltaNet-2. Its main experiments use 350M-parameter models trained on 15B tokens, with additional optimizer, learning-rate, hybrid-stack, sequence-length, scaling, and downstream evaluations. In the reported sweep, Kimi Delta Attention with Muon achieves the lowest final validation loss, while a pure Gated DeltaNet stack with AdamW delivers the highest normalized training throughput. Hybrid stacks generally trade throughput for lower loss. The proposed Cross-Layer Value Routing improves matched DeltaNet and Gated DeltaNet runs modestly, but no empirical inference-speed benchmark is provided.
No heat snapshots are available in the last 24 hours.