This preprint identifies a numerical failure in ALiBi positional encoding: linearly increasing distance biases can underflow floating-point precision, forcing many attention weights to zero and making affected heads partially blind. Using pretraining experiments with 148M-parameter decoder models, the authors separate this effect from ordinary out-of-context degradation. They report substantial harm to passkey-style token retrieval but only minor changes on standard decoder benchmarks. Four training-time mitigation strategies are evaluated, with log-scaled distances providing the most consistent passkey-retrieval improvements, although default ALiBi slopes remain a strong needle-in-a-haystack baseline.
No heat snapshots are available in the last 24 hours.