ELSAA approximates the attention score operator itself after dense Q, K, and V projections, rather than factorizing the Transformer’s learned projection or output matrices. It combines a sparse branch for selected high-similarity interactions with a low-rank branch for diffuse global context. Because the two branches may have substantially different normalization mass, ELSAA adds a denominator-aware fusion term to rescale the sparse contribution. The stated goal is longer-context Transformer training without materializing the full quadratic attention matrix.
No heat snapshots are available in the last 24 hours.