This paper proposes Self-Gating Attention (SGA), a plug-and-play attention mechanism for time-series forecasting. SGA models attention scores with a shared learnable matrix plus an input-dependent residual, avoiding the query and key projections used by standard self-attention. The authors state that this changes time and score-matrix memory complexity with respect to look-back length from quadratic to linear. Experiments cover nine public datasets spanning electricity, finance, weather, medical monitoring, human activity, and climate records. The abstract reports improved inference efficiency and competitive forecasting performance, but does not provide exact metrics, hardware settings, or speedup values.
No heat snapshots are available in the last 24 hours.