This paper proposes an exactly solvable mechanism for delayed generalization in linear models trained with full-batch heavy-ball optimization and weight decay. It identifies a population-active component of the empirical null space, termed the grokking subspace, where training predictions remain unchanged and weight decay drives slow relaxation. The resulting discrete- and continuous-time laws predict grokking times, including the weak-regularization scaling $(1-β)/(ηλ)$. The analysis distinguishes coupled L2 regularization from decoupled weight decay, extends locally to nonlinear networks, and reports parameter-free verification in a synthetic model plus scaling agreement in modular addition.
No heat snapshots are available in the last 24 hours.