DeepLoop studies residual scaling in looped Transformers, where the same physical blocks are revisited across multiple computational rounds. The paper introduces a first-order perturbation bound governed by a visit-alignment coefficient, κ_R. Under a conservative aligned-visit assumption, the residual-scaling exponent increases from 1/4 to 1/2 as unrolled depth grows at fixed physical depth. DeepLoop uses Post-LN DeepNorm with α=(2N)^{1/2} and β=(8N)^{-1/2}. Experiments on GPT-2 small- and medium-scale looped language models report neutral behavior without block revisitation and improved validation loss and downstream accuracy when recurrent depth is active.
No heat snapshots are available in the last 24 hours.