The paper introduces gate-zero growth, a function-preserving operator for continual learning that adds residual blocks behind zero-initialized gates. Under a transversality condition, the functional Jacobian separates old directions, newly added weight directions, and gate directions: new weights are initially flat, while gates are the only first-order source of functional change. The reported 300M-to-857M Transformer experiment, adapted from WikiText-103 to BookCorpus, achieves near-zero old-domain forgetting with ΔA < 0.1 under both Isolation and Freeze-Nothing settings. A non-function-preserving stacking control shows an order-of-magnitude larger forgetting.
No heat snapshots are available in the last 24 hours.