This paper proposes Isospectral Optimization (ISO), an RLVR-native framework that keeps a model’s singular-value spectra fixed while optimizing the associated input and output singular frames. Its offline component, ISO-Merger, combines frame changes from specialists sharing a base model without post-merge data, rollouts, gradients, or on-policy distillation. Its online component, ISO-Optimizer, applies optimizers such as AdamW and Muon to frame variables. Across 1.5B–8B models on reasoning and coding tasks, the authors report higher accuracy or comparable scores with fewer training steps. On Qwen3-8B-Base, ISO-AdamW matches AdamW’s 0.495 aggregate accuracy after 100 versus 270 steps, then reaches 0.509 after 210 steps.
No heat snapshots are available in the last 24 hours.