The paper introduces Subspace-Aligned Rewiring (SAR), a post-hoc method for editing reinforcement-learning updates in large language models. SAR keeps update components aligned with the base model’s spectral space and removes orthogonal components that may suppress reasoning or increase cross-domain interference. The authors report that compact reasoning cores can use as little as approximately 0.58% of total parameters while preserving over 99% of post-training performance. They also report improved high-k exploration in mathematical reasoning, gains on six of seven in-house agentic-coding benchmarks, purification of mixed-domain updates, and stronger cross-domain model merging.
No heat snapshots are available in the last 24 hours.