This paper directly compares merged specialists with the joint multi-task reinforcement-learning model that model merging is often presented as a substitute for. Using LOOP-trained Qwen3-8B specialists at AppWorld difficulty levels 1 and 2, the authors evaluate TIES, RAM+, and related merges against a jointly trained model on the same data. All merge variants are statistically indistinguishable from joint RL on task-goal completion. Task vectors have low cosine similarity, 0.06–0.10, despite approximately 65% support overlap. The authors argue that this decoupling between direction and support makes sign- and support-based methods behave similarly to near-uniform averaging.
No heat snapshots are available in the last 24 hours.