This paper asks whether large language models with substantially different parameter spaces can be merged through direct weighted averaging without distillation or semantic alignment. It introduces training-free dimensional adaptation: expanding the smaller checkpoint into the larger space for union-style merging, or truncating the larger one for intersection-style merging. Experiments on Qwen-family pairs across reasoning, code, language understanding, commonsense, knowledge, and instruction-following tasks suggest that small-ratio interpolation can transfer complementary capabilities and sometimes outperform source checkpoints. Near-balanced interpolation often collapses, however, and gains remain task-dependent.
No heat snapshots are available in the last 24 hours.