This paper connects pseudo-gradient aggregation in local SGD and DiLoCo with task-arithmetic model merging. Evaluating several merging methods, the authors identify Iso-C as a promising aggregation rule: DiLoCo SGD with Iso-C outperforms both simple pseudo-gradient averaging and momentum-based DiLoCo. They then introduce IsoLoCo, which combines Iso-C with Nesterov momentum. Experiments on language-model pretraining across different worker counts, model sizes, and local-step counts reportedly show consistent gains over DiLoCo, with the advantage increasing as the number of workers grows.
No heat snapshots are available in the last 24 hours.