This paper examines which evaluation protocols reliably compare federated pre-trained models. Using centralized and federated versions of a 16M-parameter Transformer trained on identical client data, the authors compare GLUE downstream fine-tuning, including full, head-only, and reduced-data settings, with next-token prediction on GLUE text. Downstream fine-tuning does not consistently preserve the ranking established by pre-training test perplexity, while the intrinsic next-token signal shows a strong correspondence. The results caution against relying on downstream fine-tuning alone when evaluating federated pre-training quality.
No heat snapshots are available in the last 24 hours.