Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

Evaluating Federated Pre-Training: On the Reliability of Downstream Fine-Tuning and Intrinsic Evaluation

First seen · 7/20/2026, 11:14 PMLatest activity · 7/20/2026, 11:14 PM

This paper examines which evaluation protocols reliably compare federated pre-trained models. Using centralized and federated versions of a 16M-parameter Transformer trained on identical client data, the authors compare GLUE downstream fine-tuning, including full, head-only, and reduced-data settings, with next-token prediction on GLUE text. Downstream fine-tuning does not consistently preserve the ranking established by pre-training test perplexity, while the intrinsic next-token signal shows a strong correspondence. The results caution against relying on downstream fine-tuning alone when evaluating federated pre-training quality.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorarXiv7/20, 11:14 PMnot independentRepresentative
    Evaluating Federated Pre-Training: On the Reliability of Downstream Fine-Tuning and Intrinsic Evaluation