This paper reports a joint model-and-data scaling law for generic Transformers pretrained on collider jets. Fitted only on small models spanning three orders of magnitude in training compute, the law predicts the loss of models trained later with more than 100 times the compute to within 1%. Lower pretraining loss is systematically associated with lower fine-tuning loss and higher background rejection on two standard tagging benchmarks. The authors release five pretrained models across multiple sizes, the full training recipe, and code. The frontier model matches published results from physics-aware foundation models on accuracy, AUC, and quark/gluon rejection, with a remaining advantage for the physics-aware model in the high-purity tail of top tagging.
No heat snapshots are available in the last 24 hours.