WorldCupArena introduces a dynamic benchmark for evaluating language models and deep-research agents on football forecasting. Before each match, systems either use a shared evidence package or conduct their own research, then predict the result, exact score, likely players and events, match statistics, and competition outcomes. The first evaluation covers 104 matches and 13 systems. The authors report that the strongest system achieves only small gains over betting-market and human-fan baselines on result and exact-score accuracy, while showing a clearer advantage on a graded Scoreline metric. Code, prompts, predictions, and evaluation scripts are open sourced.
No heat snapshots are available in the last 24 hours.