PACE proposes a proxy-evaluation framework for estimating expensive LLM agent performance from a compact set of instances drawn from non-agentic benchmarks. It combines target-relevance local selection with globally informative selection, then fits a regression from atomic capability scores to agentic benchmark scores. Across 14 models, 4 agentic benchmarks, and 19 non-agentic benchmarks, PACE-Bench achieves leave-one-out cross-validation mean absolute error below 4%, Spearman correlation above 0.80, and pairwise model-ranking accuracy around 85%, while costing less than 1% of a full agent evaluation.
No heat snapshots are available in the last 24 hours.