Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

PACE: A Proxy for Agentic Capability Evaluation

First seen · 7/3/2026, 12:00 PMLatest activity · 7/3/2026, 12:00 PM

PACE proposes a proxy-evaluation framework for estimating expensive LLM agent performance from a compact set of instances drawn from non-agentic benchmarks. It combines target-relevance local selection with globally informative selection, then fits a regression from atomic capability scores to agentic benchmark scores. Across 14 models, 4 agentic benchmarks, and 19 non-agentic benchmarks, PACE-Bench achieves leave-one-out cross-validation mean absolute error below 4%, Spearman correlation above 0.80, and pairwise model-ranking accuracy around 85%, while costing less than 1% of a full agent evaluation.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorHuggingFace Daily Papers7/3, 12:00 PMnot independentRepresentative
    PACE: A Proxy for Agentic Capability Evaluation