Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

Benchmarks Are Not Monolithic: Sample-Level Auditing and Orchestration for LLM Evaluation

First seen · 7/31/2026, 03:51 AMLatest activity · 7/31/2026, 03:51 AM

The paper proposes a dataset-centric meta-evaluation framework that audits individual benchmark samples across five latent dimensions: cognitive and knowledge demands, language and content quality, task properties, context, and ethics, safety, and fairness. It applies the framework to five influential benchmarks—MMLU, ARC, WinoGrande, HellaSwag, and TruthfulQA—and reports substantial internal heterogeneity that aggregate accuracy does not capture. The annotations are then used to orchestrate composite benchmark subsets targeting capabilities such as reasoning depth and ethical sensitivity.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorarXiv7/31, 03:51 AMnot independentRepresentative
    Benchmarks Are Not Monolithic: Sample-Level Auditing and Orchestration for LLM Evaluation