The paper proposes a dataset-centric meta-evaluation framework that audits individual benchmark samples across five latent dimensions: cognitive and knowledge demands, language and content quality, task properties, context, and ethics, safety, and fairness. It applies the framework to five influential benchmarks—MMLU, ARC, WinoGrande, HellaSwag, and TruthfulQA—and reports substantial internal heterogeneity that aggregate accuracy does not capture. The annotations are then used to orchestrate composite benchmark subsets targeting capabilities such as reasoning depth and ethical sensitivity.
No heat snapshots are available in the last 24 hours.