Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

Partition, Prompt, Aggregate: Statistical Self-Consistency in Language Models

First seen · 7/17/2026, 12:00 PMLatest activity · 7/17/2026, 12:00 PM

This paper tests whether language-model estimates obey the law of total probability. Using binary trees, the authors recursively partition populations, prompt models with verbalized subgroup descriptions, aggregate subgroup estimates with their priors, and compare the result with direct population-level estimates. Across domains and frontier models, the study reports widespread violations of statistical self-consistency. In persona-prompting experiments, it identifies a “macro fallacy”: estimates reconstructed from fine-grained subgroup responses are often closer to human reference data than direct aggregate estimates. The authors propose statistical self-consistency as a reference-free evaluation criterion and report that implicit prompting can partially recover the effect.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorHuggingFace Daily Papers7/17, 12:00 PMnot independentRepresentative
    Partition, Prompt, Aggregate: Statistical Self-Consistency in Language Models