Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

Are LLMs Ready for Scientific Discovery? A Capability-Oriented Benchmark for AI Scientists

First seen · 7/15/2026, 12:00 PMLatest activity · 7/15/2026, 12:00 PM

The paper introduces SDABench, a capability-oriented benchmark for scientific data analysis. It evaluates six capabilities: descriptive, exploratory, inferential, predictive, causal, and mechanistic analysis, across Biology, Chemistry, Environment, Geography, and Physics. The benchmark contains 527 real-data instances and 6,000 synthetic instances, each available in multiple-choice and open-ended formats. Evaluations of 15 representative LLMs show that models perform relatively well on descriptive analysis but degrade substantially when they must select assumptions, model latent processes, choose suitable procedures, or provide mechanistic explanations. The authors also propose a five-stage error analysis framework.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorHuggingFace Daily Papers7/15, 12:00 PMnot independentRepresentative
    Are LLMs Ready for Scientific Discovery? A Capability-Oriented Benchmark for AI Scientists