Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

Information-Seeking Failures of Large Language Models in Agentic Clinical Reasoning

First seen · 7/11/2026, 08:10 PMLatest activity · 7/11/2026, 08:10 PM

This study evaluates 32 frontier models in hematologic oncology using a three-round agentic framework. Models must proactively request clinical information before committing to a diagnosis and treatment plan. The best overall accuracy was only 68%. Information utilization, the fraction of available data requested, was the strongest predictor of diagnostic accuracy (R = 0.69, P < 0.001), but fell from 57% to 26% in the final round, leaving molecular and cytogenetic evidence unexamined. High-scoring reasoning traces were not correlated with correctness. Search satisficing, anchoring, and premature closure dominated the errors.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorarXiv7/11, 08:10 PMnot independentRepresentative
    Information-Seeking Failures of Large Language Models in Agentic Clinical Reasoning