This study evaluates 32 frontier models in hematologic oncology using a three-round agentic framework. Models must proactively request clinical information before committing to a diagnosis and treatment plan. The best overall accuracy was only 68%. Information utilization, the fraction of available data requested, was the strongest predictor of diagnostic accuracy (R = 0.69, P < 0.001), but fell from 57% to 26% in the final round, leaving molecular and cytogenetic evidence unexamined. High-scoring reasoning traces were not correlated with correctness. Search satisficing, anchoring, and premature closure dominated the errors.
No heat snapshots are available in the last 24 hours.