This study evaluates an evidence-sufficiency prompt across 1,200 paired cells from Real-POCQi, HealthBench, and MedRBench, using GPT-5.5, Claude Opus 4.8, Gemini 3.5 Flash, and Grok 4.3. The primary judge, GPT-5.4-nano, scored unsafe overconfidence at 49.3% with the standard prompt versus 24.7% with the wrapper, a 24.7-point paired reduction. Claude Sonnet 5 agreed on direction but estimated an effect of only 13.1 points. Correct diagnosis fell from 80.3% to 50.3%, with costs varying sharply by model. Clinician review found the primary judge highly sensitive but insufficiently specific.
No heat snapshots are available in the last 24 hours.