This paper stress-tests open-ended medical conversations by deleting the latter half of the final user turn in HealthBench cases. It evaluates Claude Opus 4.8, GPT-5.5, Grok 4.3, and Gemini 3.5 Flash with four provider-based LLM judges and blinded human raters. Judge agreement is only moderate (Fleiss’ kappa = 0.65), and a positive same-provider association remains after leniency adjustment. On a blinded 50-item sample, LLM judges credited appropriate uncertainty on 66–84% of cases, compared with 52% for an independent clinician, suggesting that evaluator design can materially change conclusions about medical safety.
No heat snapshots are available in the last 24 hours.