Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

Evaluating Medical AI Under Missing Information: Same-Provider Judges and Human Raters Change Apparent Safety

First seen · 7/21/2026, 04:05 PMLatest activity · 7/21/2026, 04:05 PM

This paper stress-tests open-ended medical conversations by deleting the latter half of the final user turn in HealthBench cases. It evaluates Claude Opus 4.8, GPT-5.5, Grok 4.3, and Gemini 3.5 Flash with four provider-based LLM judges and blinded human raters. Judge agreement is only moderate (Fleiss’ kappa = 0.65), and a positive same-provider association remains after leniency adjustment. On a blinded 50-item sample, LLM judges credited appropriate uncertainty on 66–84% of cases, compared with 52% for an independent clinician, suggesting that evaluator design can materially change conclusions about medical safety.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorarXiv7/21, 04:05 PMnot independentRepresentative
    Evaluating Medical AI Under Missing Information: Same-Provider Judges and Human Raters Change Apparent Safety