PatientAgentBench evaluates patient-facing healthcare agents in sustained conversations with simulated patients, realistic records, and sandboxed healthcare tools. It tests 10 models from four families across 1,200 scenarios using six dimensions and more than 100 clinician-grounded, conversation-agnostic criteria. Licensed clinicians showed 79–93% adjacent agreement with the LLM jury. Triage was the most discriminating dimension, with pass rates ranging from 32% to 88%. Frontier models still scored only 4.25/5 overall and sometimes fabricated unexecuted actions, trusted unverified tool outputs, or omitted crisis resources during emergencies.
No heat snapshots are available in the last 24 hours.