Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents

First seen · 7/28/2026, 05:24 PMLatest activity · 7/28/2026, 05:24 PM

PatientAgentBench evaluates patient-facing healthcare agents in sustained conversations with simulated patients, realistic records, and sandboxed healthcare tools. It tests 10 models from four families across 1,200 scenarios using six dimensions and more than 100 clinician-grounded, conversation-agnostic criteria. Licensed clinicians showed 79–93% adjacent agreement with the LLM jury. Triage was the most discriminating dimension, with pass rates ranging from 32% to 88%. Frontier models still scored only 4.25/5 overall and sometimes fabricated unexecuted actions, trusted unverified tool outputs, or omitted crisis resources during emergencies.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorarXiv7/28, 05:24 PMnot independentRepresentative
    PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents