EarlyDx introduces a benchmark for open-ended emergency-department diagnosis generation using 154,834 encounters from MIMIC-IV. It restricts inputs to evidence available at admission time t0 and supervises with diagnoses documented during the ED encounter, avoiding discharge labels that reflect the full inpatient course. An LLM auditor marks generated labels as fully supported, partially supported, or unsupported, with primary scoring limited to fully supported diagnoses. Evaluated frontier, medical-specialized, and in-domain post-trained systems remain unreliable at synthesizing evidence. Post-training improves recall for diagnoses requiring inference from 3–31% to 56%, but no system achieves a clinician-like sensitivity–precision balance for time-critical conditions.
No heat snapshots are available in the last 24 hours.