The paper evaluates Claude Opus-4.7, GPT-5.4, and Gemini-3.1-Pro when asked for medical diagnoses from chest X-rays, brain MRIs, or dermatology cases that contain no image, only demographic descriptors. The models do not reliably abstain: their diagnoses shift systematically with age, race, and sex. Claude reportedly predicts melanoma for a 65-year-old white man asking about a mole and sarcoidosis for a 32-year-old Black woman asking about a chest X-ray. GPT-5.4 shows broader fabrication. The study also identifies a mismatch between cautious prose and disease-bearing structured fields, plus sensitivity to probe wording.
No heat snapshots are available in the last 24 hours.