Wearable health metrics provide a continuous record of physiological signals, yet assessing how well language models interpret longitudinal sensor data remains difficult. WearableQA introduces a benchmark of 4,084 ten-option multiple-choice questions derived from up to 500 days of sensor readings, blood biomarkers, and demographics across 200 real users. Designed around both data-level computation and physiological interpretation, the benchmark preserves authentic device noise and variance. Across evaluations of 14 proprietary and open-source models, accuracy ranged between 19.6% and 72.9%, with most failing to cross 60%.
No heat snapshots are available in the last 24 hours.