WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data
Wearable health metrics provide a continuous record of physiological signals, yet assessing how well language models interpret longitudinal sensor data remains difficult. WearableQA introduces a benchmark of 4,084 ten-option multiple-choice questions derived from up to 500 days of sensor readings, blood biomarkers, and demographics across 200 real users. Designed around both data-level computation and physiological interpretation, the benchmark preserves authentic device noise and variance. Across evaluations of 14 proprietary and open-source models, accuracy ranged between 19.6% and 72.9%, with most failing to cross 60%.
Why it's worth reading
By linking noisy, multi-month wearable sensor data with clinical blood biomarkers, WearableQA provides a realistic diagnostic for evaluating whether AI models can interpret continuous personal physiological records.