Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

Read original
arXiv·Ji Soo Lee·Sep 4, 2026, 5:52 PM

WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data

Papers78

Wearable health metrics provide a continuous record of physiological signals, yet assessing how well language models interpret longitudinal sensor data remains difficult. WearableQA introduces a benchmark of 4,084 ten-option multiple-choice questions derived from up to 500 days of sensor readings, blood biomarkers, and demographics across 200 real users. Designed around both data-level computation and physiological interpretation, the benchmark preserves authentic device noise and variance. Across evaluations of 14 proprietary and open-source models, accuracy ranged between 19.6% and 72.9%, with most failing to cross 60%.

Why it's worth reading

By linking noisy, multi-month wearable sensor data with clinical blood biomarkers, WearableQA provides a realistic diagnostic for evaluating whether AI models can interpret continuous personal physiological records.

Tags

WearableQALLM BenchmarkDigital HealthTime SeriesHealthcarePhysiological ReasoningWearables

Also reported by

  • HuggingFace Daily Papers — WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data

Score breakdown

  • Novelty79
  • Impact76
  • Practicality75
  • Credibility80
  • Timeliness82