Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data

First seen · 9/4/2026, 04:00 AMLatest activity · 9/5/2026, 01:52 AM

Wearable health metrics provide a continuous record of physiological signals, yet assessing how well language models interpret longitudinal sensor data remains difficult. WearableQA introduces a benchmark of 4,084 ten-option multiple-choice questions derived from up to 500 days of sensor readings, blood biomarkers, and demographics across 200 real users. Designed around both data-level computation and physiological interpretation, the benchmark preserves authentic device noise and variance. Across evaluations of 14 proprietary and open-source models, accuracy ranged between 19.6% and 72.9%, with most failing to cross 60%.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorHuggingFace Daily Papers9/4, 04:00 AMnot independent
    WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data
  2. AggregatorarXiv9/5, 01:52 AMnot independentRepresentative
    WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data