🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
← All papers  /  Sep 7 – Sep 7, 2026
Evaluation · Reasoning

WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data

First page
WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data
The curator’s take

Ji Soo Lee and colleagues at Meta and KAIST build WearableQA from the wearable time series, blood biomarkers, and demographics of 200 real users with up to 500 days of daily measurements each.

Ask this paper

Key points
01

Real distributions retained. The benchmark preserves device noise and inter-individual variability rather than using cleaned or synthetic signals.

02

Two reasoning axes. Data versus health reasoning separates computation over longitudinal measurements from physiological interpretation; single- versus cross-signal reasoning separates one signal from integration across several. 16 question types span the grid.

03

Dual grounding for question construction. Literature-grounded physiological findings are combined with statistically validated population-grounded patterns, which is how the authors build 4,084 ten-option questions at scale without hand-authoring each.

04

Discriminative. 14 proprietary and open-source LLMs span 19.6% to 72.9% against a 10% chance baseline.

05

Not close to solved. Most models score below 60%.

Abstract

Recent advances in wearable sensing enable continuous monitoring of physiological and behavioral signals, yet existing benchmarks rarely evaluate whether AI systems can reason over a real user's longitudinal wearable record. We introduce WearableQA, a benchmark comprising 4,084 10-option multiple-choice questions constructed from the wearable time series, blood biomarkers, and demographics of 200 real users, each with up to 500 days of daily measurements. WearableQA preserves authentic wearable distributions that include device noise and inter-individual variability. To evaluate distinct reasoning capabilities, we introduce 16 question types organized along two complementary axes: data versus health reasoning, which distinguishes computation over longitudinal measurements from physiological interpretation; and single- versus cross-signal reasoning, which separates reasoning about individual signals from the integration of multiple signals. To construct reliable questions at scale, we adopt a dual-grounding framework that combines literature-grounded physiological findings with statistically validated population-grounded physiological patterns. This enables the capture of meaningful relationships observed in real-world wearable data. Evaluation of 14 proprietary and open-source LLMs demonstrates that WearableQA effectively differentiates model capabilities, with performance ranging from 19.6% to 72.9% against a 10% chance baseline. Moreover, WearableQA remains far from solved: most models achieve accuracies below 60%. Overall, WearableQA provides a realistic and diagnostic benchmark for evaluating LLM reasoning over real-world wearable data.

Every Monday
Get next week’s papers.
Subscribe on Substack