The dataset you already own
If you wear a WHOOP, an Oura ring or a modern smartwatch, you are generating a continuous physiological record: nocturnal resting heart rate, heart-rate variability, sleep duration and timing, respiratory rate, skin temperature, activity and, on some devices, blood oxygen. For most people that record lives inside an app and is glanced at over coffee. Clinically it is something more useful: a longitudinal measurement of how your body is coping, collected on the days you are not in a clinic. Which is nearly all of them.
Why a clinic visit misses things
A consultation samples one moment, usually a moment when you have traveled, rushed, and are mildly anxious about being measured. But physiology moves. Recovery collapses during a stressful quarter. Sleep erodes gradually across a year. Resting heart rate drifts upward for weeks before anything feels wrong. Alcohol, illness, a new medication, a heavy training block and a time-zone change all leave signatures. Trend data makes those visible, and visible early. That is the honest case for integrating it.
How accurate is any of this
This is where most wellness content stops asking questions, so it is worth being specific. A validation study published in Physiological Reports compared five consumer devices against an ECG reference across 536 nights of sleep. Accuracy varied significantly between devices: for nocturnal resting heart rate, Oura's third and fourth generation rings showed the strongest agreement, with WHOOP moderate and Polar poorer. For heart-rate variability the spread was wider still. The useful conclusion is not about which brand wins: the metric and the device together determine whether a number deserves clinical weight.
Sleep stages are the weakest link
Consumer sleep staging is the metric patients trust most and the one that deserves it least. Photoplethysmography-based approaches reach roughly 60 to 72 percent accuracy for four-stage sleep classification, with systematic underestimation of REM and overestimation of deep sleep. The physiological reason is straightforward: both REM and light sleep show elevated heart-rate variability, so cardiac signals alone struggle to separate them. Total sleep time and sleep timing consistency are far more trustworthy than the colored bar chart of stages, and clinically they matter more anyway.
The three patterns we read
First, sleep regularity (whether you sleep at consistent times), which predicts outcomes at least as well as duration. Second, heart-rate-variability trend, read as a multi-week direction rather than a nightly verdict; a downward drift alongside rising resting heart rate is a meaningful stress-load signal. Third, resting heart rate response to specific inputs: alcohol, late meals, travel, illness, a training overload. None of these require precision to three decimal places. They require weeks of data and a clinician who reads context.
Where the numbers mislead
Devices estimate. Sensors slip, skin contact fails, motion creates artifact, and an artifactually low HRV reading can send a healthy person into a spiral of concern. There is also a documented behavioral cost: for some people, optimizing a readiness score becomes a source of the very stress the score is measuring. We have told patients to take the ring off for a month, and it was the right clinical advice. Data serves the plan; it is not the plan.
How we use it at Organic Well
With your consent, wearable data is integrated through Heads Up Health and reviewed by a licensed provider alongside your labs and symptoms. It is used to answer specific questions: is the sleep intervention working, did recovery improve after we treated the iron deficiency, is training load compatible with the recovery you are actually getting. You can disconnect data sharing at any time. Wearable data is not used to diagnose disease, and it never replaces clinical evaluation.
Evidence
What the research says
We cite the studies directly, including where the evidence is thin. Every link goes to the primary source.
For HRV, Oura devices provided the highest accuracy… WHOOP showed moderate accuracy, followed by poor agreement from both Garmin and Polar.
Dial MB, et al. Validation of nocturnal resting heart rate and heart rate variability in consumer wearables. Physiological Reports, 2025.
Thirteen adults, 536 nights, five devices measured simultaneously against an ECG reference. Device choice materially changes how much weight a metric deserves.
Photoplethysmography (PPG)-based heart rate variability (HRV) is the dominant approach in current wearables, achieving 60–72% accuracy for four-stage sleep classification.
Systematic review of wearable sleep monitoring, medRxiv, 2025.
Both REM and light sleep raise heart-rate variability, which is why cardiac signals alone cannot cleanly separate them. Treat stage breakdowns as indicative, not diagnostic.
Key issues as wearable digital health technologies enter clinical care.
Ginsburg GS, Picard RW, Friend SH. New England Journal of Medicine, 2024;390:1118–1127.
The framing paper on what has to be true before consumer sensor data can responsibly inform medical decisions: validation, context and interpretation, in that order.


