Bottom line: your tracker knows whether you slept. It is guessing at almost everything else. Sleep-versus-wake detection clears 90% accuracy on every device tested. Sleep stages land somewhere between 50% and 86%. Calorie burn is the worst number on the screen, “active minutes” carries 45% to 55% error, and the recovery score you plan your training around has almost no published validation at all. If you are reading a “deep sleep” number as a precise measurement, you are reading it wrong. This page collects what the peer-reviewed validation research actually found, study by study, so you can judge for yourself.
Why nobody had done this

Manufacturers make accuracy claims. Researchers test those claims in sleep labs against polysomnography, the clinical standard that reads brain activity directly. Then they publish the results in journals that almost nobody shopping for a watch will ever open. The evidence exists. It is just filed somewhere useless.
Nobody had collected them in one place. So we did. Every figure below is drawn from a named, peer-reviewed study with the journal, year, sample size, and device generation stated. We have not tested any of these devices ourselves, and we do not claim to have. This is a synthesis of published research, and where studies disagree with each other, we say so rather than picking the flattering number.
What the lab measures that your watch cannot
Polysomnography records brain waves, eye movement, and muscle activity. It is how sleep stages are clinically defined. A wearable has none of that. It infers stages from movement, heart rate, and heart-rate variability, then runs a proprietary algorithm the manufacturer does not publish. That gap is the entire reason accuracy varies.
Two measures matter in the tables below. Sensitivity is how often the device correctly identifies a given state. Cohen’s kappa measures agreement with the gold standard adjusted for chance, on a scale where 1.0 is perfect. Kappa values between 0.21 and 0.40 are generally read as fair agreement, 0.41 to 0.60 as moderate.
The study everyone quotes, and the detail they leave out
Citation: Robbins R, Weaver MD, Sullivan JP, Quan SF, Gilmore K, et al. Accuracy of Three Commercial Wearable Devices for Sleep Tracking in Healthy Adults. Sensors. 2024;24(20):6532.
Method: 35 participants aged 20 to 50 without a sleep disorder, single-night inpatient study at Brigham and Women’s Hospital Center for Clinical Investigation. Participants wore an Oura Ring Gen3, a Fitbit Sense 2, and an Apple Watch Series 8 simultaneously while monitored with polysomnography.
| Device | Sleep vs. wake sensitivity | Sleep-stage sensitivity | Cohen’s kappa |
|---|---|---|---|
| Oura Ring Gen3 | ≥95% | 76.0–79.5% | 0.65 |
| Apple Watch Series 8 | ≥95% | 50.5–86.1% | 0.60 |
| Fitbit Sense 2 | ≥95% | 61.7–78.0% | 0.55 |
What it found: All three devices identified sleep versus wake with sensitivity of 95% or higher. Stage classification was far weaker, ranging from 50% to 86% across devices and stages. The Oura Ring did not differ significantly from polysomnography in its estimation of wake, light sleep, deep sleep, or REM. The Apple Watch showed the widest spread, performing well on REM (86.1%) and poorly on deep sleep (50.5%).
A conflict of interest worth knowing about. The lead author, Dr. Rebecca Robbins, is a paid medical advisor to ŌURA, the company that makes the device this study found most accurate. This is disclosed in the paper and we are not suggesting misconduct. But when a study funded or advised by a manufacturer finds that manufacturer’s product wins, that is context you deserve before you treat the ranking as settled. Read the next study with that in mind.
Six devices, one lab, a completely different ranking
Citation: Schyvens AM, Peters B, Van Oost NC, Aerts JM, Masci F, Neven A, Dirix H, Wets G, Ross V, Verbraecken J. A performance validation of six commercial wrist-worn wearable sleep-tracking devices for sleep stage scoring compared to polysomnography. SLEEP Advances. 2025;6(2):zpaf021.
Method: Laboratory-based validation of six wrist-worn devices against polysomnography, with no manufacturer affiliation disclosed among the authors.
| Device | Cohen’s kappa |
|---|---|
| Apple Watch Series 8 | 0.53 |
| Fitbit Sense | 0.42 |
| Fitbit Charge 5 | 0.41 |
| Withings ScanWatch, Garmin Vivosmart 4, WHOOP 4.0 | Within the overall 0.21–0.53 range |
What it found: Every device detected more than 90% of sleep epochs, confirming the pattern from Study 1. But specificity, the ability to correctly identify wake, ranged from just 29.4% to 52.2%. In plain terms: these devices are strongly biased toward calling you asleep. If you lie awake in bed, there is a meaningful chance your tracker counts it as sleep.
Where it disagrees with Study 1. This study put the Apple Watch Series 8 at kappa 0.53. Study 1 put the same device at 0.60. Same watch, different labs, different answers. That spread is not a flaw in either study, it is the honest reality of measuring something this noisy, and it is the strongest argument against treating any single ranking as definitive. The authors concluded that all devices “can benefit from further improvement for multistate categorization” while noting the higher-kappa devices could reasonably track prolonged changes in sleep architecture.
Eleven trackers, including ones you do not wear
Citation: Accuracy of 11 Wearable, Nearable, and Airable Consumer Sleep Trackers: Prospective Multicenter Validation Study. PubMed ID 37917155, 2023.
Method: A multicenter study comparing 11 commercially available consumer sleep trackers against in-lab polysomnography. Devices included five wearables (Google Pixel Watch, Galaxy Watch 5, Fitbit Sense 2, Apple Watch 8, Oura Ring 3), three nearables (Withings Sleep Tracking Mat, Google Nest Hub 2, Amazon Halo Rise), and three app-based trackers.
What it found: Performance varied by stage rather than by device. Among the wearables, the Google Pixel Watch and Fitbit Sense 2 performed best specifically on deep sleep. The authors concluded that some trackers showed substantial agreement with polysomnography while others were only partially consistent. Notably, this study included device classes the others did not, and found that a non-wearable app-based tracker outperformed the wrist devices on wake and REM detection.
WHOOP across 86 nights, with one big asterisk
Citation: A validation study of the WHOOP strap against polysomnography to assess sleep. PubMed ID 32713257.
Method: Twelve healthy adults, mean age 22.9, in a 10-day laboratory protocol. 86 separate sleeps scored in 30-second epochs against polysomnography. Bedtimes were entered manually by researchers, which matters for interpreting the result.
What it found: WHOOP overestimated total sleep time by 8.2 minutes on average, a difference the authors reported as non-significant. For simple sleep-versus-wake it reached 89% agreement, 95% sensitivity to sleep, and Cohen’s kappa of 0.49. For four-stage classification, agreement dropped to 64% with kappa of 0.47, and specificity for wake was 51%. The authors concluded WHOOP is reasonable for estimating sleep in field settings where polysomnography is impractical, if bedtimes are entered accurately.
Your tracker flatters good nights and misses bad ones
Citation: Kainec KA, Caccavaro J, Barnes M, Hoff C, Berlin A, Spencer RMC. Evaluating Accuracy in Five Commercial Sleep-Tracking Devices Compared to Research-Grade Actigraphy and Polysomnography. Sensors. 2024;24(2):635.
Method: Fifty-three young adults monitored for one night with five consumer devices, research-grade actigraphy, and polysomnography simultaneously.
What it found, and why this is the most practically useful result on this page: the errors are not random, they are proportional. Every device overestimated sleep on nights with shorter wake times and underestimated it on nights with longer wake times. In plain terms, your tracker flatters good nights and is least accurate on bad ones, which is precisely when you are most likely to be looking at it. The Oura Ring underestimated light sleep at any duration. REM bias was low across all devices. Every device except the Garmin Vivosmart estimated total sleep time comparably to research-grade actigraphy.
The tiebreaker: 24 studies, 798 people, one answer
Citation: Lee YJ, Lee JY, Cho JH, Kang YJ, Choi JH. Performance of consumer wrist-worn sleep tracking devices compared to polysomnography: a meta-analysis. J Clin Sleep Med. 2025;21(3):573–582.
Method: A meta-analysis pooling 24 studies covering 798 participants and devices including Fitbit, WHOOP, Garmin, Apple Watch, Jawbone, Basis B1, Zulu Watch, Huami Arc, the E4 wristband, Fatigue Science Readiband, and Xiaomi Mi Band 5. Databases searched through March 2024.
This is the highest tier of evidence available on the question, because it aggregates across individual studies rather than relying on any single lab, single night, or single manufacturer relationship. It is the appropriate reference point when a single study’s ranking looks surprising.
What every study agrees on
- Sleep versus wake detection is genuinely good. Above 90% sensitivity in every study, every device. If your tracker says you slept seven hours, that number is probably close.
- Sleep-stage detection is substantially weaker. Roughly 50% to 86% depending on device and stage. Deep sleep and REM numbers should be read as rough trends, not measurements.
- Every device over-calls sleep. Specificity for detecting wake ran as low as 29%. Time you spend lying awake is frequently logged as sleep.
- No device wins across the board. Different studies rank the same devices differently, and devices that lead on one stage trail on another.
Your tracker is least accurate on your worst nights. Kainec et al. found the bias is proportional, not random: devices overestimate sleep when wake time is short and underestimate it when wake time is long. The night you sleep badly and check your score in the morning is the night the number is most likely to be wrong.
So should you actually buy one?
Buy a sleep tracker if you want to see trends in your sleep duration and consistency over weeks and months. Every device tested is good enough for that, and the cheaper ones are close to the expensive ones on the measure that actually works.
Do not buy one expecting a precise deep-sleep or REM number. That is the metric marketing leans on hardest and the one the research supports least. And no consumer wearable can diagnose a sleep disorder. If you suspect sleep apnea, insomnia, or another condition, that is a conversation with a doctor, not a purchase.
Heart rate accuracy: wrist sensors versus chest straps
Sleep is one claim. Heart rate is the one that actually changes how you train. The pattern in the research is consistent and it is not what most buyers assume: a chest strap is close to clinical-grade, and a wrist device is good at rest and degrades as you move.
The reason is mechanical. A chest strap, like the Polar H10, reads the heart’s electrical signal, which is what an ECG measures. A wrist device shines light into your skin and infers pulse from blood-volume changes, a method called photoplethysmography. Motion, skin tone, sweat, strap tightness, and rapid heart-rate changes all interfere with that optical reading. None of them interfere with an electrical one.
Study: Pasadyn et al., Cardiovascular Diagnosis and Therapy. Fifty healthy athletic adults ran on a treadmill at six speeds from 4 to 9 mph while wearing a three-lead ECG, a Polar H7 chest strap, and two randomly assigned wrist devices from a set including the Apple Watch Series 3, Fitbit Ionic, Garmin Vivosmart HR, and TomTom Spark 3. The Polar H7 showed the highest agreement with ECG, with the Apple Watch the strongest of the wrist devices. Measurement error rose with treadmill speed across every wrist device tested.
Study: Etiwy et al., reported in Global Heart. Eighty cardiac rehabilitation patients were measured against ECG limb leads during cycling and treadmill work while wearing a Polar chest strap, Apple Watch, Fitbit Blaze, Garmin Forerunner 235, and TomTom Spark Cardio. The Polar chest strap correlated with ECG at r = 0.99. The Apple Watch reached r = 0.80 overall and improved to r = 0.89 during cycling, which involves less arm movement than running.
Study: Heart rate measurement under transient states, validated against 12-lead ECG. This one is the most useful for training. It tested the Fitbit Charge 5, Fitbit Sense 2, Garmin Vivosmart 4, WHOOP 4.0, and Withings ScanWatch during a protocol built around rapid heart-rate change rather than steady effort. A research-grade chest device performed strongly across all dynamic conditions. Wrist devices were evaluated specifically on how they handled transitions, which is exactly where optical sensors struggle.
| What you are doing | How much to trust a wrist reading |
|---|---|
| Resting, sleeping, sitting | High. Close to chest-strap accuracy. |
| Steady cardio (cycling, easy running) | Good. Cycling reads better than running because the arm is still. |
| Intervals, sprints, rapid changes | Low. This is where optical sensors lag and miss. |
| Lifting, rowing, anything gripping | Low. Wrist flexion and forearm tension disrupt the signal. |
The practical takeaway: if you train by heart-rate zones, a chest strap is not a luxury upgrade, it is the difference between training in the zone you think you are in and one you are not. Our heart rate monitor picks cover which strap fits which setup, and the heart rate zone calculator will tell you what those zones should be for you. If you only want resting heart rate, sleep data, and general trends, your wrist device is fine and a strap adds nothing.
Step counting and active minutes: the gap nobody mentions
Step counting is supposed to be the easy one. It is, right up until you walk slowly, at which point every device on the market falls apart together.
Study: the CADENCE-adults study, a catalog of validity indices for step-counting wearables during treadmill walking. Across normal walking speeds, 15 devices including the Apple Watch Series 1, Fitbit Ionic, Fitbit One, Fitbit Zip, Garmin vivoactive 3, and Garmin vivofit 3 came in under 5% mean absolute percentage error. Excellent. But the same devices averaged 40% error at slow walking speeds, against 7% at normal speeds. Shuffling around a kitchen is not the same measurement problem as walking to work, and no device solves it.
Study: a free-living comparison of the Apple Watch 2, Fitbit Charge 2, and Fitbit Alta published in PLOS One. Forty-eight participants wore the devices for 24 hours of normal life against a criterion pedometer, an ActiGraph, and a Polar H7 chest strap. Step correlations were strong at 0.84 to 0.95, but mean absolute percentage error ran 17% to 36%. That is the gap between a laboratory treadmill and an ordinary day.
The finding that should change how you read your app: in that same study, moderate-to-vigorous activity minutes, the metric most apps present as your daily “active minutes” goal, showed errors of 45% to 55% against the research-grade reference. Heart rate in the same test ran 4% to 16% error. So the metric people optimise hardest for is the least trustworthy number on the screen, by a wide margin.
Study: a 2025 evaluation in the Journal for the Measurement of Physical Behaviour tested the Apple Watch Ultra, COROS Vertix 2, Garmin Fenix 6, and Polar Grit X against hand-tallied steps in a lab and against GoPro footage during a 3.2 km trail run. All four were accurate and reliable on the treadmill and outdoors while running. All four were least accurate during activities of daily living, potentially underestimating total daily steps.
| What you are doing | How much to trust the step count |
|---|---|
| Walking at normal pace, or running | High. Under 5% error in controlled testing. |
| A normal day of mixed activity | Moderate. Expect roughly 17% to 36% error. |
| Slow or shuffling walking | Low. Around 40% error across every device tested. |
| “Active minutes” / MVPA | Very low. 45% to 55% error. Treat as a rough prompt, not a measurement. |
None of this means step counting is useless. It means the number is a consistent relative signal rather than an exact count. Ten thousand steps on your watch today versus six thousand yesterday tells you something real. Ten thousand steps on your watch versus ten thousand on someone else’s, or against an actual tally, does not.
HRV and resting heart rate: the raw signal is good
Citation: Dial MB, et al. Validation of nocturnal resting heart rate and heart rate variability in consumer wearables. Physiological Reports. 2025;13:e70527.
Method: Thirteen healthy adults wore a single-lead ECG chest strap reference alongside multiple wearables during sleep, across 536 nights. Devices tested: Oura Gen 3, Oura Gen 4, WHOOP 4.0, Garmin Fenix 6, and Polar Grit X Pro. Agreement was measured with Lin’s concordance correlation coefficient, which penalises both weak association and systematic offset, so it is a stricter test than a simple correlation.
| Device | Resting heart rate (CCC / error) | HRV (CCC / error) |
|---|---|---|
| Oura Gen 4 | 0.98 / 1.94% | Highest agreement of devices tested |
| Oura Gen 3 | 0.97 / 1.67% | 0.97 / 7.15% |
| WHOOP 4.0 | 0.91 / 3.00% | 0.94 / 8.17% |
| Garmin Fenix 6 | Lower concordance | 0.87 / 10.52% |
| Polar Grit X Pro | 0.86 / 2.71% | 0.82 / 16.32% |
What it means: resting heart rate is measured well by all of these. Overnight HRV is where they separate, with the Oura rings closest to ECG, WHOOP acceptable, and the two watches less consistent. Two caveats the authors themselves raise: the Garmin Fenix 6 is several hardware generations old and current models may perform differently, and the sample was thirteen healthy adults, which is small even at 536 nights.
Why the Apple Watch is missing: it does not measure overnight HRV continuously the way a ring or band does. It takes periodic spot readings instead, which is a different measurement, not a worse one. It cannot be compared like for like.
The finding that matters most: your recovery score is not validated
Everything above is about the raw signal. Almost nobody looks at raw HRV, though. People look at a score: WHOOP Recovery, Oura Readiness, Garmin Body Battery, Fitbit Daily Readiness, Samsung Energy Score. That number is what drives whether you train hard or take the day off.
Citation: Doherty C, Baldwin M, Lambe R, Burke D, Altini M. Readiness, recovery, and strain: an evaluation of composite health scores in consumer wearables. Translational Exercise Biomedicine. 2025;2(2):128–144.
The researchers catalogued 14 composite health scores across 10 manufacturers, including Fitbit Daily Readiness, Garmin Body Battery and Training Readiness, Oura Readiness and Resilience, WHOOP Strain and Recovery, Polar Nightly Recharge, Samsung Energy Score, Suunto Body Resources, and Ultrahuman Dynamic Recovery. The most common inputs were heart rate variability (86% of scores), resting heart rate (79%), physical activity (71%), and sleep duration (71%).
Their conclusion is the single most useful sentence on this page: none of the manufacturers disclosed their exact algorithmic formulas, and few provided empirical validation or peer-reviewed evidence supporting the accuracy or clinical relevance of their scores. They also found significant discrepancies between brands in data collection windows and metric weighting.
So there are two layers here and they are not the same thing. The bottom layer, heart rate and HRV, has a measurable ground truth you can check against an ECG, and the good devices do well on it. The top layer, the score, is a proprietary interpretation with no shared standard and little published validation. That is why a WHOOP and a Garmin worn on the same body on the same night can tell you opposite things. Neither is broken. They are running different undisclosed formulas over similar inputs.
Practical version: trust your resting heart rate and your HRV trend. Treat the recovery score as one brand’s opinion rather than a measurement, and never let it override how you actually feel. If you are choosing between devices on this basis, the recovery tracker comparison and our WHOOP versus Garmin breakdown go deeper on what each brand actually gives you for the money.
Calories burned: the least accurate number on the device
Citation: Shcherbina A, et al. Accuracy in Wrist-Worn, Sensor-Based Measurements of Heart Rate and Energy Expenditure in a Diverse Cohort. Journal of Personalized Medicine. 2017;7(2):3.
Method: Sixty volunteers of deliberately diverse age, height, weight, skin tone, and fitness level wore seven devices (Apple Watch, Basis Peak, Fitbit Surge, Microsoft Band, Mio Alpha 2, PulseOn, Samsung Gear S2) while being measured simultaneously with continuous telemetry and indirect calorimetry, the laboratory gold standard for energy expenditure, across sitting, walking, running, and cycling.
What it found: the same devices that measured heart rate well were substantially unreliable for calories. Error was lowest for cycling and highest for walking. This is the study most commonly cited for the claim that wearable calorie estimates carry error in the tens of percent, and it remains the clearest demonstration that heart-rate accuracy and calorie accuracy are unrelated problems.
Why calories are harder than everything else on this page. Heart rate is measured. Steps are counted. Calories are estimated, through a chain of assumptions: your basal metabolic rate from a population equation, then activity on top of that, then a conversion from movement and heart rate to energy. Each step adds error, and basal metabolic rate alone varies meaningfully between individuals of identical height, weight, age, and sex.
It also gets worse for some people than others. A more recent analysis of Apple, Garmin, Samsung, and Fitbit against indirect calorimetry found that calorie estimation error increased with body fat percentage across every brand tested, and that bias differed significantly by device. So the number is not just imprecise, it is imprecise in a way that varies with who is wearing it. Treat that finding as preliminary, since it is a preprint rather than a peer-reviewed publication.
Practical version: never eat back the calories your watch says you burned. If you are tracking intake against expenditure, the expenditure side of that equation is the weakest number in the whole system. Use the trend across weeks, and adjust based on what actually happens to your weight, not on what the watch claims.
VO2 max: a usable trend, an unusable number
Citation: Investigating the accuracy of Apple Watch VO2 max measurements: a validation study.
Method: Thirty adults in Dublin wore an Apple Watch for five to ten days to generate a VO2 max estimate, then completed a maximal treadmill test with indirect calorimetry as the reference.
What it found: the Apple Watch underestimated VO2 max by an average of 6.07 mL/kg/min. The limits of agreement were wide, running from roughly 6 below to 18 above, meaning individual estimates could land well off in either direction even though the average error pointed one way. As a rough fitness-trend indicator it is usable. As a number to compare against published fitness norms or against another person, it is not. Our VO2 max calculator estimates it from an actual timed effort instead, which is a different and more transparent method.
How we built this page
We searched for peer-reviewed validation studies comparing consumer sleep-tracking devices against polysomnography, and included every study we could find that named its devices, sample size, and metrics. We report the figures as published. We did not test any device ourselves and make no claim to have. Where a study has a manufacturer relationship, we say so. Where studies contradict each other, we show both numbers rather than averaging them into a false consensus.
For a plainer walkthrough of the same question without the study-by-study detail, see how accurate fitness trackers are. This page will be updated as new validation research is published. If you know of a study we have missed, we want to hear about it.
Frequently asked questions
Which sleep tracker is the most accurate? There is no single answer, and any site giving you one is overselling. The Oura Ring Gen3 posted the highest overall agreement in the 2024 Sensors study, but that study was led by a paid Oura advisor, and an independent 2025 study ranked the Apple Watch Series 8 highest among the devices it tested. Different labs, different results.
Are sleep trackers accurate enough to be useful? For total sleep time and night-to-night trends, yes. Every device tested detected sleep versus wake with above 90% sensitivity. For stage-by-stage breakdowns, treat the numbers as directional.
Why does my tracker say I slept when I was lying awake? Because specificity for wake detection ran as low as 29% in laboratory testing. Devices infer sleep partly from stillness, and lying still while awake looks like sleep to an accelerometer.
Can a wearable diagnose sleep apnea? No. Some devices now flag possible breathing disturbances, but diagnosis requires clinical testing. Treat any wearable alert as a prompt to see a doctor, not as a result.
Does a more expensive tracker sleep-track better? Not reliably. On the measure that works best, sleep versus wake, all tested devices performed similarly regardless of price.