What's the Most Accurate Wearable? 30+ Studies, 6 Devices, Ranked (2026)

Updated: Sep 9
Last Updated: September 9, 2026

The most accurate wearable depends on what you are tracking. We analyzed peer-reviewed studies from 2023 to 2026 comparing Oura Ring, Apple Watch, WHOOP, Garmin, Fitbit, and others against gold-standard medical measurements across sleep staging, HRV, heart rate, SpO2, step counting, VO2 max, and more. Below is everything we found, organized by metric, with study funding and study population flagged so you can evaluate the data for yourself.
We also built a free interactive comparison tool based on this research that lets you pick your devices and the metrics you care about to see them side by side: wearable accuracy comparison tool.
Most Accurate Wearable: Master Summary by Metric
Most accurate wearable by metric (2026):
Sleep staging: Apple Watch in independent healthy adults (κ=0.53). Oura leads only in its own funded study. Every device drops sharply in people with sleep complaints.
Deep sleep detection: WHOOP (69.6%, independent). REM: Apple Watch (68.6%).
Nocturnal HRV and resting heart rate: Oura Ring 4 (CCC 0.99 and 0.98).
Daytime heart rate: Fitbit Charge 6 and Garmin Vivoactive 5 in the top tier of a ten-device test. Model matters more than brand.
Blood oxygen (SpO2): Apple Watch.
Step count: Apple Watch and Garmin, both under 5% error in the lab. Apple holds up best in free living.
Calories: nobody. Oura's 13% daily error is the best daily figure; Apple pools at about 28%.
VO2 max: Garmin, at roughly 7% error across three independent studies.
Skin temperature: Oura (own study).
No single device wins everywhere. Pick by the metric you care about most.
This table compiles findings across the peer-reviewed studies analyzed. Where a value comes from a different study than its neighbours, the study is named. Each metric section below includes the full data, study details, and funding disclosures.
Biometric | Winner | Second | Third | Worst |
Sleep staging, healthy adults (independent, Antwerp 2025) | Apple Watch S8 (κ=0.53) | Fitbit Sense (κ=0.42) | Fitbit Charge 5 (κ=0.41) | Garmin Vivosmart 4 (κ=0.21) |
Sleep staging, healthy adults (Oura-funded, Robbins 2024) | Oura Gen 3 (κ=0.65) | Apple Watch S8 (κ=0.60) | Fitbit Sense 2 (κ=0.55) | not reported |
Sleep staging, clinical sample (independent, Lee 2023) | Fitbit Sense 2 (κ=0.42) | Oura Gen 3 (κ=0.35) | Apple Watch S8 (κ=0.30) | not reported |
Deep sleep detection (independent, healthy) | WHOOP 4.0 (69.6%) | Fitbit Sense (50.9%) | Apple Watch S8 (50.7%) | Garmin Vivosmart 4 (47.5%) |
REM detection (independent, healthy) | Apple Watch S8 (68.6%) | WHOOP 4.0 (62.0%) | Fitbit Sense (61.3%) | Garmin Vivosmart 4 (33.1%) |
Wake detection (independent, healthy) | Apple Watch S8 (52.2%) | Fitbit Charge 5 (42.7%) | Fitbit Sense (39.2%) | Garmin Vivosmart 4 (27.6%) |
Total sleep time bias | Oura Gen 3 (-3 min, meta) | Fitbit Sense (+6.3 min) | Fitbit Charge 5 (+11.1 min) | Withings ScanWatch (+39.9 min) |
Nocturnal HRV | Oura Ring 4 (MAPE 5.96%) | Oura Gen 3 (7.15%) | WHOOP 4.0 (8.17%) | Garmin Fenix 6 (10.52%) |
Resting heart rate (nocturnal) | Oura Ring 4 (CCC 0.98) | Oura Gen 3 (0.97) | WHOOP 4.0 (0.91) | not reported |
Daytime heart rate (ten-device test, tiers not ranks) | Fitbit Charge 6 (MAPE 5.5%) | Garmin Vivoactive 5 (6.3%) | Apple Watch SE (7.3%) | Fitbit Inspire 3 (16.5%) |
SpO2 | Apple Watch (MAE 2.2%) | Garmin Fenix (about 4.5%) | Withings (about 4.8%) | Garmin Venu (5.8%) |
Step count, free living | Apple Watch S6 (+2.1%, 3 weeks) | Garmin (10 to 18%) | Fitbit Sense (+18%) | Oura Gen 2 (50.3%) |
Calories, daily error | Oura (13%) | Apple Watch (about 28%) | not reported | not reported |
VO2 max | Garmin (about 7%, 3 studies) | Fitbit Charge 2 (about 10%, retired algorithm) | Apple Watch (13 to 16%, underestimates) | not reported |
Skin temperature | Oura (r² above 0.99 lab, own study) | not reported | not reported | not reported |
Sleep Staging Accuracy (4-Stage Classification)
Sleep staging is the most studied and most contested metric in wearable accuracy research. Four studies from 2023 to 2025 produced different rankings, and two things explain most of the difference: who funded the study, and whether the participants were screened healthy sleepers or people with sleep complaints. The same device can score twice as well in the first group as in the second.
Brigham and Women's Hospital Study (2024), Oura-Funded, Healthy Adults
Robbins et al. compared Oura Ring Gen 3, Fitbit Sense 2, and Apple Watch Series 8 against polysomnography (PSG) in 35 healthy adults screened for sleep disorders, over multiple nights.
Device | Overall (κ) | Deep Sleep Sensitivity | Deep Sleep Bias |
Oura Ring Gen 3 | 0.65 (Substantial) | 79.5% | No significant bias |
Apple Watch Series 8 | 0.60 (Moderate) | 50.5% | -43 min (underestimates) |
Fitbit Sense 2 | 0.55 (Moderate) | 61.7% | -15 min (underestimates) |
Funding: This study was funded by Oura Ring Inc. Lead author Dr. Rebecca Robbins is an Oura scientific advisor. The method is the same as the independent studies (attended in-lab PSG, epoch-by-epoch, Cohen's kappa); the higher number comes from screening out anyone with a sleep disorder, not from a method trick.
University of Antwerp Study (2025), Independent, Healthy Adults
Schyvens et al. tested six wrist devices against PSG in 62 healthy adults, one night each. Funded by VLAIO (Flanders Innovation and Entrepreneurship), no device manufacturer funding. Oura was not included, so this study cannot rank rings against watches.
Device | κ (4-stage) | TST bias | Deep sleep | REM | Wake | Notes |
Apple Watch Series 8 | 0.53 | +19.6 min | 50.7% | 68.6% | 52.2% | Best κ, best REM, best wake |
Fitbit Sense | 0.42 | +6.3 min | 50.9% | 61.3% | 39.2% | Lowest bias |
Fitbit Charge 5 | 0.41 | +11.1 min | 43.3% | 47.5% | 42.7% | |
WHOOP 4.0 | 0.37 | +24.5 min | 69.6% | 62.0% | 32.5% | Best deep sleep |
Withings ScanWatch | 0.22 | +39.9 min | not re-verified | not re-verified | 29.4% | |
Garmin Vivosmart 4 | 0.21 | +38.4 min | 47.5% | 33.1% | 27.6% | Oldest hardware |
Note: All six devices misclassified wake, deep sleep, and REM as light sleep, a conservative approach shared across consumer wearables, and all significantly underestimated wake after sleep onset by 12 to 48 minutes. Correction, September 2026: the Fitbit Sense and Garmin deep and REM percentages in an earlier version of this post were shuffled cells from the study's error matrices (Fitbit deep read 48.3% and REM 55.5%; Garmin deep 32.1% and REM 28.7%). The κ, TST and wake columns were always correct.
Korean Multicenter Study (2023), Independent, Clinical Sample
Lee et al. (an earlier version of this post credited it to Park) tested 11 sleep trackers against PSG in 75 adults across 2 centers, 349,114 epochs. This was a clinical sample: every participant had a subjective sleep complaint and the mean apnea-hypopnea index was 18, which is moderate sleep apnea. Several authors are affiliated with Asleep Co., whose SleepRoutine app was the top performer; the values below are for competitor devices, so that conflict cuts against them, not for them.
Device | κ (4-stage) | Deep sleep | REM |
Fitbit Sense 2 | 0.42 | 67.1% | 68.1% |
Oura Ring Gen 3 | 0.35 | 77.8% | 71.2% |
Apple Watch Series 8 | 0.30 | 41.3% | 42.8% |
Note: These numbers are a population effect, not a disagreement with the other studies. Apnea-fragmented sleep is harder to stage, so every device drops: Apple from 0.53 (healthy) to 0.30, Oura from 0.65 (healthy, funded) to 0.35. Fitbit is the most population-robust device at 0.42 in both settings. A separate independent study at Charité Berlin (Herberger et al., 2025) put Oura Gen 3 at 53.2% four-stage accuracy in 45 sleep-clinic patients, against roughly 79% in Oura's own healthy cohort.
Deep Sleep, REM and Total Sleep Time
Per-stage sensitivity is the share of true deep or REM epochs the device caught. No single study tested every brand, so the study is named on each value. Healthy-adult values first, clinical in brackets.
Deep sleep: Oura Gen 3 79.5% (Robbins, Oura-funded) [77.8% clinical, Lee]. WHOOP 4.0 69.6% (Schyvens). Fitbit Sense 50.9% (Schyvens) [67.1% clinical]. Apple Watch S8 50.7% (Schyvens) [41.3% clinical]. Garmin Vivosmart 4 47.5% (Schyvens).
REM: Oura Gen 3 76.0% (Robbins) [71.2% clinical]. Apple Watch S8 68.6% (Schyvens) [42.8% clinical]. WHOOP 4.0 62.0% (Schyvens). Fitbit Sense 61.3% (Schyvens) [68.1% clinical]. Garmin Vivosmart 4 33.1% (Schyvens).
Total sleep time bias: Oura Gen 3 -3.0 min (Khan 2025 meta, 6 studies, N=388). Fitbit Sense +6.3, Fitbit Charge 5 +11.1, Apple Watch S8 +19.6, WHOOP 4.0 +24.5, Garmin Vivosmart 4 +38.4, Withings +39.9 (all Schyvens). Within about 30 minutes is generally treated as clinically acceptable.
Wake detection is universally poor. Every device detects sleep well (89 to 95%) and wake badly (27 to 52% specificity), so all of them overstate how long you slept.
The full sleep breakdown, including sleep/wake agreement and manufacturer claims, is in our most accurate sleep tracker comparison and the sleep tracker accuracy tool.
Nocturnal HRV (Heart Rate Variability) Accuracy
An Ohio State University / Air Force Research Lab study (Dial et al., 2025) validated nocturnal HRV across 13 participants and 536 nights using a Polar H10 ECG chest strap as reference. No industry funding disclosed.
Device | CCC | MAPE | Rating |
Oura Ring 4 | 0.99 | 5.96% ± 5.12% | Nearly Perfect |
Oura Gen 3 | 0.97 | 7.15% ± 5.48% | Substantial |
WHOOP 4.0 | 0.94 | 8.17% ± 10.49% | Moderate |
Garmin Fenix 6 | 0.87 | 10.52% ± 8.63% | Poor |
Polar Grit X Pro | 0.82 | 16.32% ± 24.39% | Poor |
CCC Scale: >0.99 = Nearly Perfect, 0.95–0.99 = Substantial, 0.90–0.95 = Moderate, <0.90 = Poor
Note: Garmin Fenix 6 is 2+ generations behind current hardware, and the study authors acknowledged this. Sample size was 13 participants, though 536 total nights of data were collected (470 on Gen 3, 138 on Ring 4). Oura Ring 5 shipped in June 2026 with a redesigned sensor and has no validation of any kind yet, so the Ring 4 figures do not transfer to it.
Resting Heart Rate Accuracy
From the same Ohio State study (Dial et al., 2025):
Device | CCC | MAPE | Rating |
Oura Ring 4 | 0.98 | 1.94% ± 2.51% | Nearly Perfect |
Oura Gen 3 | 0.97 | 1.67% ± 1.54% | Substantial |
WHOOP 4.0 | 0.91 | 3.00% ± 2.15% | Moderate |
Polar Grit X Pro | 0.86 | 2.71% ± 2.75% | Poor |
Note: Garmin Fenix 6 was excluded from RHR analysis due to timestamp reporting issues that prevented alignment with the Polar H10 reference data.
Daytime and Exercise Heart Rate Accuracy
Correction, September 2026: an earlier version of this section used WellnessPulse figures (Apple Watch 86.3%, Fitbit 73.6%, Garmin 67.7%). WellnessPulse is a consumer aggregator, and its accuracy percentages are pooled correlation coefficients expressed as percentages, not accuracy rates. They have been replaced with peer-reviewed data. The largest single-protocol test is Gielen et al. (2026, KU Leuven, no conflicts): ten devices, one Zephyr chest strap reference, 45 participants, seated rest, a stress task, treadmill walking and intermittent walking. Only two watches were worn per session, so the range is quotable and the exact order is not. Read it as tiers.
Tier | Device | Median MAPE | Agreement (CCC) |
1 | Fitbit Charge 6 | 5.5% | 0.93 |
1 | Garmin Vivoactive 5 | 6.3% | 0.83 |
1 | Google Pixel Watch 2 | 6.7% | 0.87 |
2 | Apple Watch SE | 7.3% | 0.70 |
2 | Garmin Vivosmart 5 | 8.1% | 0.78 |
3 | Polar Ignite 3 | 11.2% | 0.63 |
3 | Xiaomi Watch 2 | 11.9% | 0.69 |
3 | Polar Pacer | 13.1% | 0.66 |
4 | Oura Ring Gen 3 | 15.0% | 0.61 |
4 | Fitbit Inspire 3 | 16.5% | 0.45 |
Three things stand out. No device clears the 5% line cleanly. The best and worst devices are both Fitbits, so model matters more than brand. And the Apple device tested is the SE, the budget model, not a flagship.
Two other independent tests fill the gaps. Van Oost et al. (2025, 12-lead ECG, 24 healthy adults, all devices worn at once) put WHOOP 4.0 at 8.5% MAPE, second worst of five; the Garmin Vivosmart 4 posted the lowest error (4.4%) but dropped 49.6% of its readings to get there, so the two Fitbits (Sense 2 5.6%, Charge 5 5.7%) are the honest winners of that study. Kim et al. (2023) ran Apple Watch 7 and Galaxy Watch 4 against a hospital 12-lead ECG during a treadmill stress test in 44 cardiac patients: both under 2% error, with the Galaxy Watch losing agreement above 160 bpm.
The failure mode is irregular arm movement, not intensity. Steady running is one of the easiest conditions for a wrist device; badminton, soccer and rowing are the worst. Night is easy and day is hard, which is why every manufacturer claim (WHOOP 99.7%, Oura 99%) comes from sleep data. Two free fixes with measured effect sizes: wear the watch three finger-widths above the wrist bone and tighten the strap.
Full device tiers and activity-by-activity data are in the heart rate accuracy tool, the ten-device heart rate ranking and heart rate accuracy by activity type.
Blood Oxygen (SpO2) Accuracy
Garmin Venu 2s underestimated SpO2 in 67.4% of readings. None of these SpO2 features are FDA-cleared for medical use. They are classified as wellness features.
Device | MAE | MDE | Within Range | Missing Data |
Apple Watch Series 7 | 2.2% | -0.4% | 58.3% | 11% |
Garmin Fenix 6 Pro | ~4.5% | not reported | ~44% | 28% |
Withings ScanWatch | ~4.8% | not reported | ~38% | 31% |
Garmin Venu 2s | 5.8% | 5.5% | 18.5% | 14% |
Sources: PLOS, Nature, various validation studies.
Step Count Accuracy
Correction, September 2026: the earlier table here (Garmin 82.6%, Apple 81.1%, Fitbit 77.3%) was WellnessPulse data, which are pooled correlations rather than accuracy rates. Replaced with peer-reviewed error rates. Lab and free-living are different tests, so both are shown.
Device | Lab or treadmill (MAPE) | Free living | Source |
Apple Watch | 0.9 to 3.4% normal gait; 9.3% slow and shuffle walking (Series 5) | +2.1% over 3 weeks (Series 6, Miwa 2026); 6.4% MAPE (Series 6, Kim 2024) | Rowe 2025; Miwa 2026; Kim 2024 |
Garmin | 4.6% (Vivoactive 4, de Leon 2026); about -15% (Fenix 6, Rider 2025) | 10 to 17.8%; Fenix 6 equivalent to criterion in the field | de Leon 2026; Rider 2025 |
Fitbit | 3.6% (Inspire 2, Cheung 2025) | +18.0% over 3 weeks (Sense, Miwa 2026); 17 to 35% (Charge 2, Alta) | Cheung 2025; Miwa 2026; Giurgiu 2023 |
Oura Ring | No lab step data exists | 50.3% MAPE vs pedometer, +2,124 steps a day (Gen 2, Kristiansson 2023, corrected); +1,416 a day vs ActiGraph (Niela-Vilén 2022) | Kristiansson 2023; Niela-Vilén 2022 |
Samsung Galaxy Watch | r=0.82 vs ActivPAL (Watch 4); no MAPE published | No MAPE published | limited |
WHOOP | No published validation | No published validation | none |
In the lab, Apple, Garmin and Fitbit all land under 5% on a normal walk. Free living separates them: Apple stays within a few percent, Fitbit overcounts by about 18%, and both studies that measured Oura found it overcounting by 1,400 to 2,100 steps a day. Oura's 2025 Real Steps algorithm undercounts instead and has no published validation. No Fitbit newer than the Sense or Inspire 2, no Pixel Watch, and no WHOOP has any peer-reviewed step validation.
Full breakdown in the step count accuracy tool and which wearable has the most accurate step count.
Whichever device wins for you, Kygo connects its data to your nutrition and finds your personal correlations. Get it free on iOS or Android.
Energy Expenditure (Calories) Accuracy
All wearables are weak at calorie estimation. The cleanest multi-device figure is the Stanford study (Shcherbina et al., 2017): across seven devices, energy expenditure error ran from 27% to 93% while heart rate on the same devices was under 5%. Correction, September 2026: the earlier table here used WellnessPulse percentages and an Oura 87% figure that was simply 100 minus its 13% error, a different scale that made Oura look top-ranked. Only two devices have a daily-level error rate from a peer-reviewed source.
Device | Daily energy expenditure error | Source |
Oura Ring | 13% free living (21.1% in the lab); underestimates, and error grows with intensity | Kristiansson 2023 (corrected) |
Apple Watch | about 28% (56-study meta-analysis); overestimates in women, underestimates in men | Choe and Kang 2025 |
Fitbit | No daily error rate. Pooled bias 0.19 kcal/min with limits of -5.3 to +5.7 kcal/min | Chevance 2022 |
Garmin | No daily error rate. Steady cardio about 6.7%, light activity 16.5% (Firstbeat engine, PulseOn device) | Parak 2017 |
WHOOP | The 18.4% figure that circulates has no primary publication | unverifiable |
Samsung | 9 to 21%, one small study | limited |
Bout-level figures (kcal per minute during one activity) are not comparable with daily totals, so Garmin and Fitbit do not get a daily rank. Activity-specific errors: Apple resistance training about 52%, Garmin resistance training 57%, Apple walking about 20% and running 24%, Garmin walking 32% and running 22%.
Note: None should be treated as precise calorie counters. Oura's daily figure beats Apple's, but a ring has almost no motion signal when the hand is still, so it undercounts cycling, lifting and hard efforts. Our calorie burn accuracy calculator shows the likely real range for your device and activity.
See our full breakdown of how accurate wearable calorie burn is by device.
VO2 Max Estimation Accuracy
Method beats brand. Estimates built from a real outdoor run have about zero group-level bias against the lab, while estimates built from resting heart rate overestimate by about 2 mL/kg/min (INTERLIVE meta-analysis, 2022). Either way individual error is wide. Accuracy degrades in highly trained users, and the one elite study that avoided this used a chest strap. See the full VO2 max accuracy comparison for how each device estimates it.
Device | MAPE | Bias | Source |
Garmin Fenix 6 | 7.05% | CCC 0.73 vs 30-second lab averages | Carrier 2025 |
Garmin Fenix 6 + chest strap | 6.85% | No degradation in a 95th-percentile group | Carrier 2023 |
Garmin Forerunner 245 | 6.7% overall | 9.4 to 10.4% in highly trained vs 2.8 to 4.1% in moderately trained | Engel 2026 |
Fitbit Charge 2 (retired algorithm) | about 10% | Small overestimate (+1.6 to +2.6 mL/kg/min) in both studies | Freeberg 2019; Klepin 2019 |
Apple Watch Series 10 | 13.2% | -6.25 mL/kg/min, underestimates | Mayo Clinic Proceedings: Digital Health 2026 |
Apple Watch Series 9 / Ultra 2 | 13.3% | -6.07 mL/kg/min, underestimates | Lambe 2025 |
Apple Watch Series 7 | 15.8% | Underestimates | Caserman 2024 |
Polar | 13.7% | -1.0 mL/kg/min | Neudorfer 2025 |
Samsung, WHOOP, Oura, Coros | No independent study | Vendor claims only | none |
Garmin has the most independent validation of any brand. Apple has three independent studies and underestimates in all of them. Fitbit's two studies both found a small overestimate, but on an algorithm Google retired in May 2026 when it moved VO2 max to outdoor GPS runs only. WHOOP's VO2 max exists only on the 5.0 and MG, not the 4.0, and has no independent study.
For a device-by-device breakdown, see which wearable is most accurate for VO2 max, and what actually affects your VO2 max in the first place.
Skin Temperature Accuracy
Oura’s internal validation study (2024) tested temperature sensing across 16 individuals over 1 week (93,571 data points):
r² > 0.99 in lab conditions, r² > 0.92 in real-world conditions, with precision of ±0.13°C per minute.
⚠️ Funding: This is Oura’s own study, not independently peer-reviewed. However, Oura’s temperature data has been validated in independent menstrual cycle tracking studies (Maijala et al., 2019). Apple Watch, Garmin, WHOOP, and Samsung all track skin temperature, but limited independent comparative data exists.
FDA-Cleared Features
Most wearable metrics are wellness estimates. A few features have earned FDA authorization:
Device | Feature | Status |
Apple Watch (Series 4+) | ECG / Atrial Fibrillation Detection | FDA Cleared |
Samsung Galaxy Watch (4+) | ECG / Atrial Fibrillation Detection | FDA Cleared |
Apple Watch (Series 9+, Ultra 2) | Sleep Apnea Notification | FDA Authorized |
Samsung Galaxy Watch | Sleep Apnea Detection | FDA Authorized (Feb 2024) |
Apple Watch | Blood Oxygen (SpO2) | Wellness feature (NOT FDA cleared) |
Fitbit | Irregular Rhythm Notification | FDA Cleared |
Important Caveats
Before drawing conclusions from any of this data, keep these limitations in mind:
No single device wins everywhere. The best device depends on which metric matters most to you.
Healthy vs clinical is the biggest gap. The same device scores far higher in screened healthy young sleepers than in people with sleep complaints: Oura Gen 3 goes from κ=0.65 to 0.35, Apple from 0.53 to 0.30. Numbers from healthy cohorts do not describe most people reading this.
Study funding matters. The primary sleep study (Robbins et al.) was Oura-funded. Independent studies (Schyvens, Lee, Herberger) found different rankings. The funded study used the same method; its higher number comes from its screened population.
Current hardware is barely studied. Peer review lags hardware by 2 to 4 years. Oura Ring 5, WHOOP 5.0 and MG, Apple Watch Series 10, 11 and Ultra 3, Galaxy Watch 7 and 8, Pixel Watch 3 and 4, and every Garmin flagship since the Fenix 6 have no independent validation for any metric here. Validation of one generation does not transfer to the next.
Small sample sizes. The HRV/RHR study had 13 participants (536 nights). Antwerp had 62 participants, 1 night each. Brigham had 35 participants over multiple nights. The ten-device heart rate test had 10 sessions per device.
Mean bias is the wrong metric. A bias near zero means overestimates and underestimates cancelled. Read the error rate and the limits of agreement.
All wearables are estimates. None are medical devices (except specific FDA-cleared features listed above). Data should inform, not diagnose.
Individual variation. Accuracy varies with skin tone, tattoos, BMI, wrist fit, and activity level. Most validation studies have predominantly Caucasian participants, a documented research gap.
PSG is imperfect too. The gold standard polysomnography has interrater reliability of about κ=0.75, meaning even human experts disagree about a quarter of the time on sleep staging.
Common device failure mode. All consumer devices tend to misclassify wake, deep sleep, and REM as light sleep, a conservative approach that inflates light sleep totals.
Why Accuracy Matters for Understanding Food-Biometric Patterns
If you’re trying to understand how nutrition affects your sleep, recovery, or energy levels, the accuracy of your wearable data is the foundation everything else builds on. When measurement error is high, real patterns between what you eat and how your body responds get harder to detect. When accuracy is high, the data can surface connections, like how meal timing affects your overnight heart rate, or whether a supplement is actually changing your HRV, that you’d never spot manually.
This is part of the reason we built Kygo Health to integrate with multiple wearable platforms. Different devices bring different strengths. Connecting them to nutrition data in one place gives you a more complete picture to work with.
Using Multiple Wearables Together
Many people in the biohacking and quantified self communities wear multiple devices simultaneously to capture different metrics from different strengths, Oura Ring for sleep plus Apple Watch for workouts, or WHOOP plus Garmin for different contexts.
The challenge is getting that data to talk to each other. We wrote a detailed guide on this: How to Centralize Health Data from Multiple Devices.
If you’re specifically using Oura for sleep and want to connect that with food tracking, check out: How to Combine Oura Ring with Food Tracking.
Want to compare devices yourself? Explore all the data from these studies in our free Wearable Accuracy Comparison Tool.
Now you know which device is most accurate for each metric. The harder question is what your own numbers actually respond to. Kygo connects your wearable data to what you eat and surfaces your personal correlations. Get it free on iOS or Android.
Sources
Robbins R, et al. (2024). Accuracy of Three Commercial Wearable Devices for Sleep Tracking in Healthy Adults. Sensors, 24(20), 6532. DOI: 10.3390/s24206532. Funded by Oura Ring Inc.
Schyvens AM, et al. (2025). Performance of six consumer sleep trackers in comparison with polysomnography in healthy adults. Sleep Advances, 6(2), zpaf021. DOI: 10.1093/sleepadvances/zpaf021. Independent (VLAIO-funded).
Lee T, et al. (2023). Accuracy of 11 Wearable, Nearable, and Airable Consumer Sleep Trackers. JMIR mHealth and uHealth, 11, e50983. DOI: 10.2196/50983. Independent; clinical sample; several authors affiliated with Asleep Co.
Herberger S, et al. (2025). Oura Ring Gen 3 in sleep-clinic patients. Scientific Reports, 15, 9461. Independent.
Khan S, et al. (2025). Oura Ring sleep meta-analysis, 6 studies, N=388. OTO Open, 9, e70181. Independent.
Dial MB, et al. (2025). Validation of nocturnal resting heart rate and heart rate variability in consumer wearables. Physiological Reports, 13(16), e70527. DOI: 10.14814/phy2.70527. Air Force Research Laboratory funded.
Gielen W, et al. (2026). Ten-device wrist heart rate validation. JMIR Formative Research, e85186. KU Leuven, no conflicts.
Van Oost L, et al. (2025). Five consumer wearables versus 12-lead ECG. Sensors, 25(20), 6319. VLAIO-funded.
Kim J, et al. (2023). Apple Watch 7 and Galaxy Watch 4 versus 12-lead ECG during cardiopulmonary exercise testing. Annals of Rehabilitation Medicine. Independent.
Shcherbina A, et al. (2017). Accuracy in wrist-worn, sensor-based measurements of heart rate and energy expenditure in a diverse cohort. Journal of Personalized Medicine, 7(2), 3.
Choe J, Kang M (2025). Apple Watch accuracy meta-analysis, 56 studies. Physiological Measurement.
Chevance G, et al. (2022). Accuracy and precision of energy expenditure, heart rate, and steps measured by Fitbit. JMIR mHealth and uHealth, 10(4), e35626.
Kristiansson E, et al. (2023). Validation of Oura ring energy expenditure and steps in laboratory and free-living. BMC Medical Research Methodology, with correction DOI: 10.1186/s12874-023-02029-w.
Niela-Vilén H, et al. (2022). Oura ring versus ActiGraph in free living. Computers, Informatics, Nursing, 40(12), 856-862.
Miwa M, et al. (2026). Apple Watch Series 6 and Fitbit Sense step counts over three weeks versus ActiGraph GT9X. PLOS ONE.
Kim, et al. (2024). Apple Watch Series 6 free-living step count versus ActivPAL. Sensors.
Rowe, et al. (2025). Apple Watch Series 5 step count at slow and shuffle walking speeds. PLOS ONE.
de Leon, et al. (2026). Garmin Vivoactive 4 treadmill step count and heart rate. Applied Sciences, 16(3), 1286.
Rider BC, et al. (2025). Four sports watches in laboratory and field. Journal for the Measurement of Physical Behaviour, 8(1).
Cheung YS, et al. (2025). Validity of Fitbit Inspire 2 in step count during treadmill walking. Physiotherapy Practice and Research, 46(2), 95-102.
Parak J, et al. (2017). Energy expenditure estimation with a wrist device using Firstbeat modeling. JMIR mHealth and uHealth.
Molina-García P, et al. (2022). Validity of estimating VO2 max by consumer wearables: INTERLIVE systematic review and meta-analysis. Sports Medicine, 52(7).
Carrier B, et al. (2025). Validation of aerobic capacity and pulse oximetry in wearable technology. Sensors, 25(1), 275.
Carrier B, et al. (2023). Validation of aerobic capacity and lactate threshold in wearable technology for athletic populations. Technologies, 11(3), 71.
Engel FA, et al. (2026). Validity of VO2max estimates from the Forerunner 245 in highly vs moderately trained endurance athletes. European Journal of Applied Physiology, 126, 591-603.
Caserman P, et al. (2024). Assessing the accuracy of smartwatch-based estimation of maximum oxygen uptake using the Apple Watch Series 7. JMIR Biomedical Engineering, 9, e59459.
Lambe R, et al. (2025). Investigating the accuracy of Apple Watch VO2 max measurements. PLOS ONE, 20(5), e0323741.
Apple Watch Series 10 VO2 max validation (2026). Mayo Clinic Proceedings: Digital Health.
Freeberg KA, et al. (2019). Assessing the ability of the Fitbit Charge 2 to accurately predict VO2max. mHealth, 5, 39. Klepin K, et al. (2019). Medicine and Science in Sports and Exercise, 51(11), 2251-2256.
Neudorfer, et al. (2025). Polar VO2 max estimate versus cardiopulmonary exercise testing.
Christakis, et al. (2025). A guide to consumer-grade wearables in cardiovascular clinical care. npj Cardiovascular Health, 2, 82.
Oura Internal Validation (2024). Temperature sensor validation study. 16 participants, 93,571 data points.
Maijala A, et al. (2019). Nocturnal finger skin temperature in menstrual cycle tracking. BMC Women's Health, 19, 150.
Lanfranchi, et al. (2024). Samsung Galaxy Watch SpO2 validation. Journal of Clinical Sleep Medicine.
FAQ: Wearable Accuracy Questions
Which wearable is the most accurate for sleep tracking?
It depends on the study and on the population. In the Oura-funded Brigham study (2024, healthy adults), Oura led with κ=0.65 and 79.5% deep sleep sensitivity. In the independent Antwerp study (2025, healthy adults), Apple Watch led overall (κ=0.53) while WHOOP led deep sleep detection (69.6%). In the independent Korean study (2023, people with sleep complaints), Fitbit Sense 2 led at κ=0.42, with Oura at 0.35 and Apple at 0.30. Fitbit is the most consistent across populations; Apple is best in healthy adults.
How accurate is Oura Ring HRV compared to medical devices?
Oura Ring 4 achieved a 0.99 concordance correlation coefficient with a Polar H10 ECG chest strap in an independent 536-night study (Dial et al., 2025), the highest HRV accuracy among the consumer wearables tested. That is nocturnal HRV only; Oura's daytime heart rate ran 15% error in a separate ten-device test. Oura Ring 5 has not been validated.
Is WHOOP accurate for HRV tracking?
WHOOP 4.0 showed a CCC of 0.94 and MAPE of 8.17% for nocturnal HRV in the Dial et al. (2025) study, rated Moderate on the concordance scale. WHOOP's 99.7% marketing figure comes from a one-night sleep study on the WHOOP 3.0 and does not describe daytime accuracy, where WHOOP 4.0 ran 8.5% error against a 12-lead ECG.
Does skin tone affect wearable accuracy?
PPG sensor accuracy is affected by skin pigmentation. Most validation studies have predominantly Caucasian participants, which is a known research gap. Accuracy data may not generalize equally across all skin tones.
Can I use multiple wearables together?
Yes. Many people wear multiple devices to capture different metrics from each device’s strengths. The challenge is correlating data across platforms, which typically requires a third-party platform or manual comparison.
Which wearable is best for tracking how food affects sleep?
For nutrition-sleep correlation analysis, you need accurate sleep and HRV data paired with consistent food logging. The studies above show which devices perform strongest for each metric. The best choice depends on which signal you care about most: a ring for nocturnal HRV and resting heart rate, a watch for sleep stage labels and daytime heart rate.
Are wearable calorie estimates reliable?
No wearable tracks calories with high precision. In the Stanford seven-device study, calorie error ran from 27% to 93%. The best daily figure from a peer-reviewed source is Oura at 13% error in free living; Apple Watch pools at about 28% across 56 studies. All devices should be treated as rough estimates, and accuracy drops further during high-intensity or multi-modal exercise.
Why do different studies show different accuracy rankings?
Study funding, sample size, population demographics, device firmware version, number of nights tested, and PSG scoring protocols all affect results. This is why we include multiple studies and flag funding sources throughout this article.
Disclaimer: Kygo Health LLC is a personal data aggregation and insights platform designed for informational purposes only. The information provided does not constitute medical advice, diagnosis, or treatment. Always consult a licensed healthcare provider with any questions regarding medical conditions.
Have questions about wearable accuracy or data you think should be included? Reach out directly at Ryan@kygo.app. If you have sources or credible data that isn’t listed here, share it and we’ll review and update accordingly.