Heart Rate Accuracy by Activity Type: What Breaks Your Watch
- Ryan - Kygo Health

- Aug 5
- 13 min read
Updated: Aug 9
Last Updated: August 4, 2026

Wrist heart rate is most accurate during steady running and least accurate during racquet sports, rowing, interval work and ordinary household movement. On one device across 77.5 hours of real training, running produced 1.2% error and badminton produced 16.2%. On another, the same bike and the same people produced 11.8% error at maximal steady effort and 26.0% on intervals. That ordering surprises most people, because the popular assumption is that hard efforts break the sensor. They don't. What breaks it is unpredictability, whether that is your arms moving erratically or your heart rate changing faster than the device can follow.
I went looking for a per-activity table and found nine independent studies that had each built part of one, almost none of which cited the others. Lined up, they agree on the shape and disagree in two interesting places. If you want the device-by-device picture instead, that is in how accurate your heart rate monitor is. This post is about when your device is wrong rather than which one you own.
Heart rate accuracy by activity, one device, one sample
The cleanest per-activity data comes from Ceugniez and colleagues, who put a Fitbit Charge 4 on 26 people for 77.5 hours across 55 genuine training sessions, referenced against a Polar H10 chest belt. Using one device across all activities removes the usual confound.
Activity | Error (MAPE) | Bias | Spread (limits of agreement width) | Agreement (ICC) |
Running | 1.2% | +0.1 bpm | 21.6 bpm | 0.90 |
Cycling | 8.1% | +4.8 bpm | 64.1 bpm | 0.66 |
Tennis | 8.9% | -6.2 bpm | 62.0 bpm | 0.42 |
Orienteering run | 9.5% | -8.6 bpm | 69.4 bpm | 0.80 |
Badminton | 16.2% | -16.5 bpm | 103.5 bpm | 0.37 |
Soccer | 17.5% | -16.5 bpm | 86.3 bpm | 0.49 |
The authors name the mechanism directly. Badminton and tennis involve sharp movements and rotations of the non-dominant arm, and soccer creates sensor instability through random arm and wrist actions.
Watch the bias column, because the direction matters more than the magnitude. During badminton and soccer the watch under-read by 16.5 bpm. It did not produce noise around the truth, it systematically told those athletes they were working easier than they were. If you train by heart rate zones in a field or racquet sport, your watch is quietly parking you a zone below where you actually are, and every downstream number inherits that error. Your calorie burn, your training load and your recovery score are all built on top of it.
It replicates, and arm movement is measurably the cause
Vermunicht and colleagues tested a different device on a different population, cardiac rehabilitation patients, against the same chest-strap reference. Same ordering.
Activity | Error (MAPE) | Mean absolute error | Agreement (ICC) |
Walking | 3.8% | 3.8 bpm | 0.96 |
Cycling | 6.9% | 8.7 bpm | 0.81 |
Running | 8.5% | 12.1 bpm | 0.79 |
Rowing | 13.4% | 19.8 bpm | 0.44 |
Rowing was significantly worse than everything else in that protocol at p less than 0.001. Worth being precise, though: it is the worst modality in that study, not the worst number in this post. Soccer at 17.5%, badminton at 16.2% and interval cycling at 26.0% are all worse. Rowing is not especially high intensity. It is, however, almost entirely arms.
That study then measured the mechanism instead of inferring it. Activities involving intensive arm movement, specifically rowing and arm biking, produced accurate readings only 56.6% of the time, against 78.2% for activities without intensive arm movement. That is the clearest causal evidence available, and it is why the ranking holds across devices and populations.
Two honest notes on that table. This cohort was cardiac patients, and running here reads 8.5% against 1.2% in the healthy sample above, so population moves the absolute numbers even when the shape survives. And walking beating running is the reverse of the first table. That reversal shows up again by age: in a study splitting young adults from over-65s, walking error rose from 3.77% to 7.06% in the older group while running stayed flat at about 2.5%. That is a low-amplitude pulse problem, not a motion one.
Everything above describes averages. What your own device does on your own wrist during your own sport is a different question, and the only way to answer it is your own data. Kygo syncs heart rate and recovery from Oura, Apple Health, Health Connect, Fitbit, Garmin and WHOOP and lines them up against what you actually did and ate, so you can see which of your sessions produce numbers worth trusting. Free on iOS and Android.
The cleanest proof: one machine, hands on or off
Almost every comparison above changes the sport, the intensity and the population at once. One study changed only whether the participant's hands were on the moving levers. Same elliptical machine, same session, same person, same workload setting.
Condition | Apple abs % difference | Apple CCC | Garmin abs % difference | Garmin CCC |
Elliptical, no arm levers | 3.2% | 0.94 | 9.7% | 0.54 |
Elliptical, with arm levers | 6.5% | 0.75 | 13.7% | 0.31 |
Engaging the levers roughly doubled the error on both watches and dropped agreement hard: Apple from 0.94 to 0.75, Garmin from 0.54 to 0.31. The Polar chest strap worn in the same sessions held 0.99 in every condition, so this is the wrist sensor failing, not the protocol.
Two limits decide how far this goes. Participants wore two watches, one per wrist, and the Apple and Garmin groups were largely different people, so the gap between the two brands is between-subjects and is not a like-for-like comparison. What survives cleanly is the within-device change when the levers engage, and that is the finding. The paper also prints no heart rate values for either elliptical condition, so equal intensity is inferred from the machine setting, and pulling the levers recruits more muscle, which plausibly raises heart rate somewhat.
The same pattern at population scale
Zhang and colleagues pooled 44 articles, 738 effect sizes and 15 brands. Different devices, different labs, same shape.
Activity | Pooled mean difference | 95% CI |
Rest | -0.01 bpm | -0.02 to 0.00 |
Sleep | -0.40 bpm | -1.64 to +0.83 |
Treadmill, walking to running | -0.51 bpm | -1.60 to +0.58 |
Cycling | -4.55 bpm | -7.24 to -1.87 |
Resistance training | -7.26 bpm | -10.46 to -4.07 |
Treadmill work sits essentially at zero across the entire literature, including at running pace. Cycling and lifting, the two activities where the hands grip something and the arms carry load, are where it falls apart.
There is also a dose-response effect during resistance training: the error grew by about 3 bpm for every 10 bpm rise in true heart rate. So the harder the set, the further off the number, which is the one place the intensity story does hold.
One caveat we would rather surface than have you find: the rest row's confidence interval of plus or minus 0.01 bpm is two orders of magnitude tighter than every other row in that table and is not credible as a precision estimate. It is quoted as published. Do not use it to argue that wrist sensors are essentially perfect at rest.
Cycling is the activity everyone gets wrong, including the researchers
If you cycle, you have probably read that your watch is fine on the bike. The evidence does not support that, and it does not support the opposite either.
Zhang's meta-analysis makes cycling the second worst activity at -4.55 bpm. Chevance's pooled Fitbit analysis found cycling worse than daily living, treadmill and overground walking. Ceugniez measured 8.1% and Vermunicht 6.9%. But the Stanford study, still the most cited work in the field, found cycling the single most accurate activity it tested at 1.8% median error, better than walking at 5.5%.
Those results genuinely conflict, and the reason is probably that nobody has isolated the variable that matters. Gripping handlebars fixes the wrist in extension and presses the sensor against the wrist bones, which is mechanically very different from a wrist swinging freely. That mechanism is well supported for rowing and racquet sports. For cycling it has never been tested directly, so the conflicting numbers may be measuring different bike positions rather than different truths.
The practical read: treat steady cycling as unresolved, note that every cycling figure in this research except one pooled result comes from an indoor ergometer, and if the number matters to you on the bike, wear a strap and compare for a week.
Intervals break it worse than maximal effort does
This is the finding that complicates the arm-movement story, and it produces the single worst number in the research.
Reddy and colleagues ran 20 people through six conditions with a Garmin and a Fitbit worn simultaneously, one on each wrist, against a Polar H7 chest strap.
Condition | Garmin MAPE |
Maximal treadmill | 5.8% |
HIIT treadmill | 9.0% |
Maximal cycle ergometer | 11.8% |
HIIT cycle ergometer | 26.0% |
Resistance circuit | 10.6% |
Activities of daily living | 13.0% |
Same machine, same participants, same session structure. Going from maximal steady effort to intervals more than doubled the error on a bike and raised it by half on a treadmill. Note what that does to the story: interval cycling is the worst number anywhere in this research, and cycling is not an arm-heavy activity. So the honest mechanism is not purely arm movement. It is unpredictability. Erratic arm motion and rapidly changing heart rate both defeat the smoothing these devices depend on.
One caveat on the absolute numbers. The Garmin contributed roughly 28% as many data pairs as the Fitbit in every condition and the paper never reconciles that, which is the same completeness objection that reverses device rankings elsewhere. Read this table as one device compared across conditions, not as a Garmin error level.
A second study found the same shape on different hardware. A Garmin Fenix 6 and a Polar Grit X worn simultaneously against a Holter ECG held agreement of 0.954 and 0.553 during maximal rucking, then fell to 0.589 and 0.243 during a Tabata circuit. Chest straps in the same sessions held 0.909 to 0.997 throughout, circuit included.
Lifting: the biggest hole in the research
Resistance training is where the evidence splits hardest, and where you should trust nobody's confident answer.
Zhang's pooled figure makes lifting the worst activity in the entire meta-analysis at -7.26 bpm. Polar's own white papers, which are unusual in publishing unflattering numbers, put strength training at 5.6 to 9.1 bpm mean absolute error against 1.0 to 2.4 bpm for running and cycling.
Then Lee and colleagues tested 62 men against ECG across four watches and found resistance agreement better than endurance, with correlations of 0.96 to 0.97. Three things belong with that result: the sample was 62 men and zero women, three of the four devices systematically over-read by 2.5 to 4.5 bpm, and the study was funded by a health care company with an employed co-author. Correlation is also flattered by the wide heart rate range inside a lifting session, which makes agreement look stronger than the beat-by-beat reality.
Here is the part worth knowing. Per-brand figures do exist for individual lifts and circuits: Garmin 10.6% on a resistance circuit, Samsung 6.2% and Fitbit 15.7% on squats, Garmin 10.2% and Polar 14.2% on a Tabata circuit. What does not exist anywhere is per-brand error across a conventional multi-lift resistance session. So if you find a table online ranking watches for lifting accuracy, it is not sourced from research on lifting as you actually do it.
Burpees break every device and every placement
Moghaddam and colleagues ran a 30-second burpee condition across 28 participants to test whether moving the sensor somewhere better could rescue ballistic movement. It mostly cannot.
Only a WHOOP worn on the upper arm held bias under 1 bpm, and even that fell to an agreement coefficient of 0.46. Every other placement scored below 0.35, with limits of agreement spanning roughly 40 to 60 bpm in each direction. A Garmin Forerunner 55 on the wrist was off by 33 bpm. Most striking, a Polar Verity Sense on the forearm that is near perfect on a treadmill, at 0.03 bpm bias and 0.997 agreement, degraded to 6.48 bpm during burpees.
If your training is CrossFit-style, plyometric or circuit-based, no wrist or forearm device is going to give you a reliable in-session number. That is a hardware limit, not a settings problem.
Swimming, where the evidence nearly runs out
Only two studies report usable per-device swimming figures, and part of one of them fails outright. Over 500 m against a Polar H10, a Garmin Forerunner 945 posted 3.29% error and a Polar Ignite 8.61%. Look past the means, though: the correlations were 0.49 and 0.08, which are poor even where the average error looks acceptable.
The same study's 1000 m figures, 55% and 51%, are a criterion failure rather than a device failure. Both devices collapse identically, which is the signature of a broken reference, and the authors themselves question whether a chest strap stays valid underwater over that duration. Those numbers should never be quoted as wearable accuracy, and they occasionally are.
No swimming validation with usable per-device agreement statistics exists for Apple, Fitbit, Samsung, WHOOP or any ring.
The condition nobody tests: ordinary daily life
Every study above tested exercise. The most interesting finding sits outside it.
Nelson and Allen instrumented one person with ambulatory ECG for 24 hours and captured 102,740 heartbeats. An Apple Watch Series 3 posted 3.01% error during running and 13.70% during ordinary activities of daily living, and daily living was the only condition in the entire study to exceed the 10% threshold. That is an N of 1, so treat it as a well-documented case rather than a population estimate.
It is not an isolated signal, though. Chevance's meta-analysis of 32 Fitbit studies found heart rate more accurate during moderate to vigorous activity than during light activity, which is the exact reversal of the intensity model and fits the arm-movement one precisely. Light activity is where reaching, carrying, typing and washing up live.
So the resting heart rate your watch reports while you are awake and pottering around is likely its least reliable number of the day, which is one more reason overnight resting heart rate, measured while you are still, is the metric worth watching. We compare devices on that in the most accurate sleep tracker.
Nobody can rank wearables by activity, including us
It is worth being blunt about how thin this evidence is. Lay every activity in this post against every major brand and you get 119 cells. Thirty-two carry a number. Seventy-seven are empty.
Some specifics. Rowing has been measured on exactly one device, a Fitbit. Racquet and team sport, also one device, also a Fitbit. WHOOP has no validation during any named sport or structured training modality at all, only sleep, treadmill walking, burpees and unstructured daily life. No current-generation smart ring has any published exercise data. Skiing has never been tested on anything, which matters because cold drives peripheral vasoconstriction. Outdoor cycling has never been separated from indoor on any brand.
There is also a population trap. The same lab ran the same protocol on healthy adults and on cardiac rehabilitation patients and got opposite answers: in healthy adults the treadmill was the easier condition, in patients it was the cycle. Any best-device-for-your-sport claim built on healthy volunteers does not transfer to the people most likely to care.
So when a site publishes a table ranking watches for your sport, check whether the underlying study exists. For most sports it does not.
What to actually do with this
Match the tool to the activity rather than trusting one number everywhere.
For running, steady cardio and anything on a treadmill, your wrist device is fine. The error is small enough that zone training works.
For racquet sports, field sports, rowing, arm biking and circuit work, wear a chest strap or an upper-arm optical sensor if the number matters. Moving the sensor off the wrist is the single biggest available improvement, and an upper-arm sensor measured 1.35% error against 6.82% for a wrist watch in a direct comparison.
If a strap is not an option, two free adjustments help. Move the watch further up your forearm, roughly three finger-widths above the wrist bone, which cut error by 11.4 percentage points in a controlled test. And tighten the strap, because contact pressure affected signal quality more than exercise intensity did in a study that varied it directly.
For lifting, assume the number is soft in both directions and judge the session by what you lifted rather than by what your watch said your heart did.
You can filter these findings by device and activity in our heart rate accuracy tool rather than working from static tables.
Common questions
Which activity is wrist heart rate most accurate for?
Steady running. It produced 1.2% error in the best real-world study available, and treadmill work sits essentially at zero bias across a 44-article meta-analysis. Continuous, predictable arm motion is the easiest condition an optical sensor sees.
Which activity is it worst for?
Interval cycling at 26%, then soccer at 17.5%, badminton at 16.2% and rowing at 13.4%. Racquet and field sports involve sharp arm movement; intervals defeat the device a different way, by changing heart rate faster than its smoothing can follow. Ballistic work like burpees is worse still.
Is my watch accurate for weight lifting?
The evidence conflicts. A pooled meta-analysis puts lifting worst of all activities at -7.26 bpm, while one industry-funded ECG study of 62 men found it better than endurance work. Per-brand figures exist for single lifts and circuits, roughly 6% to 16%, but nothing covers a conventional multi-lift session. Treat the number as soft.
Why does my heart rate read low during tennis or soccer?
Because wrist sensors under-read during racquet and field sport rather than reading noisily. One study measured a 16.5 bpm systematic under-read in both. Sharp rotations of the non-dominant arm disrupt the optical signal.
Is heart rate accuracy worse at higher intensity?
Generally no, and often the opposite. A meta-analysis of 32 studies found readings more accurate during moderate to vigorous activity than during light activity. What does hurt is change: the same bike and the same people produced 11.8% error at maximal steady effort and 26.0% on intervals. Resistance training is the other exception, where error grew about 3 bpm for every 10 bpm rise in true heart rate.
Is my watch accurate for cycling?
Unresolved for steady riding. One meta-analysis makes cycling the second worst activity, while the Stanford study found it the most accurate one it tested, and handlebar grip has never been isolated as a variable. Interval cycling is a separate and clearer case at 26% error. Note also that every cycling figure here except one pooled result comes from an indoor ergometer.
Does a chest strap fix this?
For arm-heavy activities, largely yes. A strap is not perfect either, running 2.28% to 3.86% error against a 12-lead ECG, but its advantage grows sharply exactly where wrist sensors fail.
The bottom line
Wrist heart rate is not one accuracy number, it is a range from about 1% to 26% depending entirely on what you are doing. Running, walking and steady cardio are near the top of that range. Rowing, racquet sports, field sports, intervals and burpees are near the bottom, and ordinary daily pottering may be worse than either.
Intensity is not the variable. Unpredictability is, whether that is your arms moving erratically or your heart rate changing fast. Which means the cheapest accuracy upgrade available to you is not a new device, it is moving the one you own further up your arm, tightening the strap, and putting on a chest strap for the two or three activities where you actually care.
The other half is knowing what your numbers respond to once you trust them. Kygo ties your heart rate, HRV and recovery data to what you eat and how you train, so the patterns that emerge are yours instead of a study average. Get it free on iOS or Android.
Key sources: Ceugniez 2025 (JMIR mHealth, six sports, no funding statement), Vermunicht 2025 (Eur Heart J Digit Health, rowing and arm movement), Gillinov 2017 (MSSE, the elliptical arm-lever isolation), Reddy 2018 (JMIR mHealth, intervals vs maximal effort), Merrigan 2023 (Sensors, Tabata circuit vs rucking), Budig 2022 (Sensors, swimming), Zhang 2020 (J Sports Sci, 44-article meta-analysis, abstract-verified), Chevance 2022 (JMIR mHealth, intensity reversal), Chow and Yang 2020 (JMIR mHealth, the age axis), Nissen 2022 (JMIR Form Res, ten tasks vs Holter), Etiwy 2019 (Cardiovasc Diagn Ther, the cardiac-population inversion), Shcherbina 2017 (J Pers Med, the Stanford study), Moghaddam 2026 (Sensors, ballistic movement), Nelson 2019 (JMIR mHealth, N=1 daily living), Lee 2026 (Sensors, resistance, industry-funded), Schweizer 2025 (JMIR Cardio, upper arm vs wrist), Scardulla 2020 (Sensors, contact pressure), Polar (self-published strength training data).
Disclaimer: Kygo Health is a personal data aggregation and insights platform designed for informational purposes only. The information provided by Kygo, including correlations, patterns, and trends identified in your data, does not constitute medical advice, diagnosis, or treatment. Always consult a licensed healthcare provider with any questions regarding medical conditions.
If you have compared your watch against a strap during a real session, I would want to know which activity showed the biggest gap, and whether it read high or low.