Sleep & Recovery Science explainer
Why Two Sleep Trackers Can Give Different Scores for the Same Night
Updated
Sleep score differences are common when two gadgets share one bed. A ring, a watch, and a mat do not share one scale. This is a source-reviewed guide, not a hands-on lab test.
People notice sleep score differences on travel nights, after wine, or after they buy a second gadget. One app says 91. The other says 74. The body feels the same. The numbers do not.
Product pages for best sleep trackers stay on features and prices. They should not settle a clinic question. A contactless sonar tool such as SleepScore is a different instrument from a wrist optical sensor. A snore recorder such as SnoreLab is another class again.
What a sleep score actually is
In a sleep lab, scorers read brain waves, eye movements, and muscle tone. They mark 30-second epochs. That method is polysomnography. It is the reference most validation papers still use.
A consumer score is downstream of a model. The model sees motion, pulse, temperature, or a bed sensor. Software then labels sleep, wake, and often light, deep, and REM. A second model blends those labels with heart-rate variability, schedule regularity, or a readiness mix. The badge you see is that blend.
Two badges can disagree even when both devices logged the same clock time in bed. They are not measuring one shared quantity. They are publishing two proprietary summaries.
The American Academy of Sleep Medicine’s 2018 position statement is still the clinical line. Consumer sleep technologies that lack FDA clearance cannot diagnose or treat a sleep disorder. They may support a conversation. They do not close the case.
Why sleep score differences appear on the same night
Sensor class is the first split. Wrist and ring devices lean on accelerometers and optical pulse. Under-mattress mats lean on pressure and ballistocardiography. Phone tools may use sonar or a microphone. Those signals are proxies. They are not EEG.
Stage maps are the second split. Brands do not share one epoch length. They do not share one rule for “light.” A paper that collapses N1 and N2 into “light” is already on a different map from a five-stage lab score.
Score recipes are the third split. One brand may weight total sleep time. Another may punish late bedtime. A third may fold in resting pulse. If those inputs move for reasons that do not match how you feel, the badge moves too.
Firmware is the fourth split. Algorithms update after a paper is printed. Chinoy and colleagues compared seven consumer devices with PSG in healthy adults (Sleep, 2021). Later commentary in SLEEP Advances (2025) warned that device generation and firmware lag can make an old paper a poor certificate for a 2026 build.
What lab papers measured on the same night
The useful studies put several gadgets on one person and keep PSG as the referee. They still do not validate a marketing score. They validate sleep–wake and stage labels that later feed a score.
| Question or claim | Evidence source | Study type | Population | Reference standard | Outcome | Key finding | Limitation | Applicability |
|---|---|---|---|---|---|---|---|---|
| Two consumer devices should match each other and PSG on stages | Lee et al., JMIR 2023 | Prospective multicenter validation | 75 adults in Korea; clinic and community recruits with sleep discomfort | Lab PSG, epoch-aligned | Macro F1 and Cohen’s κ for four stages | Macro F1 ranged from 0.69 to 0.26 across 11 tools. Wearables often called wake “light.” Nearables often called REM “light.” | One-country sample; missing nights from battery and account errors | Explains sleep score differences when a watch and a mat share a night |
| Wrist devices detect wake as well as they detect sleep | Schyvens et al., SLEEP Advances 2025 | Lab validation, one night | 62 adults (52 men, 10 women); mean age 46 | PSG | Sensitivity, specificity, κ, TST bias | Sleep-epoch sensitivity above 90%. Wake specificity 29.39% to 52.15%. κ 0.21 to 0.53. Most devices overestimated TST versus PSG. | Male-heavy sample; one night; firmware dated to the study year | A high “asleep” rate can hide quiet wake and inflate a score |
| Stage minutes from five commercial tools match PSG | Kainec et al., Sensors 2024 | Lab comparison vs PSG and actigraphy | 53 healthy adults aged 18–30; 50 nights after exclusions | PSG | Bias in light, deep, and REM minutes | Staging accuracy versus PSG was poor. Light-versus-deep errors were large. Some stage estimates differed by up to 250 minutes. | Young healthy sleepers; not a clinic OSA sample | A deep-sleep swing can move a score even when time in bed is stable |
| A consumer score can replace a sleep study | Khosla et al., AASM 2018 | Position statement | Clinical practice | None | Policy | Uncleared consumer tools cannot diagnose or treat sleep disorders. | Statement age; some later devices sought clearance for narrow uses | A badge is not an AHI and not a treatment plan |
| Public health sleep advice requires a wearable score | CDC About Sleep; NHS insomnia pages | Public guidance | General population | None | Duration and symptoms | CDC copy still centres adult sleep duration of 7 hours or more. NHS insomnia copy starts with symptoms and GP care, not an app grade. | Pages are not device trials | Use official duration and function first |
Score recipes are not a shared scale
Even if two devices agreed on total sleep time, their badges can still split. One formula may treat 7 hours as a ceiling. Another may treat regularity as the main lever. A third may drop the score after a late heart-rate bump that you never felt.
Manufacturer blogs sometimes publish internal accuracy claims for pulse or HRV. Those claims are not a shared sleep-score standard. They also change with firmware. We do not copy unaudited “percent accurate” lines into this page as fact.
Lee’s 2023 matrices are the clearer public picture. Devices did not fail in the same way. Wearables leaned toward calling wake “light.” Nearables leaned toward calling REM “light.” Airable phone apps mixed light and deep. If those labels feed a score, the morning number will drift by brand even on one mattress.
That is why sleep score differences are the expected output, not a glitch you can “calibrate away” by wearing both gadgets forever.
Quiet wake, firmware and missing minutes
Still wakefulness is the classic miss. You lie still and think. Motion is low. Pulse may look sleep-like. Many wrist tools then mark light sleep. Schyvens 2025 said wake-within-sleep is hard for that reason. Specificity stayed low even when sensitivity looked excellent.
Missing wear time is the next miss. A loose watch, a dead ring, or a mat that loses the sleeper creates gaps. Some apps fill gaps. Some drop the night. A filled gap can look like extra sleep. A dropped night can look like a “crash” the next day.
Alcohol, a warm room, a cold hand, and a tattoo under an optical path can all move the pulse signal. Those are ordinary nights. They are not proof the second device is “more scientific.”
Short sleep also changes next-day hunger. That is a separate problem from a badge fight. See sleep and weight regulation for the appetite side, not for a stage verdict.
How to read two devices without chasing a grade
Pick one instrument for two weeks. Hide the second score if you can. Log bedtime, wake time, and a simple 0–2 morning alertness note. Compare weeks, not Wednesdays.
Ask a question the stronger metrics can answer. “Did I give myself under 7 hours in bed four nights running?” is a duration question. CDC adult copy still treats 7 hours or more as the usual need. “Did my REM share drop 4%?” is a weaker question.
If you already own an Android alarm tool such as Sleep as Android, treat it as a clock and a log. Do not add it to a three-way score tribunal. Generated audio such as Pzizz is not a referee either. It is sound, not EEG.
If sleep score differences appear and you feel fine, believe function. If both badges look “perfect” and you are unsafe to drive, believe function. The paper trail above is why.
When a score is the wrong tool
Loud snoring with pauses, gasping, morning headaches, and crushing sleepiness belong with a GP or sleep clinic. NHS sleep apnoea pages start with those symptoms. They do not start with a 0–100 badge.
Insomnia that lasts for months needs a person, and often a structured therapy path, not a new ring. Chasing sleep score differences can become its own arousal. Clinicians have described that loop as orthosomnia. Turning stage charts off is a fair experiment.
This page does not name a winner. The evidence says the disagreement is structural. Sensors differ. Maps differ. Recipes differ. Use one calm diary, or use none, and keep the clinic door for red flags.
Frequently asked questions
Why do two trackers disagree on the same night?
Which score should I trust?
Can a higher-priced device remove sleep score differences?
Do sleep scores diagnose sleep apnoea?
How long should I wear a tracker before judging a pattern?
Should I take screenshots of REM percentages to a GP?
Sources
- 1. Accuracy of 11 Wearable, Nearable, and Airable Consumer Sleep Trackers: Prospective Multicenter Validation Study
- 2. A performance validation of six commercial wrist-worn wearable sleep-tracking devices for sleep stage scoring compared to polysomnography
- 3. Evaluating Accuracy in Five Commercial Sleep-Tracking Devices Compared to Research-Grade Actigraphy and Polysomnography
- 4. Performance of seven consumer sleep-tracking devices compared with polysomnography
- 5. Consumer Sleep Technology: An American Academy of Sleep Medicine Position Statement
- 6. About Sleep
- 7. Insomnia
Guidance changes. Figures were checked against the sources above at the time of review; always confirm current advice with your GP, pharmacist or clinician.
Image credits
- Photo: Photo by Airam Dato-on on Pexels / Openverse
- Photo: Photo by Elina Volkova on Pexels / Openverse
- Photo: Photo by Pexels contributor on Pexels / Openverse
- Photo: Photo by MART PRODUCTION on Pexels / Openverse
- Photo: Photo by Andrea Piacquadio on Pexels / Openverse
Was this useful?
Anyone can react — no account needed.
Discussion
0 comments · Name and email only · Email is never shown