What does the lawsuit actually allege?
The complaint's one-line version is that the marketing outran the sensors. Surber v. Oura Inc. was filed on 20 August 2026 in the U.S. District Court for the Northern District of California by Clarkson Law Firm on behalf of a California buyer, and it asks to represent a class. It quotes Oura's marketing as promising rings "built for accuracy", with "unparalleled accuracy", able to see "what only a hospital sleep lab can see", and it cites accuracy figures it says Oura has published over the years: 79 percent, and more recently "95% Sleep Staging Accuracy compared to clinical sleep lab". The complaint's core argument is a hardware one. Sleep stages are defined by brain activity, eye movement and muscle tone, the ring records none of those, and the firm says the real figure for staging is nearer to 50 percent, which the filing calls a coin flip. It asks the court to stop the marketing and to refund purchasers.
Those are allegations. Nothing has been decided and the class has not been certified. I am not a lawyer, so nothing here is a view on the filing's legal merits. What I can read is the science both sides are pointing at, because it is published.
Oura's answer is the more interesting document. In its statement to the press the company said it stands behind its science and its accuracy claims, that the ring "estimates sleep stages using multiple physiological signals, including heart rate, heart rate variability, movement, breathing patterns, and temperature", and that there is peer-reviewed evidence that sleep stages "are associated with distinct, measurable, reproducible changes in physiology". It added that the ring "is not a medical device or a substitute for a clinical sleep study", and that its staging "has been validated and compared favorably in multiple studies against polysomnography, the gold standard". Two days later its blog called the suit baseless and put the position more plainly: the autonomic signal is one of the main signals wearables use to infer sleep stages, and "it is not a coin flip; it is physiology", read through a different window.
Read the defence carefully. It rests on the validation studies, which the next two sections go through, and on one phrase. Associated with is true, and it is a different claim from measured. The company's own verbs are estimate and infer. Everything below is about the distance between those verbs and the word accuracy.
What does a ring actually sense?
Pulse timing, movement, skin temperature, breathing, and on some models a blood-oxygen trend. That is Oura's own list, and it is the whole of the raw material. The instrument reads those signals well. In Cao and colleagues' 2022 comparison against a reference ECG in 35 sleepers, an Oura ring tracked overnight heart rate at r = 0.99 and the standard beat-to-beat variability measure at r = 0.92. I have no quarrel with the sensor. A wrist band starts from the same short list, which is why sorting any wearable's spec sheet into sensed and computed is the most useful thing a buyer can do.
Sleep stages are defined in different terms. A sleep lab records brain electrical activity from scalp electrodes, eye movement from electrodes beside the eyes, and muscle tone from the chin, and a technician scores the night epoch by epoch against the American Academy of Sleep Medicine's manual. Wake shows fast, low-amplitude brain activity. Light sleep is scored on sleep spindles and K-complexes, deep sleep on slow delta waves. REM looks like waking on the brain trace, with the eyes moving and the muscles switched off. None of that is available at a finger.
So the ring is measuring something else and handing it to a model. The model was trained on nights where a ring and a sleep lab ran at the same time, and it learned which patterns of pulse, movement and temperature tend to go with each label a scorer wrote. Oura's 2022 post says the current algorithm also folds in "mathematical modeling that incorporates well known sleep patterns", which means it carries a prior about what a night usually looks like. A model with a good prior will look right on an ordinary night. The question is what it does on the night that is not ordinary, which is the night the person bought the ring for.
That is what a stage label from a ring is: a prediction of what a technician would have written, made from signals the technician does not use. It can be a good prediction. It is still a prediction, and an accuracy figure for it has to be read that way.
Is the Oura Ring's sleep staging accurate?
Neither a coin flip nor a sleep lab. It separates sleep from wake well when you are moving, and badly when you are lying still. The independent study the complaint appears to be leaning on is de Zambotti and colleagues' 2019 comparison of an earlier Oura ring against polysomnography in 41 adolescents and young adults. The ring identified sleep with 96% sensitivity and identified wake with 48% specificity, so about half of the awake epochs were scored as sleep. Stage agreement was 65% for light sleep, 51% for deep and 61% for REM, and the ring underestimated deep sleep by about 20 minutes a night and overestimated REM by about 17.
Those numbers describe an older ring and an older algorithm, and Oura has updated both. Its own validation, published in 2021 by two authors affiliated with Oura Health, used 440 nights from 106 people and reports 96% accuracy for sleep versus wake and 79% across four stages, up from 57% when the model used movement alone. In a 2022 independent comparison of six wearables against polysomnography, one night each in 53 adults, the second-generation ring scored 89% on sleep versus wake and 61% across four stages, which become kappa 0.51 and kappa 0.43 once agreement by chance is removed. Oura's August 2026 post lists further studies it says found 76.3 and 76.4 percent four-stage agreement and around 92 percent for sleep versus wake. I have not opened those studies, so I report them as Oura's summary.
| Study | Ring and sample | Sleep vs wake | Four stages |
|---|---|---|---|
| de Zambotti 2019, independent | Earlier ring, 41 adolescents and young adults | 96% sleep sensitivity, 48% wake specificity | 65% light, 51% deep, 61% REM |
| Altini and Kinnunen 2021, Oura-affiliated | 440 nights, 106 people | 96% | 79% |
| Miller 2022, independent, six devices | Generation 2, 53 adults, one night | 89% (kappa 0.51) | 61% (kappa 0.43) |
| Studies listed in Oura's August 2026 post, not opened here | Various | About 92% | 76.3% and 76.4% |
Put together, recent studies put four-stage agreement between three-fifths and four-fifths, the older independent one put single stages nearer half, and sleep-versus-wake sits in the high eighties and nineties for a reason the next section explains. The failure that matters is the still, awake one. Lying in the dark with your eyes open looks a great deal like sleep to an accelerometer, and the people who most need a sleep tracker are the people who do most of that. What cortisol is doing through those hours is the subject of cortisol and sleep.
What does an accuracy figure need beside it?
Five things, and the marketing rarely supplies any of them. A single percentage on a staging task is close to meaningless on its own.
- How many classes. Sleep versus wake is a different problem from wake, light, deep and REM, and the two get quoted in the same breath. Oura's own paper reports 96 percent for the first and 79 percent for the second, from the same nights.
- The base rate. Someone in bed is asleep for most of the night, so a model that never reports waking scores well on sleep versus wake before it has found anything. That is why sleep researchers report agreement corrected for chance. In the 2022 six-device study, the Oura ring's 89 percent sleep-versus-wake agreement becomes a kappa of 0.51 once chance is removed, and its 61 percent four-stage agreement becomes 0.43.
- What it was compared against, and how. Simultaneous polysomnography on the same nights, scored epoch by epoch? Or agreement on totals, where a night can match in aggregate while nearly every boundary inside it sits in the wrong place? Menghini and colleagues published a step-by-step framework for exactly this, with discrepancy analysis, Bland-Altman plots and epoch-by-epoch comparison, and open-source code to run it.
- Who was in the study. Validation in healthy young adults says little about the people who buy a tracker to solve a problem. The 2019 sample was adolescents and young adults. The people with the most at stake are usually the ones a sample excluded.
- Which firmware. These models update over the air. Oura shipped a new staging algorithm in November 2022. The model somebody validated is often not the model running on the ring today, and the industry has no convention for tying an accuracy claim to the code that earned it.
If a company will not put those five beside its figure, the figure is a headline rather than a measurement.
Why can't a recovery or readiness score have an accuracy figure at all?
Because there is nothing to compare it against. Sleep staging at least has a reference standard to be wrong against, which is why it is possible to quote a figure and argue about it. Recovery, readiness and stress scores do not. There is no clinical gold standard for recovery, and every stress feature on the market is scored against your own baseline for the same reason. It is a proprietary composite of pulse timing, temperature, movement and last night's sleep with undisclosed weightings, so it cannot be validated, only asserted. An accuracy claim for it is undefined rather than unproven, and the two kinds of claim fail in completely different ways. That seemed worth saying about a story on accuracy, because the lawsuit is about the claim that can be tested, and most of what a ring shows you in the morning is the kind that cannot.
The inputs also move for reasons unconnected to the state they claim to describe. Breathe slowly for a few minutes and heart-rate variability rises, which is ordinary respiratory physiology and not recovery. Posture moves the heart. Temperature moves it, and a fever accelerates it. Chronic drinking lowers variability, and it comes back when the drinking stops. Variability also falls as health declines. Add which hours of the night happened to get sampled, and a metric that responsive to conditions is a thin foundation for a precise-sounding number.
None of that makes the score useless. A trend against your own baseline is a reasonable thing to look at, and I look at mine. It makes the score unfalsifiable, which is a different property, and one that should keep the word accuracy out of the sentence. What the score is good for, and the ordinary things that move it overnight, is the subject of why a recovery score reads low.
Where I'm standing
I am building a sweat cortisol sensor, so I am not a neutral party here. It is not finished, and it is harder than it looks from the outside. What the work has mostly taught me is restraint in the copy, because I now know how much sits between a raw signal and a number a person can act on.
One example. The Auromone app's test suite contains a test that fails if any of the copy strings it covers contains "normal range", "abnormal", "consistent with" or "risk of". It is a failing test rather than a style guide or a review step, and it only sees the strings it is pointed at, so it is a floor rather than a guarantee. Review catches that sort of thing unevenly and late. A failing test catches what it covers every time, including on the Friday afternoon when the copy was written in a hurry.
The device itself, the Auromone Curve, measures cortisol continuously from a trace of sweat at the wrist, so the thing it reports is the hormone rather than a state inferred from pulse timing. It does not stage sleep or compute a readiness score, and it will not tell you whether a reading is normal, because that is a clinician's question. The first 500 units ship in Q4 2026. If you want the longer version of why a heartbeat-derived score and a hormone are different measurements, that is the subject of cortisol vs HRV, and which watches and rings measure cortisol answers the device-by-device question. The same distinction between a prediction and a measurement runs through the whole category, which is what hormone tracking devices sorts out.
What a dispute like this does to trust
What it costs the category is the reader's willingness to sort. Once wearables earn a reputation for overclaiming, a carefully hedged number and a confident guess get discounted at the same rate, and the discount lands hardest on whoever was disciplined about language. That is the part of this case the rest of the industry should be watching, whatever the court decides.
The mistake runs the other way too. Baron and colleagues coined the term orthosomnia for patients who arrived in clinic seeking treatment for sleep problems they had diagnosed from tracker data, and observed that the tracker often felt more true to them than polysomnography did. Overclaiming and over-trusting come from the same misreading, which is taking a model's output for a measurement. The fix is the same in both directions: say what the sensor touches, and tie every accuracy claim to the firmware that earned it and the confusion matrix behind it. I would want that from any wearable, and I am trying to hold mine to it. What we publish about our own sensor is on the science page.
This guide is for general wellness education only. The Auromone Curve is a general wellness device, not a diagnostic, and does not replace medical advice or clinical testing. Oura is a trademark of its owner. The description of the lawsuit reflects the public complaint and the parties' published statements as of September 2026, and allegations in a complaint are not findings. If you have symptoms that concern you, talk to a healthcare provider.
References
- Surber v. Oura Inc. et al., No. 3:26-cv-08686 (N.D. Cal., filed 20 August 2026). Class action complaint, PDF hosted by plaintiff's counsel. (The marketing phrases quoted: "built for accuracy", "Unparalleled Accuracy", "what only a hospital sleep lab can see", "79%", "95% Sleep Staging Accuracy compared to clinical sleep lab"; "a coin flip's chance"; injunctive relief and restitution.)
- Clarkson Law Firm. Class action lawsuit filed against Oura for false advertising of its smart rings' sleep tracking capabilities. 21 August 2026. (Case caption and court; the 95% figure; what the firm says the sensors do and do not record; the coin-flip characterisation; relief sought.)
- ClassAction.org. Oura lawsuit claims rings cannot measure sleep as advertised. (Filing date of 20 August 2026; the 79-to-95-percent marketing figures; the complaint's "roughly 50 percent" characterisation; what the complaint says the ring measures.)
- TechCrunch. Oura faces lawsuit accusing it of misleading consumers about sleep-tracking accuracy. 21 August 2026. (The marketing phrases quoted in the complaint; Oura's statement in full, including "estimates sleep stages", "associated with distinct, measurable, reproducible changes in physiology" and "not a medical device or a substitute for a clinical sleep study".)
- Oura. Standing Behind Our Science: How Oura Measures Sleep and Validates Accuracy. 23 August 2026. (The company's response post; "infer sleep stages"; "it is not a coin flip; it is physiology"; the studies Oura cites at 76.3%, 76.4% and 91.7 to 91.8%; not intended to replace a clinical sleep assessment.)
- Oura. Oura's New Sleep Staging Algorithm: More Accurate Than Ever Before. 16 November 2022. (79% agreement with polysomnography on four stages; signals used, including "mathematical modeling that incorporates well known sleep patterns"; the algorithm update itself.)
- Altini M, Kinnunen H. The Promise of Sleep: A Multi-Sensor Approach for Accurate Sleep Stage Detection Using the Oura Ring. Sensors. 2021;21(13):4302. (Authors affiliated with Oura Health; 440 nights from 106 people against polysomnography; 94% and 96% for two-stage, 57% and 79% for four-stage, with and without autonomic and circadian features.)
- de Zambotti M, Rosas L, Colrain IM, Baker FC. The Sleep of the Ring: Comparison of the ŌURA Sleep Tracker Against Polysomnography. Behav Sleep Med. 2019;17(2):124-136. (n = 41 adolescents and young adults; 96% sleep sensitivity, 48% wake specificity; stage agreement 65% light, 51% deep, 61% REM; deep sleep underestimated by about 20 minutes and REM overestimated by about 17.)
- Miller DJ, Sargent C, Roach GD. A Validation of Six Wearable Devices for Estimating Sleep, Heart Rate and Heart Rate Variability in Healthy Adults. Sensors. 2022;22(16):6317. (53 adults, one night each; Oura Ring Generation 2 among six devices against polysomnography; Oura two-state 89%, kappa 0.51; four-stage 61%, kappa 0.43; two-state agreement 86 to 89% across devices.)
- Menghini L, Cellini N, Goldstone A, Baker FC, de Zambotti M. A standardized framework for testing the performance of sleep-tracking technology: step-by-step guidelines and open-source code. Sleep. 2021;44(2):zsaa170. (Discrepancy analysis, Bland-Altman plots and epoch-by-epoch analysis against polysomnography, with open-source R functions.)
- Patel AK, Reddy V, Shumway KR, Araujo JF. Physiology, Sleep Stages. StatPearls. Updated January 2024. (Polysomnography records EEG, electrooculogram, electromyogram, ECG, oximetry, airflow and respiratory effort; the EEG signatures of wake, N1, N2, N3 and REM.)
- American Academy of Sleep Medicine. The AASM Manual for the Scoring of Sleep and Associated Events, Version 3, 2023. (The scoring rulebook a sleep lab stages against.)
- Cao R, Azimi I, Sarhaddi F, et al. Accuracy Assessment of Oura Ring Nocturnal Heart Rate and Heart Rate Variability in Comparison With Electrocardiography. JMIR Mhealth Uhealth. 2022;10(1):e27487. (n = 35; heart rate r = 0.99, RMSSD r = 0.92.)
- Shaffer F, Ginsberg JP. An Overview of Heart Rate Variability Metrics and Norms. Front Public Health. 2017;5:258. (Respiratory sinus arrhythmia; slow breathing increases it, with paced breathing at 6 breaths a minute producing a peak at 0.1 Hz; time-domain HRV declines with decreased health.)
- Fatisson J, Oswald V, Lalonde F. Influence diagram of physiological and environmental factors affecting heart rate variability: an extended literature overview. Heart Int. 2016;11(1):e32-e40. (Heart rate rises from lying to sitting to standing; fever accelerates the heart; chronic alcohol use lowers HRV, reversibly.)
- Baron KG, Abbott S, Jao N, Manalo N, Mullen R. Orthosomnia: Are Some Patients Taking the Quantified Self Too Far? J Clin Sleep Med. 2017;13(2):351-354.