Home/Measuring cortisol/The Oura lawsuit

Measurement guide

The Oura lawsuit, and what a sleep tracker actually measures.

On 20 August 2026 a proposed class action accused a smart-ring maker of overstating how accurately it stages sleep. I build a wearable sensor, so I am not a neutral reader of that filing. What follows is what I think it gets right, where it reaches, and the question I would put to any wearable, including the one on my own bench.

The short answer

Surber v. Oura Inc., a proposed class action filed on 20 August 2026, alleges that Oura's 79% and 95% sleep-staging accuracy claims overstate what a ring can measure. Oura disputes this, and nothing has been decided. A ring senses pulse timing, movement, temperature and breathing. Sleep stages are defined by brain waves, eye movement and muscle tone, so a stage label from a ring is a prediction of what a sleep technician would have written down. Oura's own study reports 79% four-stage agreement with a sleep lab, and an independent study found 61%. Neither figure means much without the class count, the base rate, the comparison method, the sample and the firmware beside it. Recovery and readiness scores cannot carry a figure at all, because there is nothing to compare them against.

What does the lawsuit actually allege?

The complaint's one-line version is that the marketing outran the sensors. Surber v. Oura Inc. was filed on 20 August 2026 in the U.S. District Court for the Northern District of California by Clarkson Law Firm on behalf of a California buyer, and it asks to represent a class. It quotes Oura's marketing as promising rings "built for accuracy", with "unparalleled accuracy", able to see "what only a hospital sleep lab can see", and it cites accuracy figures it says Oura has published over the years: 79 percent, and more recently "95% Sleep Staging Accuracy compared to clinical sleep lab". The complaint's core argument is a hardware one. Sleep stages are defined by brain activity, eye movement and muscle tone, the ring records none of those, and the firm says the real figure for staging is nearer to 50 percent, which the filing calls a coin flip. It asks the court to stop the marketing and to refund purchasers.

Those are allegations. Nothing has been decided and the class has not been certified. I am not a lawyer, so nothing here is a view on the filing's legal merits. What I can read is the science both sides are pointing at, because it is published.

Oura's answer is the more interesting document. In its statement to the press the company said it stands behind its science and its accuracy claims, that the ring "estimates sleep stages using multiple physiological signals, including heart rate, heart rate variability, movement, breathing patterns, and temperature", and that there is peer-reviewed evidence that sleep stages "are associated with distinct, measurable, reproducible changes in physiology". It added that the ring "is not a medical device or a substitute for a clinical sleep study", and that its staging "has been validated and compared favorably in multiple studies against polysomnography, the gold standard". Two days later its blog called the suit baseless and put the position more plainly: the autonomic signal is one of the main signals wearables use to infer sleep stages, and "it is not a coin flip; it is physiology", read through a different window.

Read the defence carefully. It rests on the validation studies, which the next two sections go through, and on one phrase. Associated with is true, and it is a different claim from measured. The company's own verbs are estimate and infer. Everything below is about the distance between those verbs and the word accuracy.

What does a ring actually sense?

Pulse timing, movement, skin temperature, breathing, and on some models a blood-oxygen trend. That is Oura's own list, and it is the whole of the raw material. The instrument reads those signals well. In Cao and colleagues' 2022 comparison against a reference ECG in 35 sleepers, an Oura ring tracked overnight heart rate at r = 0.99 and the standard beat-to-beat variability measure at r = 0.92. I have no quarrel with the sensor. A wrist band starts from the same short list, which is why sorting any wearable's spec sheet into sensed and computed is the most useful thing a buyer can do.

Sleep stages are defined in different terms. A sleep lab records brain electrical activity from scalp electrodes, eye movement from electrodes beside the eyes, and muscle tone from the chin, and a technician scores the night epoch by epoch against the American Academy of Sleep Medicine's manual. Wake shows fast, low-amplitude brain activity. Light sleep is scored on sleep spindles and K-complexes, deep sleep on slow delta waves. REM looks like waking on the brain trace, with the eyes moving and the muscles switched off. None of that is available at a finger.

So the ring is measuring something else and handing it to a model. The model was trained on nights where a ring and a sleep lab ran at the same time, and it learned which patterns of pulse, movement and temperature tend to go with each label a scorer wrote. Oura's 2022 post says the current algorithm also folds in "mathematical modeling that incorporates well known sleep patterns", which means it carries a prior about what a night usually looks like. A model with a good prior will look right on an ordinary night. The question is what it does on the night that is not ordinary, which is the night the person bought the ring for.

That is what a stage label from a ring is: a prediction of what a technician would have written, made from signals the technician does not use. It can be a good prediction. It is still a prediction, and an accuracy figure for it has to be read that way.

Is the Oura Ring's sleep staging accurate?

Neither a coin flip nor a sleep lab. It separates sleep from wake well when you are moving, and badly when you are lying still. The independent study the complaint appears to be leaning on is de Zambotti and colleagues' 2019 comparison of an earlier Oura ring against polysomnography in 41 adolescents and young adults. The ring identified sleep with 96% sensitivity and identified wake with 48% specificity, so about half of the awake epochs were scored as sleep. Stage agreement was 65% for light sleep, 51% for deep and 61% for REM, and the ring underestimated deep sleep by about 20 minutes a night and overestimated REM by about 17.

Those numbers describe an older ring and an older algorithm, and Oura has updated both. Its own validation, published in 2021 by two authors affiliated with Oura Health, used 440 nights from 106 people and reports 96% accuracy for sleep versus wake and 79% across four stages, up from 57% when the model used movement alone. In a 2022 independent comparison of six wearables against polysomnography, one night each in 53 adults, the second-generation ring scored 89% on sleep versus wake and 61% across four stages, which become kappa 0.51 and kappa 0.43 once agreement by chance is removed. Oura's August 2026 post lists further studies it says found 76.3 and 76.4 percent four-stage agreement and around 92 percent for sleep versus wake. I have not opened those studies, so I report them as Oura's summary.

Study Ring and sample Sleep vs wake Four stages
de Zambotti 2019, independentEarlier ring, 41 adolescents and young adults96% sleep sensitivity, 48% wake specificity65% light, 51% deep, 61% REM
Altini and Kinnunen 2021, Oura-affiliated440 nights, 106 people96%79%
Miller 2022, independent, six devicesGeneration 2, 53 adults, one night89% (kappa 0.51)61% (kappa 0.43)
Studies listed in Oura's August 2026 post, not opened hereVariousAbout 92%76.3% and 76.4%

Put together, recent studies put four-stage agreement between three-fifths and four-fifths, the older independent one put single stages nearer half, and sleep-versus-wake sits in the high eighties and nineties for a reason the next section explains. The failure that matters is the still, awake one. Lying in the dark with your eyes open looks a great deal like sleep to an accelerometer, and the people who most need a sleep tracker are the people who do most of that. What cortisol is doing through those hours is the subject of cortisol and sleep.

What does an accuracy figure need beside it?

Five things, and the marketing rarely supplies any of them. A single percentage on a staging task is close to meaningless on its own.

If a company will not put those five beside its figure, the figure is a headline rather than a measurement.

Why can't a recovery or readiness score have an accuracy figure at all?

Because there is nothing to compare it against. Sleep staging at least has a reference standard to be wrong against, which is why it is possible to quote a figure and argue about it. Recovery, readiness and stress scores do not. There is no clinical gold standard for recovery, and every stress feature on the market is scored against your own baseline for the same reason. It is a proprietary composite of pulse timing, temperature, movement and last night's sleep with undisclosed weightings, so it cannot be validated, only asserted. An accuracy claim for it is undefined rather than unproven, and the two kinds of claim fail in completely different ways. That seemed worth saying about a story on accuracy, because the lawsuit is about the claim that can be tested, and most of what a ring shows you in the morning is the kind that cannot.

The inputs also move for reasons unconnected to the state they claim to describe. Breathe slowly for a few minutes and heart-rate variability rises, which is ordinary respiratory physiology and not recovery. Posture moves the heart. Temperature moves it, and a fever accelerates it. Chronic drinking lowers variability, and it comes back when the drinking stops. Variability also falls as health declines. Add which hours of the night happened to get sampled, and a metric that responsive to conditions is a thin foundation for a precise-sounding number.

None of that makes the score useless. A trend against your own baseline is a reasonable thing to look at, and I look at mine. It makes the score unfalsifiable, which is a different property, and one that should keep the word accuracy out of the sentence. What the score is good for, and the ordinary things that move it overnight, is the subject of why a recovery score reads low.

Where I'm standing

I am building a sweat cortisol sensor, so I am not a neutral party here. It is not finished, and it is harder than it looks from the outside. What the work has mostly taught me is restraint in the copy, because I now know how much sits between a raw signal and a number a person can act on.

One example. The Auromone app's test suite contains a test that fails if any of the copy strings it covers contains "normal range", "abnormal", "consistent with" or "risk of". It is a failing test rather than a style guide or a review step, and it only sees the strings it is pointed at, so it is a floor rather than a guarantee. Review catches that sort of thing unevenly and late. A failing test catches what it covers every time, including on the Friday afternoon when the copy was written in a hurry.

The device itself, the Auromone Curve, measures cortisol continuously from a trace of sweat at the wrist, so the thing it reports is the hormone rather than a state inferred from pulse timing. It does not stage sleep or compute a readiness score, and it will not tell you whether a reading is normal, because that is a clinician's question. The first 500 units ship in Q4 2026. If you want the longer version of why a heartbeat-derived score and a hormone are different measurements, that is the subject of cortisol vs HRV, and which watches and rings measure cortisol answers the device-by-device question. The same distinction between a prediction and a measurement runs through the whole category, which is what hormone tracking devices sorts out.

What a dispute like this does to trust

What it costs the category is the reader's willingness to sort. Once wearables earn a reputation for overclaiming, a carefully hedged number and a confident guess get discounted at the same rate, and the discount lands hardest on whoever was disciplined about language. That is the part of this case the rest of the industry should be watching, whatever the court decides.

The mistake runs the other way too. Baron and colleagues coined the term orthosomnia for patients who arrived in clinic seeking treatment for sleep problems they had diagnosed from tracker data, and observed that the tracker often felt more true to them than polysomnography did. Overclaiming and over-trusting come from the same misreading, which is taking a model's output for a measurement. The fix is the same in both directions: say what the sensor touches, and tie every accuracy claim to the firmware that earned it and the confusion matrix behind it. I would want that from any wearable, and I am trying to hold mine to it. What we publish about our own sensor is on the science page.

This guide is for general wellness education only. The Auromone Curve is a general wellness device, not a diagnostic, and does not replace medical advice or clinical testing. Oura is a trademark of its owner. The description of the lawsuit reflects the public complaint and the parties' published statements as of September 2026, and allegations in a complaint are not findings. If you have symptoms that concern you, talk to a healthcare provider.

References

Keep reading

More measurement guides

Straight answers

Oura lawsuit FAQ

What does the Oura lawsuit claim?

Surber v. Oura Inc., filed on 20 August 2026 in the U.S. District Court for the Northern District of California, is a proposed class action alleging that Oura marketed its rings as accurate at tracking sleep stages, with figures of 79 percent and later 95 percent against a clinical sleep lab, when the ring lacks the sensors that define those stages. The complaint calls the staging no more reliable than a coin flip and asks the court to stop the marketing and refund purchasers. Oura disputes the allegations, says its staging has been validated against polysomnography in multiple studies, and says it will defend its work in the appropriate legal forums. Nothing has been decided.

Is Oura Ring sleep tracking accurate?

It depends on which question you ask. Telling sleep from wake, the ring agreed with a sleep lab 96 percent of the time in Oura's own 2021 study, partly because most of any night is sleep. Across four stages, that study reports 79 percent agreement, an independent 2022 comparison found 61 percent four-stage agreement, a chance-corrected kappa of 0.43, for the second-generation ring, and a 2019 study of an earlier ring found 51 to 65 percent agreement on individual stages and scored about half of awake time as sleep. Oura itself says the ring is not a substitute for a clinical sleep study.

Can any ring or wristband measure sleep stages?

No. Sleep stages are defined by brain electrical activity, eye movement and muscle tone, recorded in a sleep lab with electrodes on the scalp, beside the eyes and on the chin. A ring or wristband records pulse timing, movement, skin temperature and breathing, and a model predicts which stage a technician would have scored from those. Some devices predict well. The output is still an estimate, which is the word Oura uses for it.

What is a readiness or recovery score based on?

A proprietary combination of resting heart rate, heart-rate variability, skin temperature, movement and last night's sleep estimate, compared against your own recent baseline, with weightings the manufacturer does not publish. There is no clinical gold standard for recovery, so the score cannot be validated the way sleep staging can, only asserted. It can still be a useful trend against your own history. An accuracy figure for it has no meaning.

Does the Auromone Curve track sleep stages?

No. The Auromone Curve measures cortisol continuously from a trace of sweat at the wrist. It does not stage sleep or compute a readiness or recovery score, and it does not tell you whether a reading is normal or abnormal. It reports the hormone and shows you the shape of your day. It is a general wellness device, not a diagnostic, and the first units ship in Q4 2026.

A prediction is not a measurement.

The argument of this page is about sensors, not about you. The Auromone Curve is designed to read cortisol itself from a trace of sweat on your wrist, continuously, and to show you the curve as it moves. It does not stage sleep and will not tell you whether anything is wrong. The first 500 units ship in Q4 2026, and joining the waitlist is free.

Join the waitlist