Health AI Ranked: Reality vs Marketing (Apple Watch, Oura, Whoop, Hims/Hers...)
an AI health app, wearable, or any sort of smart device where you rely on it for health advice, I want you to know how to evaluate it before you actually trust it. I'm going to show you how to do that on a ranked tier list, and the difference between the top and the bottom is not small. And all these products are on the market.
The apps at the top have FDA clearance, so they've been under intense scrutiny, while some of the bottom just have marketing language. Most of the apps that you'll use will be somewhere in the middle. My name is Nass, I build AI software in health for a living, and I've spent the last several weeks evaluating different AI health products.
To actually evaluate or validate it, I ask three questions. First, was the product tested independently, so not by the company itself? Second, how closely related are the claims of the company to what was tested? And third, was the test a prospective one or an observational one? I'll show you what that means in a second.
If a tool scores high on all three, then it's S tier. If it scores low or scores zero on all three, then it's near F tier. So, to show you what I mean, let me give you a quick example of something S tier. So, the only tools that score high on all three dimensions are prescription digital therapeutics.
Think like Rejoyn for depression, Sleepio RX for insomnia, and Daylight RX for anxiety. All of these three have been cleared by the FDA or their intended outcome and prospectively and independently evaluated. On the opposite end of the spectrum, you've got F tier, and examples of this are commercial AI fitness coaches as a category.
So, for example, a 2025 systematic review by Montalegre and colleagues in Artificial Intelligence Review, they went looking for peer-reviewed interventional studies of AI virtual assistants for physical therapy. Across the entire published literature, they only found Uh and here's a part that matters the most. These eight studies are academic research and they're not the products that have been released in the wild.
It's just researchers building and testing ideas in controlled settings, not the wild. None of them test the commercially available fitness coaching apps that consumers actually are buying now. So, notice the difference between those two? On one hand, you got STR and that's prospectively evaluated. So, they made a prediction and then they tested it and they randomized who was going to go in what group and it worked out.
And then, they actually measured the outcome in STR. So, for example, in Sleep IO, they were looking at sleep. That was the outcome. And finally, it was independently tested. So, it's not the company doing all the research, it's actually third parties doing the research. And in most companies, you'll still find the companies a little bit involved because it's hard to incentivize researchers to do research on something um that's commercial and still be independent.
So, before we go into the products, I want to be clear about what I mean by AI. So, it's not like a smart watch heart rate converting to calories burnt. It's more like what takes what black box takes the raw data and converts it into an output, right? It detects whether you have a fibrillation. It detects whether it's a contraceptive green day or red day and that kind of thing.
Another thing that I want to mention is that the FDA does have its own standard. They look at safety and they look at efficacy or how effective it is. My framework is pretty similar. I just also added the independent testing review. In some cases, you know, I'll give it a points if it was reviewed by the FDA because I consider that to be independent, especially for a specific type of classification, which I'll get into later in the video.
So, let's start with something eight-tier. You've probably heard of Apple, it's a rapidly growing startup. And what I want to evaluate is the Apple watches. Not all the features, but specifically the AFib detection. Right, so Apple has the calorie burn tracker, sleep stages, activity rings, and stress notifications.
I want to focus on the AFib detection. And that one feature is what gets eight-tier. Let's go into why. So, for AFib, Apple submitted a 600, approximately 600, participants study to the FDA. FDA granted de novo authorization in 2018. De novo authorization is sort of like the best authorization that you can get from the FDA, in my opinion, because you're sort of breaking new ground.
And so, they're going to give it a lot of scrutiny. They don't usually do that for like very critical stuff, they'll do it for something like an Apple Watch. So, Apple got 98. 3 sensitivity for AFib detection and 99. 6 specificity for sinus rhythms. Sensitivity is how well it can detect something that exists, and specificity is that if it detected something, what are the chances that it's the right thing.
After that study, Stanford grant independent study, and they evaluate 419,000 Apple Watch users they enrolled in a study. And that's the Apple Heart Study in the New England Journal of Medicine. So, this was a prospective study. They made predictions about the users, and then they watched what happened.
It wasn't a randomized controlled trial where they told users what to do, but it also wasn't observational. So, it's good, but it's not S-tier. That said, it is independent and pretty large, so it does get an A tier. And again, that is the AFib detection on the Apple Watch. It's an A tier. So, that's Apple Watch's FDA detection.
It's unbiased because FDA reviewed it with a de novo classification, but also you've got the prospective study from Stanford. It's also not S tier because they actually just looked at AFib, not the actual heart outcomes. Now again, I will give Apple credit and not deduct them because they don't claim uh that it reduces stroke outcomes or health stroke outcomes.
They strictly focus on AFib detection. So, their claim is backed by their study. The next A tier thing we have earns its place in a very different way. So, it's called Natural Cycles. It's an app that tracks your basal body temperature and then tells you which days you can and cannot get pregnant. This is the first software-only contraceptive cleared by the FDA in the United States.
2018, also de novo authorization number DEN170052. So, the FDA had to review pregnancy outcome evidence sufficiently to define an entire new class of software. That is the highest bar on the unbiased axis besides having somebody else do the whole study for free. So, here's what the study did. It looked at a prospective cohort, so again, future-looking, uh 22,000 women.
They found that it was comparable to condoms, less effective than hormonal contraceptives, and the framing is in the FDA clearance. They did not overclaim. The company's published algorithm is a basal temperature formula, so it's not AI in the conventional sense, it's still formula, but it's auditable, and they have the data to back it up.
The app is basically math of the calendar, uh which uh turns out to be more trustable than a black box AI readiness score. So, that was Natural Cycles. It's A-tier. De novo authorization is the unbiased input. Large prospective cohort is the prospective input. And then, the pregnancy outcomes are about as high proximity as contraceptive evidence gets.
Now, when we get into the B-tier, things get a bit more complicated. These are the products that most of you use every day. The first example that we've got is the Oura Ring for sleep detection. Now, the rating is score. So, Vincent et al. in 2024 in Sleep Medicine, there were uh five independent academic authors at the University of Tokyo and other institutions.
So, what they did is a prospective multi-night comparison against polysomnography. And that's the clinical gold standard for sleep staging. They had about 96 participants, and they found that the Oura Ring had a 91. 7 overall sleep accuracy. It was the best in class for consumer wearables. But, here's the disclaimer that I have to give.
This is for the sleep categorization. But, Oura mostly markets its readiness score, for which I couldn't find validation evidence that satisfied me. But, for the sleep claim, I'll give it a B-tier. And Whoop has the same pattern. But, the independent study goes a little deeper. So, for Whoop, we're looking at the recovery score.
In 2022, Moreno et al. published in the Sensors journal. So, there were three academic authors at the Appleton Institute for Behavioral Science at Central Queensland University. The study was mostly funded by the Australian Institute of Sport. The study compares six wearable devices against polysomnography for sleep and heart rate.
Whoop Sleep Onset Latency Agreement scored 86% on the two-stage sleep, and about 50 to 65% on the three-stage sleep. But again, like Oura, this is for the sleep-wake agreement or the sleep cycle detection. But, what they try to sell is the readiness score. So I give Whoop a B tier on the sleep detection because of the three axes again, it was validated on the detection layer.
But I couldn't find validation for the recovery score or the validation that would satisfy me according to the three axes that I mentioned in this video. The next one is in B tier for a different reason. So Lumen has their own study. There is a published peer-reviewed study in the Interactive Journal of Medical Research.
It was a prospective study, so they predicted and then looked at what happened with a sample of about 33 people. They compared the outcome or what they were measuring to metabolic cart readings at rest and after a glucose load or having sugar. And the directional change detection was statistically significant.
So it was able to detect the change in utilization of fuel. There are two things to note here. The authors work for the parent company, MetaFlow, so they had the study themselves, but they did publish it and it was peer-reviewed, so that says a lot and it's hard to find researchers that are, you know, going to do things for free.
It's usually some, you know, incentive to do the work. Another thing is that the study itself says to use the device to track individual change over time, not to replace metabolic cart precision. So the company's marketing says measuring metabolism in real time and personal metabolic coach. That is not the same claim.
It is a much wider claim than what the study found. So Lumen is in B tier. There is published evidence. It's vendor authored. It's for a narrower claim than the marketing. And the paper's own authors tell you what it doesn't do. So that's actually the honest version. Most apps don't even publish the caveats so they deserve to be credited for that.
Now the seat here is where the framework does something interesting. So, the evidence is good. It's just that the product doesn't pass the test that it's being sold for. So, we're talking about non-diabetic continuous glucose monitors or CGMs. You've got Levels, you've got Stelo by Dexcom. So, these are products that let healthy people see their blood sugar in real time.
So, in 2025, Amdenal published a systematic review in the Cureus Journal. It was a review of continuous glucose monitors or CGMs in non-diabetic individuals for cardiovascular prevention. So, there were eight authors across NHS hospitals, Saudi Arabia, Kuwait, UAE, and Sudan. And all of them declared no industry conflicts.
So, fully independent. Now, what the review actually found is that yes, CGM is useful for seeing your glucose patterns. It can help you time exercise to reduce your blood sugar spike after a meal. It can monitor mean physical activity. Those findings are real. But, what the review also found is that evidence of a direct impact on hard cardiovascular endpoints remain limited.
And that's what they were looking at. And the evidence on a direct impact on classic cardiovascular risk factors is still nascent or new. So, the Dexcom sensor hardware is FDA cleared, including the OTC version. Stelo cleared March 2024. The FDA cleared it for one thing, understanding how diet and exercise affects your blood sugar, like that study.
That exact language. Not for metabolic health upgrades that Levels and Stelo actually market. The FDA's own intended use is narrower than their marketing. And like that research paper showed, the evidence of a direct impact on those risk factors for metabolic health is still new. So, for seat here, the evidence is there.
The sensor measures what it tells you it's going to measure, but the outcome that they sell you on is still not there yet. The strongest independent review went looking for that and found that the evidence was limited. So, high-quality research confirming the limitation is our framework working correctly.
If the evidence actually suggested that the direct relationship is strong, then it would have been bumped up from a C tier according to the framework. Now, before we get to the F tier, there is a quick note on something that I keep. So, I keep a running list of AI health creators and AI health tools, and I evaluate them using the same frameworks that I did in this video.
When something new launches, I try to put it on the list as soon as possible. And when there's news that changes that placement, I update that list. So, I'm going to put a link in the description and pinned comment below if that's something you're interested in. Now, let's go back to the F tier. Now, in the open, we talked about AI fitness coaches.
Now, we're also going to add something else. Specifically, the Hims & Hers AI weight loss. And I can be specific here because there is a documented public record. So, Hims & Hers markets an AI-driven personalized weight loss program. Their MedMatch system is described as a proprietary custom-built technology that uses AI and machine learning.
So, that was the marketing claim. What does the research say? So, in 2025, Luo and al. published in the Journal of Medical Internet Research. It was a scoping review of AI medical questionnaire validation. And they screened 49,091 publications. They found that 14 met inclusion criteria, which I'll tell you about in a second, but that's 0.
03% of the ones they screened. Now, of those 14, only 21% had entered the clinical validation phase. The remaining 79% were still exploratory. Now, Hims & Hers own MedMatch press release describes the initial beta deployment as being for anxiety and depression, not weight loss. The AI claim being marketed for weight loss is ahead of where the company's own deployment is.
But, here's the thing that matters. The drug works. Semaglutide and tirzepatide have FDA approval and phase three clinical trial evidence behind them. The Those So, those drugs are real. The AI claim being scored is whether the AI questionnaire layer is adding validated clinical value on top of that human clinician prescription.
That's the thing with no evidence. This is an F-tier because there is no independent perspective outcome evidence for the AI claim. So, the drug is real. It's just that the AI layer on top is not validated. And that brings us back to the S-tier where we appreciate why the prescription only designation is the one that brings it up to the S-tier.
So, the only tools that score high in all three dimensions for an outcome claim are the ones that require that FDA prescription or clearance. So, we've got Rejoice for depression, Sleepio RX for insomnia, and Daylight RX for anxiety. So, let's start with Rejoice. K231209 The FDA cleared it in April 2024.
The pivotal trial was a phase three double-blind sham-controlled randomized controlled study. There were 386 participants, 13 weeks, and they measured actual change in depression symptom scores. That's unbiased, perspective, and proximity to the outcome. So, it hits all three on the framework. Now, for Sleepio RX, that was FDA cleared in August 8, 2024.
Prefrontal 2025 and then JMR Mental Health journal. There were 336 participants. There was a two-arm randomized controlled trial against active control. The primary outcomes were insomnia severity, time to fall asleep, time awake at night, and they measured them up to 24 weeks post randomization, so after they assigned them to a group.
Again, it's good on all three. As for Daylight RX K233872, FDA cleared in September 4, 2024. Forces in all in 2025 in the JAMA Network Open published with 251 participants. It was a single-blind parallel group randomized control trial, and the primary endpoint was that at 10 weeks they would look at remission via blinded clinician, so the clinician did not know rating plus self-reported anxiety.
So, all of these STR ones are prescription only. Like, you cannot download them without a clinician authorizing access. So, that prescription requirement is not a limitation. That's telling you what the real bar for evidence actually looks like. So, the FDA requires clinical proof before allowing the companies to make those claims.
That standard is what most consumer apps do not meet. If your AI health app is not prescription only, that doesn't mean it's worthless. It just means it hasn't cleared that bar that prescription ones meet in terms of evidence. By the way, if you like this video and you like how I analyze these products, uh you might also like another video that I did where I ranked AI doctors by how believable they are and therefore how dangerous they are to people in general.
I'm just going to put a link for that right here.