One morning, a new alert showed up in our charts. Just a colored banner. Nobody told us what data it learned from, who it had been tested on, or what we were supposed to do when it went off at 3 a.m. on a patient who looked fine. We were expected to act on it anyway.
We were trained to pick apart a study and question a drug rep. Then software showed up and we stopped asking those questions. We should keep asking them. Checking the tools we use on patients is part of taking care of patients.
The word “innovation” covers three very different things. The proof behind each one is different too. Drugs have the highest bar. It is still lower than most of us think. Some cancer drugs reach patients early through accelerated approval, which means the drug got cleared based on a stand-in measure, like a tumor shrinking, instead of proof that people live longer. Researchers looked at 46 of them with more than 5 years of follow-up. In 43 percent, the follow-up trials showed patients actually lived longer or felt better. But 63 percent were upgraded to full approval. More drugs got upgraded than got proven.
Next are devices and software the government reviews. The U.S. Food and Drug Administration (FDA) keeps a public list of artificial intelligence (AI) tools it has authorized. As of its June 2026 update, that list held 1,524 of them, and about 3 out of 4 are radiology tools. Most got in through a path called 510(k) clearance.
Real physician voices, twice a week
Free, and one click to unsubscribe.
Clearance is not approval. It means the company showed its product is close enough to something already being sold. FDA says it asks for clinical data on fewer than 10 percent of these submissions. So clearance answers one question. Is this a lot like the product before it? It does not answer the question we care about. Do my patients do better?
A 2025 study in JAMA Network Open went through all 903 AI devices on FDA’s list through August 2024. Just over half, 55.9 percent, had any clinical performance study at all. Another 24.1 percent said plainly that no study was done. Of the studies that did exist, 8.1 percent followed patients forward in time instead of digging through old charts. Only 2.4 percent randomized anyone.
The third group is the one that surprises people. Much of it never goes to FDA at all. A law called the 21st Century Cures Act pulled certain clinical decision support software out of the device rules. Scheduling, billing, note writing, and most electronic health record (EHR) features sit outside those rules too. If your hospital builds a risk score and switches it on, usually nobody outside your hospital ever checks it.
That is not rare. Using 2023 national survey data, a Health Affairs study found 65 percent of U.S. hospitals were using prediction models. Of those, 79 percent got the models from their EHR company. Only 61 percent tested a model on their own patients to see if it worked there. Only 44 percent checked it for bias.
The clearest example is a sepsis warning tool that ran at hundreds of hospitals. Researchers in Michigan tested it across 38,455 hospital stays. The company had reported it could separate the patients who would get sepsis from the ones who would not about 76 to 83 percent of the time. The measured number was 63 percent. It missed 67 percent of sepsis cases. It fired on 18 percent of everyone admitted. Doctors had to work up 8 patients to find 1 who would actually develop sepsis. People had been acting on it for years.
Now the fair part. The company rebuilt the model. A study published this February tested the new version at 4 health systems, and it performed much better. It still flags roughly 4 to 8 people for every real case. Tools can improve. They improve when somebody measures them on local patients and says the result out loud.
Ambient AI scribes are the newest version of the same gap. They listen to the visit and write your note. A randomized trial in NEJM AI tested 2 of them against normal work across 238 clinic physicians. One cut time spent in the note by 9.5 percent. The other showed no real change at all. Both are sold on saving you time. Neither needed FDA review.
So what do you do when a vendor makes a claim? Read it the way you read a paper. What was the product compared against, a real alternative or nothing? Who paid for the study and who wrote it? Did they look backward at old records or follow patients forward, and did anyone get randomized? Was the result something a patient would feel, or a stand-in like alert accuracy? Who were the patients in the study, and do they look like yours? And the question people skip: What happens after it goes live, when your patients change and nobody has rechecked the numbers?
You also have a right most physicians do not know about. Under a federal health technology rule that took effect at the end of 2024, certified EHRs must show 31 specific facts about any prediction tool built into them. That covers what data it learned from, whether outsiders tested it, how well it performed, and how often it gets updated. You can ask your informatics team for that list by name. One catch. A rule proposed in December 2025 would remove that requirement, and it has not gone through. The right is there right now.
My own version of this lesson came from the investing side. A colleague sent me a company that tracks hospital supplies and recalls, and I understood the problem in about a minute because I had lived it. Expired medications. Crash carts. The small failures that never show up in a sales deck. Before I put money behind it, I ran it past seventeen frontline people, including nursing managers and supply chain leads, in under a month. Their read was the whole decision. That habit came off the wards, not out of a finance course.
None of this turns you into a regulator. It asks you to treat a sales claim the way you would treat a weak study. The people selling to your hospital do this for a living, and the committee signing the contract usually has nobody in the room who will use the thing at 3 a.m.
You will. When the tool is wrong, it is wrong in front of you, with your patient in the bed. That is why this is worth learning.
Harsha Moole is an internal medicine-trained physician-scientist with more than 100 peer-reviewed publications, including work featured in the New England Journal of Medicine. After years of clinical practice and gastroenterology outcomes research, he made an unconventional transition from the bedside to the boardroom by founding PhysicianEstate, a health care-focused venture capital firm.
Over the past seven years, Dr. Moole has made 22 early-stage health care investments across digital health, medical devices, biotech, and therapeutics. He has also built a network of more than 200 physicians from institutions such as Johns Hopkins and Stanford who help source opportunities and provide clinical diligence before capital is deployed. His core thesis is that physician-scientists with firsthand clinical experience are uniquely positioned to identify health care investments that generalist investors often miss.
His research background is reflected in his publication record on Google Scholar, and he shares professional updates on LinkedIn.


