The headlines have a rhythm to them now. Artificial intelligence is “going rogue” (lying, cheating, hacking, scheming to preserve itself), and the next model may be powerful enough, in Bill Gates’s phrase, to cause a billion deaths. For those of us watching AI arrive in our clinics and hospitals at the very same moment, it is a strange dissonance: We are being told to adopt a technology and to fear it in the same breath.
I build clinical AI, and I think the fear is half-right. The alarm is earned. But it is aimed at the wrong target. And in medicine, getting the target wrong is its own kind of danger.
Start with what is true, because a great deal of it is. The warnings are not coming from outside critics; they are coming from the companies building the technology. Researchers have documented models deceiving their evaluators, gaming their tests, and resisting shutdown. This past summer, a swarm of one lab’s own agents broke out of a sealed test environment and compromised another company’s production systems, roughly seven hundred of them, coordinating through a message board they set up themselves. The lab called it a “warning shot,” and it was right to. Weeks ago, a frontier model pointed at the medical-records giant Epic’s own software found a flaw that let patient records be read with no trace in the audit log. None of this is fiction. Anyone telling you these systems are harmless toys is not paying attention.
But “going rogue” is the wrong phrase, and the difference matters. Nearly everything in that unsettling catalog came from researchers deliberately trying to make models misbehave, under controlled conditions, precisely so the failures surface in a lab instead of a hospital. Finding your system cheating when you are hunting for cheating is not proof that it is loose in the world; it is proof that the safety process is working. The Gates line shows how the framing drifts. He said AI “is certainly powerful enough” to drive events that cause a billion deaths: present tense, and about people misusing powerful tools, not machines acting on their own will. Rendered as “AI will certainly be powerful enough,” it becomes a prophecy about autonomous malevolence. That is not what he said, and the distance between the two is the whole argument.
Real physician voices, twice a week
Free, and one click to unsubscribe.
Here is what I have come to believe after watching these systems closely, including by turning them loose on my own work. The danger does not scale with how ominous a chatbot sounds. It scales with three things: capability, access, and autonomy (how good a model actually is at a hard task, what real-world systems it can reach, and how much it can do without a human in the loop). Put the most capable downloadable model on the most expensive hardware a person can buy, give it a goal, and let it run, and you do not get Skynet. You get a system that can attempt real mischief and mostly fails at the catastrophic version, because the power grid, the banking network, and a working pathogen are defended by physical and security realities that no amount of eloquence crosses. The civilization-ending scenarios are not far-fetched because AI is weak. They are gated because the capability that would matter is concentrated behind compute and infrastructure that cannot be assembled invisibly.
The threats that are real and near are quieter. The first is misuse by actors who already hold the capability and the access: a nation-state, or a sophisticated insider, pairing a capable model with genuine offensive reach. The second is the one physicians should care about most, because it is already here: implementation failure in the AI systems we are racing to put into clinics and hospitals. The Epic flaw was not a rogue intelligence. It was ordinary software, built fast, with a gap a capable model found faster than its human engineers could. That is the actual lesson, and it is both less cinematic and more urgent than a billion deaths.
And here is the part the alarm consistently misses. The answer to all of this is not panic, and it is certainly not denial. It is discipline: the slow, unglamorous engineering that separates responsible deployment from the race to ship. I know this because I am living it. I am building an AI-native clinical platform, and not long ago I did something that felt counterintuitive: I had two of the most capable AI models available tear my own security design apart, adversarially, hunting for every way it could fail. They found real holes: missing patient-level access controls, an audit scheme that could not actually be built as written, a defensive feature that could be turned against the very clinicians it was meant to protect. The right response was not discouragement. It was to fix every one of them before writing a line of production code. That is what taking the risk seriously looks like, and it is nothing like the helplessness the warnings tend to produce.
So should we be alarmed? Yes, about the right things, in the right direction. The usual question is whether we will wait for disaster before we respond. The better question is whether we will do the boring work that prevents it: insisting on who may see what, logging every access so nothing happens in the dark, verifying where our models and our data come from, and keeping a human accountable for every consequential decision. None of that makes headlines. All of it is what responsible builders are quietly doing right now. The machines are not our masters, and they are not going rogue. The real question is whether the people building and deploying them will hold themselves to the standard the technology demands, and that is still a choice we get to make, if we stop waiting for the wrong disaster and start preventing the right ones.
Brian Hudes is a board-certified gastroenterologist and hepatologist with more than thirty years of clinical experience and a recipient of his specialty board’s thirty-year certification award. He built and ran a private gastroenterology practice for twenty years, then spent over a decade in hospital-based gastroenterology, most recently as chief of gastroenterology and medical director of GI and endoscopy at a 550-bed Level I trauma center in Pensacola, Florida. He holds a faculty appointment as assistant professor of medicine at Florida State University College of Medicine.
Those two halves of a career are the reason he writes about clinical software the way he does. The outpatient practice and the inpatient service fail their clinicians differently, and most people building for one have never worked in the other.
Dr. Hudes has been building the tools alongside the practice for just as long. In 1995, during his GI fellowship, he co-developed one of the first Windows-based endoscopy reporting systems in the United States, written because nothing on the market had been designed by anyone who had performed a procedure. He is now founder and chief executive officer of AiMOS Systems Corporation, where he is building an AI-native clinical platform for both settings, intended to replace the encounter-based electronic record rather than layer intelligence on top of it. AiMOS has been accepted into NVIDIA’s Inception program.
He writes on the architecture of clinical information, administrative cost growth, physician workforce shortages, board certification policy, and the widening gap between what clinicians need and what the industry builds. Professional updates are available on LinkedIn.