For many physicians, the first truly useful encounter with generative AI has been remarkably mundane. It takes notes.
Ambient AI listens to the encounter, organizes the conversation, drafts the clinical note, and gives the physician something increasingly scarce in medicine: time. That is a genuine accomplishment. But it may also have created a misleading picture of where clinical AI is headed.
The next generation of AI will not simply document what the physician decided. It will increasingly help assemble the evidence, interpret it, recommend an action, and initiate parts of the clinical workflow. We are moving from AI as scribe to AI as assistant, from assistant to adviser, and from adviser toward agent.
That raises a question medicine has not adequately confronted: At what point does assisting physician judgment become replacing physician judgment?
Human in the loop is not enough
The conventional safeguard is reassuringly simple: Keep a human in the loop. But imagine a physician seeing a medically complex patient. An AI retrieves the relevant history, summarizes laboratory results, compares the case with guidelines, ranks possible diagnoses, recommends testing and treatment, and drafts the order. The physician reviews the screen and clicks Approve.
There is unquestionably a human in the loop. But how much human judgment was actually in the loop?
The Food and Drug Administration (FDA) framework for certain clinical decision-support software emphasizes that physicians should be able to independently review the basis for a recommendation rather than rely primarily on the software. The American Medical Association is confronting the same issue. Its recent guidance asks whether AI can replace clinical decision-making and how physicians know when they are relying too heavily on it. These are no longer philosophical questions.
The agent is coming
In September, the Advanced Research Projects Agency for Health (ARPA-H) announced contracts under a program intended to develop an FDA-authorized clinical agentic AI system for cardiovascular care, ultimately functioning around the clock as a digital member of the clinical care team. The program is ambitious, and success is not guaranteed. But the direction is unmistakable.
Clinical AI is moving beyond recording what happened in the examination room. It is moving closer to the decision itself. That could be enormously beneficial. AI can process volumes of information no physician can reasonably keep in working memory. It can identify patterns across years of laboratory results, compare a patient’s presentation with clinical literature, and flag something the physician might otherwise miss.
But there is another possible trajectory. AI retrieves the information, interprets it, frames the alternatives, recommends the decision, and prepares the action. The physician’s role gradually moves downstream until the principal remaining cognitive task is deciding whether to accept the machine’s recommendation.
That may not be augmentation. It may be cognitive outsourcing.
The real risk is cognitive atrophy
Medicine understandably concentrates on whether AI makes mistakes. We should also ask what happens when AI is right most of the time.
A system that is frequently wrong will be challenged. A system that is right 98 times out of 100 creates a different problem. The physician learns that checking it carefully usually produces the same conclusion the machine already reached. Review becomes quicker. Independent reasoning becomes less frequent. Approval becomes habitual. Then comes case 99.
The danger is not simply automation bias. It is the gradual weakening of the independent capability required to recognize when the machine is wrong.
The question therefore cannot simply be whether a physician remains legally responsible for the decision. The more important question is whether the physician remains cognitively capable of making it independently.
We should automate the right things
None of this is an argument for preserving inefficient work. Physicians should not spend valuable cognitive capacity searching through screens for laboratory results, reconstructing conversations for documentation, or performing administrative tasks simply because medicine has always required them. AI should eliminate as much cognitive drudgery as possible.
But there is a crucial difference between outsourcing work and outsourcing judgment. The capacity released by AI should be reinvested higher in the value chain: understanding the patient, questioning assumptions, considering competing diagnoses, recognizing unusual circumstances, and exercising judgment where evidence is incomplete.
I have described this broader approach as Human-AI Cognitive Co-evolution, or HACC. The objective is not merely to make humans faster by transferring more tasks to machines. It is to structure interaction with AI so that human cognitive capability develops as well.
The alternative is Human-AI Cognitive Outsourcing, or HACO. The AI becomes increasingly capable while the human becomes increasingly dependent on it. Medicine should have a strong preference between those two futures.
Measure cognition, not clicks
Health systems therefore need another metric when evaluating clinical AI. We already measure accuracy, time saved, adoption, clinician satisfaction, and patient outcomes. We should also ask whether physicians retain the ability to reason without the system.
Can they identify when its recommendation conflicts with the clinical picture? Can they explain why they accepted or rejected its conclusion? Can they solve an unfamiliar clinical problem without AI assistance? Can they recognize when the model has confidently framed the wrong question? Those are harder to measure than minutes saved per note. They may also be considerably more important.
A physician clicking Approve tells us almost nothing about whether meaningful human judgment occurred.
The physician should become more capable, not less
The first generation of clinical generative AI offered medicine an appealing bargain: Let the machine do some of the clerical work so the physician can spend more time being a physician. We should take that bargain. But the next bargain is more complicated.
As AI moves from documenting decisions toward participating in them, medicine will have to decide what forms of cognition should remain human and what can safely be delegated. The objective should not be to protect physicians from AI. It should be to build clinical systems in which physicians become more capable because AI is present.
That requires more than keeping a human somewhere in the workflow. It requires keeping human judgment alive.
The question is no longer whether physicians will practice medicine with AI. They will.
The question is whether AI will make physicians better thinkers or gradually become the thinker whose decisions physicians approve.
Matt Hasan is an economist, AI strategist, and founder of aiRESULTS. He advises health systems, payers, and life sciences organizations on the strategic implications of artificial intelligence, digital transformation, and emerging technologies. Over a career spanning more than four decades, he has held leadership and advisory roles with organizations including AT&T, IBM, Deloitte, Capgemini, and Citigroup, and previously served on the faculty of New York University’s Stern School of Business.
Dr. Hasan’s work focuses on the intersection of technology, institutions, and human decision making, with particular emphasis on how AI is reshaping medicine, governance, leadership, and professional practice. He is the founder of The AI Humanist Movement and an advocate for Human-AI Synergy, a framework that views AI not merely as a tool, but as a cognitive partner capable of extending human capabilities.
His writing includes “A Profession at the AI Frontier: Medicine Must Reinvent Itself or Cede Ground,” published in Health Affairs Forefront. He shares updates on LinkedIn and Medium.














![Why patients stop trusting doctors who listened to them [PODCAST]](https://kevinmd.com/wp-content/uploads/listening-isnt-enough-podcast-190x100.png)



