At the hospital where I operate, most of the pediatric subspecialists never walk through the door. Endocrinology, hematology, infectious disease: The consult happens by phone. They read the chart, they ask questions, they leave a plan. I would like to tell you the plan is usually right. I have no way to know that. The hospitalist examined the child. The specialist did not, and there is no second version of the case to compare against. I can judge how the reasoning sounds, and that is all.
We do not think this is a good arrangement. We think it is a poor compromise we stopped arguing about years ago.
I have started to wonder what that says about the argument we keep making about artificial intelligence.
The usual defense of the profession runs through the body. A machine cannot lay hands on a patient, cannot feel the belly, cannot register the particular quality of a child who looks worse than the numbers say. I believe that. But if it is the load-bearing argument, we have conceded more than we admit. The examination still happens. It just gets written down, and the specialist reasons from the writing. Once the findings have become text, the question of who reads that text is wide open, and reading it is exactly what these systems do best.
The machine is a better talker than we are
The finding that unsettled people arrived in 2023. Researchers took patient questions from a public forum, paired each physician answer with one generated by ChatGPT, and had clinicians grade both blind. The graders preferred the machine most of the time. On empathy it was not close.
The study has been picked apart since, and fairly. The physician answers were volunteer replies typed between patients, a few dozen words, while the machine’s ran four times longer. But length is not the explanation. When researchers held the two to the same word count, the gap held. What closes it is changing the humans. Set the machine against answers from a forum staffed by board-certified specialists rather than doctors dashing off a favor, and the difference disappears. Researchers who put the original answers before a lay audience found it still ahead, by a far smaller margin, and coded the language to see why. Nothing mystical. The machine validated, reassured, withheld judgment, and never sounded rushed. That is not artificial intelligence. That is what any of us sounds like when we are not carrying a pager.
We did not lose that contest on understanding. We lost it on manner. And that is still not the part that worries me.
The autothrottle again
I wrote here a couple of months ago about Asiana Flight 214, which struck the seawall at San Francisco on a clear day in 2013. The crew believed the automation was holding their airspeed. It was not, and because they believed it was, they stopped watching. I told it as a warning about surgical robots. It turns out to describe something already happening with no robot in the room.
Researchers at Mass General Brigham asked six radiation oncologists to answer messages from cancer patients. First the doctors wrote their own replies. Then they were handed drafts written by GPT-4 and asked to edit them. Some of those drafts were dangerous, and the danger was not in the biology. The machine misjudged how sick the patient was, which is the error that kills people.
Here is the finding I cannot put down. When the physicians were done editing, their final answers resembled the machine’s draft more than what those same physicians had written minutes earlier on their own. They had already formed a judgment. Then they read a fluent, confident paragraph and let it quietly replace theirs.
Nobody was overpowered and nobody was lazy. Six experienced oncologists did what the Asiana crew did, which is trust a system that sounded like it had things handled.
The same pattern turned up in a trial that gave physicians a large language model for hard diagnostic cases. The doctors with the model did no better than the doctors without it. The model working alone did better than either group. The tool was good, the doctors were good, and the combination added nothing, because nobody has taught us how to use it.
What the consult is actually for
The comfortable story among physicians has been that machines will take the pattern recognition while we keep the human part. The evidence points the other way. Patient, unhurried, validating conversation is what these systems already do well, and what we do worst at four in the afternoon with sixty messages waiting. Judging acuity, knowing this particular child is sicker than the chart says, is where they fail, and it is the thing we are least able to explain when someone asks how we knew.
We are defending the wrong border with the wrong argument.
Which brings me back to the phone. When the endocrinologist calls with a plan, I am not the one taking the call. The pediatrician is. What arrives at the bedside is not an examination. It is a trained judgment, a license standing behind that judgment, and a person who will pick up when the child gets worse at two in the morning. Those three things came bundled together for so long that we stopped noticing they could come apart. They can come apart now.
And the reason that unsettles me is the same reason the phone consult does. When a plan arrives from somewhere I cannot see, I have no independent way to check it. I can only judge how it sounds. That was tolerable when the voice on the other end had trained for a decade and could be paged at midnight to answer for it. It is a different kind of weakness when the thing on the other end has been optimized to sound right.
Aviation did not settle this by arguing about whether pilots would be replaced. It named the phases of flight that require a human paying attention, wrote them down, and protected them on purpose. We have not done that work. We are finding out one consult at a time.
Colin G. Knight is a board-certified pediatric surgeon practicing on the Treasure Coast of Florida at HCA Florida Lawnwood Hospital. He is a clinical assistant professor of surgery at the Florida State University College of Medicine and at the Florida International University Herbert Wertheim College of Medicine.
He earned his undergraduate degree at Yale University and his medical degree at the University of Virginia. Before his surgical training, he served four years on active duty in the United States Air Force as a flight surgeon, work that shaped his interest in operating-room safety. He completed his general surgery residency at Allegheny General Hospital and his pediatric surgery fellowship at Children’s Hospital of Michigan.
His research spans minimally invasive and robotic pediatric surgery as well as the management of pediatric appendicitis, with work appearing in the Journal of Pediatric Surgery, the Journal of Laparoendoscopic and Advanced Surgical Techniques, and Archives of Surgery. He can be found at ped-surg.com and shares updates on LinkedIn, Instagram, and X.



















