Not long ago, I wrote a Sensible Medicine article about AI called “I invented a headache to test AI and here is what I found.” I concluded that for diagnosing my headache (and a few other conditions), AI got things wrong quite often. When given full history of my problem, it did well, but a typical patient doesn’t know the pertinent negatives and positives to provide the GPT. False-positive assessments would have unnecessarily sent patients to the ER and false negatives may have worrisomely kept a serious illness unaddressed (e.g., subarachnoid hemorrhage). Human vs. AI? The verdict was that humans currently win.
But, here is a scenario where I believe AI wins, or at least we shouldn’t feel embarrassed as doctors if we use it or it can be a productive part of the article-writing process. A few weeks ago, Substack announced that it was adding an AI checker, Pangram, to the Substack website, where readers could see how much of the article was AI vs. human generated. Substack mentioned that they weren’t against AI, but rather wanted readers to have full transparency.
Seemed reasonable. So, as an inquiring mind, I tested the AI checker on a recent Substack article of mine. Its verdict: 100 percent AI. That was worrisome and a wake-up call. It is true I use Claude or ChatGPT to write an initial draft of the article. But, here is the critical point. I feed into the GPT the transcript of my podcast (Live Long and Well With Dr. Bobby) that contained all of the points I wished to include in the article. The article comes after developing the podcast.
Next step on my journey, I tested my actual podcast transcript. And the new verdict: 100 percent human. How can it be that my podcast was judged human but the article based solely on that human information was judged AI? To create the podcast, I humanly created the premise, the flow of ideas to explore, the personal vignettes to describe, the key studies to cite and discuss, and the action points for the reader. A conundrum resulted. Was my article truly AI, or did the concept of AI testing miss important considerations?
A brief background on AI detectors (and I am not an expert on them). They look at sentence structure, stylistic approaches like asking a series of questions, paragraph transitions, multiple summarizations of ideas, and likely many other elements. These programs don’t look at the idea provenance or the origin of the arguments or any of the other features that were part of my podcast creation.
So, was it appropriate or immoral that I used AI to generate a draft of the article? Let’s recall how we got here. To find key studies, we used to peruse tables of contents for key journals. Then PubMed helped find them. After that, Google searches added depth by including non-academic articles. Today, no one seems upset when a doctor uses OpenEvidence to find the studies that do or don’t support a clinical topic. And, if you have used OpenEvidence (or Claude or other GPT) to find studies, there is wraparound text summarizing the findings. No one maligns doctors for using this streamlined approach.
But what if a physician writes an article using AI? Many would object. Stepping back once again, we have always had “help” writing articles, whether it was our mentor, a writing coach, the journal editor, a copy editor, or even someone who wrote the initial draft for us (e.g., research assistant). In my early career, my initial drafts received vast amounts of red ink from others with “helpful” suggestions for me to incorporate. Was the final article mine or a composite of many? And, did readers care about how that final article came to be?
Today, a professional writer likely beats AI in quality of text. But, for more than 90 percent of the rest of us, AI likely writes better than we do. I believe that not only can AI often do better than I, it has taught me the stylistic ways to make my writing more enjoyable to read. Those repetitive questions in an AI-generated text sound good, so why not use that approach as I write?
If we have always had support with writing and the key element is the provenance of ideas, then how can a reader know whether the ideas came from the author or AI? I assert that we can demonstrate idea provenance by providing a link in the article to the GPT chat that generated the text. That audit trail would provide evidence for how the article came to be, whether the author loaded into the chat a full set of ideas, or whether it began as an open-ended query with the AI providing both the wording and the underlying ideas.
The concept of an audit trail follows directly from concerns about published database analyses. How many variations on analytic approach and tested variables led to the final “striking” headline? In general, the reader cannot know. But, there have been efforts towards requiring an analytic plan before initiating work and having that “locked box” plan available to the journal editors or anyone with methodologic interest. Offering the chat “conversation” would analogously assure the reader where the ideas originated.
I assert that if the author generated the ideas but used AI to develop a draft (and assuming that the author carefully modified that draft), this approach does not greatly differ from what we have always done. AI has become our amazingly fast research assistant, making the writing process more efficient. Some might argue that the written words do reflect the author. That may be true, and AI-generated text would miss that element. For me, the ideas matter most, and the stylistic elements, whether mine or with assistance, just help the reader absorb and hopefully enjoy the text more.
As physicians, we need to responsibly work with AI. My prior article had humans winning for diagnoses (at this point). For article writing, AI holds an important partner status. I will use this approach as it increases efficiency, and all key ideas that I share came from my humanoid brain. I will not shy away from candor about my approach. I hope that others consider it, and then decide what fits best for their reading and writing enjoyment.
Bobby Dubois is a physician-scientist, podcaster, and multiple-time Ironman triathlete whose career has centered on using scientific evidence to understand what works and what does not in helping patients, and then doing everything in his power to move them down a path toward better health. In recent years he has turned his attention toward guiding others to live longer and with renewed energy as they age, work he shares on his website.
He is board-certified in internal medicine, received his MD from Johns Hopkins, his PhD in health policy from the RAND Corporation, and his undergraduate degree from Harvard College. He worked with managed care for a decade, the electronic health records industry for seven years, and the pharmaceutical industry for ten years, and he guided a D.C.-based think tank.
Dubois has published roughly 180 peer-reviewed articles spanning quality of care measurement, the appropriateness of medical and surgical procedures, comparative effectiveness research, and the value of health care spending, with work appearing in JAMA, the New England Journal of Medicine, Annals of Internal Medicine, Health Affairs, and Value in Health. His podcast, Live Long and Well With Dr. Bobby, focuses on evidence and health and is available on Apple Podcasts and Spotify. He shares updates on Instagram.



















