Should You Ask AI or Your Doctor? 2026 Evidence — Jul 27, 2026
Listen & watch
Show notes
Chatbots beat doctors on paper. Then real patients used them and it fell apart.
Run time: 6:50
In today's episode:
- Chatbot answers rated more empathetic than doctors, preferred 79% of the time
- GPT-4 alone beat physicians by ~15 points on diagnostic vignettes
- Google's AMIE matched or beat primary care docs in text consultations
- Microsoft's system solved 85% of hard NEJM cases vs doctors' 20%
- Oxford RCT: patients using LLMs did no better than a web search
- Half of consumer chatbot medical advice carried potential for harm
- ~1 in 6 US adults now use an AI chatbot for health advice monthly
TL;DR:
- On clean written cases, AI genuinely out-diagnoses doctors — that part is peer-reviewed, not hype.
- On real patients doing their own typing, that edge vanishes: an Oxford RCT (~1,300 people) found LLM users did no better than web search, because patients omit the detail that matters and models mix right advice with wrong.
- What's actually deployed is the opposite of a doctor replacement: clinician tools like OpenEvidence (used by ~65% of US doctors) keep a licensed human between the model and the prescription. No LLM is FDA-cleared to diagnose you solo.
Sources cited:
- JAMA Internal Medicine, Ayers et al.
- JAMA Network Open, Goh et al.
- Nature, Google DeepMind
- Microsoft AI
- Oxford, Nuffield Dept of Primary Care (Payne et al.)
- NBC News on OpenEvidence
Subscribe: YouTube
medAI Times is for educational and informational purposes only. The content does not constitute medical advice, diagnosis, treatment recommendation, or professional clinical guidance. Consult qualified healthcare professionals and refer to official sources before making clinical, research, regulatory, or business decisions.
Transcript
Auto-generated from the episode audio. Click any timestamp to jump the player there.
Chatbots beat doctors on paper. Then real patients used them, and it fell apart. Welcome to MedAI Times podcast, your daily update on medical AI. Don't forget to like and subscribe. Here are today's beats. On paper, the machine already wins.
A chatbot's answers to patient questions were rated more empathetic than real physicians and preferred nearly eight times in 10. One experimental AI system solved 85% of the hardest published cases, where a panel of doctors managed 20.
And then the twist that reframes everything. When ordinary people use those same models on their own symptoms, they did no better than a plain web search. One callback before we start. Back on June 6th, we covered the White House push to build a regulatory pathway for autonomous AI doctors.
The evidence meant to justify that pathway is exactly what we are weighing now. So, should you ask AI or your doctor, this is not a hypothetical anymore. By late 2025, roughly one in six American adults told the Kaiser Family Foundation they use an AI chatbot for health advice every single month.
And other surveys put the share who tried it in the last 30 days closer to a quarter. People are already doing this, quietly, between the appointment they waited six weeks for and the one they cannot get. The honest version of this question is not whether AI will replace your doctor someday.
It is whether the thing already sitting in your pocket gives you a better answer than the human with the license. So let us look at what the head-to-head studies actually found and be strict about what each one measured.
Start with the case for the machine, because it is stronger than skeptics admit. In 2023, a study in JAMA Internal Medicine took nearly 200 real patient questions from an online forum, put physician answers next to chat GPT answers, and had licensed clinicians grade them blind.
The chatbot won on quality and on empathy, and the graders preferred it 79% of the time. Then in 2024, a randomized trial in JAMA Network Open pitted GPT-4 against practicing physicians on tough diagnostic vignettes.
GPT-4 working alone scored about 15 points higher than doctors using their usual resources. It gets sharper. Google built a conversational system called AIMEE, tuned specifically for taking a medical history through text chat.
In a study published in Nature, specialist physicians and patient actors judged AIMEE equal to or better than primary care doctors on the large majority of measured axes, including awkwardly empathy.
And Microsoft's experimental diagnostic orchestrator paired with a reasoning model correctly cracked 85.5% of 304 of the hardest New England Journal of Medicine cases. The physicians it was benchmarked against solved about 20%.
On the clean written case, the machine is not close. It is ahead. Now the case against, and this is where the headline numbers quietly collapse. Researchers at Oxford ran the experiment almost nobody else did.
Instead of feeding a tidy case file to the model, they gave nearly 1,300 real people, realistic symptom scenarios and a chat bot and asked them to figure out what was wrong and what to do. The people using large language models did not make better decisions than people using an ordinary web search or their own judgment.
The models knew the medicine. The humans could not extract it. Patients left out the detail that mattered and the model handed back answers that mixed a correct recommendation with a wrong one with no signal about which was which.
That is the finding that should stay with you. A separate risk assessment published this year in the International Journal for Quality in Healthcare found that roughly half of the medical advice from consumer chat bots carried some potential for harm, usually not because the fact was wrong, but because the model missed the emergency or reassured when it should have escalated.
The benchmark studies test the model on a perfect transcript. Your kitchen at 11 at night is not a perfect transcript. You do not know which symptom is the load-bearing one and the chat bot will not tell you that you forgot to mention it.
So what is actually deployed today in real clinics right now? The AI that is winning in medicine is not the patient-facing doctor replacement. It is the doctor's tool, a system called open evidence, which grounds its answers in the medical literature and cites sources.
It was used by about 65% of United States physicians across nearly 27 million clinical encounters in a single month this spring, according to NBC News. Notice the shape of that.
The machine is answering the doctor's questions, not the patient's. And a licensed human is still standing between the output and the prescription. No large language model is cleared by the FDA to diagnose you on its own.
Here is the spotlight worth remembering. The most counterintuitive result in this whole file is that doctors given GPT-4 did not beat GPT-4 working alone. The human plus AI team underperformed the AI.
That is not a story about brilliant machines. It is a story about how hard it is for any person, expert or patient, to actually pull the right answer out of a model. The interface is the bottleneck, not the intelligence.
So the verdict with the evidence tiers stated plainly that AI can out-diagnose doctors on written cases is peer-reviewed and real, in JAMA and in nature, that the 85% superintelligence figure is a vendor preprint on cherry-picked hard cases, not a clinical trial, and it has not been tested on the runny nose and the vague ache that fill an actual clinic,
and that patients using AI do not get better outcomes is also peer-reviewed from Oxford, and it is the single most relevant finding to the question you came in with. Use the chatbot to prepare for your appointment, to draft the questions, to understand the words on your lab report.
Do not use it as the doctor. And if it tells you something is fine while your gut says otherwise, believe your gut and make the call. Full sources and every study are linked in the description. Here is the one I want your answer to.
If a chatbot and your own doctor gave you different answers about a worrying symptom tomorrow, which one would you actually act on first and why? Thanks for listening. Find us on YouTube and your favorite podcast app.
See you tomorrow.