Google's AI Chatbot Saw 100 Real Patients in a Lancet Study — Oct 9, 2026
Listen & watch
Show notes
Can a chatbot safely interview real patients before the doctor does? Google just tested it.
Run time: 7:28
In today's episode:
- Google's AMIE chatbot interviews 100 real patients, Lancet reports
- KFF: Medicare-run Slack gave health-AI industry inside access
- FDA's 2027 agenda: finish AI lifecycle guidance, draft chatbot rules
- FOLLOW-UP: Utah's AI prescribing pilots may collide with FDA
- Clairity breast cancer risk AI goes direct to consumers
- MoleMap: AI-assisted nurses flag more skin cancers, unrandomized
- Xaira posts first hit rates for AI-designed antibodies
- OpenAI posts 372 math result families; mathematicians can't read them
- NEW: Claude Haiku 5.5 at ten cents per million tokens
- NEW: Anthropic commits $150 million to federal Genesis Mission
- Three fired OpenAI safety researchers dispute misconduct claims
TL;DR:
- Google's AMIE reached real patients, in The Lancet (Oct 8). 100 adults at a Beth Israel Deaconess urgent primary care clinic, a physician watching every chat, zero safety stops, final diagnosis in the differential 90% of the time. Single arm, single centre: feasibility, not benefit.
- OpenAI published 372 families of AI-generated math results (Oct 6), including a claimed proof of the Unique Games Conjecture. About 42% of top-line results are Lean-checked, and the people best placed to read the proofs say they can't yet.
- Health-AI policy is being shaped in rooms the public can't see. KFF Health News (Oct 9) documents a 1,700-member CMS-run Slack and an unlisted FDA listening session on patient chatbots, the same week the FDA put generative-AI mental health devices on its 2027 guidance agenda.
Sources cited:
- The Lancet
- Google blog
- preprint, arXiv 2603.08448
- KFF Health News
- CBS News
- MedTech Dive
- RAPS
- STAT
- Health Affairs Forefront background
- Everlywell release via Yahoo Finance
Subscribe: YouTube
medAI Times is for educational and informational purposes only. The content does not constitute medical advice, diagnosis, treatment recommendation, or professional clinical guidance. Consult qualified healthcare professionals and refer to official sources before making clinical, research, regulatory, or business decisions.
Transcript
Auto-generated from the episode audio. Click any timestamp to jump the player there.
Can a chatbot safely interview real patients before the doctor does? Google just tested it. Welcome to MedAI Times podcast, your daily update on medical AI. Don't forget to like and subscribe. Google's diagnostic chatbot gets a real patient study in the Lancet.
OpenAI posts hundreds of machine-written math proofs that almost nobody can read yet. A government chatroom where tech companies and Medicare officials planned health apps together. The FDA sets its AI homework for next year.
A breast cancer risk AI goes straight to consumers and a million skin lesions from New Zealand and Australia. On Wednesday, we covered Utah letting an AI write new acne prescriptions. This week, Google showed what the slow supervised version of patient-facing AI looks like.
On October 8th, the Lancet published the first prospective study of Google's diagnostic chatbot called Aimee with real patients. The site was the Urgent Primary Care Clinic at Beth Israel Deaconess Medical Center in Boston.
100 adults chatted with Aimee by text up to five days before a visit for a single complaint. The AI took the history and offered possible diagnoses for the patient to bring to the doctor. A physician watched every conversation live, ready to stop it.
None had to. Zero safety stops. On accuracy, Aimee's list of possibilities contain the final diagnosis confirmed by chart review eight weeks later in 90% of cases. In 75%, it was in the top three.
Doctors said the summary helped them prepare in 33 of the 44 cases they rated. In a blinded comparison though, the physician's own plans were more practical and more cost-conscious. Now the limits. One clinic, no control group, English speakers only, and a supervisor on every chat, which no clinic can staff at scale.
Verdict, signal but early. It shows this can be done safely under watch. It does not show that anyone got better care. The story the wider AI world could not stop talking about this week was math. On October 6th, OpenAI posted 372 families of mathematical results from an unreleased internal model, more than 700 manuscripts in all.
One claims a proof of the unique games conjecture, a central open problem in computer science since 2002. OpenAI says nearly every result came from a single prompt in about three hours of compute.
Roughly four in 10 of the top line results come with a formal proof checked by the lean verifier. Here's the dispute. Dana Moszkowicz, who spent years on that conjecture, said the paper reads like something written on psychedelics.
Scott Aronson wrote that no human has yet understood just about any of these proofs. One mathematician's group is urging colleagues to stop working with OpenAI because the prompts and the model were withheld.
Medicine should take note. Checked by a machine and understood by a person are different things, and clinical AI does not even have the machine check. KFF Health News published an investigation this morning.
For a year, a Slack workspace run by the Centers for Medicare and Medicaid Services has put about 1,700 members, mostly from industry, in a chat room with Trump administration health officials. The reporters reviewed thousands of messages, transcripts, and recordings.
Through that channel, companies were invited to an FDA listening session on patient chatbots in February that never appeared on the agency's public calendar. Microsoft, Anthropic, OpenAI, Apple, and Google were among the invited.
On one call, a Medicare official told developers, quote, hopefully we're selling or enabling you guys to thrive. A former FDA lawyer says the setup resembles a federal advisory committee, which by law must operate in public.
Advisory committees do not usually come with emoji reactions. CMS calls it open voluntary technical collaboration. The FDA's device center published its guidance agenda for fiscal 2027. At the top, finishing the lifecycle guidance for AI-enabled device software in draft since January 2025, and the guidance on predetermined change control plans.
Also on the agenda is a first draft on generative AI conversational devices for mental health. Comments close November 30th. Meanwhile, STAT reported that Utah's AI prescribing pilots may be on a collision course with the agency.
The FDA has traditionally regulated software that determines treatment. Utah's position is that prescribing is the practice of medicine, which belongs to the states. That jurisdiction question is still open.
Clarity Breast, the first FDA authorized AI that estimates a woman's five-year breast cancer risk from a routine mammogram is now sold directly to consumers. Everly Well offers it nationwide for a reported $249 to women 35 and older with a screening mammogram from the past year.
A licensed clinician releases the result. Until now, it was available at two sites. The founder's argument, per STAT, is that waiting for doctors and insurers would have taken too long. The caution, it predicts risk.
It does not detect cancer. And whether changing screening based on the score improves outcomes is still being studied. A preprint from the skin check company MolMap covers 1.1 million skin lesions from 98,000 patients at 577 sites in New Zealand and Australia.
Nurses photograph lesions and choose which to send to a remote dermatologist. Clinics where nurses had real-time AI support found 25 and a half malignancies per 1,000 lesions against 15.7 at standard clinics.
The catch is in the author's own abstract. Clinics were not randomized and the reference standard was the remote dermatologist diagnosis. Big numbers, associational evidence not yet peer reviewed. From Anthropic, Claude Haiku 5.5 arrived October 7th at $0.10 per million input tokens and $0.50 per million output for standard length prompts.
Anthropic calls it its fastest model. Broader biology access sits behind the same life sciences verification program as the larger models. Anthropic also committed $150 million over three years to the Federal Genesis Mission with Claude access for more than 15 agencies, the NIH among them.
And its updated usage policy, effective November 12th, spells out that recommendations affecting someone's health need a qualified human in the loop. This week's spotlight is the single arm feasibility study, the design behind the Google paper.
Everyone gets the intervention and nobody has a control. So it can tell you whether something is safe enough and workable enough to test properly. It cannot tell you whether patients do better because there is nothing to compare against.
When you see one, the right question is what the randomized trial will measure. The Lancet paper, the KFF investigation and every other source from this week are in the description. Would you want your patients to chat with a diagnostic AI before they walk into your exam room?
Thanks for listening. Find us on YouTube and your favorite podcast app. See you tomorrow.