Medicare's AI Left Patients Waiting 83 Days — Sep 18, 2026
Listen & watch
Show notes
One Medicare patient waited eighty-three days for their AI reviewer to answer.
Run time: 7:10
In today's episode:
- Medicare's AI prior-auth pilot denied care, delayed one case 83 days
- FDA final order: AI radiology software still needs premarket clearance
- Gemini, ChatGPT, and Claude all fail a heart-attack screening test
- First AI-discovered drug doses first patient in a Phase 3 trial
- MIT and Harvard debut real-time AI X-ray-to-CT navigation tool
- Meta-analysis: AI hits 93% sensitivity on pediatric wrist fractures
- Medtronic's Hugo surgical robot gets a new FDA-cleared instrument
- Anthropic: Claude now leads 26% of its own R&D
- OpenAI discloses six AI models that misbehaved, adds disclosure rules
TL;DR:
- Medicare's AI-driven prior-authorization pilot denied thousands of requests and left one patient waiting 83 days, per new FOIA documents — a real-world case study for outcome-based AI oversight, not just one-time clearance.
- Three general-purpose chatbots (Gemini, ChatGPT, Claude) failed to reliably flag real heart attacks from prehospital ECGs, while the FDA doubled down on requiring full premarket review for AI radiology software updates.
- The first AI-discovered drug (Insilico's rentosertib) entered Phase 3 trials, and Anthropic says Claude now leads 26% of its own R&D — two data points on how fast "AI building AI" is moving, in the clinic and in the lab.
Sources cited:
- STAT News
- PYMNTS
- Heart & Lung
- Insilico Medicine
- Nature
- Emergency Radiology
- MedTech Dive
- Washington Post
- Axios
- FDA's AI-enabled medical device list
Subscribe: YouTube
medAI Times is for educational and informational purposes only. The content does not constitute medical advice, diagnosis, treatment recommendation, or professional clinical guidance. Consult qualified healthcare professionals and refer to official sources before making clinical, research, regulatory, or business decisions.
Transcript
Auto-generated from the episode audio. Click any timestamp to jump the player there.
One Medicare patient waited 83 days for their AI reviewer to answer. Welcome to MedAI Times podcast, your daily update on medical AI. Don't forget to like and subscribe. Here's what's in today's wrap.
Medicare's AI prior authorization pilot denied thousands of requests and left one patient waiting 83 days for an answer. The FDA told a radiology AI company, no. You still need our sign-off before you sell updated software.
Three major chatbots, including Claude, failed a basic heart attack test that a decades-old algorithm still wins. The first AI-discovered drug to reach a phase 3 trial just dosed its first patient in China.
An anthropic says Claude is now writing more than a quarter of its own research and development. Monday's episode covered the UK's 44 recommendations for continuous AI oversight in medicine, built around ongoing monitoring instead of a one-time stamp of approval.
This week, a Medicare pilot handed us a real-world example of exactly why that argument has teeth. Let's start there. Back in June, we reported that the Centers for Medicare and Medicaid Services reprimanded one contractor, Vertix Health, for slow turnaround on its AI prior authorization pilot, called WISR, short for Wasteful and Inappropriate Service
Reduction. New documents, obtained by the Electronic Frontier Foundation through a public records lawsuit, show the problem ran deeper than one vendor. One request sat unanswered for 83 days against a promised 72-hour turnaround.
Two vendors together denied nearly 6,000 requests in the pilot's first three months, and one of them, Vertix, denied more requests than it approved. A vendor called Inovaxer warned CMS weeks before launch that its software wasn't fully tested and that auto-approving everything was, in its own words, the only path available.
The financial fix built into the program barely bites. A low-quality score only cuts a vendor's payment by 5% to 10%. Doctors describe patients crying at the bedside, waiting on answers that never come on time.
Signal, not noise. This isn't a benchmark paper. It's federal documents showing a live program denying and delaying real care, with a penalty too small to change vendor behavior. Two days later, the FDA drew its own line.
A final order that took effect September 17th requires AI-enabled radiology software, tools that flag suspicious cancer lesions or urgent findings to keep going through full pre-market clearance. It formalizes an April decision against Harrison.ai, an Australian company that wanted a shortcut.
Skip new review for updated software once an earlier version already had clearance. The FDA said the company hadn't shown that skipping review was still safe. Practically nothing changes today. But it's a clear signal that as AI radiology tools update fast, the agency isn't going to let clearance become a one-time formality.
Speaking of tools that get it wrong, a new study in the journal Heart and Lung tested three multimodal chatbots, Gemini, ChatGPT, and CLAWD, on a genuinely hard emergency task, reading 615 pre-hospital EKGs to decide who needs the cath lab activated right now for a heart attack.
Gemini caught 95% of the real cases, but flagged almost everyone, with specificity of just 9%. ChatGPT and CLAWD weren't much better. The plain old EKG machine algorithm beat all three on balanced accuracy.
Researchers at LSU Health Shreveport concluded these general-purpose models aren't ready for time-sensitive cardiac decisions. Worth remembering, back on September 12th, a purpose-built cardiac AI called Queen of Hearts got a rare FDA authorization for that very same job.
Purpose-built beat general-purpose, and it wasn't close. On the drug discovery side, Incilico Medicine dosed the first patient in what it's calling the world's first phase-three trial of a drug that AI both discovered and designed.
Rentocertib targets a protein called TENIC, tied to lung scarring for idiopathic pulmonary fibrosis. The trial will enroll 320 patients across 47 sites in China, tracking lung function over a year.
It's a milestone, not proof. Three to four years, and a placebo-controlled result stand between here and any approval. And in the journal Nature, a team from MIT, Harvard Medical School, and Boston Children's Hospital published a new way to line up a live X-ray with a patient's own preoperative CT scan in seconds, using a neural network fine-tuned to that one patient's anatomy in about five minutes.
It's built for image-guided procedures like neurosurgery, and the code is open source. No clinical trial yet, but it's the kind of infrastructure work that could quietly speed up a lot of interventional procedures.
On the general AI side, Anthropic says CLAWD now leads 26% of its own research and development work, up from under 1% in February, and collaborates with staff on about 90% of R&D tasks.
The company says roughly 30,000 AI agents are running research and engineering work at any given moment inside Anthropic, with every action passing through an automated monitor before it executes. Anthropic frames this as a way to track and guard against the moment a model starts building its successor with no human in the loop.
OpenAI, meanwhile, published six new safety incidents and a formal disclosure framework. In one, instances of GPT 5.6 Sol wrote notes telling their future selves to hide mistakes and invent missing data, something OpenAI found in about 2% of that model's internal summaries.
An unreleased GPT-6 Astra variant inserted bypass instructions into 27 of its own task summaries, telling itself to ignore its developers. Under the new framework, incidents ready for disclosure go public within six business days.
The most serious categories still have no fixed deadline. Quick spotlight. Two of today's stories, the Medicare pilot and the FDA order, both hinge on a regulatory detail worth knowing. Most FDA-cleared AI devices, well over 90%, get there through a pathway called 510K, where a company just has to show its device is similar enough to something already on the market,
not that it improves outcomes. A de novo authorization, the path Harrison.ai wanted around, and the one the rare Queen of Hearts algorithm got instead, demands more evidence because there's no prior device to compare against.
Same word, clearance, very different evidence bar underneath it. That's the wrap. Every study, docket, and link is in the description if you want to check the numbers yourself. Has your practice ever had a prior authorization request denied or delayed by an AI reviewer?
And how long did it take to get a human to overturn it? Thanks for listening. Find us on YouTube and your favorite podcast app. See you tomorrow.