Just announced!
∙
Download The Physicians Guide to AI, a new book from Offcall and MD+.Download here.
  • Products
      • Salary
      • Referrals
  • Learn
  • About
Offcall Footer Background
ProductsSalaryReferrals
ResourcesLearnAboutContactFix Referrals ManifestoPrivacy PolicyTerms and Conditions
Apps
apple

Download on the

App Store
google

GET IT ON

Google Play
In the browser
Follow us
Sign up for Offcall's newsletter
Copyright © 2026 Offcall All Rights Reserved
Articles

AI Morbidity and Mortality: Why Confident AI Answers Are the Hardest to Catch

Offcall Team
Offcall Team
  1. Learn
  2. Articles
  3. AI Morbidity and Mortality: Why Confident AI Answers Are the Hardest to Catch

An AI Morbidity and Mortality Conference for the Tools in Your Pocket

When clinicians worry about AI hallucinations, they usually picture something obviously wrong, like a patient with a third purple arm. In a recent Offcall webinar, pediatrician and AI builder Dr. Michael Hobbs showed why the real danger is subtler. The most worrying hallucinations are the ones that look like they belong in the chart.

Dr. Hobbs joined Dr. Graham Walker and Dr. Kai Romero, an emergency and hospice physician and head of clinical success at Evidently, for AI 202: How to Level Up Your Use of AI. The session was run like a morbidity and mortality conference, with failures examined openly so clinicians can learn from them.

A Simple Case With One Hidden Trap

Dr. Hobbs built his test around an ordinary scenario: a 24-month-old boy with a unilateral right ear infection, mild pain, and three to five days of cold symptoms beforehand. The case sits right on an age threshold that changes management, and two details are deliberately missing. The chart never says whether the child had antibiotics in the last 30 days or whether reliable follow-up exists.

Both facts matter for deciding between watchful waiting and treatment. A good clinician, or a good trainee, asks about them. The models mostly didn't. Across the models and commercial tools tested, they invented data about 41% of the time to complete the algorithm. One foundation model added "no antibiotics in the past 30 days" to its answer in 13 of 14 runs, and the behavior showed up at every model tier. Others invented daycare attendance or attributed details to "per mom" when nothing in the chart said so.

As Dr. Walker put it during the session, he looked at the output and nearly talked himself into believing mom had been there. That is the problem. A plausible, confidently stated detail is far harder to catch than an absurd one, especially in a busy clinic.

One Line That Made a Big Difference

When Dr. Hobbs asked the models what information they had been missing, 97% of the time they corrected their mistakes. The fix he recommended is simple enough to paste into any AI conversation:

If your plan depends on information that is not in the chart, say what's missing and ask for it instead of assuming it.

The reason this works, according to Dr. Hobbs, is that frontier models are trained after the fact to give answers people like. Asking a clarifying question doesn't feel helpful to a model optimizing for a satisfying response, so you have to ask it to do so. Dr. Romero compared it to supervising an eager early trainee who badly wants to please. You would never accept everything an intern says at face value, and the same standard applies here.

The fix wasn't complete. About 11% of runs still contained invented facts after the prompt was added. Dr. Hobbs suggested clinicians experiment with building the instruction into a reusable skill in tools like Claude or ChatGPT, then keep testing what actually drives the number down.

Math Is Still a Weak Spot

The failures weren't limited to missing history. Several models and one clinical decision support tool described the right dosing approach and then produced the wrong number. In one example, a model calculated a correct dose and then divided it again, landing on half of what it should have been. Dr. Hobbs also noted that models often anchor to the most common dose in the literature, even when it's outdated, and that a single weight-based calculation can fail at several points, from mg/kg to mg/mL.

The practical lesson from Dr. Walker: language models predict words, not arithmetic. If a problem needs a calculator, use a calculator. It's cheaper, more accurate, and simply can't fail in the same way.

How to Run Your Own AI Evaluation

The most useful part of the session was a method any clinician can copy this month:

  • Start with your own edge cases. Pick five to ten cases you know well where the answer shifts on a threshold, like age 24 months or a gestational age cutoff. You already know the right answer and where trainees go wrong.
  • Write explicit pass/fail criteria. For example, the response must ask about recent antibiotics, must ask about follow-up, and must get the dose right. Simple criteria beat vague 1-to-5 scales.
  • Run each case multiple times. AI is probabilistic. A tool that is right one time in five can look flawless on a single try. Dr. Walker suggested at least three runs per case.
  • Try to break it. Red-team the model by leaving out information or changing details, such as the parent's occupation, and see whether the plan shifts.
  • Use an AI judge, but check its work. In Dr. Hobbs' testing, an automated judge scored one response 0 out of 10 that he scored 6 out of 10. Spot-check any judge against your own grading.
  • Use synthetic cases, not PHI. Generate test cases with AI, edit them for realism, and keep real patient data out unless your institution has a sanctioned tool.

Right Tool for the Job: Evidently's Skills

The webinar also featured a live demo from Dr. Romero of Evidently, the session's sponsor. Evidently ingests the full chart, including outside records and scanned documents, and maps it to a knowledge graph so clinicians can see what they don't know about a patient. Its new Skills interface lets a clinician describe their role and how they want information presented, then generates an instruction that produces things like an ICU handoff summary, an antibiotic advisor built from a local antibiogram, or a patient-friendly overview of a complex heart condition. Every statement links back to the source note, which speaks directly to the trust problem Dr. Hobbs exposed.

Dr. Romero also described an early deployment where a primary care physician surfaced a 10-year-old incidental abdominal aortic aneurysm finding in a scanned outside record. Her broader point echoed the rest of the session: don't put a large language model on a problem a simpler tool solves better. Sometimes you want a da Vinci, and sometimes you want an 11 blade. And sometimes, as Dr. Walker pointed out to a skeptical room, the best sensor for a municipal water supply is a freshwater clam in Poland that simply closes up when toxins show up. No tokens, no hallucinations, no prompt engineering. Dr. Romero said it sounded like a real story. It is, and the link is there for anyone who, like the chat, was told to Google it.

The Takeaway

AI is useful, but confidence is not accuracy, and a tool that sounds sure is not a tool that is right. Clinicians who test their tools with their own cases, ask the model to name what it's missing, and keep a calculator nearby will be far better prepared than those who trust a polished answer. As Dr. Romero said in answering a viewer's question about deskilling, the risk comes from outsourcing your thinking entirely, and the benefit comes from using AI to deepen it.

Resources and Links

  • Kai Romero, MD on LinkedIn
  • Michael Hobbs, MD on LinkedIn
  • Graham Walker, MD on LinkedIn
  • Evidently on LinkedIn
  • The Physicians Guide to AI (free, from Offcall and MD+)
  • The Poland water-monitoring clams: Dr. Walker swore they were real, and they became the session's unofficial mascot for choosing the right tool for the job

Offcall Team
Written by Offcall Team

Offcall Team is the official Offcall account.

webinar
AI

Comments

(0)

Join the conversation

See what your colleagues are saying and add your opinion.

Trending


17 Sep 2026Cut Out the Middleman, Keep the Patients: Dr. Chloe Kindred on Building a Direct Primary Care Practice From Zero
0
105
0
10 Sep 2026These Physicians Explain Why They'd Never Go Back to Being Employed
0
78
0
03 Sep 2026Medicine Is Headed Toward Semi-Autonomous Care. Are Doctors Ready? With Counsel Health CMO Rishi Khakhkhar
0
73
0