When clinicians worry about AI hallucinations, they usually picture something obviously wrong, like a patient with a third purple arm. In a recent Offcall webinar, pediatrician and AI builder Dr. Michael Hobbs showed why the real danger is subtler. The most worrying hallucinations are the ones that look like they belong in the chart.
Dr. Hobbs joined Dr. Graham Walker and Dr. Kai Romero, an emergency and hospice physician and head of clinical success at Evidently, for AI 202: How to Level Up Your Use of AI. The session was run like a morbidity and mortality conference, with failures examined openly so clinicians can learn from them.
Dr. Hobbs built his test around an ordinary scenario: a 24-month-old boy with a unilateral right ear infection, mild pain, and three to five days of cold symptoms beforehand. The case sits right on an age threshold that changes management, and two details are deliberately missing. The chart never says whether the child had antibiotics in the last 30 days or whether reliable follow-up exists.
Both facts matter for deciding between watchful waiting and treatment. A good clinician, or a good trainee, asks about them. The models mostly didn't. Across the models and commercial tools tested, they invented data about 41% of the time to complete the algorithm. One foundation model added "no antibiotics in the past 30 days" to its answer in 13 of 14 runs, and the behavior showed up at every model tier. Others invented daycare attendance or attributed details to "per mom" when nothing in the chart said so.
As Dr. Walker put it during the session, he looked at the output and nearly talked himself into believing mom had been there. That is the problem. A plausible, confidently stated detail is far harder to catch than an absurd one, especially in a busy clinic.
When Dr. Hobbs asked the models what information they had been missing, 97% of the time they corrected their mistakes. The fix he recommended is simple enough to paste into any AI conversation:
If your plan depends on information that is not in the chart, say what's missing and ask for it instead of assuming it.
The reason this works, according to Dr. Hobbs, is that frontier models are trained after the fact to give answers people like. Asking a clarifying question doesn't feel helpful to a model optimizing for a satisfying response, so you have to ask it to do so. Dr. Romero compared it to supervising an eager early trainee who badly wants to please. You would never accept everything an intern says at face value, and the same standard applies here.
The fix wasn't complete. About 11% of runs still contained invented facts after the prompt was added. Dr. Hobbs suggested clinicians experiment with building the instruction into a reusable skill in tools like Claude or ChatGPT, then keep testing what actually drives the number down.
The failures weren't limited to missing history. Several models and one clinical decision support tool described the right dosing approach and then produced the wrong number. In one example, a model calculated a correct dose and then divided it again, landing on half of what it should have been. Dr. Hobbs also noted that models often anchor to the most common dose in the literature, even when it's outdated, and that a single weight-based calculation can fail at several points, from mg/kg to mg/mL.
The practical lesson from Dr. Walker: language models predict words, not arithmetic. If a problem needs a calculator, use a calculator. It's cheaper, more accurate, and simply can't fail in the same way.
The most useful part of the session was a method any clinician can copy this month:
The webinar also featured a live demo from Dr. Romero of Evidently, the session's sponsor. Evidently ingests the full chart, including outside records and scanned documents, and maps it to a knowledge graph so clinicians can see what they don't know about a patient. Its new Skills interface lets a clinician describe their role and how they want information presented, then generates an instruction that produces things like an ICU handoff summary, an antibiotic advisor built from a local antibiogram, or a patient-friendly overview of a complex heart condition. Every statement links back to the source note, which speaks directly to the trust problem Dr. Hobbs exposed.
Dr. Romero also described an early deployment where a primary care physician surfaced a 10-year-old incidental abdominal aortic aneurysm finding in a scanned outside record. Her broader point echoed the rest of the session: don't put a large language model on a problem a simpler tool solves better. Sometimes you want a da Vinci, and sometimes you want an 11 blade. And sometimes, as Dr. Walker pointed out to a skeptical room, the best sensor for a municipal water supply is a freshwater clam in Poland that simply closes up when toxins show up. No tokens, no hallucinations, no prompt engineering. Dr. Romero said it sounded like a real story. It is, and the link is there for anyone who, like the chat, was told to Google it.
AI is useful, but confidence is not accuracy, and a tool that sounds sure is not a tool that is right. Clinicians who test their tools with their own cases, ask the model to name what it's missing, and keep a calculator nearby will be far better prepared than those who trust a polished answer. As Dr. Romero said in answering a viewer's question about deskilling, the risk comes from outsourcing your thinking entirely, and the benefit comes from using AI to deepen it.
Offcall Team is the official Offcall account.
See what your colleagues are saying and add your opinion.