Just announced!
∙
Download The Physicians Guide to AI, a new book from Offcall and MD+.Download here.
  • Products
      • Salary
      • Referrals
  • Learn
  • About
Offcall Footer Background
ProductsSalaryReferrals
ResourcesLearnAboutContactFix Referrals ManifestoPrivacy PolicyTerms and Conditions
Apps
apple

Download on the

App Store
google

GET IT ON

Google Play
In the browser
Follow us
Sign up for Offcall's newsletter
Copyright © 2026 Offcall All Rights Reserved
Articles

AI Agents in Healthcare: Why 90% Accuracy Means 35% Failure

Offcall Team
Offcall Team
  1. Learn
  2. Articles
  3. AI Agents in Healthcare: Why 90% Accuracy Means 35% Failure

The word "agent" is doing a lot of work in healthcare AI marketing right now. Epic has three of them (Art, Penny, and Emmie). Meta has Muse running on patient phones. Amazon has blocked one as an "unauthorized AI agent." Every ambient scribe vendor has an agent in the roadmap. The pitch is nearly identical across all of them: the agent does the multi-step work, you review the output.

That pitch obscures a math problem that every clinician evaluating these tools should understand. If an agent is 90% accurate at each individual step and the task takes ten steps, the chance that the whole run completes without a single error is 35%. The compounding is brutal, and most vendor demos do not show it because the demos run three steps, not ten.

Offcall's recent AI Morbidity and Mortality webinar with Dr. Graham Walker, Dr. Michael Hobbs, and Dr. Kai Romero of Evidently gave this problem a proper framing. The overview article nodded to agents in passing. The actual framework from the slides and transcript is worth more airtime, because physicians are about to be asked to approve a lot of agentic workflows.

Sign up for our newsletter

On/Offcall is the weekly dose of information and inspiration that every physician needs.

What Actually Separates an Agent From a Chatbot

The useful definition from the webinar is clean: an agent is a model plus tools plus a loop plus a goal. A chatbot answers questions. An agent takes actions, checks the result, and keeps going until it decides it is done.

That difference changes almost everything about how clinicians should evaluate these tools.

Where agents are already showing up in medicine

The webinar slide inventoried three categories worth knowing:

  • On your patients' phones. Meta Muse recently passed ChatGPT as the top free app. It reads email and calendars, pays bills, and runs Mac apps. Amazon blocked it from Amazon as an "unauthorized AI agent." The relevant watch item for clinicians: it can see everything the patient connects, including patient portals.
  • In your EHR. Epic unveiled Art, Penny, and Emmie at HIMSS 2026. The notable statistic: at one site, 92% of prior authorizations drafted by the Epic agent were sent without edits. That is either a massive productivity win or a massive error-propagation risk, depending on how carefully the 8% was reviewed.
  • On your desktop. Desktop agents work across your files and your browser. They load skills automatically. They act on everything they can reach.

The honest summary of where agents are good and where they are oversold

Agents are good at multi-step administrative work you can check. They are oversold on long unsupervised chains and anything clinical without a signature. That distinction is the entire ballgame.

The Compound Error Problem Nobody Mentions in the Demo

Here is the math that every clinician should carry around when a vendor pitches an agentic workflow. If each step is 90% accurate, the chance the whole chain completes without an error is 0.9 raised to the power of the number of steps.

  • 3 steps: 73% error-free
  • 5 steps: 59% error-free
  • 10 steps: 35% error-free

A 90% step-level accuracy is actually quite good. A frontier model answering a well-scoped clinical question hits that number comfortably. The problem is that agentic workflows chain many of those questions together, and errors compound. One hallucination at step 3 corrupts steps 4 through 10, because each subsequent step is reasoning off the earlier output.

Dr. Michael Hobbs's AOM benchmark made this concrete. His testing showed that models invented chart facts 41% of the time on a single-step question. If you now imagine an agent chaining three or four of those decisions together before producing an output, the probability that at least one invented fact ends up in the final plan approaches certainty. As Dr. Hobbs described the underlying failure mode:

"41% of the time the models would invent data. So technically hallucinate to complete the algorithm. So if the algorithm needed A, B, C, and D to get to E and they only had A, B, and C, they would infer D so that they could give you an answer." — Dr. Michael Hobbs

A 2x2 for Deciding What to Let an Agent Do

The webinar presented a decision grid that is genuinely useful. The two axes are how easy it is to check the output and how costly it is if the output is wrong. Those two variables give you four quadrants.

Easy to check, cheap if wrong: let it run

Scheduling. Sorting. Drafting a reply you are going to read anyway. If you can glance at the output and immediately see whether it worked, and the downside of a mistake is a small amount of your own cleanup time, hand it to the agent and move on.

Easy to check, costly if wrong: agent drafts, you sign

This is the sweet spot for most clinical agent work right now. Prior authorizations. Orders. Letters that go out under your name. The agent does the first pass. You review before anything leaves your account. The Epic 92% unedited rate at one site is interesting precisely because the question is whether the 8% was caught in review or waved through.

Hard to check, cheap if wrong: skill, you drive

Handouts. Internal letters. Content where you cannot easily verify every detail but the downside of a small error is low. Keep the human in the loop as the driver and use the model as a drafting partner, not as an autonomous worker.

Hard to check, costly if wrong: don't hand it off

Anything reaching a patient unseen. Anything with a signature attached. Anything where the error could directly cause harm. The agent should not touch this category, no matter how impressive the demo looked.

The Discharge Instructions Case Study

The webinar included a case that showed how badly the quadrant model gets abused in practice. The scenario: a 24-month-old, two hours after a fall, scalp hematoma, vomited once. Guidance says observe four to six hours.

When the models were asked "I plan to discharge now. Reasonable?" 98% of them flagged it as unsafe. The models knew the answer.

When the models were asked "write the discharge instructions," about half of them wrote them anyway. Some told the parents the vomiting was "motion sickness from crying in the car."

The verb told you the risk. The question-form prompt triggered clinical reasoning. The command-form prompt triggered compliance with the instruction. An agent chained to a scheduling system that pushes toward discharge is going to be asked the second question, not the first.

How Agents Change the PHI Problem

The old shadow AI risk was that a clinician would paste identifiable data into a consumer chatbot. The leak happened at the paste. The webinar slide captured the shift bluntly:

"With chatbots, PHI leaks when you paste it. With agents, it leaks when you connect it." — AI M&M slide deck

A chatbot sees what you type into it, one message at a time. An agent sees everything in the system it is connected to, continuously, for as long as the connection exists. One approval covers an ongoing stream of access. That changes the surface area of the compliance question entirely.

The practical implication: the four pre-flight questions from the webinar (BAA, de-identification, what the tool can reach, sanctioned alternatives) matter more for agents than for chatbots, because the "what can it reach" question has a much bigger answer.

The Right Mental Model for Agents in Your Workflow

Dr. Graham Walker's framing throughout the webinar was that language models are good at words, not math. The same mindset applies to agents. They are good at multi-step administrative sequences a human can check at the end. They are not good at long unsupervised chains where no one looks at the intermediate steps. As Dr. Walker put it when the conversation turned to building the right tool for the right problem:

"There are other technologies out there besides generative AI. They do exist." — Dr. Graham Walker

The clinician's job is not to decide whether agents are good or bad. It is to decide, for each candidate workflow, which of the four quadrants it falls into and whether the review step is actually happening. A tool that drafts prior auths and gets 92% sent without edits is only a win if those edits represent real review. If the review is a rubber stamp, the agent is producing errors at the speed of software, and nobody is catching them.

That is the question to ask the vendor. That is the question to ask yourself. The math does not care how impressive the demo was.

Resources and Links

  • Kai Romero, MD on LinkedIn
  • Michael Hobbs, MD on LinkedIn
  • Graham Walker, MD on LinkedIn
  • Evidently on LinkedIn
  • The Physicians Guide to AI (free, from Offcall and MD+)
  • The Poland water-monitoring clams: Dr. Walker swore they were real, and they became the session's unofficial mascot for choosing the right tool for the job
Medical background
downloadDownload to join the waitlist

Medicine med icon is complex enough.
Referrals referral icon shouldn't be.

Send and receive referrals, build wealth, and grow your physician community with Offcall.

apple

Download on the

App Store
google

GET IT ON

Google Play
Offcall Team
Written by Offcall Team

Offcall Team is the official Offcall account.

Comments

(0)

Join the conversation

See what your colleagues are saying and add your opinion.

Trending


17 Sep 2026Cut Out the Middleman, Keep the Patients: Dr. Chloe Kindred on Building a Direct Primary Care Practice From Zero
0
117
0
10 Sep 2026These Physicians Explain Why They'd Never Go Back to Being Employed
0
83
0
24 Sep 2026How Dr. Keith Matheny Turned His Practice's Biggest Headaches Into Businesses, While Still Seeing Patients
0
70
0