Just announced!
∙
Download The Physicians Guide to AI, a new book from Offcall and MD+.Download here.
  • Products
      • Salary
      • Referrals
  • Learn
  • About
Offcall Footer Background
ProductsSalaryReferrals
ResourcesLearnAboutContactFix Referrals ManifestoPrivacy PolicyTerms and Conditions
Apps
apple

Download on the

App Store
google

GET IT ON

Google Play
In the browser
Follow us
Sign up for Offcall's newsletter
Copyright © 2026 Offcall All Rights Reserved
Articles

Medical AI Benchmarks: Which Models Perform Best?

Offcall Team
Offcall Team
  1. Learn
  2. Articles
  3. Medical AI Benchmarks: Which Models Perform Best?

The Urgent Need for Better Benchmarking

As medical artificial intelligence evolves at breakneck speed, evaluating model performance has become a rapidly moving target. A model that was considered state-of-the-art six months ago may now be vastly outpaced by new releases.

To address this fragmented evaluation landscape, Dr. Liam McCoy, a neurology resident and incoming Harvard faculty member, recently previewed new research from the Stanford and Harvard Arise group during an Offcall webinar. The Arise group has built a centralized, highly rigorous benchmark suite designed specifically to keep pace with frontier models.

"We realized that all these benchmarks are getting written and they just get stranded," Dr. McCoy explained. "What we knew about AI in 2024 tells us almost nothing about these new models." The Arise suite solves this by providing dynamic infrastructure to test new models almost immediately upon release, testing them across billions of tokens and over 600,000 model responses.

Inside the Arise Benchmark Suite

The Arise suite moves beyond simple multiple-choice questions, recognizing that AI models possess distinct knowledge and reasoning characteristics. To capture a holistic view, the suite utilizes ten diverse evaluation sets spanning multiple clinical domains:

  • Clinical Pathologic Cases (CPCs): Based on the New England Journal of Medicine's 120-year archive of complex case series.
  • Medical Imaging: Large-scale benchmarks testing the interpretation of chest X-rays and dermatologic images.
  • Agentic Reasoning: Evaluating how well tools like Claude Code can operate autonomously within simulated clinical environments.
  • Harm Avoidance: A critical safety benchmark testing a model's ability to avoid making dangerous or toxic clinical recommendations.
  • Script Concordance Testing: Assessing how well a model updates its answers and diagnostic reasoning when faced with medical uncertainty.

General AI vs. Clinical AI: Who Wins?

When analyzing the massive dataset, Dr. McCoy's core finding was striking: "no one model is completely dominant across the board."

Different foundation models excel in entirely different niches, which means the "best" model depends entirely on the use case. According to the research:

  • Google's Models: These models "seem to knock everybody else out of the water on imaging," likely due to Google's vast proprietary imaging data.
  • Anthropic's Claude: While weaker on imaging, Claude proved to be "excellent at agentic execution," making it ideal for coding and multi-step reasoning tasks.

However, the most critical takeaway for healthcare systems and clinician-builders is the performance of specialized platforms. The Arise group tested generalist frontier models against purpose-built clinical tools from partners like Glass Health, Amboss, and Doximity.

The results clearly validated the hard work of specialized clinical engineering. Dr. McCoy revealed that "clinical-specific products... consistently outperformed general-purpose chatbots on clinical tasks." Glass Health, in particular, swept the board in the clinical realm. This data proves that bolting specialized medical architecture, verified clinical literature, and localized reasoning onto an AI backend yields significantly better results than relying on a raw, one-size-fits-all chatbot.

Where Human Clinicians Still Lead

The inevitable question in medical AI research is always: how does the machine compare to the doctor?

The Arise data showed that AI models actually outperformed human benchmarks in the harm-avoidance category. The models proved highly adept at navigating risk and strictly avoiding dangerous clinical options.

However, physicians shouldn't hang up their stethoscopes just yet. Humans still consistently outperform AI models in highly nuanced visual and reasoning tasks. Specifically, human doctors scored better on radiology, dermatology, and script concordance testing.

Interestingly, the benchmark highlighted a unique flaw in newer "thinking" models (models that outline their internal reasoning before answering). Under conditions of deep medical uncertainty, Dr. McCoy noted that these models often "think themselves into an extreme answer that kind of drives things off." For now, the nuanced ability to sit with uncertainty, update hypotheses based on subtle clinical context, and visually interpret complex imaging remains a distinctly human advantage.

Offcall Team
Written by Offcall Team

Offcall Team is the official Offcall account.

Comments

(0)

Join the conversation

See what your colleagues are saying and add your opinion.

Trending


23 Jul 2026Vibe Coding Session: Git, GitHub, Permissions, and What Heidi's CEO Builds Himself
0
622
0
29 Jun 2026Announcing The Physician's Guide to AI: A Free Resource for Physicians Across Every Specialty
0
231
0
25 Jun 2026Your Patient Trusts ChatGPT More Than You Now: The New Yorker's Dr. Dhruv Khullar on Medical Authority in the Age of AI
0
118
0