As medical artificial intelligence evolves at breakneck speed, evaluating model performance has become a rapidly moving target. A model that was considered state-of-the-art six months ago may now be vastly outpaced by new releases.
To address this fragmented evaluation landscape, Dr. Liam McCoy, a neurology resident and incoming Harvard faculty member, recently previewed new research from the Stanford and Harvard Arise group during an Offcall webinar. The Arise group has built a centralized, highly rigorous benchmark suite designed specifically to keep pace with frontier models.
"We realized that all these benchmarks are getting written and they just get stranded," Dr. McCoy explained. "What we knew about AI in 2024 tells us almost nothing about these new models." The Arise suite solves this by providing dynamic infrastructure to test new models almost immediately upon release, testing them across billions of tokens and over 600,000 model responses.
The Arise suite moves beyond simple multiple-choice questions, recognizing that AI models possess distinct knowledge and reasoning characteristics. To capture a holistic view, the suite utilizes ten diverse evaluation sets spanning multiple clinical domains:
When analyzing the massive dataset, Dr. McCoy's core finding was striking: "no one model is completely dominant across the board."
Different foundation models excel in entirely different niches, which means the "best" model depends entirely on the use case. According to the research:
However, the most critical takeaway for healthcare systems and clinician-builders is the performance of specialized platforms. The Arise group tested generalist frontier models against purpose-built clinical tools from partners like Glass Health, Amboss, and Doximity.
The results clearly validated the hard work of specialized clinical engineering. Dr. McCoy revealed that "clinical-specific products... consistently outperformed general-purpose chatbots on clinical tasks." Glass Health, in particular, swept the board in the clinical realm. This data proves that bolting specialized medical architecture, verified clinical literature, and localized reasoning onto an AI backend yields significantly better results than relying on a raw, one-size-fits-all chatbot.
The inevitable question in medical AI research is always: how does the machine compare to the doctor?
The Arise data showed that AI models actually outperformed human benchmarks in the harm-avoidance category. The models proved highly adept at navigating risk and strictly avoiding dangerous clinical options.
However, physicians shouldn't hang up their stethoscopes just yet. Humans still consistently outperform AI models in highly nuanced visual and reasoning tasks. Specifically, human doctors scored better on radiology, dermatology, and script concordance testing.
Interestingly, the benchmark highlighted a unique flaw in newer "thinking" models (models that outline their internal reasoning before answering). Under conditions of deep medical uncertainty, Dr. McCoy noted that these models often "think themselves into an extreme answer that kind of drives things off." For now, the nuanced ability to sit with uncertainty, update hypotheses based on subtle clinical context, and visually interpret complex imaging remains a distinctly human advantage.
Offcall Team is the official Offcall account.
See what your colleagues are saying and add your opinion.