AI vs. Traditional Diagnosis: Which is Better?
- For decades, the medical field has employed artificial intelligence (AI) too improve diagnoses, utilizing diagnostic decision support systems (DDSSs).
- The research, published in JAMA Network Open, indicates that DXplain performed marginally better than the LLMs in diagnosing patient cases.
- "Amid all the interest in large language models, it's easy to forget that the first AI systems used successfully in medicine were expert systems like DXplain," said Dr.
Expert systems Slightly Outperform LLMs in AI Diagnosis Study
Updated June 08, 2025
For decades, the medical field has employed artificial intelligence (AI) too improve diagnoses, utilizing diagnostic decision support systems (DDSSs). Massachusetts general Hospital (MGH) developed its DDSS, DXplain, in 1984. This system uses thousands of disease profiles and clinical data points to help clinicians generate potential diagnoses. Now, MGH researchers have compared DXplain’s diagnostic capabilities to those of modern large language models (LLMs) like ChatGPT and Gemini.
The research, published in JAMA Network Open, indicates that DXplain performed marginally better than the LLMs in diagnosing patient cases. Though, the LLMs also demonstrated strong performance. The investigators propose combining dxplain with an LLM to enhance clinical efficacy and improve both systems.
“Amid all the interest in large language models, it’s easy to forget that the first AI systems used successfully in medicine were expert systems like DXplain,” said Dr. Edward Hoffer, co-author and member of the LCS at MGH.
Dr. Mitchell Feldman, also of MGH’s LCS, added that these systems can enhance and expand clinicians’ diagnoses, recalling information that physicians may forget and isn’t biased by common flaws in human reasoning.He believes that combining the explanatory capabilities of existing diagnostic systems with the linguistic capabilities of large language models will enable better automated diagnostic decision support and patient outcomes.
The study assessed DXplain,ChatGPT,and Gemini using 36 patient cases,considering various demographics.Each system suggested potential diagnoses with and without lab data. With lab data,DXplain correctly diagnosed 72% of the cases,compared to 64% for ChatGPT and 58% for Gemini. Without lab data,DXplain listed the correct diagnosis 56% of the time,outperforming ChatGPT (42%) and Gemini (39%),though the results were not statistically significant.
Researchers noted that the DDSS and LLMs each identified diseases missed by the others, suggesting a synergistic potential in combining these approaches. Preliminary work suggests LLMs could extract clinical findings from narrative text, which could then be integrated into DDSSs, improving both systems’ diagnostic accuracy.
What’s next
Future research will focus on integrating LLMs with expert systems to create more robust and accurate AI diagnosis tools, ultimately improving patient care and outcomes.
