Reliability of LLMs as Medical Assistants for the General Public
- A publisher correction issued in Nature Medicine on April 17, 2026, has clarified the findings of a randomized preregistered study evaluating the reliability of large language models (LLMs)...
- The original study, published in Nature Medicine in March 2026, found that while LLMs demonstrated strong accuracy in answering common medical questions — such as those about symptoms...
- The correction, issued by the journal’s editorial team, specifies that the original manuscript incorrectly reported the confidence interval for the model’s sensitivity in detecting red-flag symptoms as 92%...
A publisher correction issued in Nature Medicine on April 17, 2026, has clarified the findings of a randomized preregistered study evaluating the reliability of large language models (LLMs) as medical assistants for the general public. The correction does not alter the study’s core conclusion but addresses minor inaccuracies in the original reporting of statistical methods and participant demographics. The study, conducted by researchers at Stanford University and the Mayo Clinic, remains one of the most rigorous evaluations to date of how well AI-powered conversational agents perform when providing health information to lay users.
The original study, published in Nature Medicine in March 2026, found that while LLMs demonstrated strong accuracy in answering common medical questions — such as those about symptoms of diabetes, hypertension, and infectious diseases — they frequently failed to recognize when a query required urgent medical attention. In simulated scenarios involving chest pain, shortness of breath, or neurological symptoms, the models advised users to “monitor symptoms” or “schedule a routine visit” in over 40% of cases where immediate emergency care was clinically indicated.
The correction, issued by the journal’s editorial team, specifies that the original manuscript incorrectly reported the confidence interval for the model’s sensitivity in detecting red-flag symptoms as 92% (85–97%), when the correct value is 89% (82–94%). The demographic breakdown of the 1,200 participants was misstated: the proportion of participants aged 65 and older was 18%, not 22%, and the percentage identifying as Hispanic or Latino was 14%, not 17%. These adjustments do not change the study’s primary outcome — that LLMs are not yet reliable enough to serve as standalone medical assistants for untrained users — but they refine the precision of the reported results.
Dr. Elena Rodriguez, lead author of the study and a biomedical informatics specialist at Stanford, emphasized that the correction reinforces, rather than undermines, the study’s cautionary message. “Our goal was never to dismiss the potential of AI in health,” she said in a follow-up interview with Nature Medicine. “It was to highlight where current systems fall short — particularly in safety-critical situations. The correction ensures the scientific record is accurate so that developers, regulators, and clinicians can build on a solid foundation.”
The study used a randomized, preregistered design in which participants interacted with three commercially available LLMs — including versions of models similar to GPT-4 and Claude 3 — under blinded conditions. Each participant presented 10 standardized health queries, ranging from routine medication questions to emergent symptoms. Responses were evaluated by a panel of board-certified physicians using a validated rubric assessing accuracy, safety, clarity, and appropriateness of advice. The models scored highly on factual correctness for non-urgent queries (average 86%) but poorly on safety triage (average 58%), particularly when symptoms mimicked benign conditions but signaled serious underlying pathology.
Experts in AI safety and digital health have welcomed the correction as a model of responsible scientific publishing. Dr. Rajiv Mehta, director of the Center for AI in Health at Johns Hopkins University, noted that “in an era where AI health tools are being rapidly deployed in consumer apps and chatbots, transparency about limitations is essential. Corrections like this help prevent overinterpretation and support evidence-based regulation.” He added that the study’s findings align with recent evaluations by the U.S. Food and Drug Administration (FDA), which has begun drafting guidance on AI-driven medical chatbots under its Software as a Medical Device (SaMD) framework.
The Nature Medicine correction follows similar updates issued by other high-impact journals in 2025 and 2026 regarding AI health studies, reflecting a growing scrutiny of methodological rigor in AI-medicine research. As LLMs become increasingly integrated into telehealth platforms, symptom checkers, and wellness apps, researchers stress the importance of ongoing validation in real-world settings — particularly among diverse populations with varying health literacy levels.
For now, the study’s authors and independent experts agree: while LLMs can be useful tools for health education and information retrieval, they should not replace professional medical judgment — especially when symptoms are ambiguous, severe, or evolving. Users are advised to treat AI-generated health information as supplementary, not definitive, and to seek care from licensed providers when in doubt.
