Google AMIE study tests medical AI with patients under doctors
- Primary-care patients tested Google’s conversational diagnostic system, AMIE, before urgent medical visits in a study published on Oct.
- The study moved artificial intelligence evaluation out of the exam room and into direct patient interactions ahead of scheduled appointments.
- Patients chatted with AMIE to take their medical history prior to urgent visits while a clinician observed every exchange.
Primary-care patients tested Google’s conversational diagnostic system, AMIE, before urgent medical visits in a study published on Oct. 8 in The Lancet, marking a shift for large language models from licensing exam questions into real clinical workflows. Alphabet funded the research, conducted with Beth Israel Deaconess Medical Center, which involved 98 patients and required supervising physicians to monitor every exchange.
How Primary-Care Patients Tested AMIE Before Urgent Visits
The study moved artificial intelligence evaluation out of the exam room and into direct patient interactions ahead of scheduled appointments. None of the patient-AI conversations had to be stopped during the trial.
Patients chatted with AMIE to take their medical history prior to urgent visits while a clinician observed every exchange. Among the 98 patients who completed both the chat and the in-person appointment, AMIE included the final diagnosis within its first seven suggestions in 90% of cases, while its top suggestion matched the final diagnosis 56% of the time.
What Supervising Physicians Found During Patient Conversations
Physicians reviewed the AMIE transcripts before meeting the patients in 44 cases during the trial. Those doctors reported that the system helped them prepare 75% of the time and potentially changed how they handled the appointment in 57% of cases, according to Beth Israel Deaconess Medical Center.
Supervising doctors identified one hallucination and separately added clinical clarification in five cases across the evaluations. Patients also reported improved attitudes toward artificial intelligence in healthcare after completing the chats, and those positive attitudes persisted after their in-person physician visits.
The Performance Gap Between Autonomous AI and Human Users
While systems like AMIE perform well in controlled testing and structured chat environments, performance drops significantly when people use the models independently. A randomized Oxford experiment showed that large language models tested alone identified relevant medical conditions in 94.9% of written scenarios, but participants using those same models achieved success in fewer than 34.5% of cases—no better than individuals using standard sources of their choice.
Additional research demonstrates how sensitive these systems remain to minor adjustments in input data. A Mount Sinai study found that models altered clinical recommendations when researchers changed only a patient’s demographic details.
Despite strong diagnostic inclusion rates in small clinic trials, experts note that validation across broader health systems is still required. The Lancet trial involved a single clinic and a limited group of 98 patients with a physician actively supervising every conversation, leaving questions about how the technology will perform across diverse health networks and larger patient populations.
