USC Researchers Test Safety of AI Health Advice

USC researchers are developing the ALICE project to identify accuracy, completeness, and safety vulnerabilities in medical advice provided by AI-based chatbots.

◉ 0 views
AI Gives Medical Advice Every Day. Who’s Checking Whether It’s Safe?

USC researchers are launching the ALICE project to test the accuracy, completeness, and safety of health information provided to users by AI chatbots.

AI and Medical Advice Risks

People are turning to AI chatbots with medical questions, and these systems respond within seconds with calm, authoritative tones. However, these systems can sometimes generate incorrect information, downplay emergencies, or omit critical warnings.

Marjorie Freedman notes that while such systems may be 90 percent accurate, the remaining 10 percent margin of error can be extremely misleading and potentially harmful.

The Goal of the ALICE Project

The ALICE project aims to develop better methods to test whether chatbot responses to medical questions are accurate, complete, and actionable.

In addition to detecting false statements and missing information, the project also examines whether the responses are clear, useful, and carry appropriate urgency for patients, parents, and healthcare professionals.

Human Experts and Accuracy Challenges

Researchers have reviewed hundreds of responses from three different chatbots, revealing that a single person checking a response once is insufficient to consider it reliable.

During initial reviews, human reviewers missed some factual errors, which were later caught by medical experts and professional fact-checkers.

Detecting Missing Information

A second working group focused on omissions in responses, with medical experts reviewing answers to real patient queries and flagging missing critical information.

The most powerful automated approach tested managed to catch over 90 percent of omissions, but also generated numerous false alarms.

Future Certification Goal

Jonathan May stated that they want to provide a certification similar to this for chatbots in high-risk contexts like medicine, so consumers can trust their outputs.

This two-year project, running through September 17, 2026, is funded by the Advanced Research Projects Agency for Health.

Share