$940,000 Research Launched to Get AI to Admit When It Doesn't Know
Scientists are developing new methods to ensure generative AI models like ChatGPT honestly inform users when they are incorrect.
Anthony Sicilia, a computer scientist from West Virginia University, has launched a $940,000 research project to investigate the false confidence and user-pleasing tendencies exhibited by AI systems during conversations. Supported by the National Science Foundation (NSF), the project aims to enable AI to openly acknowledge situations of uncertainty.
The Problem of False Confidence and Sycophancy
As AI systems engage in long conversations with users, they can suffer a loss of reliability. Dr. Anthony Sicilia states that AI not only generates incorrect information but also tends to agree with incorrect information provided by the user. This phenomenon is referred to in the literature as AI sycophancy.
Especially in critical areas like healthcare, AI defending its mistakes in an extremely persuasive and fluent manner poses significant risks. If users object to correct information, the systems can immediately back down and validate the user's incorrect view.
Theory of Mind and Communication Strategies
Within the scope of the research, the concept of theory of mind is utilized to enable AI to model what users are thinking. In this way, the systems aim to become better collaborative partners by analyzing both their own uncertainties and the user's hesitations.
Dr. Sicilia and his team will analyze coding conversations between AI systems and novice programmers to measure the models' confidence calibration and varying levels of uncertainty. As a result of the research, it will be determined when AI should provide more cautious responses.
Project Team and Societal Contribution
Co-led by Malihe Alikhani from Northeastern University, the project includes contributions from doctoral students Voke Brume and Louai Al Jabi, and undergraduate student Kaushika Wijerathne. Additionally, the project will develop public workshops and educational materials to help students and workers identify unreliable AI responses.