Methods for Evaluating Artificial Intelligence Systems in Humanitarian Aid

Serdar HocamAuthor & Editor

Deficiencies in current humanitarian aid AI evaluations are examined, and a four-stage framework guaranteeing technical accuracy, user engagement, and independent impact is proposed.

◉ 0 views
How to Evaluate Artificial Intelligence Systems for LMIC Usage - ICTworks

Analyses on the shortcomings and failures in evaluating artificial intelligence systems used in the humanitarian aid sector reveal that current approaches are fundamentally flawed. Instead of jumping directly to impact measurement, a staged evaluation model is proposed.

Flaws in Current AI Evaluations

The general approach in the sector jumps directly to impact assessments without checking whether systems are functioning properly. This leads to multi-million-dollar algorithms failing and projects ending before they even begin.

Although organizations like the Danish Refugee Council, UNHCR, and WFP have successful AI and predictive models, evaluation errors across the broader sector prevent projects from reaching their potential.

Four-Stage Systematic Evaluation Framework

The new framework proposed by the center offers four sequential evaluation levels progressing step-by-step to build evidence. This structure provides a systematic alternative to the current disarray.

The first level, model evaluation, verifies whether an AI system performs its intended function accurately and consistently before deploying it anywhere.

User Interaction and Behavior Change

The second level, product evaluation, examines whether target users abandon the system even when technical accuracy is ensured, thereby revealing actual usage rates.

In the third level, user evaluation, it is analyzed whether high engagement yields meaningful results and leads to appropriate behavior changes.

Independent Impact Assessment and Conclusion

The fourth and final level is the independent impact assessment, which measures whether investments improve the quality of life after technical reliability, user engagement, and behavior change are confirmed.

This sequential approach prevents costly mistakes, contributing to the sustainability of humanitarian aid projects and the achievement of donor objectives.