top of page

June 2026

Can Large Language Models Reliably Code Qualitative Humanitarian Data? A Benchmark Study Against Human Expert Adjudication

Study evaluating the reliability of 46 large language models (LLMs) for coding qualitative humanitarian data. Top-performing models from Anthropic, Google, and OpenAI showed 90-93% accuracy, matching average human coders. While top models excelled at identifying humanitarian needs like food and medical care, they were less accurate at detecting protection-related concerns, such as physical safety and discrimination, or needs expressed in indirect language like slang or metaphor. Findings from the study suggest that LLMs can save significant time and augment humanitarian organizations’ capacity to code qualitative data, but that LLMs are no substitute for human judgement. The article ends with a discussion of best practices, including tiered human oversight and deploying open-weights models on self-hosted infrastructure for sensitive humanitarian data. The methodology presented is reproducible and can be used to test new LLMs as they become available.

bottom of page