
Read the Q&A with AI Solutions and Delivery Manager Nayid Orozco on AI's biggest shift in humanitarian work.
Mercy Corps, which will become Prosper Global later in 2026 as part of its organizational evolution, is a global humanitarian organization running more than 200 programs across 30+ countries. Its Community Accountability Reporting Mechanism, CARM, offers the people it serves a pathway to speak up, whether that means asking a question, making a suggestion, or reporting that something has gone wrong. Claude now helps CARM teams translate, tag, and grade that feedback in a fraction of the time it once took.

Read the Q&A with AI Solutions and Delivery Manager Nayid Orozco on AI's biggest shift in humanitarian work.
Read the Q&A with AI Solutions and Delivery Manager Nayid Orozco on AI's biggest shift in humanitarian work.
Read the Q&A with AI Solutions and Delivery Manager Nayid Orozco on AI's biggest shift in humanitarian work.
Community feedback mechanisms are a baseline expectation across humanitarian work, part of the sector's accountability to the people it serves. CARM is Mercy Corps' answer. It gives participants in Mercy Corps programs and the people living around them a way to tell the organization what is working, what is not, and when something has gone wrong. That includes questions about assistance to safeguarding concerns and complaints about staff conduct. Reports arrive through hotlines, forms, help desks, and in-person conversations, very often in one of dozens of local languages, from Burmese and Ukrainian to Spanish and French. Before Claude, someone had to transcribe each report, translate it, read it, decide which theme it fell under, grade its severity against the organization's guidance, and route it for action.
Most reports are ordinary program feedback. Some are not. "The whole point of accountability is that the person who raised the concern trusts that we hear it, act on it and report back in a timely manner," said Nayid Orozco, AI Solutions and Delivery Manager at Mercy Corps. "If a serious report sits in a queue unread or gets logged as routine when it should have been escalated, we have not just missed a data point. We have let down someone who took a risk to speak up, and we may have left a protection issue unaddressed."
The strain showed up in three places: volume, language, and consistency. Teams had more feedback than they could process quickly, local dialects were hard to translate and respond to, and two people could grade the same case differently. The cost was staff time and delay. "Speed and correct severity grading are the difference between a mechanism that protects people and one that just collects forms," Orozco said.

Turn limited resources into lasting impact. Generate grant proposals, track program outcomes, and free your team to focus on serving your community.
Turn limited resources into lasting impact. Generate grant proposals, track program outcomes, and free your team to focus on serving your community.
Turn limited resources into lasting impact. Generate grant proposals, track program outcomes, and free your team to focus on serving your community.
For triage that could touch safeguarding, the first questions "were not about speed," Orozco recalled. "They were about data, responsibility and control." Mercy Corps needed a provider that would not train on its data, that held inputs for as short a window as possible, and that could show real certifications rather than promises. Anthropic's commercial terms met both requirements, and together with ISO 27001, ISO 42001, and SOC 2 certifications, they "are what let us even start the conversation," Orozco said.
"What would have disqualified a model was any ambiguity on data use, or any sign that it would confidently invent things on sensitive content," Orozco explained. "We tested by running real cases through it and checking the output, not by taking it on trust."
A CARM colleague ran 20 real feedback cases through Claude, graded against the team's standard guidance, then compared the results to the original human grades. Claude agreed about 75% of the time. "The most useful part was the disagreements," Orozco noted. Some traced to genuine ambiguity in the rubric, and a few prompted the team to look again at whether the original human grade was right. "It gave us a baseline for automated grading, and it worked as an audit of our own consistency," he said. "The reason it gave us confidence to continue was that the divergences were explainable and instructive, not random." The team initially found that Claude sometimes changed its grade when the same case was run again. However, after refining the grading definitions and guidance provided to Claude, the system produced more consistent and accurate classifications.
Trust also had to be earned on the procurement side, not just in the model's answers. Mercy Corps works with Anthropic through the Claude Enterprise interface for day-to-day work and the Claude API for production work such as automated grading, under one set of commercial terms and one Data Processing Agreement. The team completed its vendor forms and a formal privacy impact assessment from Anthropic's published documentation. "For a nonprofit that does not have a large procurement or security team, that consolidation genuinely shortened the review," he added.
In the current test flow, a report comes in, often in a local language. Direct identifiers and the most sensitive content are screened and redacted, then Claude translates it into clear English, tags it by theme, and proposes a severity grade against the team's six-level grading matrix with a short rationale. A CARM staff member reviews that output against the original and confirms or adjusts the grade. The work runs through the Claude Enterprise interface with grading guidance loaded as shared project knowledge.
"Claude never makes the final call," Orozco said. "It speeds up translation, tagging and a first-pass grade, but a trained staff member always reviews the output and owns the decision. Communities can trust the process because accountability still rests with a person, not a model."
"We are still in a testing phase here, not a fully integrated production system," Orozco emphasized. "CARM is deliberately piloting Claude and checking its results case by case before we trust it with anything live." The model the team is working toward is a shared Claude Project holding the CARM grading guidance and standard prompts, so the staff member who handles feedback in any country, what Mercy Corps calls a focal point, can get consistent translation, theme tagging, and a suggested grade without building the setup themselves. In the coming months, Mercy Corps plans to connect this to Zendesk, its CARM feedback management system.
"The rollout is less about the technology and more about people," he noted. "It means training focal points on how to prompt and, just as importantly, how to sanity check what comes back." The team is expanding only as the testing gives it confidence.
CARM team members report that feedback analysis that used to take up to a week now takes about an hour, and a report that took 8 hours to produce now takes 30 minutes. The speed ensures serious reports no longer sit unread in a queue, and the people who raised them hear back sooner.
The gains extend well beyond one team. Emily Joy, a director on Mercy Corps' corporate partnerships team, pointed to "the work that Claude has enabled that just wouldn't have been possible before." On the monitoring, evaluation, and learning (MEL) side, a research design estimated at roughly 280 person-hours was completed in about 20, and a market dataset analysis for Mercy Corps' Sudan research went from about a week to a day and a half, with its processing script now in production. Staff without developer skills, she said, "are now able to run reports and get those questions answered that they would not have been able to self-serve before."
In Mercy Corps' pilot midline survey of 42 users across 13 teams, 98% reported faster workflows, the average time reduction across tracked tasks was 88%, and more than 540 hours were saved across 29 tracked tasks in a single month.
"The reaction that stuck with me was a MEL colleague saying their biggest worry was that Claude might go away after the pilot," Orozco recalled. "When the loudest complaint is fear of losing the tool, that tells you something about how embedded it had become."
Next comes moving CARM from testing into a real rollout, taking the shared project to feedback staff in more countries. "Underneath all of it, we are building an AI adoption and governance framework so that expansion stays deliberate and responsible rather than ad hoc," Orozco said. Moving the heavier production work onto the Claude API is, in his words, what makes that scale "realistic rather than theoretical."