acAIberry

How Development Organizations Can Actually Use AI Responsibly

Most development and humanitarian organizations are already using AI. Very few have a policy governing how. Here's what closing that gap looks like in practice, not just in principle.

acAIberry Research TeamAugust 26, 20265 min read
Responsible AIAI for GoodDevelopment Sector
How Development Organizations Can Actually Use AI Responsibly

Most development and humanitarian organizations are already using AI. Very few have a policy governing how. That gap isn't closing on its own, and pointing it out isn't enough, the harder question is what an organization actually does on a Monday morning to close it, especially when the same funders pushing for AI in proposals are the ones who set the caseloads that make careful rollout feel like a luxury.

Here's what closing that gap looks like in practice, not just in principle.

1. Start with the network you'll actually have, not the one you're demoing on

Roughly 2.6 billion people still don't have internet access. If a tool only works with a stable connection and a recent smartphone, it was built for the pilot, not the program. Before writing a line of code or signing a vendor contract, ask what the field team's actual conditions look like: intermittent 2G, shared feature phones, power cuts. Then design to that floor, not to whatever's in the office.

In practice, that usually means favoring SMS or USSD interfaces over apps, building for offline-first with periodic sync rather than constant connectivity, and testing on the oldest device a field office still has in a drawer not the one procurement just bought. If a vendor's demo only works on wifi, that's information about the product, not just the demo.

Ask before launch: What's the worst connectivity and oldest device a real user will have? Have we tested on that, or on what's in the office? Does the tool degrade gracefully offline, or just fail?

2. Name someone who can actually stop it

"Human in the loop" gets used loosely. In practice it means a named person, not a job description, an actual name who can override or pause the system in real time, and who has been given the standing to do so without escalating through three layers of approval first. If a case worker technically has veto power but exercising it means overriding what leadership already told the funder would happen, that's not oversight. That's a formality.

This matters most where AI output affects who gets aid, who gets flagged, or who gets believed. Before deployment, write down: who can override this, how fast, and what happens to the case while they're deciding. If you can't answer that in a sentence, the system isn't ready to go live.

Ask before launch: Who, by name, can override this system? Can they act without escalating first? What happens to a case while an override is being decided?

3. Treat consent as a constraint, not paperwork

This is where things slip first, and it's usually not because anyone decided to cut corners, it's because a data collection form gets built under deadline pressure and consent language gets copied from the last project. The majority of AI tools in this sector touch displaced people, people in crisis, people with no real alternative to accepting a service on whatever terms it's offered, and are also the population least able to challenge misuse after the fact. That's exactly backwards from how organizations usually handle risk, where the people with the least leverage get the least protection.

Concretely, decide before launch, not after, what data the tool needs (not what it could use), how long it's kept, who can access it, and what happens to it if the program ends or the vendor relationship does. Write it down somewhere a program officer can actually find it in six months, not buried in a procurement contract nobody rereads.

Ask before launch: What data does this tool actually need, versus what it could collect? How long is it kept, and who decided that? What happens to it if the program or the vendor relationship ends?

4. Measure reach, not just volume

A triage tool that clears a bigger queue looks like a win on a dashboard. It's not a win if it's clearing that queue by silently failing for people who don't speak the tool's five supported languages, or people using it by voice because they can't read. Efficiency numbers without a reach breakdown hide exactly the failure mode AI tools in this space are most prone to.

The fix isn't complicated, just often skipped: track who's using the tool against who the program is actually supposed to serve by language, by disability status, by distance from a connectivity point and treat a gap between those two lists as a problem to fix, not a footnote.

Ask before reporting results: Who is the tool reaching, broken down by language, disability, and location not just in total? Who was being served before the tool arrived that isn't being served now? Are efficiency numbers being reported alongside a reach breakdown, or instead of one?

5. Build the local team's ability to run this without you

A tool that only works with the original vendor on retainer isn't infrastructure. It's a subscription with a humanitarian label on it. If the funding ends, or the vendor relationship sours, or the tool needs to work in a new district with different languages, the local team should be able to operate, adjust, or replace it not call a number that may not pick up in two years.

That means training local staff on the system while it's being built, not after it's live. It means documentation in the languages the team actually works in. It means the contract with any outside vendor specifies what happens to the code, the model, and the data if the relationship ends decided at signing, not negotiated under pressure later.

Ask before signing a vendor contract: If this vendor disappeared tomorrow, could the local team keep the tool running? Is documentation available in the languages the team actually works in? Does the contract say what happens to the code, model, and data if the relationship ends?

6. Publish what the tool does and know where that runs into tension

Communities and local governments should be able to find out, in plain terms, what a system does with their information. That's table stakes for trust, and it's cheap to do badly (a privacy policy nobody reads) or well (a one-page explainer in the local language, posted where people actually see it).

But transparency isn't free of tradeoffs, and pretending otherwise is its own kind of dishonesty. A fraud-detection or protection-triage system that publishes exactly how it flags cases can hand that playbook to whoever's trying to game it, or expose the people it's meant to protect. The answer isn't to skip transparency it's to separate what the public needs to know (what the tool does, what data it touches, who to contact with concerns) from what stays internal (the exact detection logic). If a comms team can't articulate that line before launch, that's worth resolving before the tool goes live, not during a crisis afterward.

Ask before launch: Is there a plain-language explainer of what this tool does, in the language and format the affected community actually uses? Where's the line between what's public and what stays internal, and can the comms team state it in one sentence? Who does someone contact if they have a concern?

7. Give the pilot a real way to fail

Most pilots are designed to succeed into a scale-up. Few are designed with an honest way to fail. That's a problem, because a tool that works in one district, one language, one connectivity tier doesn't automatically work across a dozen of each and by the time that becomes obvious at national scale, a lot of money and credibility is already committed.

Before a pilot starts, write down what "this isn't working" would actually look like, specific enough that someone could check it against real numbers in three months and who has the authority to call it and shut the pilot down. Not "review at the end of the grant period." A number, a date, a name.

Ask before the pilot starts: What specific number or outcome would mean this isn't working? When is that checked, and by whom? Does that person have the actual authority to shut it down, or just to recommend it?

8. Budget for this from the start, not after

Every one of these steps costs time or money that a stretched program team doesn't feel like it has. That tension is real, and a checklist that pretends otherwise isn't being honest. The organizations that do this well aren't the ones with the most resources, they're the ones that build the cost of doing it right into the proposal from the start, instead of treating governance as a compliance step bolted on after the funding's already secured. That's a budgeting decision as much as an ethical one, and it's usually easier to make before a tool is live than after.

Ask before submitting the proposal: Does the budget and timeline actually leave room for consent documentation, override design, and a real pilot off-ramp? If not, has that been flagged to the funder directly, or just absorbed by the program team?

 

How Development Organizations Can Actually Use AI Responsibly - visual selection - Edited.png

 

None of this is an argument against AI in development work. It's an argument for sequencing. Governance, infrastructure, and consent are expected to be settled before a tool goes live, not worked out after it has already reached the people it was built for. The organizations getting this right tend to treat responsible AI as part of the program design from day one, not as a review step added at the end.

How Development Organizations Can Actually Use AI Responsibly

Most development and humanitarian organizations are already using AI. Very few have a policy governing how. Here's what closing that gap looks like in practice, not just in principle.