Most of what goes wrong in development is not a shortage of evidence. It is that the evidence never reaches the person deciding, or reaches them in a form they cannot use.
A minister has forty minutes. A project team has a deadline on Friday. The paper that would help them is ninety pages long and sits behind a search box nobody opens. That is the problem I think AI is actually good for. Not writing the paper. Getting it read, and used, by the right person at the moment it matters.
I came to this from behavioral science, and it shows in what I ask of a tool. The first question is never what the model can do. It is who needs to act, what they need to decide, and what is stopping them. Only then does the tool have a job. Sometimes the job is a short answer with a page reference. Sometimes it is a message tested on the people it is meant to move, before anyone spends money sending it. The method comes first and the technology serves it. When it is the other way round you get demos.
Trust is earned by what a tool refuses to say. AVA, the evidence assistant we built at the World Bank, answers from a curated base of reports, cites the page, and says so when the evidence is not there. In the study we published this year, people who kept using it reported saving between two and four hours a week. The citations are not decoration. They are what lets a busy person check before they act. An assistant that is confidently wrong once is finished in a ministry. One that says I do not know survives.
Measure it, or it did not happen. That is the discipline twenty years of randomized trials leaves you with, and it applies to AI with no discount. Did people learn more? Did they decide differently? Did the outcome move? Usage is not impact. I have watched too many tools count logins and call it change. The test we use is simple to state and hard to pass: what can this person decide on leaving that they could not decide walking in.
Scale is a design choice, not a hope. Most of the world's policymakers will meet these tools on an ordinary phone, in a language other than English, inside a government system with its own rules. Building for that from the start is harder, and it is the whole point. A tool that works beautifully for a hundred people in Washington has not solved anything yet.
Behavior is the product. Every one of these systems exists to change what a person does. That is a behavioral problem before it is a technical one, and the field that has spent decades on how people actually decide has a great deal to offer the people building the models. I would like the two to work together more than they do. That is most of what I am trying to do now.
If you are working on any of this, I would be glad to compare notes.