Business Growth

How the World Bank built an AI tool people trust more than ChatGPT

Read time: 7 min
An image showing the World Bank Group team: Francesca Spagnoli, Noemi Caiazzo and Mohamad Chatila

Key takeaways

  • A small Rome-based World Bank team turned 150-page reports into training that has reached more than 60,000 learners across 195 countries
  • Their AI research assistant, Eva, was built to be more verified than ChatGPT, with source checks on both what goes in and what comes out
  • Every AI-generated podcast episode passes through five human checkpoints before it publishes, proof that fast and careful aren’t opposites

A small team at the World Bank built an AI research tool designed to be checked and verified more thoroughly than ChatGPT, not less, and used it to take research that used to sit in 150-page reports and reach more than 60,000 learners across 195 countries.

The combination of AI at genuine scale with more human review layered in rather than less is the part worth paying attention to. Most people assume those two things trade off against each other.  This team built the workflow specifically so they wouldn’t have to choose, and the results so far, real usage numbers rather than a pilot demo, suggest the trade-off was never as fixed as it looks from the outside.

During the Learnworlds WOL: AI summit this year, we had the pleasure of virtually meeting the World Bank Group team: Francesca Spagnoli (Senior Learning Specialist), Noemi Caiazzo (Learning Analyst), and Mohamad Chatila (Behavioural Scientist), who explained to us how they made it happen.

The problem: research nobody outside the building reads

The World Bank Group’s Development Impact Unit is a small team based in Rome, tucked inside one of the largest development institutions in the world. Their job is producing research, often running numerous pages, on topics like energy, agriculture, and job creation in developing economies.

That research is genuinely valuable. It’s also, by the team’s own description, mostly unread by the people it’s meant to help. A 150-page report doesn’t reach a policymaker in the middle of a packed agenda, and it definitely doesn’t reach the practitioner actually implementing the policy on the ground.

The team’s answer was a mission they describe as making solutions fast, far, and for all: take knowledge that already exists and build the infrastructure to actually move it to the people who’d use it, rather than producing more of it.

The unit has grown and expanded fast geographically since it started:

  • Rome, under two years ago: energy, agriculture, and job creation
  • Washington DC, added since: global policy and air quality
  • Tokyo, added most recently: resilience

The audience is deliberately broad rather than narrow, spanning academia, government, private sector, and research institutes, with the heaviest concentration of learners in Africa.

What they actually built

Two tools carry most of this work: Eva, an AI research assistant, and a weekly multilingual podcast. 

Both sit on top of the same underlying idea that scale and rigor should reinforce each other rather than compete. Both feed into an impact assessment framework the team built to keep improving the approach rather than treating either tool as finished.

EvaAccelerating Development podcast
What it doesTurns 15 years of Bank reports into summaries, footnoted citations, mind maps, slide decks, infographics, and quizzesTurns research reports into short audio briefings for busy policymakers and practitioners
LanguagesMore than 60Seven, including English, French, Spanish, Hindi, Arabic, Chinese, and Portuguese
Reach so farMore than 20,000 queries, more than 2,000 professionals in the evaluation11,100 downloads across 168 countries within a few months of launch


Neither tool is a demo, and both are already carrying real weekly and daily traffic from people outside the Bank.

The team also publishes a set of downloadable infographics alongside the courses themselves, a lower-tech companion to Eva and the podcast that serves people who’d rather skim a visual summary than query a chatbot or listen to an episode. 

Different audiences absorb research differently, and the unit builds for that rather than assuming one format fits everyone.

Why people trust it more than ChatGPT

Eva stands for AI plus verified, and that name is the whole design brief. A general-purpose tool like ChatGPT can write a fluent, confident-sounding summary of anything. What it can’t do is guarantee that summary is actually grounded in a specific, checkable Bank report.

“LLMs can generate text very well, but the issue is people can’t trust this text.”

—Mohamad Chatila, Behavioural Scientist at World Bank Group

Eva verifies on both ends:

  • On the way in: every answer is grounded in the Bank’s own research library rather than the open internet
  • On the way out: every claim carries a real citation back to the source report, so a researcher can check exactly where a number or finding came from before using it

Eva can also work in a Socratic mode, asking questions back rather than just handing over an answer. So, instead of a one-shot summary, a user gets pushed toward the specific angle their own work actually needs, and Eva can surface related reports they hadn’t thought to look for. This turns a single query into the start of new research rather than an endpoint.

Reviewing input and output is showing up in how people actually use it. Pilot participants said they trust Eva more than general-purpose tools like ChatGPT specifically for World Bank research, and roughly 80 percent said it helps them understand and synthesize that research more efficiently.

Two other numbers from the pilot back that up. 

A 96 percent satisfaction rate suggests people aren’t just trying Eva once out of curiosity, and a 63 percent follow-up rate- users coming back to keep working with an answer rather than treating it as a one-off- points the same way. Adoption like that is hard to fake with a demo that only looks good in a slide deck.

Fast doesn’t mean unreviewed

The podcast makes the same trade-off visible in a different way. Every episode is AI-generated using a tool called Wondercraft, and every episode passes through five separate human checkpoints before it gets published.

  • An AI drafts the outline first, and a producer has to approve it before anything else happens.
  • AI drafts the script next, and the producer reviews that too, checking tone as much as accuracy.
  • The report’s own author reviews the script and the first English audio cut, since they’re best placed to catch a claim that’s technically true but presented in a way the research never intended.
  • The episode is localized into the other six languages.
  • A native speaker reviews each version before it goes live.

Every step in that chain has a named human role attached to it, not a general review note tacked on at the end. Nothing moves to the next stage until the person responsible for that stage signs off.

“It’s not just about scale. It’s also about the speed through which we need to do something.”

—Noemi Caiazzo, Learning Analyst at World Bank Group

One early pattern is already showing up across the podcast’s episodes. Certain topics travel unusually well across every language the show is translated into: empowering women in growth, the middle-income trap, business readiness, jobs and work, and migration and development. 

The team calls this a directional signal rather than a firm conclusion this early, but it’s a useful clue about which research is actually landing with a global, multilingual audience.

The results, in specific numbers

Below are the confirmed results in figures, reported directly by the team:

  • More than 60,000 learners reached across 195 countries, with more than 5,000 forum and social interactions in under two years
  • Eva’s pilot tested with 116 people from 116 countries, logging more than 20,000 queries and engaging more than 2,000 professionals in evaluation, 95.4 percent of them from outside the Bank
  • Roughly 3 hours saved weekly per Eva user, with more than 80 percent saying it helps them understand and synthesize World Bank research more efficiently
  • 11,100 podcast downloads across 168 countries within a few months of launch, with the top episode ranking first in six of seven available languages

The podcast itself is public. You can listen to Accelerating Development directly, in any of the seven available languages.

What’s next, and the budget reality

The team’s next build is an AI tutor, integrated directly into their LearnWorlds courses, using the same Socratic questioning approach as Eva. It launches to 300 participants on a behavioral diagnostics course, asking questions rather than handing over answers, and nudging learners to reflect on why an answer was wrong rather than just marking it incorrect.

The stated goals for the next stage are specific: move beyond three caption languages, personalize the tutor’s response to each learner, and scale to 1 million users within two years.

The team gives three reasons for investing here specifically:

  • It’s a pedagogical bet, not just another certificate to hand out
  • It’s aimed squarely at learners who don’t normally get access to this kind of training, since the courses are free
  • It’s about closing language gaps that AI tools were previously too expensive to solve well

“We’re doing this with an insanely small amount for such a big institution.”

—Francesca Spagnoli, Senior Learning Specialist at World Bank Group

What Francesca says here matters most for a smaller training provider or nonprofit reading this and assuming a program like this is out of reach. Building Eva, the podcast, and the tutor was well within the budget and easily affordable for the team. Choosing tools deliberately and building review into the process, instead of skipping it, is what made all of this work.

Quality control, though, still takes real time:

  • Each course takes about two months to develop
  • Scripts are co-created with AI, then reviewed by the pedagogical team and subject-matter experts
  • Testing includes both topic-area users and outside reviewers, followed by a final internal test round

The AI speeds up the draft. It doesn’t skip the review.

That testing round specifically brings in people from outside the unit that produced the course, not just the same small group reviewing its own work, plus experts from other units within the Bank entirely. 

The point is to catch a blind spot that a single team, however careful, can easily miss in its own material, before it ever reaches the 60,000 learners that actually depend on it.

Scale further without losing what makes it trustworthy

The World Bank Group team’s own message is the one worth keeping: it was never just about reaching more people faster. It was about doing that without giving up the human review that makes the result trustworthy in the first place. Speed and scrutiny can coexist if the workflow is built to accommodate both.

Nothing about this required a Bank-sized budget, either. It required deciding upfront which checkpoints were non-negotiable, then building AI to speed up everything around them rather than replace them. That’s a decision any team can make regardless of size.

If you’re training staff, partners, or a global community across languages and regions, LearnWorlds’ nonprofit training platform is built for exactly this kind of multi-audience, multi-language reach. 

Start your 30-day free trial and see what your own content could reach once it’s built to travel.

Your professional looking Academy in a few clicks
Start FREE Trial

Organic Content Strategist at 

Kyriaki is the Organic Content Strategist at LearnWorlds, where she writes and edits content about marketing and e-learning, helping course creators build, market, and sell successful online courses. With a degree in Career Guidance and a solid background in education management and career development, she combines strategic insight with a passion for lifelong learning. Outside of work, she enjoys expressing her creativity through music.