What NGOs like Educate! and Youth Impact can teach us about improving labor market programs

Source: Optimizely

Even evidence-based design of policies and programs often follows a “best guess” approach to some extent. But there is a better way. A growing number of organizations use rapid learning evaluations (aka A/B testing) for iterative and adaptive programming – similar to leading tech firms. Adopting this kind of evaluation can also help increase the impact and cost-effectiveness of labor market policies and programs.

By Kevin Hempel | February 2025

A few years ago I wrote about the potential of rapid learning evaluations (also known as “Nimble RCTs”, “Rapid Fire Testing”, or “A/B testing”) in international development cooperation. The core idea behind this type of evaluation is to iteratively test design options throughout the program cycle, using small-scale rigorous evaluation methods. These evaluations do not seek to measure whether a program’s ultimate objectives are achieved. Instead, they provide evidence for optimizing specific aspects of a program’s design and delivery (e.g., outreach, curriculum, delivery channels). In doing so, they offer timely and practical insights to implementing staff on how to make programs (more) effective.

Last year, I was thrilled to see several articles and reports from organizations sharing their experience with such rapid learning evaluations. I firmly believe that we need more of these evaluations in development cooperation in general and in the field of labor market policies where I work. So, let’s take a look at what these organizations have been doing and what we can learn from them…

Moving beyond “best guess” project design

As a starting point, we must admit that we don’t have all the answers when designing a new policy or program or making changes to an existing one. Even as more rigorous evidence is emerging on what works and what doesn’t, there remain many unknowns.

Let’s take the example of a job training program, one of the most popular interventions to improve people’s access to jobs and higher incomes. We have already learned a ton about making these programs more successful, such as designing training curricula based on employer needs, providing practical on-the-job experience, and offering follow-up support, especially for disadvantaged groups. That said, we don’t have a perfect recipe for all contexts and target groups. In practice, there are countless options for program design and implementation, for example regarding outreach, training content, pedagogy, incentives for participants, and more (see Figure 1).

Figure 1: The “design space” for job training programs

How do we typically deal with all these options? We make a proposal based on what we know or believe, followed by discussions and negotiations among different stakeholders (e.g., funders, implementing agencies, consultants, etc.). The agreed-upon program design is then written up in a project document. These decisions about design and delivery become “locked in”. We may revisit them during a mid-term evaluation or when planning the next phase of the program. Sometimes we may make course corrections during a program, but usually, that is only when something is going strikingly wrong. It is a very reactive way of programming…

There must be a better way…and there is. Instead of a “best guess” approach, we could explicitly acknowledge upfront what we don’t know and put in place a structured process to test different alternatives. Lant Pritchett and colleagues at Harvard University called this process “structured experiential learning”.

If we are unsure whether it is more effective to do beneficiary outreach in person or via text messages or email, we can test it. If we are not sure whether offering transportation or childcare assistance to training participants increases take-up and reduces drop-out, we can test it. This kind of structured (proactive!) learning through rapid evaluations is not about whether or not the program should exist. It is about optimizing key ingredients of the program in our specific context to maximize effectiveness.

Case study 1: Educate!

Educate! is a youth employment organization in East Africa that offers a variety of programs for in-school and out-of-school youth. They have been championing rigorous impact evaluation for a long time, working with leading researchers to rigorously evaluate their programs. With this experience, they are also acutely aware of one of the downsides of traditional impact evaluation. As the team shared in a recent blog article: “Some evaluations can take years to return results, meaning insights come too late to make changes during critical stages of program development.”

When Educate! embarked on a new program for out-of-school youth a few years ago, they wanted to build a tailored intervention that addresses the unique barriers this demographic faces. They knew they didn’t have all the answers on what the intervention should look like. So, to ensure they were on track, they wanted to test their assumptions and key questions of program design in real-time.  

This meant embedding rapid evaluation within the program development process, which led Educate! to develop its own Rapid Impact Assessment (RIA) system. While previous evaluations took years to complete, Educate! was now able to test specific questions of program design in a matter of months, answering questions such as: “What schedule is easiest for young people out of school to participate in, especially young women and girls?

Similarly, when the COVID-19 pandemic hit, Educate! had to adapt its in-school programs to distance learning. But rather than waiting for a pilot program to end and impact evaluation results to become available, and then making updates, they set up rapid evaluations to generate real-time data that helped them make iterative improvements almost immediately. For instance, one A/B test showed that calling participants’ caregivers to tell them about Educate!’s remote learning activities could increase youth participation by 29%, so they integrated this outreach approach into their new model.

Case study 2: Youth Impact

Youth Impact is another exciting example of an organization championing rapid evaluations. Based in Botswana, the organization delivers health and education programs for young people, by young people. Youth Impact is not your average NGO. Others have described the organization in a way that it “bears as much resemblance to a data-driven tech firm as it does to a research institute and a social service implementer.”

A recent report shares some of Youth Impact’s experience with rapid experimentation in education, having used A/B testing for over 7 years. Like Educate!, Youth Impact already had experience with traditional impact evaluation, and like Educate!, they wanted to generate evidence in a more timely manner: “As we considered next steps for the programme, we knew we wanted to tweak programme components to improve impact and scalability. We wanted to test these changes rigorously, but we wanted to do it quickly, at low cost, and to be able to iterate and optimise over time.” So Youth Impact repurposed its M&E system to get it ready for A/B testing. Since 2019, they have run over 40 A/B tests, with a staggering 12 in the first half of 2024 alone. Wow!!!

For example, they ran an A/B test to improve the targeting of their remedial education program (Teaching at the Right Level). In Math, students were typically grouped based on their ability to add, subtract, multiply, divide, or do no operations at all. This was the standard program design – version A. Youth Impact wanted to know if they could improve learning outcomes by also grouping students based on their ability to recognize and interpret larger numbers (i.e. putting students in different groups based on their ability to recognize 2, 3, or 4-digit numbers). This targeting tweak became version B. Collecting data from over 1,000 students across 4 regions in Botswana, they found significant impacts on learning outcomes from version B, at minimal extra cost (just a few cents per student). 

Way forward

I fully acknowledge that proper A/B testing might not be realistic for every organization or program. For example, the program must be sufficiently large, and the M&E system must be mature enough to collect quality administrative data. In contexts where implementation capacity is weak, it can also be challenging to implement several versions of a program simultaneously.

In general, however, the above examples highlight the promise of rapid learning evaluations across sectors, including employment and labor market programs. They are not a substitute for traditional impact evaluation and other evaluation methods, but an important complement to optimize program design and delivery. These evaluations can offer practical insights in real-time. This means that program managers and implementing staff can really benefit from the evaluation to make continuous improvements, rather than just seeing the evaluation as an accountability tool.

If we are serious about maximizing the impact and cost-effectiveness of labor market policies and interventions, rapid learning evaluations must become a standard tool in our M&E toolbox.

About the author:

Kevin Hempel is the Founder and Managing Director of Prospera Consulting, a boutique consulting firm working towards stronger policies and programs to facilitate the labor market integration of disadvantaged groups. You can follow him on LinkedIn.