Rocket Learning
Conviction-buildingDiagnostics
AWW personas and journey map
Anganwadi workers are not one audience but three motivational personas, and what gates engagement is recognition and confidence: a concrete design target, not a generic "engage workers more."
3motivational personas, grounded in field data
Full study
The questionWho are AWWs, what motivates and blocks them, and which segments should Rocket Learning design for?
What we didAn ethnographic field visit plus 16 moderated interviews (Prayagraj, Bikaner), triangulated with baseline survey data, producing a motivators-and-barriers map, a journey map, and three personas.
What we foundMotivators run functional, emotional, and social; barriers are low pay, poor facilities, task overload, and weak credibility against private preschools. Three personas: Ambitious Seekers, Uncertain Achievers, Status-Quo Survivors.
Why it mattersIt gave Rocket Learning a shared language and a prioritised opportunity set, and set up the persona quantification that followed.
What would build on thisQuantifying the personas in a representative sample and confirming they predict engagement.
user research · segmentation
Source: AWW Research Findings deck
Rocket Learning
Conviction-buildingDiagnostics
AWW professional identity (WISE groundwork)
What AWWs most need, psychologically, is to be taken seriously as educators: "Can you take me seriously? Do I belong? Does my work matter?" That is exactly the identity a Wise intervention should affirm.
3core identity tensions to affirm: seen, belonging, mattering
Full study
The questionWhat are AWWs' psychological needs and identity, and which stories and prompts should a Wise intervention use?
What we didQualitative interviews and a focus group with AWWs across Prayagraj and Alibaug, exploring identity, parent perceptions, and values, and testing candidate stories and prompts.
What we foundAWWs hold a diffuse, multi-task role and rarely reflect on classroom quality. Because parents do not see them as education experts, they mimic private-preschool behaviours to signal legitimacy. They value status, growth, a stable government job, and impact.
Why it mattersIt pinpointed the identity to build, a credible early-childhood educator, and supplied the raw material for the intervention.
What would build on thisA test of whether the intervention actually shifts the targeted identity.
user research
Source: WISE Qualitative Findings
Rocket Learning
Accelerator 2025Diagnostics
AWW agency survey
AWWs report a strong sense of agency but act on it far less: nearly half had not engaged a parent in the past week, and engagement is predicted more by group dynamics than by individual traits.
180,736workers completed the survey, of 358k invited
Full study
The questionCould a large survey measure workers' sense of agency, and re-ground the personas in observable data?
What we didAnalysed the June 2025 agency surveys across five states, then re-derived personas from observable proxies for the 22,043 workers with that data.
What we foundA clear intention-behaviour gap: high reported intent, mixed behaviour, with nearly half reporting no agentic behaviour with parents in the past week. Engagement tracked group activity more than individual traits.
Why it mattersEngaging parents is the clearest place to intervene; group activity is a candidate lever; and it showed where the measure needs strengthening.
What would build on thisBehaviourally anchored items, and an experiment testing whether raising group activity raises engagement.
segmentation · measurement validation
Source: AWW Survey Analyses 2025
Rocket Learning
Accelerator 2025Diagnostics
Parent insights on enrolment and attendance
Parents see the Anganwadi as a place for care and nutrition rather than learning, and the real enrolment decisions are made by fathers and elders, so reach and regularity hinge on shifting those beliefs and reaching those decision-makers.
61parent conversations across 3 states (45 interviews, 16 focus groups)
Full study
The questionWhat drives parents' decisions to enrol children at Anganwadi Centres and send them regularly?
What we did45 one-on-one interviews and 16 focus groups with parents across Maharashtra, Uttar Pradesh, and Madhya Pradesh, structured through a five-dimension behavioural framework.
What we foundParents do not see play as learning. The mother is the daily anchor, but the father is usually the final decision-maker and elders can override. Anganwadis are chosen for being free, close, safe, and food-providing; private preschools for English and discipline.
Why it mattersIt names the specific beliefs to shift, and scoped the follow-on study on attendance regularity.
What would build on thisThe planned attendance-regularity study, triangulated with observed attendance data.
user research
Source: Parent Insights on Anganwadi Enrolment
Rocket Learning
R&DMeasurement
Remote ECD measurement (AIM-ECD): validation
A phone-based caregiver survey measures young children's development about as reliably as an in-person assessment, within 4.2 percentage points of the ground truth, making low-cost, high-frequency tracking feasible at scale.
4.2ppmean gap from in-person ground truth (not significant)
Full study
The questionCan a phone survey built on the World Bank's AIM-ECD framework measure child development reliably enough to replace in-person assessment?
What we didAdapted AIM-ECD into a phone caregiver survey for children aged 4 to 6 with Rocket Learning in Ghaziabad, surveying 1,229 parents and validating against in-person direct child assessments, triangulated with teacher reports.
What we foundCaregiver reports aligned closely with direct assessments, a mean difference of 4.2 percentage points, not significant. Literacy is the one area parents over-report; teachers over-report more than parents do.
Why it mattersProgrammes can replace costly annual in-person ECD surveys with a scalable phone survey, foundational for evaluating digital programmes at scale.
What would build on thisReplication across sites, validating the full item set in person, and refining the literacy items.
Pre-registered, AEA #0014554Co-authored: MIT, Zürich
measurement validation
Source: Remote-Scalable Measure of ECD
Rocket Learning
R&DExperiments
Caregiver-report framing experiment
Framing the survey to invite exaggeration did not move parents' answers: the AIM-ECD items resist social-desirability bias, so the tool can be deployed at scale without elaborate de-biasing.
0significant framing effects across 4 randomised arms
Full study
The questionDo parents inflate their children's development when the survey invites it, and does the framing need to be carefully controlled in the field?
What we didEmbedded a randomised survey experiment in the phone survey (N = 1,229), assigning parents to four framings: encouraging exaggeration, neutral, discouraging exaggeration, and a teacher-validation cue.
What we foundNo significant differences across framings. The specificity of the AIM-ECD items appears to suppress strategic misreporting. Teachers, by contrast, over-report more than parents.
Why it mattersThe measure is robust to how it is introduced, which simplifies large-scale administration, and it is methodological evidence that concrete items resist social-desirability bias.
What would build on thisReplicating the framing test in other populations and tying psychological items to verifiable behaviours.
Pre-registered, AEA #0014554
A/B test (survey experiment)
Source: Remote-Scalable Measure of ECD
Digital Green
Accelerator 2025Diagnostics
FarmerChat engagement and retention
FarmerChat's binding constraint is first-week retention, not acquisition: across three countries, nearly every new user goes quiet after their first week, so the highest-leverage investment is early re-engagement.
55,488farmers analysed across 496,585 questions
Full study
The questionHow engaged is FarmerChat's user base really, and do established users stay active?
What we didAnalysed Farmer.Chat product data: mapping the footprint, segmenting users by message volume, profiling tenure, measuring retention for mature cohorts, and classifying intent.
What we foundEngagement varies sharply by country: Kenya most engaged, India most one-and-done, Ethiopia best at moderate engagement. The headline is a first-week cliff: activity collapses between week zero and week one in every cohort.
Why it mattersIt redirects effort from acquisition to early retention, and the country contrasts argue for localised approaches, located from existing product data with no new collection.
What would build on thisDefining activation and retention windows prospectively, then a targeted re-engagement experiment.
segmentation · retention / cohort analysis
Source: DG FarmerChat Diagnostics
Jacaranda Health
Accelerator 2025Measurement
Can a clinic stand in for a mother's SES?
A mother's enrolling clinic cannot stand in for her individual vulnerability: most variation sits within clinics, not between them, so scalable personalisation needs a few individual questions, not a facility average.
1,116clinics, yet variation is mostly within, not between, them
Full study
The questionJacaranda's PROMPTS platform collects almost no individual data; the only consistent identifier is the enrolling clinic. Can clinic-level information proxy individual status at scale?
What we didUsing the Pathways Survey, analysed about 10,000 observations across 1,116 clinics, built a composite vulnerability index from eight binary markers, and tested how much variation sits between clinics versus within them.
What we foundThe index behaves sensibly, but most variation is within clinics rather than between them, and the same holds for counties, so clinic averages are too imprecise to substitute for an individual mother's circumstances.
Why it mattersIt rules out a tempting shortcut before it could be built into the product, and hands Jacaranda an interpretable set of onboarding questions.
What would build on thisValidating the index against health-seeking behaviour over time, and broadening sampling to poorer counties.
proxy-metric development · measurement validation
Source: Clinic Proxies, TAF and Jacaranda
Precision Development
Accelerator 2025Experiments
TarunBot: AI vs human advisories
AI-generated advisories delivered through TarunBot held farmer engagement at least as well as PxD's human-written status quo, and the strongest variant beat it, clearing the way to generate advisories with AI without losing engagement.
+3.2%pickup vs human advisories (AI + human-voice arm)
Full study
The questionDo AI-generated advisories engage farmers as well as the human status quo, and does the voice matter?
What we didA pre-registered, stratified, individually randomised A/B test among coffee farmers in Karnataka (Nov 2025 to Jan 2026), split across three arms: Human, AI advisory with AI audio, and AI advisory with a human-recorded voice. Analysed against a 3% minimum effect at 90% confidence.
What we foundBoth AI arms beat the human baseline on pickup. AI with a human-recorded voice reached 58.1% vs 56.3% for human, a +1.8pt gain (+3.2%), clearing the threshold. AI with AI audio reached 57.4%. Pairing AI-written content with a human voice performed best.
Why it mattersPxD can move advisory generation to AI without sacrificing engagement, and likely lift it: a major scale-and-cost unlock and a foundation for the planned personalisation classifier.
What would build on thisListening depth, comprehension, and recall, and whether the engagement gain carries through to farming behaviour.
Pre-registered A/Bn = 28,796 farmers
A/B test
Source: TarunBot A/B Testing
Educate!
Conviction-buildingExperiments
Does a nudge lift endline survey response?
A simple pre-survey nudge produces a small but consistent lift in endline response: promising enough to adopt and re-test, and a reminder that cheap operational levers can move the data-quality metrics evaluations depend on.
76→78%endline response, control vs nudged (directional)
Full study
The questionDoes a pre-survey nudge raise the share of SkillUp participants who complete the endline?
What we didAnalysed an A/B test embedded in SkillUp VII data collection: participants randomised to nudge or no-nudge across baseline modes, with the effect estimated by logistic regression.
What we foundA modest, positive, not statistically significant lift: about a 14% increase in the odds of responding unadjusted (roughly 76% to 78%), stable across specifications.
Why it mattersNudging is a cheap, plausibly real way to lift response and reduce attrition, worth adopting provisionally and re-testing on a partner's own data pipeline.
What would build on thisA pre-registered, adequately powered re-test.
A/B test
Source: Educate! nudge and survey-length A/B test
Educate!
Conviction-buildingExperimentsFinding in progress
Does a shorter endline form improve response?
The same trial randomised a long versus short endline form alongside the nudge, testing whether a lighter survey lifts completion without sacrificing measurement: the cheapest lever on respondent burden a phone-survey programme has.
Long vs shortendline forms randomised within the same trial
Full study
The questionDoes a shorter endline survey raise completion, and at what cost to the data collected?
What we didWithin the same SkillUp VII A/B test, participants were randomised to a long or short version of the endline form, allowing the survey-length effect to be estimated alongside the nudge.
What we foundBeing finalised. The survey-length effect is being separated from the nudge effect in the current analysis cut; figures will be added once isolated.
Why it mattersSurvey length trades respondent burden against measurement; quantifying that trade-off is directly useful to any phone-survey programme.
What would build on thisFinalising the long-versus-short comparison on both response rate and item-level completion.
A/B test
Source: Educate! nudge and survey-length A/B test
9dots
Conviction-buildingDiagnostics
Are teachers delivering the target dose?
The binding constraint for 9dots is lesson delivery, and it is falling: fewer teachers hit the 20-lesson target each year, and the shortfall sits in a growing tail of individual teachers, so support should target teachers, not just school leadership.
57.5→44%of teachers hitting the 20-lesson target, year on year
Full study
The questionAre teachers hitting the target of at least 20 lessons a year, how much does delivery vary, and does anything predict who delivers more?
What we didAnalysed teacher-delivery records across two school years: average lessons taught, share meeting the target, within- and across-school variation, and the relationship between teachers' mindset and lessons delivered.
What we foundDelivery declined: average lessons fell from 18.7 to 16.4, the share meeting the target dropped from 57.5% to 44%, and the share teaching fewer than ten rose from 20% to 30%. Variation is wide within as well as across schools.
Why it mattersThe problem is dosage and its decline, concentrated in a tail of low-delivery teachers, so target individuals. A light analysis of existing records can quantify fidelity before money goes to a fix.
What would build on thisStrengthening the mindset measure and testing a targeted intervention on the low-delivery tail.
dosage / fidelity analysis · segmentation
Source: 9dots diagnostic
Avanti Fellows
R&DExperiments
Grade 11 transitional-stories RCT
The Grade 11 intervention reliably shifts the very attitudes it is built to move, belonging and a stress-is-enhancing mindset, immediately after delivery; whether those shifts become durable exam-score gains is the open question the endline is designed to answer.
d = 0.50on stress-mindset; exam effect still null
Full study
The questionDoes a brief transitional-stories intervention improve Grade 11 students' exam performance and academic attitudes?
What we didA pre-registered RCT across Jawahar Navodaya Vidyalaya schools, individually randomised within strata. Pre/post exam composites analysed with ANCOVA, plus an immediate-post six-scale attitudinal battery, with balance, attrition, and robustness checks.
What we foundOn exam scores, the effect is small and not yet significant. On attitudes, it moved two target constructs, belonging and stress-is-enhancing mindset (d up to 0.50), both clearly significant. Pre-trends are parallel; attrition does not differ by arm.
Why it mattersThe intervention demonstrably shifts proximal attitudes; the live question is whether those translate into durable academic change.
What would build on thisThe confirmatory analysis and the endline wave, to test whether proximal shifts become lasting gains.
Pre-registered RCTWith Stanford
impact evaluation (RCT)
Source: JNV Wise Intervention RCT, preliminary (Avanti, Stanford, TAF)
Avanti Fellows
R&DExperiments
Grade 12 self-affirmation RCT
The Grade 12 self-affirmation intervention moved none of its six target attitudes and produced no detectable exam gain in this cut: an honest null that says, for older exam-focused students, this lighter touch may not be enough, and tells us where to look next.
0 of 6attitudinal scales moved; one gender signal to follow
Full study
The questionDoes a brief self-affirmation intervention improve Grade 12 students' exam performance and academic attitudes?
What we didThe Grade 12 arm of the same pre-registered RCT, individually randomised within strata, pre/post exam composites via ANCOVA plus the immediate-post attitudinal battery, with the same robustness checks.
What we foundThe exam effect is small and not significant, and none of the six attitudinal scales moved. The one signal is a gender interaction (positive for girls, roughly null for boys), to carry into the confirmatory analysis.
Why it mattersA clean null is itself a finding: it directs attention to whether the dose or framing fits an older, exam-pressured cohort, and to the gender heterogeneity.
What would build on thisThe confirmatory analysis, the endline wave, and a closer look at the Grade 12 gender interaction.
Pre-registered RCTWith Stanford
impact evaluation (RCT)
Source: JNV Wise Intervention RCT, preliminary (Avanti, Stanford, TAF)
Shamiri Institute
Conviction-buildingDiagnostics
What drives lay-provider outcomes?
Scaling a lay-provider model is an allocation-and-supervision problem, not a screening problem: training-test scores did not predict student improvement, and "like-with-like" demographic matching was unhelpful, even counterproductive for one pairing.
4,406students; test scores did not predict who improved
Full study
The questionWhich provider characteristics or provider-student matches actually improve outcomes, and how should providers be recruited and allocated at scale?
What we didA secondary analysis of Shamiri's school programme using random-forest regression to surface drivers, then multivariable regression (school-clustered SEs) to test four pre-specified hypotheses on depression and anxiety outcomes.
What we foundTraining-exam scores and fidelity did not predict improvement; attendance did not independently predict outcomes. "Like-with-like" matching was not beneficial, and female-provider/female-student pairings were associated with significantly worse anxiety (linked to co-rumination). Matching and supervision mattered more.
Why it mattersRecruiting on academic test scores and defaulting to demographic matching may be the wrong playbook; the higher-leverage levers are supervision and smarter allocation.
What would build on thisExperimentally testing the allocation rules, for example randomising strategic mismatches.
Co-authored: Harvard, OxfordWith IDinsight
driver analysis (ML + regression)
Source: Do Lay-Provider Characteristics Predict Psychotherapy Outcomes? (IDinsight, Shamiri, Harvard, Oxford, TAF)
Kabakoo Academies
Accelerator 2026Diagnostics
Where does the onboarding funnel lose learners?
The onboarding flow is not where Kabakoo loses people: the cliff is before it. More than half of invited learners never click the first button despite reminders, while those who start finish about 60% of the time, so the fix is the gap between in-person registration and the first message.
55.6%of invited learners never responded to a single onboarding message
Full study
The questionAfter a high-trust, in-person recruitment drive, only 27% of invited learners completed WhatsApp onboarding (against a 60% target). Where in the funnel are they lost?
What we didAnalysed the WhatsApp chat log from the March 2026 "Opération Missabougou" campaign in Bamako, 20,165 message events across 853 learners, building a step-level onboarding funnel and profiling the non-responders.
What we foundA single large cliff at step one: of 826 invited, only 44.4% clicked "Let's go!"; 55.6% never sent a single message, despite about five reminder templates over two weeks. Past that first click, attrition is gradual and completion runs about 60%.
Why it mattersThe flow is fine; the friction sits between face-to-face registration and the first asynchronous WhatsApp contact, pointing to motivation decay, prompt misalignment, and a trust gap.
What would build on thisLinking registration timestamps to the chat log to test the registration-to-invitation delay, and the dropout phone interviews to prioritise among the three mechanisms.
onboarding funnel analysis · segmentation
Source: Kabakoo Onboarding Funnel, Diagnostic and Proposals