play icon for videos

Impact evaluation · Practical guide

Impact Evaluation: Methods, Design and What You Can Conclude

Learn what an impact evaluation must compare against, which designs can supply that comparison, and which parts a program team can do before hiring an evaluator.

Sopact AcademyFREE PRACTICAL COURSE

Measurement and reporting: from agreement to report

Set your program’s results beside outside data that counts the same thing, and say plainly what that comparison cannot prove.

  • Match definitions before any comparison
  • Choose live query or load once
  • Keep names out of AI tools
Start the free lesson →

What is impact evaluation?

Impact evaluation estimates the change a program caused by comparing what happened to participants with a credible estimate of what would have happened to them without the program, called the counterfactual. Terms vary by field, so name the causal question and the design rather than relying on the label.

Nobody can observe the counterfactual directly, because the same people cannot both join and skip a program. Every design is a way of building a stand-in for it, and each stand-in rests on an assumption you should be able to state in one sentence.

THE SHORT VERSION

  1. An impact evaluation needs a comparison that stands in for life without the program; a rise from before to after is not one.
  2. Random assignment, difference-in-differences, matching and other designs build that comparison in different ways, each on an assumption you must check.
  3. A program team can fix definitions, keep one record per person and compare with outside data on the same definition; sizing the effect usually needs an evaluator.

How is impact evaluation different from monitoring and outcome evaluation?

Monitoring tracks agreed metrics during the year, outcome evaluation asks what changed for participants, and impact evaluation asks how much of that change the program caused. Each builds on the one before, so a weak monitoring record limits every evaluation after it.

Three kinds of evidence work, and who usually leads each
TypeQuestionEvidence it needsWho leads
MonitoringAre we on track against the agreement?Agreed metrics, collected all yearProgram team
Outcome evaluationWhat changed, and for whom?The same people measured over timeProgram team, often with a reviewer
Impact evaluationHow much change did the program cause?A comparison group or other counterfactual designEvaluator, with the program team

In a monitoring, evaluation and learning loop, you agree what to monitor from the reporting agreement and data dictionary, collect inside the work all year, evaluate what changed, for whom and why, then decide what changes next cycle. Impact evaluation is the deepest form of that third stage, and it only works if the first two kept one record per person.

Impact evaluation: Slide titled Monitoring, evaluation and learning on one record, all year. Four boxes on a loop around a center labelled One record, one ID per person: 1 Agree, what to monitor, from the agreement and dictionary; 2 Monitor, collect inside the work, all year; 3 Evaluate, what changed, for whom, and why; 4 Learn, decide what changes next cycle. Handwritten note: learning all year, not a report at the end.
An impact evaluation sits in the Evaluate stage and depends on the record the other stages keep. From the course Measurement and reporting.

The short video below separates outputs from outcomes, the distinction every impact evaluation rests on. Watch for how an outcome is worded so it can be measured the same way twice; an effect estimate is only as sound as that definition.

Video · Outputs, outcomes and how to word an outcome you can measure.
Watch on YouTube ↗

Why isn’t a before-and-after change enough?

Because other things move at the same time: the job market, who enrolled, how the outcome was counted and people’s own progress can all change a before-and-after number without any help from the program. The comparison has to separate those from the program’s part.

Sopact’s course follows a fictional regional workforce fund and its four job-training partners. This year, Partners A, B and D placed 100 people in a job within 90 days of exit. If next year’s count is higher, hiring may have picked up across the region, or a partner may have enrolled people who were already closer to work.

The count itself can shift too. Partner C reported 55 placements using a six-month window instead of the agreed 90 days; a change like that, unnoticed, looks like a rise in results. People who enroll at a low point also tend to improve anyway, a pattern called regression toward the mean.

What are the main impact evaluation designs?

Five quantitative designs cover most impact evaluations of social programs; each builds the counterfactual a different way and each depends on an assumption that can fail. The last column shows what each might look like for the workforce fund.

A starting overview, not instructions for implementing a statistical design
DesignBasic ideaAssumption or limitFor the workforce fund
Randomized comparisonA lottery decides who gets a placeAttrition and spillovers still matterOversubscribed places filled by lottery
Difference-in-differencesCompare changes over time across groupsBoth groups were on parallel trendsTrainees vs similar job seekers, before and after
Regression discontinuityCompare people near an assignment cutoffOnly the program changes at the cutoffAn eligibility score decides entry
Matching or adjustmentCompare with people alike on measured traitsUnmeasured differences can still bias itMatch on age, prior work and education
Interrupted time seriesLook for a break in a repeated seriesNothing else changed at the breakMonthly placements before and after new mentoring

Work with an evaluator who can judge feasibility, sample size and assumptions; a sophisticated method on unsuitable data does not produce a credible estimate. The World Bank’s Impact Evaluation in Practice explains counterfactuals and these designs in more depth.

The video below compares two of the designs most teams weigh, following the same people over time or comparing groups at one point, and asks which one a funder will accept as evidence that the program caused the change.

Video · Longitudinal or cross-sectional: which design shows your program caused the change.
Watch on YouTube ↗

What do interviews and theory-based evidence add?

They explain how and why an effect happened, or did not, which a single estimate cannot show. Interviews, observation and documents can reveal whether mentoring reached people, why one partner placed more graduates, or where an expected link broke.

Theory-based approaches such as contribution analysis test the program’s pathway against rival explanations, and they need the same discipline as a statistical design. A positive story assembled from supportive quotes is not evidence of effect. The guide to attribution vs contribution covers that approach step by step.

What can a program team do without an evaluator?

A program team can do the groundwork that decides whether an impact evaluation is possible at all: agreed definitions, one record per person, honest follow-up counts and comparisons on the same definition. Sizing the effect is the evaluator’s job.

Who usually does what
TaskProgram teamEvaluator
Define outcomes and time windowsLeads, in the agreement and dictionaryReviews
Keep one ID per person across wavesLeadsDepends on it
Count who is missing at follow-upLeadsAdjusts for it
Compare with earlier cohorts or outside dataLeads, labelled as contextAdvises
Choose a comparison groupContributesLeads
Estimate the effect and its uncertaintyReviewsLeads

Definitions come first. “Placed in a job” means within 90 days of exit and “retained” means the same job at 12 months; write both in a shared data dictionary before collection, or no later comparison will hold.

In Sopact Sense, a persistent unique ID from the first form links each person’s intake, placement and follow-up answers on one record, so an evaluator starts from matched records rather than merged spreadsheets. Public data loaded as an ordinary survey follows the same rules as your own.

How can outside data help, and where does it stop?

Outside data on the same definition gives your result a fair reference point, such as the local wage for the same job, but it describes other people, not your participants without the program. Treat it as context beside your results.

The fund’s agreement asks for starting wage, hourly, by track. Each track trains for an occupation with a standard code, and the Bureau of Labor Statistics publishes hourly wages by occupation and area. Since new hires start low, set graduates beside the 25th percentile as well as the median, and name the area and reference year.

Impact evaluation: Slide titled Compare with outside data, two ways. Left panel, live query, nothing copied: HubSpot and Sopact Sense both connect to Claude or ChatGPT, answering Which employer partners will hire again, and how long do our graduates stay? Right panel, pull in once, compare anytime: Bureau of Labor Statistics prevailing wage by occupation and area; Department of Labor OSHA and wage enforcement records; IRS Form 990 via ProPublica, status and finances of applicant charities; loaded like any survey, under the same rules. Footer: live for everyday questions, loaded for benchmarks.
Query outside systems live for everyday questions; load public data once for reference figures everyone reads the same way. From the course Measurement and reporting.

A comparison is fair only when both numbers count the same thing, for the same kind of people, place and period. Partner C’s single average starting wage cannot join the comparison, because it mixes tracks. The chapter Compare your results with outside data shows how to write both definitions in one row first.

Where it stops: graduates earning above the local 25th percentile is worth reporting, but people who enroll may differ from the typical worker in that job. Only a design that compares similar people can say what the program added.

How do you plan an impact evaluation with an evaluator?

Plan before enrollment opens, because most designs need the comparison group, the baseline and the outcome windows fixed before the first participant arrives. A design added after the fact usually falls back to before-and-after.

Check that the design is feasible, fair to participants and useful for a real decision; never delay needed support to create a convenient comparison. Write down outcomes and analysis choices in advance, keep negative findings, and record any change to the plan.

PROMPT · PASTE INTO CLAUDE, CHATGPT OR YOUR AI TOOL

Help me prepare a one-page brief for an evaluator. Our program: [WHAT WE DO, FOR WHOM, WHERE]. The decision the evaluation should inform: [DECISION]. Our outcome definitions, with time windows: [DEFINITIONS]. What we collect today, and when: [SOURCES AND WAVES]. Follow-up counts so far: [INVITED, ANSWERED]. No names.
1. Write the causal question in one sentence.
2. List two or three designs that might fit, with the assumption each depends on and what would make it infeasible for us.
3. List what we must collect or fix before enrollment for each design.
4. List the questions we should ask the evaluator.
5. Use only facts I gave you. Do not invent numbers, sample sizes or effects. Where information is missing, write "not in our data".

What can you conclude from an impact evaluation, and how should you say it?

State the population, the comparison, the outcome, the period and the uncertainty, and keep observed change and estimated effect in separate sentences. Say whether it would hold in another setting.

Hold the limits in view as you write. Outputs such as sessions delivered are not outcomes. A before-and-after change is not proof of cause without a comparison. Self-reported answers carry recall and courtesy bias, and the people who stop answering follow-ups may be the ones for whom the program worked least.

AI can organize source material and draft wording, but it cannot supply a missing counterfactual or rescue an unsupported claim. Check every line it writes against the records, and let a named person decide what the evaluation concludes. For the wider planning process, see program evaluation.

Start with one evaluation question this year

Pick one outcome a funder or board keeps asking about, and get it ready for an impact evaluation, whether or not you commission one this year.

  1. Write the causal question: which outcome, for whom, over what period, compared with what.
  2. Check that the outcome has one definition and time window in your agreement and dictionary.
  3. Make sure each participant carries one ID from the first form to the last follow-up.
  4. Count invited, answered and missing at each follow-up, and look at who is missing.
  5. Find one outside figure or earlier cohort on the same definition, and label it as context.
  6. Run the prompt above and take the brief to one evaluator for a conversation.

After the first cycle you have a sharp causal question, matched records, one honest comparison and a view of which design, if any, is within reach.

Frequently asked questions

What is a counterfactual in impact evaluation?

The counterfactual is what would have happened to participants without the program. It can never be observed directly, since the same people cannot both take part and not take part. Evaluations estimate it with a comparison: a randomly assigned control group, a matched group, people near an eligibility cutoff or a trend before the program started. The credibility of the whole result rests on that comparison.

Is outcome evaluation the same as impact evaluation?

No. Outcome evaluation describes what changed for participants, such as how many were placed in a job within 90 days of exit. Impact evaluation asks how much of that change the program caused, which needs a comparison. Many reports labelled impact evaluations are outcome evaluations; that is fine, as long as the report states the design it actually used.

Are randomized trials always possible?

No. Random assignment needs more applicants than places, enough people for a stable estimate and a setting where a lottery is fair. It may be wrong to withhold a service people urgently need. Where a trial does not fit, difference-in-differences, regression discontinuity or matching may, each with its own assumptions stated openly.

When is a program ready for an impact evaluation?

When delivery is stable, outcomes have agreed definitions and time windows, each participant carries one ID across waves, and follow-up reaches most people. A program still changing its model every quarter is better served by monitoring and outcome evaluation first. Evaluating too early measures an unfinished program.

Can qualitative evidence support causal claims?

Yes, when it is gathered systematically to test the program’s pathway and rival explanations. Interviews can show whether mentoring reached people and why results differed across partners. It should include people with poor outcomes and should not be reduced to selected testimonials. Many evaluations combine it with a quantitative design.

Does AI make an evaluation causal?

No. Causal credibility comes from the question, the design, the evidence and the assumptions, not from the tool that summarizes the data. AI can help organize records, theme open answers and draft a brief for an evaluator. A person still needs to check each output against the source records and decide what the evaluation can conclude.

Explore Impact Measurement →