play icon for videos

Impact Evaluation: Methods, Frameworks & AI Tools 2026

Impact evaluation methods, frameworks, and AI tools explained. See how AI-native platforms like Sopact Sense reduce evaluation analysis from months to minutes.

Updated
July 21, 2026
360 feedback training evaluation
Use Case

What is impact evaluation?

Impact evaluation is the systematic assessment of whether a program caused the changes observed in the people it served — not just whether those changes occurred. It isolates the program's contribution from everything else happening in participants' lives, using a defined method, a comparison of some kind, and evidence collected against a baseline. An outcome shows change; an impact evaluation shows the change was caused by the program.

The reason impact evaluations are hard is not the statistics; it is the attribution claim underneath them. Every evaluation makes a claim about causation, and every claim carries a debt: the evidence you must have collected, on the same people, before you can defend it. Programs that decide the method at the end, after the data is in three disconnected tools, discover the debt is unpayable.

Key takeaways

  • Impact evaluation asks whether the program caused the change, not just whether change happened — which requires a method, a comparison, and baseline evidence on the same people.
  • Every evaluation makes an attribution claim, and every claim carries a debt: the evidence you must have collected before you can defend it. Decide the claim first, then the method.
  • Sopact calls that liability the Attribution Debt: the gap between the causal claim a program wants to make and the connected, baseline-anchored evidence it actually holds on one participant record.
  • Impact evaluation is not outcome evaluation. Outcome evaluation measures what changed; impact evaluation establishes the program caused it, against a counterfactual.
  • Build the evidence architecture before you collect. The method you can defend is decided by the identity and baseline you set up at intake, not by cleverness at analysis.

Name the attribution claim, then pay the debt.

An impact evaluation starts with a claim: this program caused this change for these people. That claim decides everything downstream — the method, the comparison, the data you must have. Choosing a randomized design demands one thing; a quasi-experimental or contribution-based design demands another. The mistake is choosing the method last, after collection, when the required evidence no longer exists.

Sopact calls the liability the Attribution Debt: the gap between the causal claim a program wants to make and the connected, baseline-anchored evidence it actually holds. The debt accrues at collection, when intake, exit, and follow-up land in different tools with no shared identifier; it closes when they land on one participant record. The chain that an evaluation prices is validated on the theory of change page, and the indicator format on logframe.

Where the debt shows up is predictable: a deadweight estimate with no comparison group, a confidence score that cannot be linked pre-to-post because emails changed, a qualitative signal that never reaches the analysis. The survey mechanics that prevent it are on the survey design page; the difference from measuring change alone is below.

Impact evaluation is not outcome evaluation.

The two are constantly conflated. Outcome evaluation measures what changed for participants against a baseline — did confidence rise, did employment follow. Impact evaluation goes one step further and asks whether the program caused that change, which requires a counterfactual: what would have happened without it. The methods that do this live on outcome evaluation, and the underlying distinction on output vs outcome.

The practical rule: most programs can and should run outcome evidence continuously, and reserve full impact evaluation — with comparison groups or randomized designs — for the questions and funders that require it. Over-claiming impact from an outcome study is the fastest way to fail external review; the wider practice of measuring well is on impact measurement.

How to run an impact evaluation, step by step.

You run an impact evaluation in five steps: name the attribution claim, choose the method that claim requires, build the evidence architecture before you collect, run analysis continuously instead of in an annual batch, and convert findings into decisions. The first two steps decide whether the last three are even possible.

The method is the choice most programs get wrong by making it last. The table pairs each common method with the claim it supports and the evidence it demands, so the choice is made before collection, not after.

Impact evaluation methods, by the claim they support.

The method follows the attribution claim: a stronger causal claim demands a stronger comparison and more evidence collected upfront. Read the last column — it is the Attribution Debt each method requires you to pay.

Impact evaluation methods
MethodClaim it supportsEvidence it requires upfront
Randomized (RCT)The program caused the changeRandom assignment; baseline on both arms
Quasi-experimentalLikely caused, vs a matched groupA comparison group; matched baseline covariates
Pre / post with benchmarkChanged vs a sector expectationBaseline + follow-up on the same people; a cited benchmark
Contribution analysisPlausibly contributed, given the theoryA theory of change; evidence for each causal link
Outcome harvestingOutcomes occurred and were connectedDocumented outcomes traced back to the program

Read the last column and the point is clear: the method does not decide the evidence at analysis; the evidence you set up at intake decides which method you can defend. Closing the Attribution Debt means building that architecture before the first response, on one participant record.

An evaluation delivered at year-end is a post-mortem. The Loop runs it continuously.

A batch evaluation run once at the end reports what happened too late to change it, and discovers its evidence gaps when they can no longer be filled. Running analysis as data arrives keeps the causal picture current and surfaces a broken assumption in weeks. That is the premise of the Loop, Sopact's method for continuous impact intelligence: collect clean at the source, analyze the moment data arrives, improve while you can still act.

The Loop is also what makes an evaluation defensible. Every figure traces back to the response and the baseline it came from, so an attribution claim resolves to its evidence. That standard has its own chapter in traceability and transparency.

One method, three moves that never stop

1 · CollectClean at the source; baseline and follow-up on one participant record.
2 · AnalyzeOn arrival; the causal picture updated, not batched at year-end.
3 · ImproveIn time to act; catch a broken assumption mid-program.

Then the cycle runs again, a little sharper each cohort. Read the method: the Loop methodology →

Close the Attribution Debt this week

The fastest way to make an evaluation defensible is to name the claim and check the evidence it demands. Each prompt below pastes into Sopact Sense's Assistant, or reasons through with your team; the arrow above each links the Academy walkthrough that shows the expected output and the tips.

Academy walkthrough → Name the claim, check the evidence

For this program, name the attribution claim I want to make and the evidence it requires: [PASTE PROGRAM + FUNDER QUESTION]. Recommend the method that claim supports, list the baseline and follow-up data it needs on the same participants, and flag every piece I do not currently collect. Return a table: Claim / Method / Evidence needed / Have it?

Academy walkthrough → Trace the causal claim to evidence

For each claim in this evaluation, build a source row: the finding, the responses and baseline behind it, the comparison used, and the calculation. If a claim has no comparison or baseline, flag it as UNSUPPORTED. Return a table: Claim / Source / Comparison / Calculation. Claims: [PASTE]

Academy walkthrough → Separate outcome from impact claims

Review these claims and mark each as an OUTCOME claim (change occurred) or an IMPACT claim (the program caused it): [PASTE CLAIMS]. For every impact claim, state the counterfactual it rests on; if there is none, downgrade it to an outcome claim rather than over-claiming. Return the marked list with a reason each.

Academy walkthrough → Lock the evaluation dictionary

Turn this program's indicators into an evaluation data dictionary: [PASTE INDICATORS]. For each, give the definition, the wave schedule (baseline, mid, endline), the denominator rule, and what would invalidate a comparison between waves. Return a table.

Learn the how-to in the Academy

Each walkthrough is short and practical: what to do, the prompt to run, the output to expect, and the tips that keep it reliable.

Watch: what's changing in impact evaluation, and why the attribution claim is decided at collection, not analysis.

Frequently asked questions

What is impact evaluation?

Impact evaluation is the systematic assessment of whether a program caused the changes observed in the people it served, isolating the program's contribution using a defined method, a comparison, and baseline evidence. It goes beyond measuring change to establishing causation. In Sopact's framing, an evaluation is defensible only when it has closed its Attribution Debt — the gap between the causal claim and the connected evidence behind it.

What are the main impact evaluation methods?

The common methods are randomized controlled trials, quasi-experimental designs with a matched comparison group, pre/post with a cited benchmark, contribution analysis grounded in a theory of change, and outcome harvesting. Each supports a different strength of causal claim and demands different evidence upfront. Sopact keeps that evidence on one participant record so the method chosen at design time is defensible at analysis.

What is the difference between impact evaluation and outcome evaluation?

Outcome evaluation measures what changed for participants against a baseline; impact evaluation asks whether the program caused that change, which requires a counterfactual. Over-claiming impact from an outcome study is a common review failure. Sopact runs outcome evidence continuously and reserves full impact evaluation for the questions that require a comparison, keeping both on the same record.

What is an impact evaluation framework?

An impact evaluation framework sets the design: the attribution claim, the method, the comparison strategy, the indicators, and the evidence schedule. It is decided before collection, because the method you can defend is fixed by the identity and baseline you establish at intake. Sopact treats the framework as the architecture that closes the Attribution Debt rather than a document written at the end.

How do I choose an impact evaluation method?

Start from the attribution claim you need to make, not the data you happen to have. A stronger causal claim requires a stronger comparison and more baseline evidence collected upfront. Choosing the method last, after collection, is why evaluations fail. Sopact's Attribution Debt framing forces the claim and the method to be decided before the first response is collected.

What are examples of impact evaluation?

A workforce program using a matched comparison group to estimate employment gains net of the local labor market; a health program using pre/post readings against a sector benchmark; an education program using contribution analysis against its theory of change. Each pairs a claim with the evidence it requires. Sopact's examples keep that evidence traceable so the claim resolves to its source.

How is impact evaluation different from monitoring?

Monitoring tracks whether a program is delivering as planned, continuously; impact evaluation assesses whether it caused change, periodically and against a counterfactual. They are one system on two cadences, covered on the monitoring and evaluation page. Sopact keeps both on one participant record, so the monitoring data feeds the evaluation rather than living in a separate tool.

Do I need a control group for impact evaluation?

Not always. A randomized or quasi-experimental design needs a comparison or control group for the strongest causal claim; contribution analysis and outcome harvesting make weaker but still useful claims without one. The honest move is to match the claim to the evidence you can collect. Sopact makes the trade-off explicit by surfacing the Attribution Debt each method carries.

How long does an impact evaluation take?

A rigorous evaluation typically spans the program cycle plus a follow-up window, with months of design and data preparation if the architecture was not built in advance. Programs that set up identity and baselines at intake compress the analysis dramatically. Sopact keeps evidence connected on one record, so the evaluation is a continuous read rather than a year-end reconstruction.

Next: measure the change on the outcome evaluation page, or validate the causal chain on the theory of change page.