What is impact measurement?
Impact measurement is how an organization and its funders agree what change a program should produce, for whom, and which numbers will show it, then collect, check and report those numbers so people can act on them. It goes past counting activity to ask what changed, how sure you can be, and what to do next.
Most of the work is agreement, not arithmetic. When a funder and a funded partner (the grantee or investee) settle what each metric means before the first report, later reports can be compared, added up and checked.
THE SHORT VERSION
- Impact measurement starts with a written agreement on a few metrics and what each one means, not with a new survey or dashboard.
- Funders then check, add up and report on a portfolio, while funded partners keep one evidence base and shape it into each funder's report.
- Report what changed, what is missing and what the evidence cannot prove, and let people, not AI, decide what a number means.
How is impact measurement different from monitoring?
Monitoring counts what you delivered; impact measurement asks what changed for people and how confident you can be that the program contributed. Both matter, and the second depends on the first.
In the fictional workforce example used throughout this guide, training leads to certification, a job within 90 days of exit, the same job at 12 months and, over time, living-wage work. "Completed training" is an output; the job at 90 days and at 12 months are outcomes; living-wage work across a region is the impact. The OECD DAC calls impact the higher-level effects of an intervention, a larger claim than attendance or satisfaction.
Which impact measurement methods should you use?
Choose the method by the claim you need to make: a theory of change states the pathway, outcome tracking follows people over time, and contribution analysis or a comparison design is needed before you speak about cause. No single method is best.
Scroll horizontally to see all columns →
| Method | Use it when | What it cannot show alone |
|---|---|---|
| Theory of change | You need the pathway, assumptions and intended change written down before collecting | Whether the pathway held in practice |
| Outcome tracking | You follow people or cohorts over time on one ID | That the program caused each outcome |
| Contribution analysis | Several factors shape a result and you need a transparent account of your part | Experimental proof of cause |
| Experimental or quasi-experimental design | The stakes justify a causal design and a fair comparison is ethical and feasible | Why the result happened, in people's words |
| Qualitative and participatory inquiry | You need people's own account of what changed, and effects nobody planned | How common a result is across everyone |
Sources: OECD DAC, Applying Evaluation Criteria Thoughtfully · BetterEvaluation, Contribution Analysis.
Frameworks are layouts, not rival methods: a theory of change, logic model, logframe and results framework hold the same levels under different labels. One change model, four formats shows how to build once and re-lay.
Why are impact reports so hard to compare?
Because the funder and the funded partner rarely agree, in writing, what gets reported and what each number means, so every report arrives in its own format with its own definitions. A report is only as good as the agreement behind it.
A portfolio manager receives PDFs, spreadsheets, survey exports and emails; one partner counts a job placement within 90 days, another within six months. Meanwhile each funded partner reports to several funders, each with its own template, and rebuilds the same evidence every time.

The talk below is a Sopact introduction to impact measurement and management in the age of AI. As you watch, keep this section's point in mind: AI can only compare reports that rest on agreed definitions.
What are the steps from agreement to report?
Both sides share three steps (the reporting agreement, the change model and the data dictionary); then the funder checks, adds up and reports on a portfolio, while the funded partner maps every funder's ask to one evidence base and writes each report.
What you collect is set up earlier, when you design forms and give each person one ID, as the Foundations course covers. This journey starts at the agreement and ends at a report whose every number can be traced.

Scroll horizontally to see all columns →
| Stage | What you end up with | Chapter |
|---|---|---|
| Agree together | A reporting agreement from the call | Turn the onboarding call into a reporting agreement |
| Agree together | One change model in the funder's format | One change model, four formats |
| Agree together | One dictionary row per metric | A shared data dictionary |
| Funder road | Each report marked matches, different, missing or not split | Check each partner report against the agreement |
| Funder road | Totals of what shares a definition; the rest held | Roll up and benchmark portfolio results |
| Funder road | A board or donor report checked number by number | Write the portfolio report for your board or donors |
| Partner road | Each funder's ask mapped to a held metric | Map every funder's ask to one evidence base |
| Partner road | A short taste file per funder, plus your own | Capture each funder's taste, and your own |
| Partner road | Each report drafted, gaps flagged, checked | Write each funder's report with AI, then check it |
| Toolkit | Outside comparisons and a source per figure | Compare with outside data · Make every number match its source |
What does impact measurement look like in practice?
Follow a fictional regional workforce fund and its four job-training partners, A to D: the agreement names five metrics, one partner's report drifts from it, and only the numbers that share a definition are added up.
On the onboarding call, the fund and its partners agree five metrics: participants enrolled (unique people, not visits), completed training, placed in a job within 90 days of exit, retained in the same job at 12 months, and hourly starting wage by track.
Partner C's report shows 80 people enrolled, which matches, but counts placements within six months, leaves out 12-month retention and gives one average wage, not split by track. The fund sends one question per gap: can you count placements within 90 days, when will 12-month retention be available, and can you split wage by track?
In the roll-up, Partners A, B and D count placements within 90 days: 42, 31 and 27, a portfolio total of 100. Partner C's 55 used the six-month rule, so it is shown as held until C confirms its 90-day count.
From Partner C's side, its six-month number was written for another funder that asked for six months. C answers four templates, so the fix is one evidence base with one ID per person, and a report that says which definition each number uses.
Real teams work this way. AHA Ventures, the investment arm of the American Heart Association, is growing its portfolio from 19 to more than 40 companies; its head of impact drafts each company's logic model and data dictionary during the onboarding call with Sopact Sense. Auroville International USA funds about 80 project partners that report every quarter against shared terms.
On the partner road, Open Play Foundation's CEO, Marco Botha, put the payoff plainly: "It just took 2 prompts and my report was ready." The short video below, Turn Theory of Change Into Daily Decisions, shows a theory of change informing ongoing program decisions and evidence review, long before a report is due.
Can AI do impact measurement?
AI can draft, compare and flag gaps quickly, but only against definitions people have agreed and records it is allowed to read; it cannot decide what an outcome means or whether a program caused it.
The test is what AI does with a gap. In the course's demo dataset, a first draft flagged training hours delivered, job placements and starting wage by track as missing because the placements survey had not been selected as a source; once it was added, the draft was complete. A missing number should be flagged, never guessed.
Governance comes first; the NIST AI Risk Management Framework starts with Govern. In practice: a named owner for each definition, AI limited to declared sources, and a person who checks each number against its record.
In Sopact Sense, each person or partner keeps one ID from the first form, the AI Assistant stays locked until you pick which surveys it may use, and every line of an answer links to a record you can open. The prompt below works in any AI tool.
PROMPT · PASTE INTO CLAUDE, CHATGPT OR YOUR AI TOOL
You are helping me check one report against our reporting agreement. Agreement (metric, definition, due date): [PASTE THE AGREEMENT] Report: [PASTE THE REPORT] For each metric in the agreement, return one row: metric | agreed definition | what the report says | status (matches, different, missing or not split) Rules: - Use only numbers in the report. Do not invent, estimate or round any number. - If a metric is not in the report, write "not in our data". - If the report uses a different definition, quote its wording and mark it "different". - End with one plain question per gap.
What can impact measurement not prove?
Measurement makes numbers comparable; it does not make them true, and on its own it does not prove the program caused them. Outputs are not outcomes: 80 people enrolled says nothing yet about jobs. A before-and-after change shows movement, not cause, and a causal claim needs a comparison chosen before collection.
Self-reported wages depend on who answers, and people who stop responding often differ from those who stay, so report non-response beside every rate. A named person, not AI, decides what a result means.
How do you start impact measurement with one agreement?
Pick one funder relationship or one funded partner and run a single cycle from agreement to report before you change anything bigger.
- Record the next onboarding call, with consent, and note the change the program expects.
- Agree three to five metrics, each with a plain definition, due date and source, such as "unique people, not visits".
- Put each metric in a data dictionary and send it to the other side to confirm.
- When the first report arrives, mark each metric as matches, different, missing or not split, and send one question per gap.
- Add up only the numbers that share a definition, show what is held, and give every figure in the report a source note.
- Write down one change to the agreement or dictionary for the next cycle.
After the first cycle you have a confirmed agreement, a short dictionary, one checked report and fewer questions next time. The impact measurement and management guide shows how a funder carries this across a whole portfolio, and the impact reporting guide covers the report itself.
Frequently asked questions
What is the difference between monitoring and impact measurement?
Monitoring tracks delivery: sessions held, people enrolled, training completed. Impact measurement connects that delivery to outcomes, such as a job within 90 days or the same job at 12 months, and asks how confident you can be that the program contributed. Most teams need both, kept on the same records, because an outcome rate is only as good as the enrollment count under it.
What are the main impact measurement methods?
Common methods are a theory of change, outcome tracking over time, contribution analysis, experimental or quasi-experimental designs, and qualitative or participatory inquiry. Pick by the claim you need to make. A pathway needs a theory of change; change per person needs outcome tracking on one ID; a statement about cause needs a comparison planned before collection.
How do funders compare impact reports from different partners?
By agreeing each metric's definition before reports arrive, sharing one data dictionary with every partner, and checking each report against it. Only numbers that share a definition are added up; a number counted differently is held and clarified with one question, not added in quietly. Optional benchmarks then compare a partner with an outside number that uses the same definition.
How can we report to several funders without collecting everything twice?
Keep one evidence base with one ID per person, and map each funder's request to a metric you already hold. When a funder defines a metric differently, report their number and say how it differs, rather than changing your records. Each report takes that funder's format and words; the numbers underneath stay the same.
Do we need IRIS+ codes or an SROI ratio to measure impact?
No. IRIS+ codes are an optional standard that helps when funders compare across portfolios, so add one where a code fits your metric, such as PI4060 for total clients. SROI and other dollar values are optional too. They are worth the effort only when a reader will use the figure and your outcome evidence can carry it.

