What is the role of a theory of change in monitoring and evaluation?
In monitoring and evaluation, the theory of change is the map the plan is built from: each level gets an indicator you monitor during delivery, and each arrow between levels becomes a question the evaluation tests. Without it, M&E drifts toward counting whatever is at hand.
Monitoring follows the agreed indicators while the work runs. Evaluation asks why results moved, for whom, and how much the program contributed. Both read from the same model, so a gap in the model shows up as a gap in the evidence.
THE SHORT VERSION
- Give every output and outcome in the model one defined indicator, with a source and a due date, and agree those rows with your funder before collection starts.
- Monitor on that schedule all year and check each report against the agreement, so a changed definition or a missing number surfaces in the first quarter.
- Turn each arrow into an evaluation question, and once a year feed what you learned back into the model and the agreement.
How do you turn a theory of change into an M&E plan?
Give each output and outcome one indicator, attach its definition, source and due date, and agree those rows with the funder; that reporting agreement is the core of the M&E plan. Everything else in the plan, from baselines to review meetings, hangs off those rows.
Here is the fictional example used throughout the course Measurement and reporting: a regional workforce fund and one of its four job-training partners. On the onboarding call the partner said, “We want graduates in jobs that pay a living wage, and still there a year later.” The pathway: training and mentoring → certification → placement within 90 days → retained at 12 months → living-wage jobs.
| Level | Indicator and definition | Source | Due |
|---|---|---|---|
| Output | Participants enrolled: unique people, not visits | Enrollment form | Quarterly |
| Output | Completed training: finished the track | Completion record | Quarterly |
| Outcome | Placed in a job within 90 days of exit | Placement survey | Quarterly |
| Outcome | Retained: same job at 12 months | 12-month follow-up | Annual |
| Long-term aim | Starting wage: hourly, by track | Placement survey | Annual |
Inputs and activities carry no indicator here by choice: the funder monitors what changes for people, and the partner tracks sessions in its own records.
Each indicator then needs a full definition both sides count the same way: collection point, data type, disaggregation and source. The course chapter A shared data dictionary shows how to write that row. The rest of the plan fits in a short table.
| Element | What to record |
|---|---|
| Starting point | Baseline source, period and limits |
| Disaggregation | The splits you will report, such as by track or gender |
| Review | When results are read, and by whom |
| Responsibility | Who collects, who checks, who decides |
| Decision | Which finding would change the work |
Keep it proportionate: add a survey only where no existing record answers the question.
PROMPT · PASTE INTO CLAUDE, CHATGPT OR YOUR AI TOOL
Below is our theory of change. Draft a monitoring and evaluation plan from it. Rules: 1. List every level in order: inputs, activities, outputs, outcomes, impact. 2. For each output and outcome, propose one indicator with a definition (unit, time window, who counts), a source and how often it is due. 3. Under each arrow between levels, write one monitoring question and one evaluation question. 4. List the assumptions in our model and suggest a record that would show whether each holds. 5. Use only what is in our model and our reporting agreement. Do not invent targets, baselines or numbers. Where something is missing, write "not in our data". 6. Mark every indicator you proposed that is not in our agreement, so we can decide with our funder. Our theory of change: [PASTE] Our reporting agreement, if we have one: [PASTE, OR WRITE "NONE"]
Does M&E need a logframe instead of a theory of change?
No: a logframe is the same model laid out as a matrix, with indicator and means-of-verification columns added, so an M&E plan built from your theory of change already fills it. Many public and international donors ask for a logframe because its columns are the M&E plan.
The levels map one to one. Outputs stay outputs; the outcomes become the logframe’s purpose; the long-term change becomes its goal. The means of verification column is where each number comes from: the source column in the plan above.

Build the plan once on your own model and re-lay it for each template; the logframe guide walks through the matrix.
Which monitoring and evaluation questions come from each arrow?
Each arrow gives two questions: a monitoring question about whether the next level is appearing in the records, and an evaluation question about why, for whom and how far the program contributed. Indicators sit on the boxes; the questions sit on the arrows between them.
| Arrow | Monitoring question | Evaluation question |
|---|---|---|
| Training → certification | How many who enrolled finished the track? | Which parts of training help, and for whom? |
| Certification → placed within 90 days | How many graduates are placed within 90 days of exit? | Do employers recognize the certification? |
| Placed → retained at 12 months | How many are in the same job at 12 months? | What helps people stay, and what makes them leave? |
| Retained → living-wage jobs | What is the starting wage, by track? | Do these jobs lead to a living wage over time? |
A count of sessions tells you about delivery, not about the pathway. A handful of well-chosen questions on the arrows that matter most will teach you more than monitoring every box with equal effort.
Before you assign indicators, the introduction below is a useful refresher for the whole team. It walks a skills program from outputs to impact; notice which results get counted at exit and which need a follow-up months later.
What does monitoring against a theory of change look like during the year?
Monitoring means collecting inside the work on the agreed schedule and checking each report against the agreement, so a changed definition or a missing number shows up in the first quarter rather than at the annual report. The funder’s check is short: for each metric, does the report match, differ, leave it out, or fail to split it?
In the fictional example, Partner C’s report matched on enrollment: 80 unique people. Placements were counted “within 6 months” instead of within 90 days, retention at 12 months was missing, and starting wage came as one average instead of by track.
The fund sent three questions. “Can you count placements within 90 days, as agreed?” “When will 12-month retention be available?” “Can you split starting wage by track?” The chapter Check each partner report against the agreement shows the full check.
Monitoring also protects the roll-up. Partners A, B and D reported 42, 31 and 27 placements within 90 days, a portfolio total of 100. Partner C’s 55 used the six-month definition, so the fund held it until C confirmed its 90-day count, instead of adding it and overstating the total.
Per-person outcomes need the same person across time. In Sopact Sense, each participant keeps one ID from the first form, so enrollment, completion, the placement survey and the 12-month follow-up add to one record, and every line of an AI Assistant answer links to a record you can open.
How do you monitor the assumptions, not only the indicators?
Write each assumption as a condition you could observe, name the record that would show it, and pick the one that would most change your plan if it proved wrong. Indicators tell you whether results arrived; assumptions tell you why they did not.
In the example, “suitable jobs are within reach” matters more than which day workshops run: if it fails, placements stall however good the training is. A placement survey that asks “What got in the way?” and notes from employer contacts can show early whether it holds.
Watch for unexpected effects too: a push for quick placements may lower retention at 12 months. Evidence that challenges the model shows where the explanation or the delivery needs attention.
How does learning feed back into the theory of change?
Learning closes the loop: at a set point each year, or when evidence challenges an assumption, compare results with the pathway, agree one change, and record it in the model and the agreement. The loop runs on one record, so each cycle starts from evidence rather than from memory.

- Compare the year’s evidence with the expected pathway, level by level.
- Find the arrow with the weakest support.
- Consider delivery problems, assumptions and other explanations before deciding.
- Agree one change and the evidence you will need to judge it.
- Record the date, the reason and who approved it, and keep the earlier model with the reports that used it.
Do not quietly redraw the pathway so every result looks like success, and report misses along with gains. The CDC evaluation framework gives a broader structure for planning and using evaluation.
What can monitoring against a theory of change not tell you?
Monitoring shows whether expected results are appearing; it cannot show on its own that the program caused them. A rising count of completions says nothing about jobs, and more people placed this year than last may reflect a stronger job market, a different intake or a change in who answered.
A causal claim needs a design with a comparison, planned before collection; BetterEvaluation explains how to make the theory explicit for evaluators.
Follow-ups have gaps of their own: a graduate who has left a job may never answer the 12-month survey. Show response counts beside each outcome, and have a person check anything an AI tool drafts or flags against the records before it reaches a report.
Start with one pathway this quarter
Choose one program and its most important outcome, and run one full monitoring cycle against the model before you extend the plan.
- Write the pathway in one line, from main activity to long-term change.
- Give each output and outcome on it one indicator with a definition, source and due date.
- Agree those rows with your funder, or confirm them if an agreement exists.
- Write one monitoring and one evaluation question under the arrow you trust least.
- At quarter end, check the report against the agreement: matches, different, missing or not split.
- Hold a short learning review and record one change to the model or the plan.
After the first cycle you have an M&E plan tied to your model, one quarter of evidence read against it, and a named assumption you are watching.
Frequently asked questions
Does every box in the theory of change need an indicator?
No. Put indicators on the outputs and outcomes your decisions depend on, and on the outcome your funder cares most about. Inputs and activities can stay in program records. What you should not have is an outcome that matters with no indicator at all, because that is a claim nobody is checking. Note any important gap.
What is the difference between monitoring and evaluation in a theory of change?
Monitoring tracks the indicators on each level while the work runs, on a fixed schedule, and asks whether expected results are appearing. Evaluation works on the arrows: it asks why results moved, for whom, and how much the program contributed. Both draw on the same model and the same per-person records, so neither needs a separate dataset.
How often should we review the theory of change?
Once a year is a sensible default, timed so findings can shape next year’s plan and budget. Review sooner when evidence challenges an assumption, such as placements stalling although completion is high. Match the rhythm to when change can happen: quarterly checks suit placement within 90 days; retention at 12 months only moves once a year.
Can the theory of change change during implementation?
Yes, and it should when evidence shows an assumption was wrong. Record what changed, why and who approved it, and keep the earlier model with the reports that used it. Separate a revision driven by learning from a changed target, and tell your funder before the next report so the agreement changes with it.
Does monitoring prove impact?
No. Monitoring describes progress against the model and flags where to ask questions. Showing that the program caused a change needs an evaluation design with a comparison, such as earlier cohorts or outside data measured on the same definition, and even then the claim should say what the design can and cannot rule out.
