play icon for videos

Monitoring and evaluation · Practical guide

Monitoring and Evaluation: Differences, Plan and Practical Example

Run monitoring and evaluation as one loop on one record: monitor against the agreed metrics all year, check each report, and evaluate what changed and why.

Sopact AcademyContinue learning

Measurement and reporting: from agreement to report

FREE PRACTICAL COURSE

Check each partner report against the reporting agreement, the monitoring step most M&E plans leave vague.

  • Compare each report metric by metric
  • Turn every gap into one question
  • Keep a gap log for next year
Start the free lesson →

What is monitoring and evaluation?

Monitoring and evaluation (M&E) is how an organization tracks whether its work is happening as planned and judges what changed as a result: monitoring follows agreed measures all year, and evaluation asks what changed, for whom and why. The two work best as one loop on one set of records, not as two reports written months apart.

Funders and funded partners usually meet in M&E through a reporting agreement: the few metrics both will count, each defined and dated. Monitoring checks every report against it; evaluation uses the same records for the questions an agreement cannot answer.

THE SHORT VERSION

  1. Monitor against what was agreed: each metric with its definition, period and source, checked every time a report arrives.
  2. Evaluate with questions, not more indicators: what changed, for whom, why, and what else could explain it.
  3. Keep both on one record per person and per partner, so every finding traces back to the report and record it came from.

What is the difference between monitoring and evaluation?

Monitoring tracks agreed measures routinely to show whether the work is on track; evaluation asks focused questions at chosen points to explain the results and judge whether to continue or change the work. Monitoring shows placements dropped at one partner; evaluation asks why.

Scroll horizontally to see all columns →

QuestionMonitoringEvaluation
Main purposeSee whether delivery and results are on trackExplain results and judge their value
Example questionDid every partner count placements within 90 days?Why do placements differ between partners?
EvidenceRoutine records, reports and scheduled surveysMonitoring records plus interviews or comparisons
TimingEvery time a report or record arrivesAt points chosen around a decision
Typical useFix gaps, adjust deliveryRedesign, continue or change funding

These are practical distinctions, not rigid walls: monitoring can include outcomes, and evaluation can inform decisions mid-program. BetterEvaluation's monitoring guidance describes this broader scope, and its guidance on evaluation timing explains why the two need to work together.

What is the difference between an M&E framework and an M&E plan?

An M&E framework sets out the questions, measures and approach across the work; an M&E plan says who collects and checks what, when, and from which source. Organizations use these labels loosely, so agree on the job each document does instead of keeping several copies of one table.

Scroll horizontally to see all columns →

DocumentWhat it helps you decide
Theory of changeWhy the work should lead to the change, and which assumptions matter
LogframeHow objectives, indicators, sources and assumptions fit in one matrix
Reporting agreementThe few metrics a funder and partner report, how each is counted, when
M&E frameworkThe overall questions, measures and approach across the work
M&E planWho collects and checks what, when, from which source

The four change-model layouts carry the same levels under different labels, so build the model once and re-lay it when a funder asks for another format; One change model, four formats shows how. BetterEvaluation explains how a framework spans several evaluations while a plan covers one.

Many M&E plans still use the logframe, often filled in once and never opened again. In this short video, watch for the means of verification column: it names where each number will come from, the first thing monitoring needs.

Video · Why many logframes fail, and how to make one work.
Watch on YouTube ↗

How do you build a practical M&E plan?

Start from the decision the evidence must support and the metrics already agreed with the funder, then define each measure, name its source and schedule, and plan the analysis before collecting anything new. The six steps below use one running example.

1. Start from the decision and the agreement

The example is fictional: a regional workforce fund with four job-training partners, A to D. It must decide whether each partner's training leads to jobs that last, and its reporting agreement names five metrics: enrolled, completed training, placed in a job, retained and starting wage.

2. Choose the unit you need to follow

Enrolled counts unique people, not visits, so every person needs one ID from the first form; a partner-level check needs one record per partner. Collect personal details only where a question needs them.

3. Define each measure and its source

Write down the calculation, period and source for every metric: placed in a job means within 90 days of exit, retained means the same job at 12 months, starting wage is hourly by track. Keep these in a shared data dictionary, and keep a recorded zero apart from a missing answer.

4. Schedule collection and review

Match collection to the agreement's rhythm: enrolled is reported quarterly, retention and starting wage annually. A metric not yet due is logged as "not due", not "missing". Name who requests, checks and responds to each report.

5. Plan the analysis before collecting more

Decide which groups you will compare, such as by gender, age and disability, and how you will treat missing follow-ups. A causal question needs an evaluation design, not routine reporting.

6. Budget for the whole loop and pilot it

Allow time for partners, checks, analysis and discussion, not only collection. Run one reporting cycle end to end and revise before asking every partner to take part.

Any AI tool can draft the plan's first table from your agreement; review every row with the metric's owner.

PROMPT · PASTE INTO CLAUDE, CHATGPT OR YOUR AI TOOL

You are helping us turn a reporting agreement into a monitoring and evaluation plan.

Reporting agreement (one row per metric: name, definition, period, breakdown):
[PASTE AGREEMENT TABLE]

Our change model (theory of change, logic model or logframe):
[PASTE OR ATTACH]

For each metric, write one row: what monitoring checks, source record, who collects it, when it is due, and one evaluation question it could help answer.
Then list evaluation questions the metrics cannot answer, with the evidence each needs.

Rules:
- Use only the metrics, definitions and dates in the agreement. Do not invent targets, baselines or numbers.
- If a definition, source or due date is not stated, write "not in our data" and suggest one question to ask.
- Return a table first, then the list.

What does monitoring look like during the year?

Monitoring is the routine check of each incoming record and report against the agreement, done whenever something arrives instead of once at year end. It is one step in a four-step loop: agree, monitor, evaluate, learn.

Slide titled 'Monitoring, evaluation and learning on one record, all year.' Four boxes sit on a loop: 1 Agree, what to monitor, from the agreement and dictionary; 2 Monitor, collect inside the work, all year; 3 Evaluate, what changed, for whom, and why; 4 Learn, decide what changes next cycle. In the centre: One record, one ID per person. Handwritten note: learning all year, not a report at the end.
Monitoring and evaluation read the same record, so an evaluation question never starts from a fresh spreadsheet. From the course Measurement and reporting.

Collecting inside the work means the attendance sheet, exit form and follow-up survey feed the record the report is built from. With one ID per person from the first form, a 12-month follow-up lands on the same record as intake.

Spreadsheets can do this with discipline: one ID column, one tab per wave, no retyped names. In Sopact Sense, the ID is issued at the first form and later workflows join the same record; each partner's data can sit in its own folder, whose AI Assistant sees only that folder, while the fund sees aggregated results.

The video below covers the same shift, from a theory of change written once to one used in daily decisions. Watch for how a team reads its measures during the year instead of assembling them at the end.

Video · A theory of change used in daily decisions.
Watch on YouTube ↗

How do you check a partner report against the M&E plan?

Lay the report beside the agreed metrics and give each one a status, matches, different definition, missing or not split, before anything is added up or interpreted. Partner C's first report to the workforce fund holds one match and three gaps.

Partner C counted 80 unique people enrolled, as agreed. It counted placements within six months instead of 90 days, left out 12-month retention, and gave one average wage instead of one per track.

Slide titled 'Compare each report with the agreement. Every gap becomes one question.' A table for a fictional partner report: Enrolled, agreement unique people, report says 80 people, status matches; Placed in a job, within 90 days, report says within 6 months, status different; Retained, at 12 months, not reported, status missing; Starting wage, by track, one average, status not split. A side panel lists three questions to send: Can you count placements within 90 days, as agreed? When will 12-month retention be available? Can you split starting wage by track?
Monitoring in practice: one match and three gaps, each turned into a question the partner can answer. From the course Measurement and reporting.

Each gap is a monitoring finding, not a verdict on performance: Partner C may have placed more people than anyone, but you cannot tell until it counts the agreed way. Send the three questions together within days, while the person who compiled the report still knows its sources.

The same metrics show where monitoring stops and evaluation starts:

Scroll horizontally to see all columns →

MetricMonitoring asksEvaluation asks
EnrolledUnique people, reported on time?Who enrolled, and whom did we miss?
Placed in a jobCounted within 90 days of exit?Why do placements differ by partner?
RetainedReported at 12 months, when due?Do jobs last, and for whom not?
Starting wageHourly and split by track?Do tracks lead to living-wage work?

What evaluation questions should an M&E system answer?

Good evaluation questions come from the theory of change and the decisions ahead: whether the intended people are reached, how they experience the work, what changed and for whom, what else could explain it, and what should change. For the workforce fund, that means asking whether people furthest from work are enrolling, and whether local hiring rather than training explains a jump in placements. Write the question before choosing the method.

"How many enrolled?" is a fair monitoring question; it fails when a report uses it to answer "Did people get jobs that last?" Name that gap and choose proportionate evidence: the 12-month follow-up, interviews with early leavers, or earlier cohorts counted the same way.

Open answers often explain what numbers cannot. In Sopact Sense, an Intelligence Cell reads each exit-survey answer, such as "What got in the way?", on arrival with a prompt your team configures; a person checks the themes before they reach a report.

What can monitoring and evaluation data prove?

M&E data shows what was delivered and what changed among the people you reached; on its own it does not prove the program caused the change. Outputs such as sessions delivered are not outcomes such as a job held at 12 months, and a before-and-after difference needs a comparison before it supports a claim about cause.

Follow-up surveys are self-reported and not everyone answers, so show how many responded beside every rate and treat the rest as unknown. A monitoring check confirms that a number is defined as agreed, not that it is correct; when AI drafts a check or summary, a person reviews every row against the record.

Start with one monitoring cycle this quarter

Pick one funder agreement and one reporting period, and run the whole loop once before designing anything larger.

  1. List the agreed metrics with definition, period and breakdown in one table, and mark any that are not due this period.
  2. Name each metric's source record and owner, and check that every participant has one ID from the first form.
  3. When the report arrives, give each metric a status: matches, different definition, missing or not split.
  4. Send one message with a plain question per gap within a week, and log each answer.
  5. Write two evaluation questions the metrics cannot answer, and the evidence each would need.
  6. At the end of the period, record one change to the agreement, the plan or the program.

After the first cycle you hold checked numbers, a short gap log, two evaluation questions with an evidence plan, and one decision for next period.

Frequently asked questions

What does M&E stand for?

M&E stands for monitoring and evaluation. Monitoring is the routine tracking of agreed measures, such as enrolments each quarter or placements within 90 days of exit. Evaluation is the periodic, question-led look at what changed, for whom and why. Many teams add an L for learning, which makes the use of findings explicit.

What is the difference between an M&E framework and an M&E plan?

The framework covers the whole body of work: the questions, measures and overall approach. The plan is operational: for each measure, who collects it, from which source, when it is due and who reviews it. A framework can span several programs; a plan belongs to one, and both share one change model.

What are examples of M&E indicators?

In the fictional workforce fund on this page, the five agreed metrics are participants enrolled (unique people, quarterly), completed training, placed in a job within 90 days of exit, retained in the same job at 12 months, and hourly starting wage by track. Each needs its definition, period and source written down so every partner counts it the same way.

Does every M&E system need individual records?

No. Use records per person when a question follows the same people over time, such as retention at 12 months. Site-level, partner-level or anonymous group evidence can answer other questions. Where you do follow people, one ID from the first form saves matching names by hand later.

Does a baseline prove impact?

No. A baseline lets you compare a later condition with a starting point, which shows change among the people you reached. Claiming the program caused it needs more: a comparison group, earlier cohorts or outside data counted the same way, and an account of who did not respond.