What is monitoring and evaluation?
Monitoring and evaluation (M&E) is how an organization tracks whether its work is happening as planned and judges what changed as a result: monitoring follows agreed measures all year, and evaluation asks what changed, for whom and why. The two work best as one loop on one set of records, not as two reports written months apart.
Funders and funded partners usually meet in M&E through a reporting agreement: the few metrics both will count, each defined and dated. Monitoring checks every report against it; evaluation uses the same records for the questions an agreement cannot answer.
THE SHORT VERSION
- Monitor against what was agreed: each metric with its definition, period and source, checked every time a report arrives.
- Evaluate with questions, not more indicators: what changed, for whom, why, and what else could explain it.
- Keep both on one record per person and per partner, so every finding traces back to the report and record it came from.
What is the difference between monitoring and evaluation?
Monitoring tracks agreed measures routinely to show whether the work is on track; evaluation asks focused questions at chosen points to explain the results and judge whether to continue or change the work. Monitoring shows placements dropped at one partner; evaluation asks why.
Scroll horizontally to see all columns →
| Question | Monitoring | Evaluation |
|---|---|---|
| Main purpose | See whether delivery and results are on track | Explain results and judge their value |
| Example question | Did every partner count placements within 90 days? | Why do placements differ between partners? |
| Evidence | Routine records, reports and scheduled surveys | Monitoring records plus interviews or comparisons |
| Timing | Every time a report or record arrives | At points chosen around a decision |
| Typical use | Fix gaps, adjust delivery | Redesign, continue or change funding |
These are practical distinctions, not rigid walls: monitoring can include outcomes, and evaluation can inform decisions mid-program. BetterEvaluation's monitoring guidance describes this broader scope, and its guidance on evaluation timing explains why the two need to work together.
What is the difference between an M&E framework and an M&E plan?
An M&E framework sets out the questions, measures and approach across the work; an M&E plan says who collects and checks what, when, and from which source. Organizations use these labels loosely, so agree on the job each document does instead of keeping several copies of one table.
Scroll horizontally to see all columns →
| Document | What it helps you decide |
|---|---|
| Theory of change | Why the work should lead to the change, and which assumptions matter |
| Logframe | How objectives, indicators, sources and assumptions fit in one matrix |
| Reporting agreement | The few metrics a funder and partner report, how each is counted, when |
| M&E framework | The overall questions, measures and approach across the work |
| M&E plan | Who collects and checks what, when, from which source |
The four change-model layouts carry the same levels under different labels, so build the model once and re-lay it when a funder asks for another format; One change model, four formats shows how. BetterEvaluation explains how a framework spans several evaluations while a plan covers one.
Many M&E plans still use the logframe, often filled in once and never opened again. In this short video, watch for the means of verification column: it names where each number will come from, the first thing monitoring needs.
How do you build a practical M&E plan?
Start from the decision the evidence must support and the metrics already agreed with the funder, then define each measure, name its source and schedule, and plan the analysis before collecting anything new. The six steps below use one running example.
1. Start from the decision and the agreement
The example is fictional: a regional workforce fund with four job-training partners, A to D. It must decide whether each partner's training leads to jobs that last, and its reporting agreement names five metrics: enrolled, completed training, placed in a job, retained and starting wage.
2. Choose the unit you need to follow
Enrolled counts unique people, not visits, so every person needs one ID from the first form; a partner-level check needs one record per partner. Collect personal details only where a question needs them.
3. Define each measure and its source
Write down the calculation, period and source for every metric: placed in a job means within 90 days of exit, retained means the same job at 12 months, starting wage is hourly by track. Keep these in a shared data dictionary, and keep a recorded zero apart from a missing answer.
4. Schedule collection and review
Match collection to the agreement's rhythm: enrolled is reported quarterly, retention and starting wage annually. A metric not yet due is logged as "not due", not "missing". Name who requests, checks and responds to each report.
5. Plan the analysis before collecting more
Decide which groups you will compare, such as by gender, age and disability, and how you will treat missing follow-ups. A causal question needs an evaluation design, not routine reporting.
6. Budget for the whole loop and pilot it
Allow time for partners, checks, analysis and discussion, not only collection. Run one reporting cycle end to end and revise before asking every partner to take part.
Any AI tool can draft the plan's first table from your agreement; review every row with the metric's owner.
PROMPT · PASTE INTO CLAUDE, CHATGPT OR YOUR AI TOOL
You are helping us turn a reporting agreement into a monitoring and evaluation plan. Reporting agreement (one row per metric: name, definition, period, breakdown): [PASTE AGREEMENT TABLE] Our change model (theory of change, logic model or logframe): [PASTE OR ATTACH] For each metric, write one row: what monitoring checks, source record, who collects it, when it is due, and one evaluation question it could help answer. Then list evaluation questions the metrics cannot answer, with the evidence each needs. Rules: - Use only the metrics, definitions and dates in the agreement. Do not invent targets, baselines or numbers. - If a definition, source or due date is not stated, write "not in our data" and suggest one question to ask. - Return a table first, then the list.
What does monitoring look like during the year?
Monitoring is the routine check of each incoming record and report against the agreement, done whenever something arrives instead of once at year end. It is one step in a four-step loop: agree, monitor, evaluate, learn.

Collecting inside the work means the attendance sheet, exit form and follow-up survey feed the record the report is built from. With one ID per person from the first form, a 12-month follow-up lands on the same record as intake.
Spreadsheets can do this with discipline: one ID column, one tab per wave, no retyped names. In Sopact Sense, the ID is issued at the first form and later workflows join the same record; each partner's data can sit in its own folder, whose AI Assistant sees only that folder, while the fund sees aggregated results.
The video below covers the same shift, from a theory of change written once to one used in daily decisions. Watch for how a team reads its measures during the year instead of assembling them at the end.
How do you check a partner report against the M&E plan?
Lay the report beside the agreed metrics and give each one a status, matches, different definition, missing or not split, before anything is added up or interpreted. Partner C's first report to the workforce fund holds one match and three gaps.
Partner C counted 80 unique people enrolled, as agreed. It counted placements within six months instead of 90 days, left out 12-month retention, and gave one average wage instead of one per track.

Each gap is a monitoring finding, not a verdict on performance: Partner C may have placed more people than anyone, but you cannot tell until it counts the agreed way. Send the three questions together within days, while the person who compiled the report still knows its sources.
The same metrics show where monitoring stops and evaluation starts:
Scroll horizontally to see all columns →
| Metric | Monitoring asks | Evaluation asks |
|---|---|---|
| Enrolled | Unique people, reported on time? | Who enrolled, and whom did we miss? |
| Placed in a job | Counted within 90 days of exit? | Why do placements differ by partner? |
| Retained | Reported at 12 months, when due? | Do jobs last, and for whom not? |
| Starting wage | Hourly and split by track? | Do tracks lead to living-wage work? |
What evaluation questions should an M&E system answer?
Good evaluation questions come from the theory of change and the decisions ahead: whether the intended people are reached, how they experience the work, what changed and for whom, what else could explain it, and what should change. For the workforce fund, that means asking whether people furthest from work are enrolling, and whether local hiring rather than training explains a jump in placements. Write the question before choosing the method.
"How many enrolled?" is a fair monitoring question; it fails when a report uses it to answer "Did people get jobs that last?" Name that gap and choose proportionate evidence: the 12-month follow-up, interviews with early leavers, or earlier cohorts counted the same way.
Open answers often explain what numbers cannot. In Sopact Sense, an Intelligence Cell reads each exit-survey answer, such as "What got in the way?", on arrival with a prompt your team configures; a person checks the themes before they reach a report.
What can monitoring and evaluation data prove?
M&E data shows what was delivered and what changed among the people you reached; on its own it does not prove the program caused the change. Outputs such as sessions delivered are not outcomes such as a job held at 12 months, and a before-and-after difference needs a comparison before it supports a claim about cause.
Follow-up surveys are self-reported and not everyone answers, so show how many responded beside every rate and treat the rest as unknown. A monitoring check confirms that a number is defined as agreed, not that it is correct; when AI drafts a check or summary, a person reviews every row against the record.
Start with one monitoring cycle this quarter
Pick one funder agreement and one reporting period, and run the whole loop once before designing anything larger.
- List the agreed metrics with definition, period and breakdown in one table, and mark any that are not due this period.
- Name each metric's source record and owner, and check that every participant has one ID from the first form.
- When the report arrives, give each metric a status: matches, different definition, missing or not split.
- Send one message with a plain question per gap within a week, and log each answer.
- Write two evaluation questions the metrics cannot answer, and the evidence each would need.
- At the end of the period, record one change to the agreement, the plan or the program.
After the first cycle you hold checked numbers, a short gap log, two evaluation questions with an evidence plan, and one decision for next period.
Frequently asked questions
What does M&E stand for?
M&E stands for monitoring and evaluation. Monitoring is the routine tracking of agreed measures, such as enrolments each quarter or placements within 90 days of exit. Evaluation is the periodic, question-led look at what changed, for whom and why. Many teams add an L for learning, which makes the use of findings explicit.
What is the difference between an M&E framework and an M&E plan?
The framework covers the whole body of work: the questions, measures and overall approach. The plan is operational: for each measure, who collects it, from which source, when it is due and who reviews it. A framework can span several programs; a plan belongs to one, and both share one change model.
What are examples of M&E indicators?
In the fictional workforce fund on this page, the five agreed metrics are participants enrolled (unique people, quarterly), completed training, placed in a job within 90 days of exit, retained in the same job at 12 months, and hourly starting wage by track. Each needs its definition, period and source written down so every partner counts it the same way.
Does every M&E system need individual records?
No. Use records per person when a question follows the same people over time, such as retention at 12 months. Site-level, partner-level or anonymous group evidence can answer other questions. Where you do follow people, one ID from the first form saves matching names by hand later.
Does a baseline prove impact?
No. A baseline lets you compare a later condition with a starting point, which shows change among the people you reached. Claiming the program caused it needs more: a comparison group, earlier cohorts or outside data counted the same way, and an account of who did not respond.

