play icon for videos

Monitoring and Evaluation: Differences, Plan and Practical Example

Build a practical monitoring and evaluation plan with clear measures, a worked example, shared definitions, local flexibility and evidence-based reporting.

Updated
September 15, 2026
360 feedback training evaluation
Use Case
Impact & ESG portfolios · Practical guide

Monitoring and Evaluation: Differences, Plan and Practical Example

Build a practical monitoring and evaluation plan with clear measures, a worked example, shared definitions, local flexibility and evidence-based reporting.

Read the guide ↓

What is monitoring and evaluation?

Monitoring and evaluation (M&E) helps an organization understand what it is doing, what is changing and what it should do next. Monitoring tracks implementation, conditions and emerging results. Evaluation examines the value and results of the work through explicit questions, evidence and interpretation.

A program manager might use attendance records to notice that fewer people return after their first session. An evaluation could investigate why, whose experience differs and whether the program is still meeting the need it was designed for. Both activities matter. A well-maintained dashboard does not answer every evaluation question; a thoughtful evaluation cannot repair information that was never collected.

This guide is for teams designing a workable evidence process across programs, services, locations or partners. Start with the decision that needs support, then choose the collection and analysis required. You do not need a large indicator library to begin.

Monitoring vs. evaluation: what is the difference?

Scroll horizontally to see all columns →

QuestionMonitoringEvaluation
Main purposeUnderstand delivery, conditions and progressAssess results, value, explanations and implications
Example questionWhich sites have incomplete follow-up?Why does follow-up differ, and how does this affect results?
EvidenceRoutine records, feedback, observations and scheduled measuresRelevant monitoring evidence plus additional inquiry or comparison when needed
TimingRegularly, at a frequency useful to the workAt selected points or through an ongoing evaluative approach
Typical useAdjust delivery and identify issuesInform design, improvement, continuation or investment

These are practical distinctions, not rigid boundaries. Monitoring can include outcomes and changing conditions, not only activities. Evaluation can support decisions during implementation rather than waiting until a program ends. Choose the timing around the decision, not around the label. BetterEvaluation's monitoring guidance describes this broader scope; its guidance on evaluation timing explains why the two need to work together.

Start with useful questions

Before adding a field to a survey, write down who will use the answer. The team delivering a service, the people using it and the organization funding it may need different views of the same work.

  • Delivery: Are the intended people or locations being reached? What prevents participation?
  • Experience: What is useful, difficult or missing from the service?
  • Results: What changed, for whom, and over what period?
  • Explanation: What else may explain the change? Which assumptions need further investigation?
  • Decision: What evidence would justify changing the approach, and who decides?

A question such as “How many workshops did we deliver?” is legitimate. It becomes insufficient when the report uses it to answer “Did people's practice improve?” Make the gap explicit and decide what additional evidence is proportionate.

Framework, plan, theory of change and logframe

These documents serve related purposes. Organizations use the labels differently, so agree on the job each document will do rather than creating several versions of the same table.

Scroll horizontally to see all columns →

DocumentWhat it helps you decide
Theory of changeWhy the work might lead to the intended change, and which assumptions matter
Results frameworkHow intended results relate to one another
LogframeHow objectives, indicators, verification and assumptions fit in a matrix
M&E frameworkThe overall questions, measures and approach across the work
M&E planWho collects and analyzes what, when, with which resources and review steps

A framework can span several evaluations or programs. A specific evaluation plan adds the methods and arrangements for that inquiry. BetterEvaluation explains this distinction. The documents should inform one another as the work develops.

Build a practical M&E plan

1. Define the decision and intended change

Choose one workflow first. For example, a regional training team needs to decide whether to revise follow-up support before the next cohort. Identify who needs the evidence, when the decision occurs and what uncertainty matters most.

2. Choose the unit you need to understand

The unit might be a person, household, site, organization, project or population. Following the same participant over time requires a suitable linking method. Comparing anonymous group feedback does not require identifying every respondent. Do not collect personal details simply because the system can hold them.

3. Define each measure and its source

Record the calculation, eligible population, period, exclusions and source. Separate a recorded zero from a missing answer. Define what counts as completion or improvement before interpreting results, and retain the version used for each reporting period.

4. Schedule collection and review

Collect information when people can answer meaningfully and when the team can use it. Some measures may be useful after every activity; others need a later follow-up or a dedicated study. Assign responsibility for requesting, checking, analyzing and responding to the evidence.

5. Plan analysis before collecting everything

Decide which groups can reasonably be compared, how missing data will be handled and where comments or interviews will help explain the numbers. If the question requires a causal conclusion, design an appropriate evaluation rather than promising that routine reporting will establish it.

6. Budget for the whole process

Allow time for contributors, translation where relevant, data checks, analysis and discussion. A free form can still require substantial coordination. Pilot the process with one cycle and revise it before asking every location to participate.

A worked example: delivery is not the same as an outcome

This fictional example follows a training program. There are 100 enrolled participants. Eighty attend the first session and 60 meet the program's agreed completion rule. The completion rate is therefore 60 ÷ 100 = 60% of enrolled participants. Using attendees as the denominator would answer a different question.

Forty participants provide comparable baseline and follow-up assessments. Of those, 28 improve on the defined measure: 28 ÷ 40 = 70% of matched respondents. Matched evidence covers 40% of enrollment. The team cannot report that 70% of all participants improved, and it cannot assume the other 60 participants did not improve.

Monitoring reveals that two locations have low follow-up coverage. The next step is to investigate collection timing, access and other barriers. Evaluation might examine whether improvements are meaningful, whether they persist and what role the program played alongside other influences. The before-and-after comparison alone does not establish causation.

Comments can add context: some participants may describe useful practice opportunities while others report that their role prevents applying the skill. These accounts help identify questions to examine. They do not automatically represent everyone or prove an explanation.

Compare across locations without imposing one survey

Federated organizations often need local flexibility. A school, chapter or partner may ask questions relevant to its own setting while contributing a small shared set of measures. Start by agreeing which results genuinely need comparison or aggregation.

Keep stable registration or organizational information in the relevant record and update it when necessary. Collect period-specific activity, experience and outcome information at the appropriate time. Avoid repeatedly asking contributors to re-enter context that is already known and current.

A shared data dictionary should define the common fields, units, periods and exclusions. Local questions can remain local. Two fields should not be combined merely because their labels look similar: “people attending” and “attendance entries” count different things. Where definitions cannot be reconciled, show separate results and explain why.

Decide whether a combined rate should be weighted by respondents, locations or another appropriate unit. Publish the method alongside the result. Equal weight for every site and equal weight for every person answer different questions.

Make evidence reviewable

  • Coverage: Which expected records arrived, and which are missing?
  • Consistency: Did the question, definition or instrument change?
  • Traceability: Can a reviewer find the source behind a number or claim?
  • Access: Who may see identifiable information, comments or sensitive supporting files?
  • Interpretation: What does the evidence support, and what remains uncertain?
  • Response: Who will act on the finding and review the result?

Keep corrections and changes understandable. Replacing a figure silently makes later reports difficult to reconcile. Preserve enough history to explain which version supported a decision, without retaining personal information unnecessarily.

Where connected software and AI can help

When surveys, documents and updates arrive separately, the work often shifts to joining files and reconstructing context. Sopact's approach is to connect recurring evidence to the relevant record and make analysis and governance part of that workflow. The useful test is whether your team can manage the process, review sources and change definitions without repeatedly rebuilding the reporting chain.

AI can assist with extraction, grouping comments and drafting summaries. Check the underlying material, missing records and interpretation before using the output. A citation helps a reviewer locate evidence; it does not guarantee that the conclusion is correct. Software also cannot resolve an unsuitable study design or decide what an outcome should mean.

Evaluate a small end-to-end pilot: collect one period of evidence, review its quality, produce a summary and make a real decision. Measure the effort required to maintain that process as well as the time saved preparing a report.

Report findings people can use

A useful report includes the question, population, period, method, findings and limitations. Distinguish outputs, observed outcomes and claims about contribution. State who reviewed the evidence and what will happen next. Return an appropriate summary to contributors as well as leadership.

For the next step, use the How to Write an Impact Report guide and report examples. Keep the report tied to the evidence available, rather than filling a template with claims the data cannot support.

Watch the companion videos

These Sopact videos discuss logframes and a customer perspective. The customer account illustrates an operational practice; it is not a causal evaluation of the fictional example above.

Your Logical Framework Logframe Is Broken Here is Why

Open Play Foundation Turn Theory of Change Into Daily Decisions

Frequently asked questions

What does M&E stand for?

M&E stands for monitoring and evaluation: tracking implementation and results, then using evidence to assess the work and inform decisions.

Is evaluation only done at the end?

No. Evaluation can inform design and improvement during implementation as well as assess results later. Choose its timing around the questions and decisions it needs to support.

Does every M&E system need individual records?

No. Use individual linking when the question requires following the same person and it is appropriate to collect that information. Site-level, organizational or anonymous population evidence may be suitable for other questions.

Does a baseline prove impact?

No. A baseline supports comparison with a later condition. Explaining whether the program caused the change requires further reasoning and an appropriate evaluation design.