What is program evaluation?
Program evaluation systematically examines a program to answer questions about its design, implementation, results or value. The purpose is to support a judgment or decision using evidence. It is broader than counting activities and does not always require estimating causal impact.
Start with who needs the finding and what they will do with it. An evaluation designed for improving delivery may need different evidence from one intended to estimate the program’s effect.
Choose the type of evaluation
| Type | Typical question | Possible evidence |
|---|---|---|
| Formative | How should the program be designed or improved? | Needs assessment, interviews and pilot findings. |
| Process | Was the program delivered as intended? | Delivery records, observations and participant accounts. |
| Outcome | What outcomes were observed? | Defined measures at appropriate time points. |
| Impact | What change is attributable to the intervention? | A design that examines the counterfactual or causal explanation. |
| Economic | How do costs relate to results or alternatives? | Defined costs and outcome evidence with assumptions. |
These categories can overlap. Name the actual question and design rather than relying on a label to explain what the evaluation can conclude.
Use the current CDC framework
The 2024 CDC framework organizes evaluation around assessing context, describing the program, focusing questions and design, gathering credible evidence, generating supported conclusions and acting on findings. Collaboration, equity and use run across the work. Read the current framework.
Use it as a planning structure rather than a rigid sequence. A discovery during collection may require revisiting the design. Do not silently substitute an older six-step wording while describing it as the current framework.
A practical evaluation plan template
| Plan element | Decision to record |
|---|---|
| Purpose | Who will use the findings, for what decision and when? |
| Program | What work, population and period are in scope? |
| Questions | Which uncertainties matter most? |
| Design | What evidence and comparison can answer them? |
| Collection | Sources, timing, sampling and responsibilities. |
| Analysis | Calculations, coding and quality checks. |
| Use | How findings will be discussed and acted on. |
Include the budget, access requirements and important limitations. Keep the plan feasible: a question that requires unavailable evidence may need to be narrowed.
Worked example: a training program
An illustrative team wants to know why course completion is high but workplace application is low. It combines attendance and assessment records with follow-up accounts about opportunities to use the skill.
The evaluation may find that practice time is unavailable after the course. That finding can guide delivery changes. It does not by itself estimate how much performance would have changed without training. A separate causal question needs a suitable design.
Collect evidence that fits the question
Use existing records where they are suitable and collect new evidence for the gaps. Check definitions, coverage, timing and source quality. A satisfaction score cannot stand in for a skills assessment, and a few interviews cannot establish a population rate.
Document missingness and differences between respondents and nonrespondents where possible. Involve people affected by the program in interpreting findings, not only in supplying data.
Turn findings into action
A useful conclusion states the evidence, interpretation and limitation. Agree the action, owner and review date. Keep findings that challenge the program as visible as favorable ones.
Store sources and definitions with the analysis so a later team can understand it. AI can assist with organizing records, but accountable reviewers should verify conclusions. Continue to using a theory of change in M&E.
Frequently asked questions
Is monitoring the same as evaluation?
Monitoring routinely tracks selected measures. Evaluation investigates questions and supports judgments about the work. They inform each other.
Does every evaluation need a control group?
No. The design depends on the question. Causal effect estimation requires particular attention to what would have happened otherwise.
When should evaluation start?
Plan early enough to collect the needed evidence. Some questions can be addressed retrospectively, with stated limitations.
Who should conduct it?
Choose people with the necessary skills and appropriate independence for the purpose. Internal and external roles can both be useful.
