What is program evaluation?
Program evaluation is the systematic collection and analysis of evidence about how a program is run and what it produces, used to judge its merit and improve it. It spans four working types — formative, process, outcome, and impact — and follows a standard sequence such as the CDC's five-step framework, from engaging stakeholders to justifying conclusions and putting them to use. The point is a decision, not a document.
Most program evaluation still arrives as a post-mortem: an external evaluator is hired near the end, data is pulled from several systems and matched by hand, and a report lands months after the cohort has moved on. This guide covers the methods, the four types, and the CDC framework — and the shift that lets an evaluation inform the cycle it is measuring rather than the one after it.
Key takeaways
- Program evaluation judges a program's merit and improves it — across four types (formative, process, outcome, impact) and a standard sequence like the CDC five-step framework.
- The three method families are quantitative, qualitative, and mixed. Quant measures how much changed; qual explains why; mixed keeps the number and its reason on one record.
- Sopact calls the working version The Continuous Evaluation Record: the four evaluation types run off one participant record read on arrival, so evaluation informs the next cycle instead of arriving after the program ends.
- An evaluation delivered after the program is a post-mortem. Reading evidence during delivery is what turns a finding into a decision you can still act on.
- The methods are not the constraint — the data model is. Formative, process, outcome, and impact questions only draw on the same evidence when it sits on one record.
The four working types of program evaluation.
There are four types of program evaluation, each answering a different question: formative asks whether the design is sound before you scale it, process asks whether the program is running as intended, outcome asks whether participants changed, and impact asks whether the program caused the change and whether it lasted. A complete evaluation uses more than one, in sequence.
The types are often treated as separate studies run by separate people at separate times, which is why they rarely share evidence. Run as one continuous record, a process finding mid-cohort and an outcome finding at follow-up sit on the same participant, so the evaluation reads as one story rather than four disconnected reports. The design layer beneath them is on theory of change and logic model.
The four working types of program evaluation
| Type | The question it answers | When it runs |
|---|
| Formative | Is the design sound before we scale it? | Before and early in the program |
| Process | Is the program running as intended, at quality? | During delivery |
| Outcome | Did the participants change? | At exit and follow-up |
| Impact | Did the program cause it, and did it last? | At longer-term follow-up |
Program evaluation methods: quantitative, qualitative, mixed.
Program evaluation methods fall into three families: quantitative methods measure how much changed, qualitative methods explain why, and mixed methods keep the number and its reason together. The choice is rarely either/or — a defensible evaluation almost always needs both a measured change and the participant's account of it.
The trap in mixed methods is practical, not conceptual: when the rating and the open-ended reason come from two tools that never shared an identifier, joining them is a manual reconciliation that is out of date the moment it is built. The instrument that carries both is on mixed-method surveys, and the analysis of the two halves on survey analysis.
Program evaluation methods
| Method | What it is best for | The trap |
|---|
| Quantitative | Measuring how much changed, across a cohort | Counts without the reason behind them |
| Qualitative | Explaining why the numbers moved | Unreadable at scale without coding |
| Mixed methods | The number and its reason on one record | Two tools that never share an identifier |
The Continuous Evaluation Record: from post-mortem to next-cycle decision.
The classic evaluation is a phase that begins when the program ends: hire the evaluator, gather the data, deliver the report. That sequence guarantees the findings arrive too late to change the program they describe. The cost is not rigor — a post-hoc evaluation can be perfectly rigorous — it is timing.
Sopact calls the alternative The Continuous Evaluation Record: intake, activity, open-ended feedback, and follow-up all resolve to one participant record read on arrival, so a formative question, a process question, and an outcome question draw on the same evidence without a merge. The difference is a data-model one — an evaluation assembled from disconnected exports can only be run once, at the end, while a continuous record can be read at every stage. The practice this sits inside is on monitoring and evaluation and impact measurement and management.
The stage below runs one program cycle both ways.
Stage 1
Evaluating a program cycle
where evaluation arrives too late
TodayAn external evaluator is hired near the end · Data pulled from four systems and matched by hand · A report lands months after the cohort finished⚠ An evaluation delivered after the program ends is a post-mortem — the findings arrive when the only thing left to change is the next program, not this one.
The Loop on this stage with Sopact
Collect — clean at the source
IntakeAttendance / activityOpen-ended feedbackFollow-up survey
→ every source lands on one persistent ID
On arrival — read automatically
Intelligent Cell
Each open response is themed against your codebook on arrival, so a process problem is visible mid-cohort rather than in the final report.
Intelligent Row
Every source resolves to one participant record, so formative, process, outcome, and impact questions all draw on the same evidence without a merge.
Ask & act — the Assistant
“Which sites are underperforming this cohort, and what did participants there say?”
→ A corrective decision while the cohort is still enrolled — evaluation that informs the next cycle, not just records the last one.
The CDC framework, applied.
The CDC's framework for program evaluation is the field standard: engage stakeholders, describe the program, focus the evaluation design, gather credible evidence, and justify conclusions so they get used. It is a sequence, and the order matters — skipping the description step is how an evaluation ends up measuring the wrong thing.
Step four — gather credible evidence — is where most evaluations break, because evidence collected in disconnected spreadsheets cannot support the traceable conclusions step five requires. Keeping the evidence on one participant record is what makes step five a query rather than a reconstruction. Applying the standard to a worked cohort is walked through in analyze pre, mid and post waves.
The CDC framework for program evaluation
01Engage stakeholdersThe people who use the findings help decide what to ask
02Describe the programA theory of change or logic model — the thing being evaluated
03Focus the designChoose the evaluation type and the questions it must answer
04Gather credible evidenceCollect indicators on one participant record, baseline to follow-up
05Justify conclusions and use themTie each finding to evidence, and feed the next cycle
The CDC's five-step framework is the field standard. Step four is where most evaluations break — evidence gathered in disconnected exports cannot support the conclusions step five asks for.
A report tells you what happened. The Loop tells you in time to act.
An evaluation that reports after the program answers the last cohort; a program evaluated as it runs can answer this one. Reading evidence on arrival is what moves a finding from the post-mortem to the working period. That is the premise of the Loop, Sopact's method for continuous impact intelligence: collect clean at the source, analyze the moment data arrives, improve while you can still act.
The Loop is also what makes an evaluation defensible. Every finding traces to the participant and the response behind it, and the same codebook returns the same themes twice running, so a conclusion resolves to its evidence. That standard has its own chapter in reliability and reproducibility. Turning a finished evaluation into a report is on program report, and the live view on program dashboard.
One method, three moves that never stop
1 · CollectClean at the source; every source on one participant record.
2 · AnalyzeOn arrival; a process problem visible mid-cohort, not at year-end.
3 · ImproveIn time to act; the correction lands while the cohort is here.
Then the cycle runs again, a little sharper each cohort. Read the method: the Loop methodology →
Run your evaluation on one record
The fastest way to feel the difference is to run a type of evaluation against a cohort you already have. Each prompt below pastes into Sopact Sense's Assistant, or reasons through with your team; the arrow above each links the Academy walkthrough that shows the expected output and the tips.
Academy walkthrough → Run an outcome evaluation
Run an outcome evaluation on this cohort: [PASTE INTAKE + EXIT + FOLLOW-UP]. Match participants across waves, report the change among matched respondents versus the raw change, and disaggregate by site or subgroup. State plainly where attrition or a small cell makes a result unreliable. Return a table: Outcome / Matched change / Raw change / Subgroup gaps / Confidence.
Academy walkthrough → Explain the numbers (mixed methods)
For this evaluation result, recover the explanation from the open-ended responses: [PASTE RESULT + OPEN RESPONSES + CODEBOOK]. Give the theme distribution among the participants who moved least, the two themes most over-represented there, and three verbatims each. If the open responses cannot explain the result, say so rather than inventing a narrative.
Academy walkthrough → Run a process evaluation mid-cohort
Run a process evaluation on this in-flight cohort: [PASTE ATTENDANCE + ACTIVITY + MID-POINT FEEDBACK]. Flag where delivery is diverging from the design, which sites or facilitators are underperforming, and what participants there are reporting. Rank the issues by how much they threaten the outcomes. Return: Issue / Evidence / Sites affected / Priority.
Academy walkthrough → Focus the evaluation design
Focus a program evaluation design for this program: [PASTE PROGRAM OR LOGIC MODEL]. Recommend which of the four types (formative, process, outcome, impact) to run and in what sequence, the questions each must answer, and the indicator and instrument for each. Flag any outcome that cannot be measured with the data currently collected.
Learn the how-to in the Academy
Each walkthrough is a hands-on companion written to run on your own data: what to do, the prompt to run, the output to expect, and the tips that keep it reliable.
Watch: reading a program's evidence as it arrives, so the four evaluation types run off one record instead of four disconnected studies.
Frequently asked questions
What is program evaluation?
Program evaluation is the systematic collection and analysis of evidence about how a program is run and what it produces, used to judge its merit and improve it. It spans four types — formative, process, outcome, and impact — and follows a standard sequence such as the CDC's five-step framework. In Sopact's framing, evaluation runs off one Continuous Evaluation Record read on arrival, so it informs the cycle it measures rather than arriving after the program ends.
What are the methods of program evaluation?
Program evaluation methods fall into three families: quantitative methods measure how much changed, qualitative methods explain why it changed, and mixed methods keep the number and its reason together. Most defensible evaluations use both a measured change and the participant's account of it. Sopact's position is that the binding constraint is rarely the method — it is whether the quantitative and qualitative evidence sit on one participant record so they can be read together.
What are the types of program evaluation?
The four working types are formative (is the design sound before we scale it), process (is the program running as intended), outcome (did participants change), and impact (did the program cause the change and did it last). A complete evaluation uses more than one, in sequence. Sopact runs all four off one participant record, so a process finding mid-cohort and an outcome finding at follow-up sit on the same evidence.
What are program evaluation techniques?
Techniques are the specific procedures inside each method: pre-post and matched-cohort comparison, cross-tabulation and significance testing, codebook-based thematic coding of open-ended responses, and counterfactual or comparison-group designs for impact. The technique that most often decides an evaluation's credibility is matching participants across waves, because comparing two wave averages made of different people measures the mix, not the change. Sopact keeps one persistent ID so the matched comparison is available by default.
Can you give a program evaluation example?
A workforce training example: a 320-participant cohort with four data sources — intake survey, attendance and activity logs, open-ended mid-point feedback, and a follow-up survey at six months. The formative question was whether the curriculum matched employer demand; the process question whether attendance held across sites; the outcome question whether participants were placed in living-wage roles; the impact question whether placements lasted at twelve months. All four ran off one participant record, so each finding traced to the people behind it.
What is the CDC framework for program evaluation?
The CDC's framework is the field-standard five-step sequence: engage stakeholders, describe the program, focus the evaluation design, gather credible evidence, and justify conclusions so they get used. The order matters — skipping the description step is how an evaluation ends up measuring the wrong thing. Step four is where most evaluations break, because evidence gathered in disconnected exports cannot support the traceable conclusions step five requires; Sopact keeps the evidence on one record so that step is a query.
What is the difference between program evaluation and monitoring?
Monitoring is the ongoing tracking of indicators against a plan; evaluation is the periodic, deeper judgment of a program's merit and worth. Monitoring tells you whether you are on track; evaluation tells you whether the program is working and why. They share the same evidence base, which is why Sopact runs both off one record — the fuller treatment is on the monitoring and evaluation page.
What tools or software are used for program evaluation?
Spreadsheets and statistical packages handle the quantitative analysis; qualitative packages handle coding; survey tools collect the data. The gap most evaluation stacks hit is that these do not share a participant identity, so a formative, process, and outcome evaluation become three disconnected datasets someone merges by hand. Sopact is built to keep all four evaluation types on one participant record, which is what turns evaluation from a periodic project into a continuous read.
How do you write program evaluation questions?
Start from the type: a formative question asks whether the design will work, a process question whether it is being delivered as intended, an outcome question what changed for participants, and an impact question whether the program caused it. Each should name the indicator and the instrument that would answer it before any data is collected. Sopact ties each evaluation question to its indicator so the design is measurable from the start rather than aspirational.
How long does a program evaluation take?
A traditional end-of-program evaluation takes months, most of it spent gathering and reconciling data from disconnected systems after the program has ended. A continuous evaluation is effectively finished when the program is, because the evidence was read on arrival and each stage was evaluated as it happened. Sopact treats the gap between those two timelines as the real cost of a post-hoc evaluation — findings that arrive too late to act on.
Next: place evaluation in the wider practice on monitoring and evaluation, or describe what you evaluate on logic model.
The Continuous Evaluation Record
01FormativeIs the design sound? — read before you scale
02ProcessIs it running as intended? — read mid-cohort
03OutcomeDid participants change? — read at follow-up
04ImpactDid it cause it, and last? — all four on one record
The Continuous Evaluation Record: the four evaluation types run off one participant record read on arrival, so evaluation informs the next cycle instead of arriving after the program ends.