By Sopact Academy · Updated September 12, 2026. Training-program figures are fictional teaching examples.
How do you turn a theory of change into a practical data-collection workflow?
Turn a theory of change into a data-collection plan by defining the evidence needed for each outcome, choosing a workable collection process, establishing the right baseline and scheduling follow-up. Name the population, measure, source, timing and owner before building a form. This lesson produces a one-page plan that connects participant responses, documents and staff observations across a program.
Ask four questions: what would indicate change, how will it be observed, when should it be collected, and who will check the evidence?
Use an existing theory of change as your input. If you still need to build one, start with the practical theory-of-change lesson. For definitions and framework examples, use the theory-of-change reference guide. Here we focus on turning that framework into collection decisions.
Instead of collecting information because a form happens to ask for it, every field exists for a reason: to examine an outcome, test an assumption, support a program decision, or satisfy a reporting requirement.
Use these four planning questions within your course module:
1. Choose defensible measures for each outcome.
2. Start with the workflow that creates visible value first.
3. Establish a baseline that makes change measurable.
4. Connect every measure to the right collection moment.
The result is a one-page collection plan that connects your theory of change to applications, intake forms, mentor notes, surveys, exit assessments, and follow-up data—creating evidence that can improve programs while they are still running, not just after they end.
Course task: use the fictional training-program example below, or use a service or coaching program you know. Bring its theory, current forms and reporting requirements. Leave with a collection plan you can test on three sample records. You can design it on paper or in a spreadsheet; connected software helps implement and maintain it across repeated submissions.
Watch · 9:50 · Before Step 1
Theory of Change with AI: The 4-Step Method Funders Trust
Start here if your theory of change isn’t written yet. The four-step AI walkthrough that produces the outcomes and assumptions this chapter turns into measures — inputs to impact, outcomes and assumptions identified for review.
The one question that designs the workflow
A data-collection workflow may sound technical, but it begins with one question asked about every outcome:
What evidence would show that this outcome is occurring—and where would that evidence come from?
Your theory of change contains a series of claims: training increases confidence; confidence combined with a recognized credential improves employment prospects; placements remain stable six months later. Each claim needs a clearly defined measure, data source, and collection moment. Make those decisions for every outcome, and the structure of your workflow begins to emerge.
Important assumptions may also need evidence. If your theory assumes that employers continue to recognize a credential, for example, that assumption should become a monitoring question with its own source and collection schedule.
Starting with existing forms is convenient, but it can leave important outcomes without evidence. Review every field against a decision or learning question before carrying it into the new plan. Keep useful operational records; remove duplication only after checking their purpose and reporting requirements.
Here, the theory of change determines what gets collected. Every field must earn its place by supporting an outcome, testing an assumption, meeting a reporting requirement, or informing a decision.
Step 1 — Define the measure and its evidence
For each outcome in your framework, ask:
What observation or combination of observations would give credible evidence that this outcome is occurring?
Start with the smallest defensible set of measures. One may be enough for a simple question; employment quality may need wages, hours, stability and participant experience. The Center for Theory of Change’s indicator guidance asks you to specify who changes, the intended level of change and when it should happen. Decide the population and time window before choosing a question.
Then grade your current measurement honestly:
- AVAILABLE — A relevant measure is already being collected, and its source can be traced.
- NEEDS REVIEW — The outcome is claimed, but the available data does not yet provide sufficient evidence.
- MISSING — No defensible measure or data source has been identified.
These labels describe the collection plan, not program success. AVAILABLE means a relevant source exists; it does not establish that the source is complete, accurate or sufficient for your claim. NEEDS REVIEW means the proposed measure or existing evidence needs checking. MISSING means a necessary source or collection moment has not been identified.
Use the following prompt in an approved AI workspace, along with your theory of change, logic model, or a paragraph describing your program:
I run [describe the program]. Here are the outcomes, population, current measures and available sources: [paste them]. Propose the minimum defensible measures for each outcome. Return: outcome, indicator, population, scale or calculation, source, baseline, follow-up, owner and missing-data rule. Mark existing sources AVAILABLE, unresolved choices NEEDS REVIEW, and absent required evidence MISSING. Cite only sources I provide. Do not invent data, targets or a validated instrument. Explain what needs practitioner review before collection.
| Promised outcome | Proposed measure | Scale | Grade today |
|---|---|---|---|
| Job-ready confidence | Confidence self-rating | 1–10, same wording at start, mid, and exit | NEEDS REVIEW claimed, no baseline yet |
| Skills credential earned | Credential Y/N | Y/N at exit | AVAILABLE completion records exist; verify credential criteria |
| Living-wage employment | Hourly wage + employed Y/N | $/hr at six-month follow-up | MISSING no follow-up wave exists |
| Ongoing support quality | Support theme, classified | from mentor notes, weekly | NEEDS REVIEW notes exist, unread |
| Community-level change | — no defensible measure yet | — | MISSING honest, and stays out of the first workflow |
For living-wage employment, also name the wage threshold, location, reference date and source. A wage figure and employment status alone do not establish whether the threshold is met. Keep the wording, scale and timing comparable when measuring change. For example, fictional mean confidence scores of 4.2, 7.1 and 7.4 would be interpretable only after checking who responded at each wave and whether the question stayed the same. If different people answered, the line describes changing respondent groups, not necessarily individual progress. A self-rating does not replace an assessment of demonstrated skill.
Step 2 — Choose a useful first collection workflow
Choose a first workflow that answers a useful decision and is feasible to test. Consider data access, record quality, staff capacity and the cost of delay. Do not postpone a time-sensitive baseline while waiting to demonstrate another feature.
Start with a manageable workflow that has a clear owner, a usable source and a decision the team needs to make.
Application review can be a useful pilot when submissions already exist and the team has agreed a rubric. A backlog of notes may be a better starting point for an established service. Compare the effort of preparing records, testing the method and reviewing results; existing data still needs quality and permission checks.
Consider application review if your program has an upcoming selection round:
- The nonprofit’s training cohort opens with applications: personal statements, goal and barrier questions, sometimes an uploaded proposal. Reviewing them requires agreed criteria, reviewer availability and a way to resolve different interpretations.
- Accelerators and fellowships can face similar coordination questions, although their criteria and selection processes differ.
With a configured rubric, AI-assisted review can prepare scores and point reviewers to relevant passages. Test it on representative submissions, including different languages and formats. Review missing evidence, check inconsistent scores and record overrides. People still resolve eligibility, conflicts, borderline cases and final selection; a shared rubric does not eliminate bias or coordination.
Ask your AI to pressure-test the choice:
Here is my measure map from Step 1: [paste it]. And here are the workflows my program already runs (for example: applications, intake, mid-program survey, mentor check-ins, exit survey, follow-up).
Which bounded workflow should I test first? Compare decision value, time-sensitive baseline needs, available evidence, preparation effort and review capacity. Do not favor applications by default. Explain the tradeoff and any collection that must proceed alongside the pilot.
If applications are not the immediate need, compare intake, check-ins or existing notes. Prefer a bounded pilot whose results can be checked. Keep baseline collection on schedule even when the pilot concerns a different stage.
Step 3 — Establish a baseline before the change you want to measure
Collect the baseline before the intervention or change it is meant to precede. It may need to happen alongside your first pilot rather than afterward. The Magenta Book’s evaluation guidance emphasizes early planning for baseline and comparison data. A before-and-after record can describe change; causal questions require an appropriate evaluation design.
An application may contain usable baseline fields. Check whether it was completed before relevant services, measures the same construct, uses a comparable scale and links to the correct participant. Selection incentives can influence answers. Goals and barriers provide useful context but are not automatically baseline outcome measures. Reuse suitable fields; add or verify the rest at intake without asking people to repeat information unnecessarily.
Use baseline information to plan support. A transport barrier reported before a course gives staff a reason to follow up about access. Confirm the person’s situation and available support; an AI tag should not determine the response alone. Record what was agreed so later notes show whether the barrier changed.
Repeat comparable measures at a named later moment when change is the question. Keep the selected change measures comparable and document changes to the instrument or population. This does not require every location to use an identical form. A later score alone cannot show the earlier state; paired scores still do not prove causation. If the baseline is missing, label it missing rather than reconstructing a fictitious answer from later notes.
Here are my application, intake form and measure map: [paste them]. For each candidate baseline field, check timing relative to services, participant identity, wording, scale, source and whether selection could affect the answer. Return REUSE, VERIFY or ADD AT INTAKE, with reasons. Name the later collection moment and who will own it. Do not assume an application answer is a valid baseline or invent missing values.
Step 4 — Connect collection moments through the participant record
With the first workflow chosen and baseline requirements established, map the remaining collection moments. Work backward from the decision and the likely timing of change. The right schedule depends on the program, participants and reporting commitments.
The common menu, from which each program picks:
- Mid-program check-in — the same baseline questions re-asked, plus “what’s getting in your way right now?” Catches people drifting while support can still be adjusted.
- Mentor, coach, or staff notes — recurring, unstructured, and a source of context that needs review, if they are read on arrival instead of filed.
- LMS or attendance signals — engagement data that already exists; connected to the record, it makes an attendance change visible without assuming its cause.
- Exit — closes the before/after pair the baseline opened; completion alone is not change.
- Follow-up (3, 6, or 12 months) — can observe later outcomes such as employment status, wages or persistence. Some outcomes occur during delivery or at exit; choose timing for the actual question.
- The demand side — for the social enterprise: employer requirements, openings, and placements, so candidate supply and employer demand reconcile instead of living as two unrelated counts.
Choose against two lists: what your funder or board must see (their report defines mandatory moments), and what your team must decide (an early-warning list needs mid-program data; a staffing decision needs LMS signals). Also retain collection needed for service delivery, safeguarding or other applicable obligations. Remove a collection moment only after checking these purposes.
Design the record before automating collection. Keep a participant identifier, program enrollment, collection date, measure definition and source reference with each observation. Store repeated observations as separate dated records. A person may join more than one program; do not combine those enrollments into one undated row. A spreadsheet can support a small pilot if these relationships remain explicit.
| Moment | What's collected | Who provides it | Decision it supports |
|---|---|---|---|
| Application (if suitable) | Goals, barriers, confidence baseline, qualitative answers | Applicant | Evidence for selection; reuse baseline fields only after checking suitability |
| Mid-program check-in | Same confidence question + "what's in your way right now?" | Participant | Early warning while support can still be adjusted |
| Mentor / staff notes | Topics, progress, blockers — weekly, unstructured | Mentor | Context for a staff follow-up, with sources and human review |
| LMS / attendance signals | Engagement data that already exists | System | Flags an attendance change for follow-up; does not establish disengagement |
| Exit | Same measures as baseline + "would this have happened anyway?" | Participant | Closes the before/after pair — completion is not change |
| Follow-up (3–12 months) | Employment, wage, persistence | Participant | Later outcomes; duration requires evidence across the relevant interval |
| Employer demand (social enterprise) | Requirements, openings, placements | Employer | Supply and demand reconcile into one diagnosis |
Watch · 3 min · Workflow example
Turn Theory of Change Into Daily Decisions
Watch a Sopact explanation of connecting a theory of change to day-to-day evidence. Use the workflow as a discussion aid; the video is not proof that a particular evaluation design establishes causal impact.
Common mistakes
Delaying the baseline for a pilot. Prioritize a useful workflow without missing the pre-program observation window. Early operational value and a valid baseline serve different purposes; plan both.
Claiming change without comparable observations. Identify the population, timing and instrument used at each point. Report missing follow-up and distinguish individual change from differences between respondent groups.
Adding collection moments no one asked for. Every moment must serve the funder’s report or a real decision your team makes. Unnecessary questions add burden; check whether each serves a legitimate operational, learning, safeguarding or reporting purpose.
Flattening the history. Link people, enrollments and dated observations without overwriting earlier records. An identifier helps only when duplicate identities, permissions, definitions and amendments are also managed.
How does a theory of change guide data collection?
A theory of change guides data collection by identifying the outcomes and assumptions that require evidence. Each outcome should be mapped to an indicator, data source, collection method, responsible person, and collection moment. This prevents organizations from collecting information that is easy to count but cannot show whether stakeholders experienced the intended change.
How do you turn a theory-of-change outcome into an indicator?
Start by rewriting the outcome as an observable change in a defined population. Then ask, “What would we expect to see if this change occurred?” Select the smallest defensible set of measures, define their scales or calculations, and identify when and from whom they should be collected. Some outcomes need several measures to avoid a misleading conclusion. Use comparable measures at baseline and follow-up. Repeating suitable wording and scales is one approach; a different instrument needs a justified mapping before change can be assessed.
What should a theory-of-change data-collection plan include?
A practical plan should include the outcome, indicator, data source, collection method, responsible person, collection frequency, baseline moment, follow-up moment, and participant identifier. It should also include monitoring questions for important assumptions. Every data field should support an outcome claim, test an assumption, satisfy a reporting requirement, or inform a program decision.
Can an application form be used as a baseline?
Yes, when the application is completed before services begin and captures the same outcome measures that will be repeated later. For example, an application can establish initial employment status, wage, confidence, goals, or barriers. It should not be treated as a baseline when questions, scales, or participant identities cannot be matched reliably with later responses.
What is the difference between a data-collection plan and an M&E plan?
A data-collection plan specifies what information will be collected, from whom, when, how, and where it will be stored. An M&E plan is broader: it also defines indicators, targets, responsibilities, analysis methods, learning questions, reporting schedules, and how findings will influence decisions. The workflow in this article forms the data-collection foundation of the broader M&E plan.
How often should outcome data be collected?
Collect outcome data when meaningful change could reasonably occur and when the result can inform a decision. Common moments include baseline, mid-program, exit, and three-, six-, or twelve-month follow-up. More frequent collection is not automatically better. Every collection moment should support a comparison, reporting requirement, early intervention, or program decision.
Should every theory-of-change outcome have an indicator?
Every outcome the organization intends to manage or report should have at least one defensible indicator. Some long-term or system-level outcomes may be beyond the organization’s practical measurement capacity. Those should be marked clearly as unmeasured or outcomes with untested contribution hypotheses rather than supported with weak proxy measures.
Exercise: test your one-page collection plan
Use one outcome and three fictional or appropriately de-identified records. The aim is to find ambiguity before a real collection round.
- Write the outcome and the decision the evidence will inform.
- Define the population, measure, source, baseline, follow-up and missing-data rule.
- Assign collection and review owners; name who can access identifiable responses.
- Ask a colleague to apply your rules independently to the same records.
- Resolve disagreements, save the definition version and schedule the first review.
Fictional check: a cohort has 40 eligible starters. At follow-up, 30 have known employment status and 18 meet your definition. That is 60% of known responses, with 75% follow-up coverage; 10 remain unknown. Do not report 60% as the verified outcome of all 40 starters. Keep a count of departures and explain who remains in the reporting denominator.
Ready to continue? Another staff member can identify the same eligible records, repeat the calculation and locate each source. If not, revise the definition before building more forms.
Put the plan into a connected workflow
In a configured Sopact workflow, forms, uploaded documents and notes can be linked to the relevant participant or organization and analyzed using agreed instructions. Start by testing extraction and classifications against source records. Staff review uncertain evidence and decide what action to take. Availability of external-system imports and access controls must be confirmed for your implementation.
For baseline design, use the reference on designing an intake form that captures a usable baseline. If selection is your immediate task, use the optional application-review lesson. Explore the Case Management solution when you are ready to connect collection across staff and programs.