What is impact data?
Impact data is the evidence that a program changed the people it serves — the outcomes, the participant records, and the open-ended responses that show what changed and why. It is distinct from activity data (what a program did) because impact data ties a measured change to the person it changed. Its value depends less on having it than on being able to read it as one connected thing. The failure is rarely missing data; it is drift.
Most impact data software is a place to put the data — a survey tool, a spreadsheet, a BI dashboard — and storing it is the easy part. The hard part is that definitions drift by year and team, identifiers do not match across sources, and open-ended responses sit as unread text, so by report time the data cannot be read as one thing. This guide covers what impact data is, why it drifts, what impact data software has to do beyond store it, and the dictionary that keeps it readable.
Key takeaways
- Impact data is the evidence a program changed people — outcomes, participant records, and the open-ended reasons — as distinct from activity data, which is what the program did.
- Sopact calls the real failure The Drift Problem: impact data breaks not from missing data but from definitions, IDs, and formats drifting apart, so it cannot be read as one thing.
- Most impact data software just stores. A survey tool, a spreadsheet, a BI dashboard — storing is easy; reading the data as one connected record is the job.
- Open-ended data is where impact data is thrown away. Stored as text and never coded, the 'why' behind every outcome goes unread.
- A data dictionary is what holds impact data together. One definition per metric, one participant ID, one codebook — defined once, so the data does not drift.
The Drift Problem: the failure is rarely missing data.
Teams assume their impact data problem is gaps — data they did not collect. Far more often the data exists and cannot be used, because it has drifted: the same outcome defined three ways across three years, participant IDs that do not match between the survey tool and the program database, open-ended responses stored as text nobody coded, and units that changed without anyone noting it. At report time, reconciling it is a project.
Sopact calls this The Drift Problem, and it is why impact data software has to be built to read the data, not just store it. The difference is a data-model one — a place to put data, where drift accumulates silently, versus one record per participant against one dictionary, read on arrival, where drift cannot start. The wider measurement practice is on impact measurement and, for nonprofits, nonprofit impact measurement. The stage below shows one report pulled both ways.
Stage 1
Pulling impact data for a report
where impact data drifts
TodayData spread across surveys, sheets, docs · Definitions differ by year and team · Reconciled by hand each report⚠ The failure is rarely missing data — it is drift: the same metric defined differently across years, IDs that do not match across sources, and open text nobody coded, so the data cannot be read as one thing.
The Loop on this stage with Sopact
Collect — clean at the source
SurveysDocumentsProgram recordsOpen-ended text
→ every source lands on one persistent ID
On arrival — read automatically
Intelligent Cell
Each open-ended response is themed against your dictionary on arrival, so the qualitative data is read, not stored as unusable text.
Intelligent Row
Every source resolves to one participant under a persistent ID against one dictionary, so the impact data stays readable instead of drifting apart.
Ask & act — the Assistant
“Which outcomes moved this year, on data defined the same way as last year?”
→ A report read straight from the record — impact data that holds, not a reconciliation project.
Impact data software: built to read, not just store.
Impact data software is worth its name only if it reads the data — codes the open-ended responses, holds one participant identity across sources, and keeps every field to one definition — rather than being another place to store it. A survey tool stores responses, a BI dashboard visualizes whatever is loaded in, and neither reads the qualitative half or prevents drift.
The property to test in any impact data platform is whether an outcome can be traced from the report back to the participant response that produced it, without an export. That is the difference between a data impact platform and a data warehouse. The analytics layer is on nonprofit analytics, the tool comparison on impact measurement software, and the reporting end on impact reporting.
A traditional impact data stack vs Sopact
| Dimension | Traditional stack | Sopact |
|---|
| Where data lives | Surveys, sheets, BI, docs — separate | One record per participant |
| Definitions | Drift by year and team | One data dictionary, defined once |
| Open-ended data | Stored as text, rarely read | Coded against the dictionary on arrival |
| Reading it | Reconciled by hand each report | Read on arrival, traceable to source |
The impact data dictionary that holds it together.
An impact data dictionary is what prevents drift: one definition per metric, a persistent participant ID, a fixed codebook for open text, demographics at intake, locked units and scales, and a recorded source for every field. Defined once, it keeps impact data readable across years and teams — which is the whole of data impact management.
The rule teams break most is the first — letting a metric's definition drift between reporting cycles — which turns a trend into an artifact. Read the six rules below. Building the dictionary is walked through in how to build a data dictionary, and the framework the metrics map to on five dimensions of impact.
Six rules for an impact data dictionary that holds
| Rule | Why it matters |
|---|
| One definition per metric | The same number means the same thing every year |
| A persistent participant ID | Sources join without manual matching |
| A fixed codebook for open text | Qualitative data is comparable, not re-invented |
| Demographics captured at intake | Every later disaggregation is possible |
| Units and scales locked | A trend is real, not a format artifact |
| Source recorded for every field | Each number traces back to where it came from |
Data assembled at report time has already drifted. The Loop reads it on arrival.
Impact data reconciled once a year has already drifted by the time anyone reads it. Reading each source against one dictionary as it arrives is what keeps it from drifting in the first place. That is the premise of the Loop, Sopact's method for continuous impact intelligence: collect clean at the source, analyze the moment data arrives, improve while you can still act.
The Loop is also what makes impact data defensible. Every figure traces to the response and participant behind it, and the same definition returns the same number twice, so a report resolves to its source rather than to a reconciliation. That standard has its own chapter in traceability and transparency. Where impact data feeds management is on impact measurement and management.
One method, three moves that never stop
1 · CollectClean at the source; every source on one record, one dictionary.
2 · AnalyzeOn arrival; open text coded, definitions held, no drift.
3 · ImproveIn time to act; the report is read from the record, not rebuilt.
Then the cycle runs again, a little sharper each year. Read the method: the Loop methodology →
Find the drift in your impact data
The fastest way to see the Drift Problem is to check your own impact data against one dictionary. Each prompt below pastes into Sopact Sense's Assistant, or reasons through with your team; the arrow above each links the Academy walkthrough that shows the expected output and the tips.
Academy walkthrough → Build the impact data dictionary
Build an impact data dictionary from this program's data: [PASTE FIELDS OR OUTCOMES]. For each metric, give one definition, the unit, the source, and the codebook if it is open text. Flag any metric defined more than one way and any field with no source. Return a table: Metric / Definition / Unit / Source / Codebook.
Academy walkthrough → Trace a number to its source
For each figure in this impact report, reconstruct its source: [PASTE REPORT]. Name the responses behind it, the calculation, the definition used, and any filter. Flag any figure that cannot be traced or whose definition differs from a prior report. Return a table: Figure / Responses / Definition / Traceable? / Drift?
Academy walkthrough → Find where identity breaks
Review how our impact data is collected for whether sources join to one participant: [PASTE SOURCES + IDENTIFIERS]. Identify where a participant in one source cannot be matched to the same participant in another, and what identifier would fix it. Return a table: Source pair / Match key / Where it breaks / Fix.
Academy walkthrough → Separate impact data from activity data
From this dataset, separate impact data from activity data: [PASTE FIELDS]. Mark each field as activity (what the program did) or impact (a change in the participant), and for each impact field confirm it ties a change to a specific person. Flag anything labelled impact that is really an activity count. Return a table: Field / Activity or impact / Ties to a person?
Learn the how-to in the Academy
Each walkthrough is a hands-on companion written to run on your own data: what to do, the prompt to run, the output to expect, and the tips that keep it reliable.
Watch: reading impact data on arrival against one dictionary, so definitions, identity, and open text stay connected instead of drifting into an annual reconciliation.
Frequently asked questions
What is impact data?
Impact data is the evidence that a program changed the people it serves — the outcomes, the participant records, and the open-ended responses that show what changed and why. It differs from activity data, which is what the program did, because impact data ties a measured change to the person it changed. Sopact's framing is that impact data's value depends on being readable as one connected thing, and its usual failure is drift, not absence.
What is impact data software?
Impact data software is the system that collects, connects, and reads impact data — not just stores it. Most tools stop at storage: a survey tool holds responses, a BI dashboard visualizes whatever is loaded, and neither codes the open-ended data or prevents definitions and identifiers from drifting. Sopact is built to read impact data — coding open text, holding one participant identity, and keeping every field to one definition — which is what the word should mean.
What is a data impact platform?
A data impact platform is meant to be the single place a program's impact data lives and is read, rather than a warehouse it is dumped into. The test of a real platform is whether an outcome traces from the report back to the participant response that produced it without an export. Sopact functions as that platform by keeping one record per participant against one dictionary, so the data stays connected and traceable.
What is data impact management?
Data impact management is the discipline of keeping impact data usable over time — one definition per metric, one participant identity, a fixed codebook for open text, and a recorded source for every field. It is mostly the work of preventing drift. Sopact operationalizes it through a data dictionary defined once, so the same number means the same thing across years and teams rather than being reconciled each cycle.
Why does impact data fail if it is not missing?
Because it drifts. The same outcome gets defined differently across years, participant IDs do not match across sources, open-ended responses sit as unread text, and units change unnoticed — so at report time the data cannot be read as one thing and reconciling it becomes a project. Sopact calls this the Drift Problem and prevents it by reading every source against one dictionary on arrival, where drift cannot accumulate.
What makes impact data AI-ready?
AI-ready impact data meets three conditions before a model runs: it is connected to one participant identity, defined against one dictionary, and complete enough that the open-ended responses are present and codeable. An AI run over drifted, disconnected data produces confident nonsense. Sopact enforces the three conditions at collection, so the analysis the AI performs is reproducible and traceable rather than a guess over messy inputs.
What is the difference between impact data and activity data?
Activity data is what a program did — sessions delivered, people served, materials distributed. Impact data is what changed in the people served — a skill gained, a job secured, wellbeing improved — tied to the specific person it changed. A funder asking about impact is asking for the second. Sopact keeps impact data on a participant record so a change always traces to a person, which is what distinguishes it from an activity count relabelled as impact.
How do you keep impact data usable over time?
With a data dictionary: one definition per metric, a persistent participant ID, a fixed codebook for open text, demographics captured at intake, locked units and scales, and a source recorded for every field — all defined once and held across cycles. That is what prevents the drift that makes old impact data unusable. Sopact builds the dictionary into collection so the data stays readable rather than degrading between reports.
Next: compare the tools on impact measurement software, or build the dictionary via how to build a data dictionary.
The Drift Problem
01Not missing dataThe data exists — it just cannot be read
02Definitions driftThe same metric, defined three ways
03IDs and formats driftSources will not join; units change unnoticed
04One dictionary, one IDRead on arrival, so drift cannot start
The Drift Problem: impact data fails not from missing data but from definitions, IDs, and formats drifting apart — which is why impact data software must be built to read it, not just store it.