What is impact data?
Impact data is evidence used to understand changes in people, communities, organizations, systems, or the environment and the contribution of a program, policy, grant, investment, or business activity. It includes structured measures, demographics, services, transactions, open-ended responses, interviews, observations, documents, benchmarks, and contextual evidence. Its usefulness depends on clear definitions, relevant context, dates, appropriate permissions, calculations and sources. Personal identity is needed only where the question and authorized workflow require it.
Define impact measures once, reuse them across frameworks
Four funders can mean four spreadsheets rebuilt every reporting season. A data dictionary starts with a more useful question: what does each measure actually mean? Define the population, unit, time period, inclusion and exclusion rules, and calculation once so each report uses the same approved meaning.
Watch: One Data Dictionary, Every Impact Framework: IRIS+, SDG & ESRS Without the Rebuild.
The video explains how shared definitions support framework mappings and sector-specific vocabulary. Label each mapping Exact, Related, or Organization-Specific so a useful connection is not mistaken for an equivalent measure. Funder templates and branded reports can draw on the same governed data; each framework's evidence requirements still need to be checked.
Read the full method: build a portfolio data dictionary and map measures to standards →
Key takeaways
- Impact data is broader than metrics. Numbers can show scale and change; language and documents add evidence about experience, context and possible unintended effects.
- Activity data is not outcome evidence. Services delivered matter, but they do not establish that participants or communities changed.
- A data dictionary protects meaning. Every identifier, measure, option, segment, formula, owner, cadence, and source needs an approved definition.
- AI-ready does not mean ungoverned. Models need authorized evidence, suitable structure, clear boundaries and human review. Check generated citations against the actual source.
- Collect for a decision. Every field should support an operational, evaluation, management, or reporting need.
Impact data often becomes unusable because meaning and context separate
Organizations rarely lack data. They have surveys, spreadsheets, case systems, grant reports, interviews, notes, and PDFs. Problems arise when the same measure has different definitions, records cannot be matched, reporting periods drift, corrections are undocumented, or qualitative evidence is detached from the person or program it describes.
Later cleaning may recover some relationships when reliable identifiers and source records exist. It cannot establish an unknown date or missing outcome merely because a report needs one. The data architecture must preserve meaning and source from the moment evidence enters the workflow.
Keep definitions and context together
A field definition tells a reviewer what a value means. The context tells them how to interpret this observation. A confidence rating of 3 at baseline means something different from a rating of 3 after an earlier rating of 5. Preserve the relevant period, instrument version and record relationship.
Connect only the evidence needed for the decision. That may involve people, organizations, programs, grants or investments. Some questions need a continuing individual record; others can use anonymous feedback or aggregate evidence. Multiple source systems can remain in place if their relationships and responsibilities are clear.
Sopact's approach combines recurring collection, connected evidence, reviewed analysis and governance. Test the required relationships, changes, permissions and reporting in your own workflow rather than assuming that every integration or calculation is automatic.
Watch: why clean data can still be hard to use
The companion explores why a clean field still needs the context of the record and collection period.
How should you evaluate impact data software?
Use one reported claim and trace it through the full data workflow. Include a quantitative measure, an open-ended explanation, a document, a second period, a correction, a missing record, and the audience view.
Routine administration
Program, MEL, grant, and portfolio teams should be able to update definitions, review missingness, correct records, and answer routine questions without rebuilding exports.
Try this in a pilot: Ask an operating user to correct a definition, add a source, and reproduce a result.
Record relationships
Where evidence must connect over time or across sources, define stable identifiers for the appropriate units and check the relationships. Do not require a personal identifier for anonymous or group-level analysis.
Try this in a pilot: Open one result and trace every contributing record to the correct unit.
Volume and coverage
The platform should handle real row counts, long text, files, repeated updates, and exceptions at the required cadence.
Try this in a pilot: Use a full-volume batch and inspect latency, coverage, duplicates, and exclusions.
Longitudinal
Impact data must support change across baseline, delivery, exit, follow-up, reporting periods, and corrected history.
Try this in a pilot: Change a definition and correct a prior value; verify that both history and comparability remain clear.
Qualitative
Open-ended responses, interviews, observations and notes can help interpret a measure and reveal experiences missed by it. They do not automatically establish why a change occurred.
Try this in a pilot: Require supporting and contradictory passages for one finding.
Documents
Applications, partner reports, evaluations, plans, policies, and case documents contain material impact evidence.
Try this in a pilot: Ask a claim that depends on several files and open every cited passage.
Assisted analysis
An assistant should answer only within approved definitions and permissions and disclose included records, calculations, exclusions, and evidence.
Try this in a pilot: Ask the same question twice, change one filter, and inspect exactly how the result changes.
Reproducibility and review
Reliable impact data preserves definitions, transformations, corrections, calculations, qualitative boundaries, model configuration, review, and sources.
Try this in a pilot: Recalculate one number and inspect how one qualitative conclusion was developed from its sources. Qualitative interpretation is not necessarily a deterministic calculation.
Types of impact data
A useful evidence record combines the types required by the decision rather than treating one source as complete.
Scroll horizontally to see all columns →
| Data type | Examples | What it contributes |
|---|---|---|
| Identity and context | Participant, household, partner, site, program, grant, investment, demographics, location, dates | Defines who or what the evidence describes and supports segmentation and longitudinal follow-up. |
| Activity and service data | Enrollment, attendance, dosage, referrals, training, mentoring, funding, engagement | Shows what was delivered, to whom, when, and with what intensity. |
| Outcome and impact measures | Skills, confidence, employment, health, wellbeing, income, stability, environmental or organizational change | Shows the direction, magnitude, distribution, and durability of change. |
| Qualitative evidence | Open-ended responses, interviews, case notes, observations, stories, stakeholder feedback | Provides evidence about experience, possible mechanisms, barriers, unexpected effects and differences between groups. |
| Documents and external evidence | Applications, reports, evaluations, policies, research, benchmarks, plans, verification records | Supports context, assumptions, standards alignment, verification, and source traceability. |
Can you keep existing data systems?
Yes. Keep survey, case, grant, CRM, learning, finance, portfolio, warehouse, BI, and research tools that serve their operational purpose. Connect only authorized evidence needed for a clear decision or report.
Start with the Academy lessons on building a data dictionary, connecting quantitative and qualitative data, and governed data reliability.
Plan a shared core without imposing one form
In a federated network, each location may collect different information for its own work. Agree on the few fields needed for shared decisions, then document their meanings in the data dictionary. Let local teams keep questions that serve their own participants and services.
Collect stable registration context once where practical, and update changing information when relevant. Preserve historical context if it affects a comparison. If a person changes location, the new location should not silently replace the location used to interpret an earlier outcome.
For each proposed comparison, check the unit, population, period, definition and coverage. A mapping can connect field names; it cannot make attendance entries equivalent to unique people. Keep incompatible measures separate or explain an approved transformation.
A worked example: one result, several evidence checks
Suppose a fictional network receives reports from 20 of its 25 member organizations. Those reports describe 800 activity attendances and an estimated 500 people reached. The network should report 80% organization-level coverage, keep attendances separate from estimated people and identify how the estimate was produced.
If the reports overlap in participants, adding local people counts may double-count individuals. That does not make the activity data useless. It means the network should label what it can substantiate, such as locally reported reach, until an appropriate method supports a deduplicated total.
A supporting document from a prior quarter should remain attached to its own period. Comments about transport barriers can guide further investigation, but they do not prove that transport explains every missing report or participation gap. Preserve the source, review the interpretation and assign any follow-up.
Report what the evidence supports
Keep delivery, outcomes and causal claims distinct. A change observed after participation is not automatically an effect caused by the program. Missing evidence should remain visible rather than being filled with invented values or treated as zero by default.
Use the How to Write an Impact Report guide to organize the narrative and browse report examples. The impact dashboard guide shows how to present results with coverage and limitations.
Frequently asked questions
What is impact data?
Impact data is evidence used to understand changes in people, communities, organizations, systems, or the environment and the contribution of an intervention or activity.
What is impact data software?
It is software that helps define, collect, connect, analyze, govern, and report quantitative, qualitative, document, and longitudinal evidence.
What is the difference between impact data and activity data?
Activity data describes what was delivered, such as services, funding, sessions, or referrals. Impact data also examines outcomes, experience, context, contribution, and durability.
What is a data dictionary for impact data?
It is a governed set of definitions for identifiers, fields, measures, response options, segments, calculations, owners, cadence, sources, permissions, and standards mappings.
What makes impact data AI-ready?
Authorized evidence, appropriate record relationships, clear definitions, documented quality and human review. Check calculations and cited sources; repeated AI wording is not proof of a reliable conclusion.
Can impact data include qualitative evidence?
Yes. Interviews, open-ended responses, observations, notes, and documents are often essential for explaining mechanisms, barriers, subgroup differences, and unintended effects.
How do you keep impact data usable over time?
Preserve identity, definitions, dates, units, versions, corrections, permissions, sources, and longitudinal relationships; review quality while evidence can still be corrected.
How should AI be used with impact data?
Use AI to read language, extract fields, identify patterns, flag gaps, and support questions. Keep governance, deterministic calculations, citations, permissions, and human judgment.

