What impact reporting is, why a fast and accurate report still gets ignored, the four layers of context that make a number usable, and the failure modes that quietly break a portfolio total.
Impact reporting is the practice of explaining what changed, for whom, how the organization knows, and what it will do next — in a form the reader can act on. A report can be fast, accurate and clean and still be unusable, because a number on its own does not carry its source, its stage, the framework it maps to, the audience it was cut for, or the organization's own words. That is where reporting usually breaks: the figure is right, and the reader still cannot tell whether it is a baseline or an endline, which outcome it evidences, or what decision it was meant to inform.
Watch (6:07): why a report that is fast, clean and accurate still gets set aside — and what the number should have been carrying with it.
Speed is not the problem it is usually diagnosed as. Reports have become far quicker to produce, and the share of them that changes a decision has not moved with it. What decides whether a report gets used is whether the context travelled with the number — and context is something you design and write down, because no tool infers it.
Key takeaways
Four things have to be true before a reader can act on a figure. It has to be accurate — the number is right. Comparable — it means the same thing as the number beside it. Relevant — it answers the question this particular reader has. And credible — the reader can see where it came from. Most reporting effort goes into the first one, because it is the only one with an obvious artefact behind it. The other three decide whether the report is used.
A data dictionary defines the field: its name, its type, its allowed values, the calculation behind it, who owns it, how often it is collected. That work is genuinely load-bearing — without it the number is not right, and everything downstream inherits the error. But a dictionary describes the container, not the reading. It cannot tell you that this 3 is a baseline and that 3 is a drop from 5, because both are valid values of the same well-defined field. Build the dictionary first anyway; the Academy walkthrough on building a data dictionary covers the artefact itself.
Comparability is a fact about two programs, not about one field: it requires that both numbers were produced under the same definition at the same stage, and no dictionary you write can constrain someone else's. Relevance requires knowing the decision the reader has to make before the cut is chosen, which is knowledge about a person rather than about data. Credibility requires that the reader can open the record behind the figure, which is a property of how the evidence was collected and kept, not of how the field was named.
| Quality | The question it answers | What it actually requires | Does the dictionary give it? |
|---|---|---|---|
| Accurate | Is this figure right? | A defined field, allowed values, and one fixed calculation | Yes — this is exactly what a dictionary is for |
| Comparable | Does it mean the same as the number beside it? | A shared definition and a shared stage, agreed across both programs | No — your dictionary cannot govern someone else's field |
| Relevant | Does it answer the question this reader has? | The reader's decision, written down before the cut is chosen | No — this is knowledge about a person, not a field |
| Credible | Can the reader see where it came from? | Source record, timestamp, and a calculation trail that opens | No — traceability is a property of the record |
Context is not one thing to attach. It is four, and they fail separately: a report can carry perfect data context and no audience context, and still be unreadable to the person holding it.
The first layer is what surrounds the value on the record. Which form or interview it came from, whether it was captured at baseline, mid-point or exit, when, and what the question was trying to establish. Without it, a value is a number with no position in time, and change cannot be read at all. This layer belongs beside the field on the record, not in a separate methodology note that arrives three months later.
The second layer is the outcome, dimension or standard the field evidences: a step in a theory of change, one of the five dimensions of impact, an IRIS+ metric, or a custom framework the organization built for itself. That mapping belongs next to the field in the dictionary, not inside a prompt someone retypes each cycle, because a mapping that lives in a prompt is rebuilt — and quietly changed — every time the report is produced.
The third layer is the decision the report exists to support. A trustee committee approving a tranche, a program officer changing delivery this month, and a board setting strategy need three different cuts of the same evidence. Most reports pick one cut by habit — usually the total — and hand it to all three. Writing the reader's decision down before choosing the cut is what the Academy calls a sourced funder context profile.
The fourth layer is how the organization actually talks: what it calls a participant, which units it uses, which claims it is willing to make and which it holds back. A report written in generic language reads as though it could be about any program, and a reader who cannot hear the organization in it discounts it. This layer is the one most often left to whoever is drafting, which is why the same measure gets described four different ways in one document.
| Layer | What it attaches to the number | The question it settles | Where it has to live |
|---|---|---|---|
| Data | Source, stage, timestamp, and the purpose of the question | Is this a baseline or an endline? | Beside the field, on the record |
| Framework | The outcome, dimension or standard the field evidences | Which change does this count as proof of? | In the data dictionary, not in a prompt |
| Audience | The decision the reader has to make and the cut it needs | Why am I being shown a total? | In a written reader profile, before the cut |
| Expression | The organization's own terms, units and admissible claims | Whose words are these? | In a terminology record the drafter works from |
Audience context is the layer teams most often skip, because a single report feels more efficient than three. It is not: a report cut for nobody in particular is read by nobody in particular. The three readers below can be served from one governed evidence base without changing a single definition — what changes is the cut, the horizon and the level of detail.
| Reader | The decision in front of them | The cut they need | What makes it unusable |
|---|---|---|---|
| Trustee committee | Release, hold or decline the next tranche | Commitments against evidence, with the gaps named | A portfolio total assembled from incompatible definitions |
| Program officer | Change delivery inside this cycle | Who is drifting now, and the reason in their own words | Aggregates only, arriving after the cohort has finished |
| Board | Set or defend strategy over several years | Direction across periods, with material risks and limits | A single period with no baseline and no stated uncertainty |
These are the three failure modes that make a technically clean report unusable. None of them is a data-quality problem in the ordinary sense — every value involved is valid.
Two investees each report 40 jobs created. One counts every role filled during the year, including seasonal and replacement hires. The other counts only net new full-time positions still held at year end. Both figures are accurate against their own definition, both pass every validation rule, and the portfolio total of 80 describes nothing that exists.
The fix: agree the definition before the field is collected, not while the roll-up is being assembled, and record which definition each figure was produced under so an incompatible pair refuses to add rather than adding silently. The Academy walkthrough on giving every number one definition is the working method for this.
On a five-point confidence scale, two participants both answer 3. For one it is a baseline — the first time they have been asked. For the other it is a drop from 5 recorded six weeks earlier. Averaged together they contribute an identical value, and the one signal in the pair that should trigger a conversation this week disappears into the mean.
The fix: carry the stage and the prior value on the record rather than in the analyst's memory, and keep the participant identified across waves so the second 3 can be read as a change rather than as a level. Without a persistent identity linking wave one to wave two, this failure is not detectable at all.
In most organizations the context exists. It is in the head of the person who designed the survey, or in a tab called definitions that has not been opened since the last cycle. Because it is not attached to the data, it gets rebuilt each reporting period from memory, and it drifts a little each time — a stage boundary moves, a denominator changes, an exclusion is forgotten. Two years later nobody can say whether the trend is real.
The fix: treat context as a governed artefact with a version and an owner, sitting next to the field it describes. That is also the condition under which the same question returns the same number twice, which the Academy covers in getting stable results from governed data.
The obvious response to a slow reporting cycle is to point an AI assistant at the data and ask it questions in plain language. That works, and it is genuinely faster. It also does not solve the problem this page is about, because an assistant reads whatever context it was given and answers confidently from it — including when it was given none. Ask it for jobs created across the portfolio and it will add 40 and 40, exactly as a spreadsheet would, and it will explain the result fluently.
The useful version of the claim is narrower and more demanding: an assistant is only as good as the context attached to the data. Where stage, framework mapping, audience and terminology sit beside the field, an assistant can honour them — refuse an incompatible sum, cut the report for the reader who asked, use the organization's own words, and cite the record behind every figure. Where they do not, it produces a faster version of the report nobody could act on. Attaching that context is design work, and it happens before the question is asked. The Academy chapter on generating the audience-specific report from evidence shows what that looks like in practice.
A report produced once a year is a description of a cohort that has already finished. The context problem and the timing problem have the same root: evidence read long after it arrived cannot be corrected, and the person who could have explained an odd value has moved on. Reading each response as it lands is what makes the missing context recoverable while it still exists.
That is the premise of the Loop — collect clean at the source, analyse on arrival, improve while you can still act. The credibility half of it has its own chapter in traceability and transparency: every figure in a report resolving to the exact response, note or document it came from.
One method, three moves that never stop
Then the cycle runs again, a little sharper each period. Read the method: the Loop methodology →
Context only helps once it exists outside somebody's head. Each prompt below takes one layer and turns it into an artefact you can keep; paste it into Sopact Sense's Assistant, or work it through with your team. The arrow above each links the Academy walkthrough with the expected output and the tips.
Academy walkthrough → Give every number one definition
Write the operational definition for this measure so two different programs would count it identically: [PASTE MEASURE + PROGRAM]. Give me who is eligible, the qualifying event, the observation point, the exact calculation with its denominator, the acceptable proof, how exceptions and missing values are handled, and the decision it informs. Then list every way two teams could legitimately read it differently.
Academy walkthrough → Map each field to the framework
For each question in this reporting requirement, tell me what evidence would actually answer it, where in the workflow that evidence can be captured, who owns it, and what should happen when it is missing: [PASTE REQUIREMENT]. Mark each row decision-critical or administrative. Return a table: Question / Evidence / Source and moment / Owner / Missing rule.
Academy walkthrough → Write down what the reader needs
Build a context profile for this reader: [PASTE FUNDER / COMMITTEE / BOARD]. Separate three columns — what they have explicitly required, with the source and date; what we have observed them ask for repeatedly, with examples; and what we do not know and need to confirm, with an owner. Do not turn an observation into a rule. Then state the one decision this report has to support.
Academy walkthrough → Make every figure traceable
For every headline figure in this draft, build the trail back to evidence: the claim, the calculation, the definition version it used, the source records, and the reviewer who approved it. Where a step is missing, write MISSING and name what would close it — do not fill the gap with prose. Return a table. Draft: [PASTE]
Each walkthrough is short and practical: what to do, the prompt to run, the output to expect, and the tips that keep it reliable.
Impact reporting is the practice of explaining what changed, for whom, how the organization knows, and what it will do next, in a form the reader can act on. It combines quantitative results, qualitative evidence, documents and change over time, separates what the evidence supports from what remains uncertain, and lets the reader inspect where each important figure came from.
Because the reader cannot act on them. A report can be fast, accurate and clean and still be unusable if the number arrives without its source, its stage, the framework it maps to, the audience it was cut for, or the organization's own words. That is a context problem rather than a speed problem, and producing the same report faster does not fix it.
Impact measurement defines and observes change. Impact measurement is the workflow; impact reporting communicates the relevant evidence, its interpretation, its limits and the resulting action to a specific audience. Reporting is an output of that workflow, not a replacement for it, which is why a reporting problem is usually a measurement-design problem surfacing late.
Only partly. A data dictionary defines the field — name, type, allowed values, calculation, owner — and that buys accuracy. It cannot tell you how to read the answer, so it does not deliver comparability, relevance or credibility. Those three come from context attached beside the field: source and stage, framework mapping, the reader's decision, and the organization's terminology.
Data, framework, audience and expression. Data is the source, stage, timestamp and purpose behind a value. Framework is the outcome, dimension or standard the field evidences. Audience is the decision the reader has to make and the cut it requires. Expression is the organization's own terminology, units and admissible claims. They fail separately, and each has to be attached deliberately.
Because a shared field name is not a shared definition. Two investees can each report 40 jobs created — one counting every role filled including seasonal hires, the other only net new full-time positions held at year end — and both figures are accurate. The portfolio total of 80 describes nothing real. Comparability requires the same definition at the same stage, agreed before collection.
The formal report may be quarterly or annual, but decision-critical evidence should be read at the pace at which the team can still respond. Attendance may need same-week review, employment retention a 90-day window, a board a quarterly view. Reading on arrival is also what keeps missing context recoverable, because the person who could explain an odd value is still reachable.
An assistant can draft and structure a report from supplied material, and it is genuinely faster. It is only as good as the context attached to the data, and it will answer confidently from whatever context it was given, including none. Where stage, framework mapping, audience and terminology sit beside the field, an assistant can honour them and cite its sources; where they do not, it produces an unusable report more quickly.
Keep missing, not recorded, not applicable and negative results distinct rather than collapsing them. When sources disagree, preserve both, describe the disagreement, name the owner of the resolution and narrow the claim. A transparent partial answer is more usable than a polished unsupported one, because the reader can still decide what to do about the gap.
Next: put the structure on the page with the impact report template, or, if you are choosing tooling rather than method, compare options on the impact reporting software page.