Reproduce a result, explain a changed number and keep the evidence behind both. Use a worked example to control data snapshots, calculations and qualitative coding.
To get stable results from AI-assisted impact data, do not ask a model to reinterpret a pile of raw files each time. Give it governed context: an approved Evidence Map, data dictionary, source records, calculation rules, qualitative codebook, and version history. Then place a controlled calculation layer around AI. AI can interpret a plain-language question and prepare the query; fixed rules calculate the result; reliability checks compare it with the prior release, cite the evidence, and flag anything that still needs human judgment.
Goal: Check a reported result against known records and agreed rules.
Start with: Bring one metric definition, a small source dataset and the result you expect. Keep qualitative interpretation separate from arithmetic.
Carry forward: Leave with a repeatable test and a list of unresolved issues. Next, make the source trail clear enough for someone else to inspect.
This chapter is for program, MEL, grant, foundation, portfolio, and data leads who need an answer they can use repeatedly—not merely a fluent summary. The related lesson connected collection to the workflow. This lesson turns that evidence into a reviewed result package: the result, its scope and rule versions, its supporting records, any changes from the prior result, and the person who approved it.
The reliability rule
For these fixed counts and rates: same approved scope + same evidence snapshot + same reviewed coding decisions + same calculation rules = the same numeric result.
If evidence or a rule legitimately changes, the result may change. Reliability means the change is versioned, explained, reconciled, and approved—not silently hidden.
An AI-generated answer can vary when the context, records, model or interpretation changes. If “engaged” is undefined, a model may make an unsupported assumption or ask for clarification. Neither a fluent answer nor repeating the same prompt establishes that the count is correct. Define the question and calculate from approved records.
Imagine asking, “How many young people remained engaged, why did some drop off, and did our intervention help?” A folder may contain enrollment spreadsheets, attendance sheets, WhatsApp messages, interviews, partner PDFs, and an old report. A model can read them, but the raw material does not answer several governing questions:
A more detailed prompt may reduce ambiguity for one run, but it does not by itself create stable identities, approved definitions, persistent code versions, source priority, reconciliation, or a release history. That is why the answer begins in context—not in prompt cleverness.
The Evidence Map tells AI what question is being answered, which approved measure answers it, which sources are allowed, what context is required, and how far the organization may go in interpreting or attributing the result. The data dictionary supplies the exact language and rules behind those choices.
Together, these artifacts function like the operating instructions for the evidence. They give reviewers a way to detect when AI has treated a familiar label as a different definition. They also make different reports possible from the same governed records: a funder may need an approved indicator and disaggregation, while a program lead needs the current barrier, action owner, and next follow-up.
It is a layer that makes AI work through approved definitions, stored evidence, inspectable queries, notes about what a claim can establish, versioned qualitative metadata, and release checks. The generative model can help understand the question and evidence; it is not allowed to silently decide the population, formula, source, or attribution claim.
This diagram is a control design, not a promise that an autonomous agent can resolve every issue. Checks can be manual or configured in software. The responsible person must resolve ambiguous scope, disputed identity and evidence gaps before approving the result.
Do not ask AI to summarize the raw text and then treat its prose as data. Preserve every source excerpt and add governed metadata beside it. AI can propose that metadata; people approve sensitive or material interpretations; controlled queries then count or compare the approved records.
| Metadata stored with the excerpt | Why it matters | Reliability boundary |
|---|---|---|
| Exact excerpt + source ID | Keeps the participant’s words and location in the original message, interview, note, or document | A theme can always be checked against the source |
| Person/entity + time + program stage | Connects the statement to enrollment, first session, service, action, and later follow-up | Prevents the same excerpt from floating outside its context |
| Approved theme or barrier code | Supports repeatable comparison across many responses | Store AI proposal, human decision, override reason, and codebook version separately |
| Indicator or outcome relationship | Shows why the excerpt is relevant to a specific evidence question | Relevance does not turn a statement into proof of the indicator |
| Claim and evidence limit | Separates observation, sequence, participant attribution, plausible contribution, and causal claim | AI cannot upgrade the claim beyond the approved evidence and method |
| Review status + version | Distinguishes proposed, reviewed, approved, rejected, and superseded interpretations | Released counts use only the permitted status and version |
This creates a bridge between qualitative and quantitative evidence without pretending they are the same. A team may count approved occurrences of a theme, but the count remains connected to source excerpts, context, coding rules, attribution metadata, and limitations. For example, if a report says a theme appeared in 12 reviewed responses, retain the 12 source references and the coding rule. This is a hypothetical illustration, not an Open Play result. Decide whether you count people, responses or excerpts; one person may contribute several passages.
The released answer should show what population was counted, which evidence and rule versions were used, what the qualitative coding found, what changed after follow-up, and what cannot be attributed to the program.
Fictional continuation of the attendance workflow: A program lead asks, “Why did enrolled youth miss the first session, and what happened after WhatsApp follow-up?” Before producing a narrative, the team records the scope:
If a young person later attends, the system may accurately say that attendance occurred after the outreach. It should not say WhatsApp outreach caused the return unless the organization has an approved design and evidence capable of supporting that claim. The Evidence Map governs that boundary before the narrative is written.
Fictional continuation of the training course. The approved first release has 80 starters, 60 people with known employment status 90 days after exit and 36 employed. The known-status rate is 60%, with 75% coverage. Keep that release as a fixed record.
| Release or test | Known status | Employed | Rate among known | Explanation |
|---|---|---|---|---|
| Release A | 60 of 80 | 36 | 60% | Approved original snapshot; 20 unknown |
| Rerun A unchanged | 60 of 80 | 36 | 60% | Same records, reviewed statuses and rules must reproduce the count |
| Release B: four late confirmations | 64 of 80 | 39 | 60.94% | Three employed and one not employed; coverage now 80%; 16 unknown |
| Wrong denominator test | 64 of 80 | 39 | 48.75% of all starters | A different valid view if labeled; not the same measure as 60.94% among known statuses |
The increase from 36 to 39 confirmed employed is new evidence about the same follow-up date. It is not evidence that three people became employed between releases. Confirm the date each late response describes. The employment rate alone also hides the change in response coverage.
This is the practical record called a reviewed result package in this lesson; it is not an external certification. Reproducing the same answer does not establish that the underlying data or method is valid. A consistently wrong denominator produces a consistently wrong answer.
Keep a small reviewed test set with clear examples, ambiguous passages, multiple themes and missing context. Re-run a proposed coding change against that set. Compare source-level decisions and investigate disagreements before replacing released codes. Set the acceptance criteria for the task and its consequences; there is no universal agreement percentage that makes every qualitative analysis valid.
Once codes are approved, count the stored decisions using the same unit and rule. Re-generating themes from raw text on every request creates a new analysis, even if the underlying interviews have not changed.
A result may change when evidence, identity resolution, definitions, codebooks, or approved methods change. The system should never overwrite the prior release silently. It should produce a new version and a reconciliation that explains the difference.
| Change | Correct response | What must remain visible |
|---|---|---|
| A late attendance record arrives | Create a new evidence snapshot and rerun the same approved query | Previous value, new value, late record, and release date |
| Two participant records are confirmed as duplicates | Apply the approved identity correction and issue a reconciliation | Merge decision, owner, affected results, and reason |
| A qualitative codebook is improved | Create a new coding version; review material changes before release | Old and new codes, affected excerpts, overrides, and reviewer |
| The definition or denominator changes | Create a new metric version or restatement under an approved policy | Effective date, comparability warning, and authorized approval |
| The underlying AI model changes | Test against reviewed examples; do not silently replace released qualitative decisions | Model/rule version, test results, changed proposals, and approval |
The checks should cover meaning, scope, identity, evidence completeness, query logic, qualitative interpretation, attribution, change history, and release readiness. Their purpose is to prevent an answer from becoming more confident than the governed evidence.
Freeze the evidence snapshot, record the definition and codebook versions, save the calculation or query, preserve qualitative source excerpts, review exceptions, and issue a dated result with a change note. The method can begin in documents and spreadsheets; the difficulty is maintaining it across many files, partners, programs, and releases.
A configured Sopact Sense workflow can keep collection records, documents and agreed definitions together so teams can analyze incoming evidence and ask questions across time. Use this chapter’s checklist to test the workflow: require a reproducible calculation, traceable sources, reviewed qualitative interpretations and a documented change record. Confirm how each control is implemented and exported before relying on it for reporting. People approve definitions, resolve disputed evidence, review sensitive interpretations, decide attribution, and retain final authority.
For evidence affecting people or funding, restrict sensitive data, test qualitative interpretation across languages and groups, document overrides, define retention, monitor classification differences, and require human approval for consequential claims. Reliability is not only numerical consistency; it is disciplined, reviewable use of evidence.
Watch the 6-minute 7-second explanation of data, framework, audience and presentation context. If you watched it earlier, apply it here by checking which context is recorded with your released result.
Govern the context before asking for an answer. Use an approved Evidence Map and data dictionary, freeze the evidence snapshot, apply versioned calculation and qualitative-coding rules, preserve source citations, run identity and conflict checks, compare with the prior release, and require human approval. The same governed inputs and rule versions should reproduce the same released result.
No. Generative interpretation can vary. Reliability comes from the controlled layer around it: approved definitions, fixed evidence snapshots, inspectable queries, stored qualitative proposals and reviews, versioning, reconciliation, citations, and release controls. Use deterministic calculations for governed numeric results and treat AI-generated qualitative metadata as reviewable evidence work, not unquestionable fact.
The data dictionary defines the shared evidence language: fields, identities, events, indicators, values, formulas, sources, owners, access, and versions. The Evidence Map connects a decision or reporting requirement to those approved definitions and identifies the evidence, gaps, breakdowns, and permitted claim. Reliability requires both meaning and purpose.
AI can draft a query, but generating one does not prove it answers the intended question. Check its fields, joins, filters, scope and access rules against approved definitions. Test it on records with known answers, retain the query and result, and resolve ambiguity before release. Calculate the number from the approved query rather than accepting a count written in generated prose.
Preserve the exact source excerpt, use an approved codebook with examples and uncertainty rules, store the AI proposal separately from the reviewed decision, retain overrides and reasons, and version both codebook and model-assisted pass. Test across languages and groups. Released counts should state which review status and coding version they use.
A reliable number can change when late evidence arrives, duplicates are resolved, a source is corrected, or an approved definition or method changes. The system should show the earlier and later values, identify the changed records or rules, state whether comparisons remain valid, and record who approved the new release. Unexplained change is the problem—not change itself.
No. AI can organize evidence and test whether a proposed statement exceeds the supported claim boundary, but causal claims require an appropriate design, method, assumptions, and human judgment. A sequence—outreach followed by later attendance—does not by itself prove that outreach caused attendance. Preserve that distinction in the Evidence Map and report.
Make every released result traceable. The traceability reference connects the figure, theme, or claim to its query, definition, evidence snapshot, source records, qualitative excerpts, rule versions, exceptions, and approval. Stability answers “Will this hold when rerun?” Traceability answers “Can I inspect how it was calculated and what supports it?”
Related practice: You now have a governed, versioned result with a visible reconciliation. The traceability reference shows how to connect every result to its supporting evidence in How Do You Trace Every Result Back to Its Evidence? →
The National Academies’ report on reproducibility explains why rerunning an analysis requires its inputs, methods and conditions. The NIST Generative AI Profile addresses confidently incorrect output and the need for testing. This lesson applies those principles to routine reporting counts; it does not certify an AI system.
For your final report, use How to Write an Impact Report and report examples.
By Sopact Academy · Revised September 12, 2026. All numerical examples in this lesson are fictional practice data.
Bring a report whose numbers changed between versions. Work through the definitions, source records and review steps.
Explore Impact Measurement →