To get stable results from AI-assisted impact data, do not ask a model to reinterpret a pile of raw files each time. Give it governed context: an approved Evidence Map, data dictionary, source records, calculation rules, qualitative codebook, and version history. Then place a controlled calculation layer around AI. AI can interpret a plain-language question and prepare the query; fixed rules calculate the result; reliability checks compare it with the prior release, cite the evidence, and flag anything that still needs human judgment.
This chapter is for program, MEL, grant, foundation, portfolio, and data leads who need an answer they can use repeatedly—not merely a fluent summary. Chapter 7 produced linked evidence inside the workflow. Chapter 8 turns that evidence into a Reliability Release: the result, its scope and rule versions, its supporting records, any changes from the prior result, and the person who approved it.
The reliability rule
Same approved scope + same governed evidence snapshot + same rule versions = same released result.
If evidence or a rule legitimately changes, the result may change. Reliability means the change is versioned, explained, reconciled, and approved—not silently hidden.
Why can raw ChatGPT or Claude give a different answer?
In short: A general chat model is an excellent reasoning and drafting tool, but it is not automatically your governed evidence system. Unless you supply and enforce the context, it must infer what the files mean, which records count, how identities connect, which definition applies, and how strong a claim the evidence supports.
Imagine asking, “How many young people remained engaged, why did some drop off, and did our intervention help?” A folder may contain enrollment spreadsheets, attendance sheets, WhatsApp messages, interviews, partner PDFs, and an old report. A model can read them, but the raw material does not answer several governing questions:
- Does “enrolled” mean registered, eligibility-confirmed, or accepted into a cohort?
- Does “engaged” mean first-session attendance, any attendance, or participation above a threshold?
- Which program, site, cohort, date range, and participant population are in scope?
- Are two similar names one person, two people, or an unresolved duplicate?
- Is an empty attendance cell an absence, a late record, or missing data?
- Should a WhatsApp statement be coded as financial hardship, transport, both, or not enough information?
- Does later attendance show that outreach caused re-engagement, or only that the events occurred in sequence?
A more detailed prompt may reduce ambiguity for one run, but it does not by itself create stable identities, approved definitions, persistent code versions, source priority, reconciliation, or a release history. That is why the answer begins in context—not in prompt cleverness.
How does a well-designed Evidence Map make AI more reliable?
In short: The Evidence Map tells AI what question is being answered, which approved measure answers it, which sources are allowed, what context is required, and how far the organization may go in interpreting or attributing the result. The data dictionary supplies the exact language and rules behind those choices.
Data dictionary
Defines the person, event, indicator, population, denominator, time window, missing states, allowed values, calculation, owner, and version.
Evidence Map
Connects a reporting or decision question to the approved measure, evidence sources, gaps, breakdowns, and permitted claim.
Collection-Moment Plan
Preserves where and when the evidence appeared, its identity and source, the staff action, and what follow-up occurred.
Together, these artifacts function like the operating instructions for the evidence. They prevent AI from treating a familiar label as permission to invent a definition. They also make different reports possible from the same governed records: a funder may need an approved indicator and disaggregation, while a program lead needs the current barrier, action owner, and next follow-up.
What is the controlled evidence layer around AI?
In short: It is a layer that makes AI work through approved definitions, stored evidence, inspectable queries, operational attribution metadata, versioned qualitative metadata, and release checks. The generative model can help understand the question and evidence; it is not allowed to silently decide the population, formula, source, or attribution claim.
A reliability architecture in plain language
1 · The question
A person asks in ordinary language: “Why did enrolled youth miss the first session, and what happened after follow-up?”
↓
2 · Reliability agents interpret and guard the request
They find the approved definition, confirm the cohort and time period, select allowed sources, resolve identities, check missing or conflicting records, build the query, and apply the permitted attribution level.
↓
3A · Deterministic calculation
A saved, inspectable query applies fixed filters, joins, formulas, exclusions, and metric versions to governed records. The query remains traceable with the result.
3B · Governed qualitative reading
AI proposes themes and indicator relationships, but keeps the exact excerpt, source, codebook version, confidence, and review status.
↓
4 · Verified release
The system compares the result with the previous release, explains every change, attaches citations and limitations, and routes material exceptions to a person for approval.
“Agent” here does not mean an invisible bot free to make high-stakes decisions. It means a specialized reliability checker with a bounded job. One checks scope. Another checks identity and duplicates. Another checks definitions and query logic. Another examines qualitative evidence and operational attribution metadata. A final check prevents release when required evidence, citations, or approval are missing.
How can qualitative evidence become reliable without losing the participant’s voice?
In short: Do not ask AI to summarize the raw text and then treat its prose as data. Preserve every source excerpt and add governed metadata beside it. AI can propose that metadata; people approve sensitive or material interpretations; controlled queries then count or compare the approved records.
| Metadata stored with the excerpt |
Why it matters |
Reliability boundary |
| Exact excerpt + source ID | Keeps the participant’s words and location in the original message, interview, note, or document | A theme can always be checked against the source |
| Person/entity + time + program stage | Connects the statement to enrollment, first session, service, action, and later follow-up | Prevents the same excerpt from floating outside its context |
| Approved theme or barrier code | Supports repeatable comparison across many responses | Store AI proposal, human decision, override reason, and codebook version separately |
| Indicator or outcome relationship | Shows why the excerpt is relevant to a specific evidence question | Relevance does not turn a statement into proof of the indicator |
| Attribution level | Separates observation, sequence, participant attribution, plausible contribution, and causal claim | AI cannot upgrade the claim beyond the approved evidence and method |
| Review status + version | Distinguishes proposed, reviewed, approved, rejected, and superseded interpretations | Released counts use only the permitted status and version |
This creates a bridge between qualitative and quantitative evidence without pretending they are the same. A team may count approved occurrences of a theme, but the count remains connected to source excerpts, context, coding rules, attribution metadata, and limitations. In Sopact, this metadata is operational: the reader can move from “financial hardship appeared in 12 reviewed responses” to the relevant words, indicator relationship, attribution boundary, and review history without asking AI to reconstruct the evidence later.
Worked example: a first-session no-show result that can hold up
In short: The released answer should show what population was counted, which evidence and rule versions were used, what the qualitative coding found, what changed after follow-up, and what cannot be attributed to the program.
Illustrative continuation of the Open Play field example: A program lead asks, “Why did enrolled youth miss the first session, and what happened after WhatsApp follow-up?” The system does not immediately produce prose. It first interprets and displays the governed scope:
You asked: Why did enrolled youth miss the first session, and what happened after follow-up?
Before calculating, the reliability layer confirms
Confirmed-enrollee definition v3 · first-session event v2 · selected cohort and period · one governed participant ID · no-show, not-recorded, unreachable, and declined kept separate · WhatsApp theme codebook v2 · only reviewed theme records counted · later attendance treated as sequence, not proof of causation.
The release returns
The approved result and denominator · missing and excluded records · reviewed themes with source excerpts · WhatsApp action status · later attendance status · changes from the previous release · citations, limitations, and approval.
If a young person later attends, the system may accurately say that attendance occurred after the outreach. It should not say WhatsApp outreach caused the return unless the organization has an approved design and evidence capable of supporting that claim. The Evidence Map governs that boundary before the narrative is written.
When should a stable result legitimately change?
In short: A result may change when evidence, identity resolution, definitions, codebooks, or approved methods change. The system should never overwrite the prior release silently. It should produce a new version and a reconciliation that explains the difference.
| Change |
Correct response |
What must remain visible |
| A late attendance record arrives | Create a new evidence snapshot and rerun the same approved query | Previous value, new value, late record, and release date |
| Two participant records are confirmed as duplicates | Apply the approved identity correction and issue a reconciliation | Merge decision, owner, affected results, and reason |
| A qualitative codebook is improved | Create a new coding version; review material changes before release | Old and new codes, affected excerpts, overrides, and reviewer |
| The definition or denominator changes | Create a new metric version or restatement under an approved policy | Effective date, comparability warning, and authorized approval |
| The underlying AI model changes | Test against reviewed examples; do not silently replace released qualitative decisions | Model/rule version, test results, changed proposals, and approval |
What should Sopact AI reliability agents check?
In short: They should check meaning, scope, identity, evidence completeness, query logic, qualitative interpretation, attribution, change history, and release readiness. Their purpose is to prevent an answer from becoming more confident than the governed evidence.
- Interpret the question. Identify the decision, metric, population, period, breakdown, and comparison the person appears to mean.
- Retrieve the approved context. Select the relevant Evidence Map, data-dictionary fields, source hierarchy, calculation, codebook, and attribution rule.
- Inspect the evidence. Flag missing sources, unresolved identities, duplicates, late entries, access restrictions, contradictory records, and stale versions.
- Build an inspectable, traceable query. AI may translate the plain-language request into filters, joins, and calculations, but the governed query—not generated prose—produces the numeric result. Sopact retains the query path with its definition, evidence snapshot, and result.
- Read qualitative evidence with metadata. Propose themes and relationships, preserve exact excerpts, and route sensitive, uncertain, or material items for review.
- Test the claim. Check that the narrative distinguishes observation, output, outcome, participant attribution, plausible contribution, and causation.
- Reconcile and release. Compare with the last approved result, explain every change, attach citations and limitations, and obtain the required human approval.
What can a team do before it has Sopact Sense?
In short: Freeze the evidence snapshot, record the definition and codebook versions, save the calculation or query, preserve qualitative source excerpts, review exceptions, and issue a dated result with a change note. The method can begin in documents and spreadsheets; the difficulty is maintaining it across many files, partners, programs, and releases.
Sopact Sense makes this architecture operational at scale. It can hold the governed context with the evidence, let a user ask questions in ordinary language, prepare traceable deterministic queries, create reviewable qualitative and attribution metadata, run reliability checks, and retain the released answer with its query, source, and version history. People approve definitions, resolve disputed evidence, review sensitive interpretations, decide attribution, and retain final authority.
For evidence affecting people or funding, restrict sensitive data, test qualitative interpretation across languages and groups, document overrides, define retention, monitor classification differences, and require human approval for consequential claims. Reliability is not only numerical consistency; it is disciplined, reviewable use of evidence.
Frequently asked questions
How do you get stable results from AI-assisted impact data?
Govern the context before asking for an answer. Use an approved Evidence Map and data dictionary, freeze the evidence snapshot, apply versioned calculation and qualitative-coding rules, preserve source citations, run identity and conflict checks, compare with the prior release, and require human approval. The same governed inputs and rule versions should reproduce the same released result.
Is Sopact claiming that generative AI is deterministic?
No. Generative interpretation can vary. Reliability comes from the controlled layer around it: approved definitions, fixed evidence snapshots, inspectable queries, stored qualitative proposals and reviews, versioning, reconciliation, citations, and release controls. Use deterministic calculations for governed numeric results and treat AI-generated qualitative metadata as reviewable evidence work, not unquestionable fact.
What is the difference between a data dictionary and an Evidence Map?
The data dictionary defines the shared evidence language: fields, identities, events, indicators, values, formulas, sources, owners, access, and versions. The Evidence Map connects a decision or reporting requirement to those approved definitions and identifies the evidence, gaps, breakdowns, and permitted claim. Reliability requires both meaning and purpose.
Can AI create a database query from a plain-language question?
Yes. Sopact constrains the query to approved fields, identities, metric versions, joins, filters, and access rules. The generated query is inspectable and traceable: it is tested and retained with the Reliability Release, and the request is stopped when the question is ambiguous or evidence is insufficient. The numeric result comes from executing the governed query, not from the model estimating a count in prose.
How do you make qualitative analysis repeatable?
Preserve the exact source excerpt, use an approved codebook with examples and uncertainty rules, store the AI proposal separately from the reviewed decision, retain overrides and reasons, and version both codebook and model-assisted pass. Test across languages and groups. Released counts should state which review status and coding version they use.
Why did a number change if the system is reliable?
A reliable number can change when late evidence arrives, duplicates are resolved, a source is corrected, or an approved definition or method changes. The system should show the earlier and later values, identify the changed records or rules, state whether comparisons remain valid, and record who approved the new release. Unexplained change is the problem—not change itself.
Can AI decide whether a program caused an outcome?
No. AI can organize evidence and test whether a proposed statement exceeds the approved attribution level, but causal claims require an appropriate design, method, assumptions, and human judgment. A sequence—outreach followed by later attendance—does not by itself prove that outreach caused attendance. Preserve that distinction in the Evidence Map and report.
What happens after a result is stable?
Make every released result traceable. Chapter 9 connects the figure, theme, or claim to its query, definition, evidence snapshot, source records, qualitative excerpts, rule versions, exceptions, and approval. Stability answers “Will this hold when rerun?” Traceability answers “Can I see exactly why it is true?”
Next: You now have a governed, versioned result with a visible reconciliation. Chapter 9 shows how to connect every result to its supporting evidence in How Do You Trace Every Result Back to Its Evidence? →