How do you build data governance before an AI pilot?
Before a team asks AI to summarize a case, identify a cohort at risk, or explain an outcome, it needs a governed evidence record. Define what each measure means, who owns it, which records can be linked, which sources are authoritative, and how a reviewer can trace an answer back to its source. The result is not a bigger database. It is a practical way to decide what can be combined, compared, and used in a decision.
This exercise is for a data owner, program lead or operations team preparing an AI-assisted reporting workflow. The same decisions also arise in partner networks and commercial service organizations. Your deliverable is a one-page evidence register and a small test log, not a plan to replace every system.
- Choose one decision and the evidence it needs.
- Define the measures and appoint accountable owners.
- Set permitted uses, access and retention rules.
- Test source tracing, missing evidence and changed definitions.
- Review the results and decide whether to proceed, repair or narrow the pilot.
What does data governance mean before a nonprofit uses AI?
Data governance is the set of decisions that makes information interpretable and accountable. It assigns an owner, a definition, a permitted use, and a source to the data a team relies on.
A nonprofit may already hold relevant information in several places. Their CRM may hold relationships and gifts. A case-management system may hold service delivery. Surveys hold ratings and open comments. Spreadsheets and documents hold the report staff actually use in meetings. Each source can be useful; none automatically tells the whole story.
Ask four questions before connecting sources
Every data element should answer four questions: what is it, who owns it, what can it be linked to, and how can someone inspect it?
What does “enrolled,” “completed,” or “improved” mean?
Who can change the rule and explain it?
Which records may be joined, for what purpose and under which approved rules?
Can a reviewer find the source response, note, document, or calculation?
If one answer is missing, resolve the gap before relying on the output. These four questions organize the evidence work; they are not a complete AI risk assessment.
What should a nonprofit govern first?
Start with a decision that is already challenged in meetings: the score people distrust, the cohort that appears in several systems, or the report staff rebuild by hand.
- Name one decision. For example: “Which participants need follow-up after the next workshop?”
- List the evidence. Include scores, open comments, case notes, attendance, referrals, documents, and programme changes—not only dashboard fields.
- Set the definition. State the exact population, measure, timeframe, and exclusions.
- Assign an owner. The owner can approve changes and explain when a definition or instrument changed.
- Set linking and access rules. Document identifiers, permitted uses, access restrictions and records that must remain separate. Review disclosure risk rather than assuming that removing names makes a group anonymous.
- Test traceability. Ask a reviewer to reproduce one finding from source evidence without rebuilding the report.
Why are definitions more important than clean-looking fields?
A field can be complete and still be misleading when teams use the same word for different events.
| Team | Definition of “enrolled” | Governance decision |
|---|---|---|
| Programme team | Attended one session | Decide whether attendance is the reporting threshold. |
| Funder report | Completed intake | Document the eligible population and rule. |
| CRM report | Status equals active | Decide when and by whom the status changes. |
All three counts may be calculated correctly. They are not comparable until the organization decides which definition serves which decision and records its change history.
Assign responsibilities without making every decision an IT project
The program owner defines what a result means and how it will be used. A data steward maintains definitions, quality checks and the change log. The system administrator configures technical access and integrations. A reviewer checks whether the evidence supports the finding. One person may fill more than one role in a small organization, but record which responsibility they are exercising.
Data owners should be able to maintain ordinary collection rules within agreed boundaries. Bring in technical, privacy or measurement expertise when the proposed change requires it. Independence means knowing which decisions you can make and which require another owner, rather than bypassing the controls that protect the work.
NIST's AI Risk Management Framework Playbook is a voluntary resource for broader AI-risk work. Its core framework treats governance as an ongoing function across the AI lifecycle, not a one-time approval. Use this lesson as a focused data exercise within that wider responsibility.
Make a small evidence register
For each source, record its owner, reporting period, definition, approved uses, access roles and retention decision. Add a location or reference that an authorized reviewer can follow. If two sources disagree, identify the person who resolves the conflict rather than declaring one system authoritative for every purpose.
For example, an attendance register may be authoritative for recorded session attendance, while a corrected participant response is authoritative for that person's updated contact preference. Keep the correction and its reason. Do not erase provenance when resolving a discrepancy.
For personal information, confirm the applicable basis and conditions for the intended processing. For UK GDPR, the ICO guide to lawful basis explains that consent is one possible basis; the appropriate choice depends on the purpose and circumstances. Avoid collecting or sending unnecessary sensitive information to an AI service. Review the provider's terms and configured data handling with the appropriate owner before using real records.
To develop the register further, use reading documents as evidence and the AI access-control exercise. These lessons turn source handling and permissions into checks you can actually perform.
How should qualitative evidence be governed?
Preserve the original words, define how themes are created, and keep a path from every theme or summary back to the source passage.
A confidence result can describe responses under a defined measure. By itself, it cannot establish why confidence changed or whether the program caused the change. That context often lives in open comments, interviews, and case notes. Governance should specify who can access those sources, which permitted uses and confidentiality commitments apply, when a theme is reviewed, and what evidence supports a published finding.
How do you govern local differences and repeated cycles?
Agree the small core needed for shared decisions, then track changes to questions, populations, instruments and delivery. Different sites can keep locally useful questions. A data dictionary records which local fields map to the shared measures and which must remain separate. Reuse appropriate registration details and date changes to context such as location or role; do not ask every person to supply unchanged details every round.
Longitudinal analysis breaks quietly. A new wave may use a revised survey question, a different export, new staff, or a changed programme design. Maintain a visible change log that records what changed, why it changed, who approved it, and whether comparison with earlier periods remains appropriate.
Worked example: a community learning programme
Illustrative example: a programme director sees lower application confidence in the North cohort. Before asking AI for a cause, the team verifies that the score has a stable definition, the cohort is connected to authorized records, the programme context is documented, and source-linked comments repeatedly request more practice before real conversations.
The responsible finding is not “the North cohort is disengaged.” It is a reviewable statement: lower scores appear alongside a repeated request for more practice, and the team should test an additional practice block before the next pulse. A person still decides whether the evidence is sufficient and which action is appropriate.
Where does Sopact fit after the governance work?
Sopact helps teams run a governed evidence process repeatedly; it does not replace the organization’s responsibility to define what it collects or what a finding means.
You can begin with a data dictionary, a change log, structured files, and a manual review process. The work becomes harder when responses arrive continuously, open text grows, definitions evolve, and one report requires several systems. In a Sopact evaluation, test how the proposed setup connects new responses, uploaded documents and analysis to the correct person, organization, event and period. Ask to see source references, the handling of inconsistent definitions and the actual access controls. Confirm the integration and review arrangements needed for your workflow. People retain final judgment.
For an optional explanation of how the evidence stays connected, read Connected Data Intelligence.
What governance checks should happen before an AI pilot?
Run one real decision through the workflow before investing in a broader AI feature.
- Can the team change and document a definition?
- Does each response land on the correct authorized record?
- Can the organization identify who is missing from the evidence?
- Can repeated cycles be compared responsibly?
- Can a qualitative theme be traced to source passages?
- Are comments, documents, and notes governed by access and retention rules?
- Can a reviewer reproduce one number and its explanation?
- Are AI suggestions reviewed by people with programme and privacy context?
Test the pilot with known errors before relying on it
Use fictional records first. In this exercise, 40 people are eligible for follow-up, 32 have a recorded response and 20 of those 32 meet the agreed outcome definition. The observed rate is 62.5%; coverage is 80%. Eight people have unknown outcomes. This is an arithmetic and evidence-handling test, not a real program result.
| Test input | Expected behavior | If it fails |
|---|---|---|
| 20 outcomes among 32 responses; 40 eligible | Report 62.5% observed and 80% coverage with separate bases | Repair the calculation or reporting instruction |
| Eight missing responses | Keep unknown outcomes distinct from failures | Correct the missing-data rule |
| Document from the wrong reporting period | Flag it rather than silently using it | Review source-period checks |
| A restricted case note | Unauthorized user cannot retrieve it in a view, export or AI answer | Pause that access path and fix authorization |
| A revised outcome definition | Show its version and whether the old trend remains comparable | Repair versioning before combining periods |
Have a second reviewer reproduce the numerator and denominator and open the permitted source passages. Record the inputs, rule version, observed behavior, reviewer and correction. A source link is useful only if it actually supports the associated statement.
Decide whether to proceed, repair or narrow the use case. A summarization pilot may be appropriate while an automated decision affecting services is not. Test according to the consequence of an error; passing these examples does not certify the system for every task.
Frequently asked questions
What is data governance for a nonprofit?
It is the practical system for defining data, assigning ownership, setting access and linking rules, and documenting how information can be used. It helps a nonprofit make decisions from evidence it can explain, not merely fields it can export.
Do we need a data warehouse before we can govern data?
No. Start with a decision, a small set of sources, written definitions, ownership, and a traceability test. A warehouse cannot resolve an unclear definition, unclear permitted use, or conflicting identity record.
How do we govern open-ended survey responses?
Keep the original response available to authorized reviewers, document coding rules, apply privacy protections, and retain a reference from each theme to supporting passages.
Can AI help with data governance?
AI can help draft definitions, flag inconsistent labels, suggest themes, and surface missing information. It should not silently decide identity matches, access permissions, or high-stakes conclusions.
How often should governance rules be reviewed?
Review them whenever a programme, instrument, data source, staff workflow, or reporting requirement changes. For repeated collection, make the review part of each new cycle.
What makes an AI answer trustworthy?
An answer is more trustworthy when a reviewer can see its scope, underlying data, source passages or calculations, governing definitions, and limitations. Trust comes from inspectability and accountable human review—not fluent wording.
Explore related demonstrations
Browse the video library → Use the source, definition and access checks in this lesson when evaluating a demonstration. A video does not verify the controls in your own configuration.
Bring the governance decisions back to your workflow
Return to your course module with the evidence register, assigned owners and test log. If your pilot involves revised instruments, use changing survey questions without losing comparability. Record the next correction, its owner and the check that will show it worked.
Reviewed September 12, 2026. The community-learning scenario and pilot counts are illustrative, not customer results.