This chapter resolves check 07 Assistant of the eight checks.
A programme director sets up an AI summary of participant feedback for the board pack. She is careful about it. In the instructions she writes: do not include participant names, and do not include anything that could identify anyone. The first summary comes back clean, and she is reassured.
Three months later a colleague opens the same setup and asks a different question — why fathers in the rural sites dropped out. The answer includes a comment about a single father managing shift work and school pickups at the smallest site. No name. It does not need one. Everyone on the call knows who that is.
Nobody broke a rule. The rule was a sentence, and the person asking the second question had never read it.
How do we use AI on stakeholder data without exposing people?
You decide what the AI is allowed to see before anyone asks it anything. Almost all the worry about AI in this work is about accuracy — whether the answer is right. That is manageable, because a wrong answer can be checked. The risk that is hard to reverse is exposure: a participant recognisable in a report, a comment traced back to someone promised confidentiality. You cannot un-share that, and no amount of accuracy prevents it. What prevents it is the information not being available in the first place.
An instruction is not a control
This is the distinction most organisations have not yet made, and it is the centre of the chapter. An instruction is a sentence you write asking a system not to do something. A control is a property of the data itself: the information is not there to be revealed, regardless of who asks or how the question is worded.
| What happens when… | You wrote an instruction | The rule is attached to the data |
|---|
| Someone phrases the question differently | The instruction may not cover the new wording. It was written with one kind of question in mind. | Nothing changes. The wording of the question has no bearing on what exists to be returned. |
| A new colleague uses the same setup | They do not know the instruction is there, and may edit or replace it without knowing what it was protecting. | They inherit the protection without needing to know it exists. |
| The person who wrote it leaves | The reasoning leaves with them. The sentence survives as a line nobody can explain or defend. | The rule stays attached to the field, with an owner and a date. |
| The output is shared onward | The instruction governed the generating. It has no say over the sharing. | What was never included cannot be forwarded. |
"Please don't forward this" is a request made to a person. A file the recipient cannot open is a control. Only one of the two holds when someone is in a hurry, unaware, or asking in good faith.
Small groups: suppress the evidence, not the chart
The second common failure is a group so small that reporting it identifies people. Six participants at a site, one of them the only person in a particular category: a percentage of six is a description of individuals.
Most organisations hide the cell on the chart, which is the wrong layer. The chart is not the only way a number leaves the building — there is the export, the underlying list, the follow-up question, and the quote below the chart describing the situation precisely enough to identify the person the chart just protected. Set a minimum group size, apply it to the evidence so a below-threshold breakdown does not exist by any route, and hold quotes to the same threshold.
Who can see the report — including people who never write a prompt
Careful thought goes into who may query the data. Then a report is made from it and shared — with the board, a funder, a partner, in a link that works for anyone who has it. Everyone reading it is seeing stakeholder data, and none of them went near a prompt. A shared report has its own access question, and it is usually the harder one, because reports outlive their reason for existing. Partner staff change. A board member's term ends. The link stays live.
Two habits fix most of it: every report gets an owner and a named list of who may open it, and that list is checked when people leave — which requires it to be a list somewhere, not a set of forwarded emails.
A person decides anything consequential
AI prepares; people decide. Preparing means reading a large volume of responses, grouping them, drafting a summary, pointing to the passages behind it. Deciding means whether a participant is escalated, whether a partner's grant continues, whether a finding goes to a funder. Every consequential decision should carry a person's name and a record of what they saw, because a decision about someone's participation is the kind of thing a human has to be answerable for.
You do not have to learn prompting
None of this requires you to write good prompts or to understand how a model works. If a reliable answer depends on the phrasing skill of the person asking, that is not a solution — it is a new dependency, and it sits with whoever is most confident with AI rather than whoever knows the programme. What you do need is the ability to say, in your own words, who may see what. That is a programme judgement, and you are the only person who can make it.
How to do this without any particular software
- Sort your fields into three tiers. Identifying (names, contact details, anything unique), sensitive (health, immigration, income, disciplinary history), and open.
- Write down who may see each tier, by role rather than by person. "Case managers see identifying fields for their own caseload. Nobody outside the programme team sees them."
- Set a minimum group size and write it down. Apply it before a breakdown is produced, and to quotes as well as numbers.
- Record consent as a field on the record, including whether the person agreed to be quoted — not in a separate log nobody consults during analysis.
- Before anything goes into any AI tool, remove the identifying tier. Removing the column is the control. Asking the tool to ignore the column is not.
- Keep one list of who can open each report, with a review date and a person responsible for it.
- Log consequential decisions with a name, a date, and what the person was looking at.
Seven items, writable in an afternoon, and genuinely most of the protection.
Where it breaks
The de-identified copy that grows back. Someone makes a safe extract with names removed. A later question needs one more field, so a second extract adds it. Nobody re-checks the combination — and site plus age plus role usually identifies people in a cohort of forty.
Two safe reports that are unsafe together. One shows outcomes by site, another the same outcomes by age group. Neither identifies anyone alone. Read side by side by someone who knows the programme, they frequently do.
The narrow claim for a system: in Sopact Sense the permission travels with the field and consent status travels with the record, so what someone may see is decided by their role and the participant's consent rather than by the wording of a request. Thresholds apply to the evidence, not the chart, and reports carry their own access list. Your tiers, thresholds and human-decision rules are still yours to set.
What to ask a vendor
Listen for whether the answer is a mechanism or a reassurance.
- What leaves our system, and where does it go? Which data goes to which third party, in what form, under what agreement, and whether it trains anything.
- What does the model see when someone asks a question? The whole record, or only the fields that person is entitled to. If the answer is "we instruct it not to use the rest", you have an instruction, not a control.
- What happens when someone asks a question they are not entitled to the answer to? Ask for a demonstration, then ask them to try it a second way. A protection that only holds for the expected phrasing is not a protection.
How to test this before you trust it
Use: Your own data, and four attempts. Ask for something you are not entitled to, then ask for it again in different words. Ask for a breakdown of a group with three people in it. Ask about a participant who withdrew consent. Then have someone outside the programme team open a shared report.
Pass: All four are refused or empty, identically, regardless of phrasing. The withdrawn participant does not appear. The outside colleague sees only what their role allows, and you can see the list of who has access.
Fail: Any protection that changes with the wording of the question. Or a suppressed number sitting above a quote that identifies the person it was hiding.
Do this with a real cohort, small sites included. A demo dataset has no group of three and nobody who withdrew — the only two situations this check is about.
Frequently asked questions
Do I need to learn prompting to use any of this?
No. Ask your question in the words you would use with a colleague. If a system only performs for people who have learned to phrase things a particular way, that skill is compensating for a weakness in the tool — and the people who know your programme best are the least likely to have it.
Is our participant data used to train the AI?
Ask, and get the answer in writing. It is a straightforward question, and a vendor who cannot answer it plainly has told you something. The answer may differ between their own features and any third-party service sitting behind them.
Can't we just tell the AI to keep everything anonymous?
You can, it helps with the obvious cases, and it is not a control. It depends on the instruction being present, being read, covering the question actually asked, and surviving the next person to edit it. Remove the identifying fields from what the system can reach and none of those four conditions matters.
How small is too small to report?
Five is a common floor rather than a standard. Judge it against your context: where one attribute is rare, a group of twelve can still identify someone. Write your threshold down, apply it to quotes as well as numbers, and revisit it as your cohort shrinks.
Who should be able to open a report built from participant data?
A named list with an owner and a review date. "Anyone with the link" is how most exposure actually happens — not through a clever question, but through a report shared for a good reason in March and still open in November to people who have since left.
What if a participant withdraws consent after they have been quoted?
Withdrawal has to reach the evidence, not just the mailing list: their words stop being available for new analysis and new reports. Already-published material is a separate conversation, and a reason to be conservative about what you publish.
What decisions should never be left to AI?
Anything that changes what happens to a person or an organisation: escalation, eligibility, funding, a finding going to a funder. Preparation is useful; the decision needs a person's name against it and a record of what they were looking at.
What changes here is what confidentiality is. It used to be a promise made once, at intake, and kept by careful people — which worked when the data was read only by those who collected it. Now an answer given in confidence can be reached by a question nobody anticipated, from a person the participant never met, months after the programme ended. The promise has to become a property of the answer itself: something that travels with it for as long as it exists, and holds without depending on anyone remembering it was made.
Next: Write a cited impact narrative