When a manager, peers and a team all describe one colleague, the person answering is not the person being measured — and almost no survey tool assumes that shape. Anonymity becomes a number you choose before you collect.
Record shape B of the four in Connected Data Intelligence — several people describing one person. It leans hardest on checks 07 Assistant, 02 One record and 05 Qualitative.
A leadership programme collects feedback on each participant from four people: their manager, two peers and someone who reports to them. Everyone is told their answers are confidential. The responses come back thoughtful and candid, exactly as intended.
Then the programme tries to give each participant their results, and discovers it cannot. There is only one manager, so anything labelled "your manager said" is a named person. There is only one direct report, so that comment is attributable too. Any breakdown by group identifies two of the four raters. So the feedback sits there — collected honestly, promised confidentially, and impossible to hand over.
Nothing went wrong in the collection. The problem was decided weeks earlier, by nobody making a decision at all.
The person answering is not the person being measured. In almost every other survey those are the same — you ask a participant about themselves, and one response means one person's data. Here they come apart: four people answer, and the thing you are building a picture of is a fifth person who may not have answered at all.
That sounds like a small distinction and it is the whole difficulty. Survey tools are built to give you one row per response. What you need is one record per subject, assembled from several responses, each carrying who it came from — while never letting the subject work out which comment came from whom. Very little software assumes that shape, which is why teams end up doing it in spreadsheets and why it goes wrong so reliably.
This is not only a workplace pattern. A mentor, a teacher and a parent all describing the same young person is the same shape. So is a case worker, a clinician and a family member describing one client. If more than one person is telling you about someone else, you are in shape B.
"Confidential" is not a setting, it is a threshold: the smallest number of people whose answers you are willing to show as a group. If a participant has two peers and you show peer results, you have shown something close to two individuals' views. If they have twelve, you genuinely have a group.
So pick the number first — three and five are the common choices — and then apply it in two places most teams forget:
That second point is the one that changes response quality. People are not reassured by the word confidential. They are reassured by arithmetic they can verify.
The threshold protects you from counting problems. It does not protect you from context, and this catches people out.
A global skills competition ran feedback where each skill area had exactly one expert from each country. A competitor writing about "the expert" was, in practice, naming a specific individual — because in that skill there was only one. The team handling the reports used to go through and change "the Korean expert" to "an expert" by hand before anything was circulated. Their own description of the difficulty was honest: it is hard to know in advance which detail is the identifying one.
The general rule is that any attribute unique within a small group is an identity, whatever your threshold says. Region, job title, tenure, the programme someone attended, a disability, a language — each is a name in a small enough population. Which is why a group of five is not automatically safe, and why somebody has to read before a report is shared rather than trusting the count alone.
The most common mistake in this shape is also the easiest to make, because software offers it: taking every rating a person received and producing one number.
A manager and a direct report are not two measurements of the same thing. They are two people observing someone from different positions, with different information and different stakes. When they disagree, that is not error to be smoothed out — it is usually the most useful thing the exercise produced. Someone rated highly by their manager and poorly by their team has a specific, actionable problem, and it is exactly the finding an average erases. The single number that comes out will describe nobody at all.
So report by group, always. The same applies with a mentor, a teacher and a parent: three legitimate perspectives on one child, three different contexts, and a combined score that means nothing. Where they disagree is where the conversation should start.
Access in this shape surprises people, because it inverts the usual order. The person being described often sees more than their own manager does, and the raters see the least of all.
| Who | What they should see |
|---|---|
| The person described | Themes by group, with example wording where the group is large enough. Not a raw list of comments — that is where recognisable phrasing does its damage. |
| Their manager | Usually less, not more: the themes and the development focus, not the underlying comments. If the manager was also a rater, they must not be able to reverse-engineer the others. |
| The raters | Their own submission only. Never anyone else's, and never the final report unless the subject chooses to share it. |
| The programme team | Patterns across participants — who is progressing, where the cohort struggles — without reading individuals' feedback as a matter of routine. |
Write this down before the first invitation goes out. Every difficult conversation you will have later is a version of one of these four rows, and they are much easier to settle in the abstract than when someone is asking to see a specific person's comments.
For one cohort of fifteen or twenty people, that is completely doable by hand, and the reading step is not optional at any scale.
It breaks on the reading. Everything else scales adequately with careful spreadsheet work, but the step where a human checks each report for identifying detail is the bottleneck, and it is the step that gets skipped when a cohort grows or a deadline lands. When it gets skipped, the failure is not visible — it looks like a finished report, right up until somebody recognises themselves in it.
It also breaks across cycles. A year later a question has been reworded, the participant has a different manager, and one of last year's peers has left. The comparison still needs to be honest about all three, which is check 04.
What a system is for here is narrow: keeping every response attached to the right subject while holding the rater's group and not their identity, applying the threshold automatically so a thin group is suppressed rather than remembered, grouping written comments into themes without reproducing recognisable phrasing, and giving each of those four audiences a different view of the same underlying material. It should not be the thing that decides whether a comment is safe to show — a person still reads before sharing. It should be the thing that makes sure that person is looking at a manageable amount.
Use: Two subjects. Give the first four groups with enough people in each; give the second a group of only two. Include one comment that names a specific role, and one pair of groups that disagree sharply about the same subject.
Pass: The group of two is suppressed without anyone remembering to do it. The disagreement between groups is visible rather than averaged. The role-naming comment is flagged for a person to look at. And the four audiences each get a different view.
Fail: A single combined score appears anywhere. Or the thin group is shown because nobody caught it.
Pick a number in advance — three and five are common — and apply it when you send invitations, not when you write the report. If a group would fall below it, either invite more people or accept that the group will be suppressed. Deciding at reporting time leaves you choosing between breaking a promise and having nothing to show.
Yes, by grouping them into themes at group level rather than quoting individuals, and by having a person check for phrasing that identifies someone regardless of the count. What you cannot safely do is hand over a raw list of comments from a small group.
Because a manager and a team member are not two measurements of the same thing, and their disagreement is usually the most useful finding. Averaging removes exactly the signal the exercise was run to produce, and the resulting number describes nobody.
They see more than anyone else, but not the raw comments where groups are small. Themes with example wording is the normal boundary — enough to act on, not enough to identify who said what.
No. A mentor, teacher and parent describing one young person is the same shape, as is a case worker, clinician and family member describing one client. Any time several people describe someone else, the same rules apply.
The comparison has to disclose it. A change in who is describing someone can move the result as much as a change in the person, so a shift in group membership belongs in the report rather than hidden inside it.
No. The decisions here are judgement calls a programme lead is well placed to make — the threshold, who to invite, who sees what. None of them requires knowing how the software works, and none should require writing instructions to a chatbot.
What makes this shape hard is not the collection, and not really the privacy rules either. It is that the two things everyone wants — read what people actually wrote, and protect who wrote it — pull in opposite directions, and every design choice trades one against the other. Teams that decide the trade deliberately, at invitation time, end up with candid feedback they can hand over. Teams that leave it until the report exists end up with candid feedback nobody is allowed to read.
Start with data your teams struggle to bring together. Agree shared definitions, keep each source identifiable, and decide who can see what before asking AI for an answer.
Explore Connected Data Intelligence →