How do you organize feedback from several people about one person?
Keep the person being described separate from the person providing feedback. Link each response to the correct subject, rater role, assessment cycle and question version. Decide before collection how scores and comments will be reported, who may see them and what happens when too few people respond.
Use this reference when several people provide observations about the same person. Design a small multi-rater dataset and check a sample report before anyone receives it. The collection purpose and disclosure rules must fit your course workflow; this is not a required employee-rating model.
A leadership participant may receive feedback from a manager, peers, direct reports and themselves. Other settings include several mentors describing a learner or several service providers contributing observations. The relationship pattern is similar; the assessment, confidentiality rules and responsibilities can differ substantially.
Start with the purpose and the promise
State whether the exercise supports development, program learning or another defined purpose. Do not quietly reuse feedback collected for coaching as a promotion or disciplinary score. A different use may require a different design, review and communication process.
Explain who will see individual responses, who will see grouped findings and whether any role will be identifiable. “Confidential” describes controlled access; it does not necessarily mean that nobody administering the process can know a rater's identity.
The Center for Creative Leadership's implementation guidance emphasizes clear outcomes and attention to confidentiality and rater anonymity. Use that as a planning question: what exactly are you promising each audience in this exercise?
A single manager's feedback can be reported openly if that is the stated design. It is misleading to promise that every role is anonymous and later display a category containing one identifiable person. Resolve that difference in the invitation, not after the reports are produced.
Model the subject, rater and response separately
The subject is the person receiving feedback. The rater is the person supplying an observation. One rater may assess several subjects, and one subject may receive many responses. Self-assessment is a special relationship in which the same person holds both roles.
| Field | Purpose | Example in the exercise |
|---|---|---|
| Subject ID | Whose development record receives the observation? | S-014 |
| Rater reference and role | Who supplied it, under the approved access model? | R-082; peer |
| Cycle and observation date | Which assessment period does it describe? | Cycle 1; September 2026 |
| Instrument version and item | What behavior and response scale were used? | Version A; listens before deciding |
| Response status | Can the observation contribute to this measure? | Valid response; not enough opportunity to observe; missing |
| Reporting group and release status | How may this response contribute to an approved output? | Peer aggregate; reviewed before release |
Keep invitation details and identifying information restricted as appropriate. Removing a name from a report is not the same as removing it from the underlying administration system. Document both layers so a reviewer knows what is protected and from whom.
Do not attach every answer to the rater's personal record merely because they submitted the form. That would lose the subject relationship. Equally, do not collapse all responses into the subject record without retaining role and source references for authorized analysis.
Set reporting rules before sending invitations
Define minimum valid response counts for protected groups and what happens below them. A rule may suppress a group, combine specified groups under an approved scoring method or delay a report. These are design choices, not a universal numeric guarantee of anonymity.
Count completed, usable responses when applying the release rule. Inviting five people does not satisfy a three-response rule if only two answer. At item level, some raters may select “not enough opportunity to observe,” leaving fewer usable ratings than the group total.
CCL's published scoring example combines two peer and two direct-report responses in a labeled category. Its rater guidance also identifies categories that may be displayed individually. These are examples of a defined instrument's rules, not rules to copy into every assessment.
Plan invitations around relevant observation, not just the count needed to release a chart. Adding people who have little experience of the subject's work can make a group larger without making the feedback useful.
Worked example: a report that preserves the differences
This fictional exercise uses a five-point behavior scale. The local policy reports the manager and self-assessment openly, requires at least three valid raters for protected peer and direct-report groups, and suppresses rather than combines a group below that threshold. Comments still require review.
| Perspective | Valid ratings | Exercise values | Released result under this policy |
|---|---|---|---|
| Self | 1 | 4 | 4.0, identified as self-assessment |
| Manager | 1 | 5 | 5.0, openly attributable by design |
| Peers | 3 | 3, 3, 4 | 3.33, peer aggregate |
| Direct reports | 2 | 2, 3 | Not released as a group: below the exercise threshold |
The values in this teaching table are synthetic and fully visible so you can check the calculation. An actual participant-facing report under this policy would not display the suppressed group's individual values. Do not copy this teaching table into a real confidential report.
The visible self, manager and peer perspectives provide a starting point for discussion. They do not establish which group is “right,” diagnose the subject or prove that a particular behavior caused a business outcome.
An overall average could undermine suppression. In this example there are six ratings other than self: one manager, three peers and two direct reports. Their sum is 20. Showing an overall mean of 20/6 alongside the manager and peer totals could reveal the suppressed group's combined value. Review derived totals as well as individual cells.
Use combined scores only with a clear rationale
An overall score is not automatically meaningless, and a group breakdown is not automatically safe. Follow the assessment's documented scoring method, explain what is combined and preserve useful distinctions where reporting rules allow.
Weighting matters. An average of all individual ratings gives a group with more respondents greater weight. An average of group means gives each included group equal weight. Those answer different questions. State the rule instead of allowing the spreadsheet's default formula to decide.
For the fictional example, the manager score is 5 and the peer mean is 10/3. Averaging those four individual ratings gives 15/4 = 3.75. Giving manager and peer groups equal weight gives (5 + 10/3)/2 = 4.17. Neither should appear as an unexplained “overall score.”
Do not combine unrelated constructs or incompatible scales simply because they share a subject ID. A parent's account of home routines, a teacher's classroom observation and a mentor's attendance note may need a contextual summary rather than one numeric rating.
Read disagreement as a question to investigate
Different roles may observe different behavior, work in different settings or have different opportunities to see a skill. A gap can be useful, but it is not automatically evidence of hypocrisy, poor performance or a hidden problem.
Start with what the item asks and what each group could observe. Check completion, timing, scale interpretation and the number of valid responses. Then discuss concrete examples with an appropriate facilitator or reviewer.
Look for patterns across related items without treating every small difference as meaningful. A few ratings on a five-point scale can move noticeably when one response changes. Avoid false precision in a report intended to support a development conversation.
Keep the subject's own account in view. The purpose is not to vote one person's experience out of existence. It is to understand the perspectives and identify an appropriate next step.
Protect comments as well as score breakdowns
A comment can identify its author through a distinctive project, a unique role or recognizable phrasing. A numeric threshold does not remove that context. Review quotations, summaries and the combination of details before release.
Summarizing comments into themes can reduce unnecessary detail, but it is not a guarantee. A theme mentioned by only one person may remain recognizable. If a summary changes the meaning, it is not an acceptable confidentiality fix.
Keep source wording available to authorized reviewers, record material edits and label paraphrases appropriately. Do not invent a composite quotation and attribute it to a rater group. Explain what the summary can support and where evidence is limited.
Decide in advance how serious concerns will be handled. A developmental report is not a substitute for the organization's appropriate reporting and response process. Follow the agreed process rather than hiding a consequential concern inside an average or an automated theme.
Give each audience an appropriate view
The subject, coach, manager, raters and program administrators may need different information. There is no universal rule that the subject must always see more than everyone else, or that a manager must never see a comment. Define the purpose and access arrangement for this program.
For this exercise, the subject and assigned coach receive the reviewed development report. The manager receives an agreed development plan. Raters can confirm their own submission but cannot browse other responses. Administrators manage collection with restricted access. These are illustrative choices, not a product default.
Test the report link, downloadable file, citation and export. A protected dashboard does not protect a spreadsheet containing raw responses if that spreadsheet is shared widely. Recheck access after role changes and record what happens to previously downloaded copies.
Use the access-controls lesson for the broader acceptance test. Do not promise that a plain-language instruction to an assistant is sufficient to enforce this policy.
Keep rater changes visible across cycles
When the assessment repeats, retain the cycle, instrument version and rater-group composition. A new manager or a different peer group can change the perspective represented in the results even when the subject's behavior has not changed.
Compare the same defined measure where possible and disclose meaningful changes. Do not reveal protected rater identities merely to explain composition. An authorized analyst may need more detail than the participant-facing report.
Connect a development action to a later review without assuming the action caused a score change. For example, record the agreed behavior to practice, the review date and the relevant next-cycle evidence. The course's longitudinal-analysis lesson covers repeated observations in more depth.
Run a small acceptance test
Create two fictional subjects with separate rater lists. Give one subject enough valid responses for every protected group and the other a group below the chosen threshold. Add a distinctive comment, an “unable to observe” response and a rater who assesses both subjects.
- Confirm every response reaches the correct subject and cycle.
- Verify that the same rater can describe two subjects without merging their records.
- Apply the documented scoring and minimum-response rules, including item-level missingness.
- Check that totals, exports and summaries do not defeat a suppressed result.
- Review the distinctive comment and record the release decision.
- Open the outputs under each test role and verify the intended view.
- Repeat with a changed rater group and question version, checking the comparison labels.
For a Sopact evaluation, ask for this demonstration in the proposed configuration. Collection, linked records and AI-assisted preparation can be useful parts of the workflow; verify the reporting and access rules rather than assuming every assessment convention is built in.
The output of this lesson is a record model, a reporting policy and a tested sample report. Keep all three together so the next collection cycle does not start by rediscovering the same decisions.
Watch: why feedback needs its context
Watch the video · 2 minutes 34 seconds. This course companion explains the broader value of connecting qualitative evidence. It is not a demonstration of the scoring policy used in this exercise. Browse more videos in the video library.
Frequently asked questions
What is multi-rater feedback?
It collects observations from several people about one subject. Keep the subject, rater role, cycle and measure separate so the responses remain interpretable.
Is a manager's rating always anonymous?
No. Some designs report it openly. Explain the arrangement before collection and do not promise anonymity for a category that identifies a single person.
How many raters are enough?
Use the instrument's justified reporting rules and the context of the program. Count usable responses, not invitations. A threshold alone does not guarantee that comments or combined outputs cannot identify someone.
Can we combine rater groups?
Sometimes, under a defined scoring and disclosure policy. Explain the groups and weighting. Do not combine them silently or publish an overall result that reveals a suppressed group.
What does disagreement mean?
It may reflect different observations, contexts or interpretations. Check the evidence and discuss the gap; it is not automatically a diagnosis or proof of a performance problem.
Can AI anonymize all comments for us?
Do not assume that. An assistant can help identify details or draft themes, but context can remain identifying and summaries can alter meaning. Review the actual output under the agreed release policy.
Can we compare results next year?
Yes, with attention to instrument versions, timing and changes in rater composition. State meaningful differences so a new perspective is not mistaken for change in the subject alone.
When your report spans several programs
Continue to many programs, one picture. That reference examines evidence across several programs, where shared definitions and distinct denominators become central.
Reviewed September 12, 2026. Subjects, ratings and reporting policies in the worked exercise are fictional; they are not a validated assessment or a universal confidentiality standard.