Analyze survey results by subgroup by choosing a justified comparison, defining the outcome and group membership, and showing the response base for each result. Check whether the measures, timing and participation are comparable before interpreting a difference. Protect small or identifiable groups and investigate possible explanations. A lower average identifies a question; it does not establish that the group is being left behind or explain the cause.
This lesson is for teams comparing participant, member, customer or employee experiences. You will prepare a subgroup table and a bounded finding. Use demographic information only where its collection and analysis are appropriate to the purpose; do not infer a sensitive characteristic from names, comments or an AI guess.
- Choose a decision and relevant groups before exploring many cuts.
- Define group membership, outcome, period and population.
- Calculate results and coverage separately.
- Review comparability, uncertainty and disclosure risk.
- Use appropriate qualitative context to investigate the pattern.
- Report the limits and the next action.
Which groups should you compare?
Start with a question that matters to delivery or access. A location, language, program route or voluntarily supplied demographic category may help answer it. A site is an operational grouping, not itself a demographic characteristic. Name the type of group accurately.
Limit exploratory slicing to what the evidence and purpose can support. Repeatedly searching for the largest gap can produce an attention-grabbing result without a sound interpretation. Distinguish planned comparisons from patterns noticed during exploration.
Keep unknown and declined group responses visible in coverage. Do not assign them to a category by inference. If people can select several identities, document whether groups overlap and avoid presenting their totals as mutually exclusive.
What context belongs with group membership?
Use the profile information relevant to the observation date. A person may change location, role or program. Decide whether the question concerns their group at entry, at follow-up or at another point, and retain that rule.
Stable registration information can be reused where appropriate. Changing fields need a suitable confirmation process and history. Local forms may differ, but the compared measure needs compatible definitions and an explicit mapping. The same response label does not establish the same construct.
How should the comparison table be built?
Show the eligible population, usable responses, result numerator and denominator, and collection period for each group. If the question is change over time, identify the matched records rather than comparing two different respondent groups as though they were the same people.
| Fictional location | Eligible | Usable responses | Report the service met their need | Observed rate | Coverage |
|---|---|---|---|---|---|
| North | 100 | 80 | 56 | 56/80 = 70% | 80% |
| South | 100 | 40 | 24 | 24/40 = 60% | 40% |
The observed rates differ by ten percentage points. Response coverage also differs substantially. The table does not establish the experience of nonrespondents or why the observed rates differ. North’s higher result does not justify describing all North participants as satisfied.
The combined respondent rate is 80/120 = 66.7%. Comparing each group with that total may help describe the table, but it is not inherently better than a direct comparison or a relevant target. State the reference and why it answers the decision.
What must you check before interpreting a gap?
Review the measure, timing, eligibility, delivery context and response pattern. Ask whether the groups had comparable opportunities to use the service. Different stages or service mixes can produce different results even when the same question is used.
Separate two concerns: statistical uncertainty and the risk of revealing someone’s information. The ONS discussion of survey reliability distinguishes uncertainty in small-group estimates from disclosure protection. Its survey-specific practices are not a universal threshold for your dataset.
For formal estimates, use methods suited to the sampling design and get appropriate statistical review. An ordinary margin of error does not repair systematic nonresponse or turn a convenience sample into a representative one. For operational descriptive results, show the base and limits plainly.
What should happen with a very small group?
Apply the approved disclosure rules for the audience and consider whether combinations of tables or quotations could reveal identities. Removing a name may not be enough. Do not promise that a single minimum count makes every output safe.
Pooling periods is an option only when the question, population and collection remain compatible. It changes the period described and may hide a recent shift. Grouping categories can also hide important experiences. Record the reason and consider a restricted review or a careful non-identifying account where a public percentage would mislead or expose someone.
How can comments help explain the result?
Review relevant accounts from more than one perspective. In the fictional example, some South respondents may mention scheduling. That supports investigating the schedule, not declaring it the cause of the ten-point gap. Check how the comments were collected and whose accounts are absent.
A small number of comments can still identify an important issue. Report why the issue matters without pretending the comments estimate prevalence. Keep participant words, reviewed themes and proposed explanations separate.
What is an appropriate next action?
Choose a check proportionate to the evidence: inspect invitation delivery, review an access barrier, discuss a specific concern or improve the next collection round. Adding one question does not itself resolve the difference or improve the outcome.
For the fictional table, a useful finding is: “South respondents reported lower need fulfillment, with lower response coverage. The team will review invitation reach and service-access feedback before choosing a change.” Keep an owner and a later review point.
How does Sopact support the analysis?
You can start with a reviewed table and source register. Sopact’s connected-record approach can bring appropriate profile history, responses and comments into context, with dictionary rules available to the analysis. That helps keep definitions and coverage visible across repeated comparisons.
Test group filters, date rules, overlapping memberships and access controls in the configured workflow. AI can organize comments and draft questions for review. It should not infer sensitive group membership, expose restricted cells through a summary or turn a difference into a causal explanation.
Practice: write a comparison without a label
Use the North–South table. Calculate both rates and coverage, then list three possible explanations to investigate. Write a two-sentence finding without labeling either group as successful or failing. Add one privacy check before sharing the table.
Frequently asked questions
Does a below-average group mean inequity has been proved?
No. It is a finding to investigate in context. The measure, population, coverage and other evidence determine what can be concluded. A gap can matter without its cause already being established.
Is there one safe minimum subgroup size?
No universal number is supplied here. Precision and disclosure depend on the data, design, audience and other available information. Use appropriate review and the organization’s approved rules rather than a generic threshold.
Can missing demographic fields be inferred?
Do not infer sensitive characteristics from names or narrative to complete a table. Retain unknown or declined status and explain how missing information limits the analysis.
Should we always pool small groups across time?
No. Pooling may be unsuitable if measures, populations or conditions changed. It also changes the reporting period. Consider the decision and disclosure implications, and document the choice.