Compare two cohorts by checking who is included, what was measured, when it was measured and how complete the evidence is before interpreting the difference. A common data dictionary helps establish compatible definitions. It does not remove differences in participants, delivery or opportunity to achieve the outcome.
This lesson produces a comparison brief: the result, reporting bases, known differences and a conclusion the evidence can support. Use it for training groups, service locations, member activities or other groups only where the intended comparison makes sense.
Name the comparison and the decision
Are you describing this year’s results, deciding where support is needed or asking whether a program change caused an improvement? These are different questions and require different evidence.
Write the group definitions and review date. An intake cohort is not automatically the same population as the people who completed the program or answered follow-up. Keep exclusions and missing observations visible rather than changing the group halfway through the analysis.
Check measurement before calculating a difference
| Check | Question to resolve |
|---|---|
| Outcome definition | Does the result mean the same thing in both groups? |
| Unit and denominator | Are you counting people, episodes or events, and who is eligible? |
| Timing | Did both groups have a comparable observation window and opportunity? |
| Instrument | Did wording, response options, assessment criteria or collection mode change? |
| Coverage | Who has usable evidence and who is missing? |
| Delivery and context | Were there relevant differences in support, setting or participant characteristics? |
A percentage accounts for its stated denominator; it does not automatically make two groups comparable. Dividing an outcome by program length is not a general solution either. Some outcomes develop unevenly, and a shorter follow-up window may answer a different question.
Work through a composition example
The following figures are fictional. Two cohorts use the same application criterion and review point. The only two categories in this simplified example are prior experience at entry and no prior experience.
| Group | Prior experience | No prior experience | Overall |
|---|---|---|---|
| Cohort A | 9 of 10 meet the criterion: 90% | 4 of 10 meet it: 40% | 13 of 20: 65% |
| Cohort B | 18 of 20 meet the criterion: 90% | 4 of 10 meet it: 40% | 22 of 30: about 73.3% |
The overall percentage is higher in B, but the percentage within each experience category is unchanged. B contains a larger share of people with prior experience. Calling the overall difference a demonstrated improvement in delivery would miss what the table shows.
This example illustrates one known composition difference. In real data, there may be other differences, incomplete evidence and uncertainty. Matching one category does not prove that all relevant explanations have been addressed.
Keep useful shared profile fields without imposing one survey
Different locations can use local questions while collecting the few shared fields needed for a justified comparison. In the example, prior experience has a defined meaning at entry. It should not be replaced later by someone’s experience after the program.
Reuse suitable registration information and retain its timing. Define categories and missing states in the dictionary. Collect a characteristic only when it serves an appropriate purpose; more personal information is not automatically a better comparison.
Keep incompatible measures separate. If one location records self-confidence and another observes task performance, a shared field name cannot turn those into the same outcome. Explain the difference instead of manufacturing a conversion.
Distinguish comparison problems
A changed question can create a measurement difference. Missing follow-up can change who is represented. Different entry characteristics can affect the group result. A change in delivery may be part of the explanation you want to investigate. Do not call all of these one generic “confound” and assume a single adjustment fixes them.
List the relevant differences you know, the source for each and what remains unknown. Do not ask AI to name every possible factor or claim its direction without evidence. The purpose is to make the comparison inspectable, not to produce an exhaustive-looking list.
For a causal question, choose an evaluation design and analysis appropriate to that question. A descriptive table, a linked record or a list of caveats does not by itself establish that the program caused the difference.
Review coverage alongside the result
If one cohort has follow-up for nearly everyone and the other has follow-up for only part of the group, show that difference. Compare the eligible count, usable observations and missing states. Where appropriate, investigate what is known about the people missing from the evidence without inventing their outcomes.
Distinguish a comparison of different cohorts from change within the same people. Both can be useful, but their denominators and interpretations differ. A matched pre/post analysis should explain who has the required observations and who was excluded.
Make the analysis reproducible and useful
Save the group definitions, observation window, data snapshot, measure version and calculation. Keep the source records available to authorized reviewers. If a later correction changes a numerator or denominator, explain the change rather than presenting it as a new program effect.
In Sopact, continuing records and shared definitions can help assemble and inspect these inputs. AI can assist with a proposed comparison summary and evidence gaps. The team still decides whether the measures can be combined and what the result supports; the platform should not be treated as automatically correcting every design difference.
Write the conclusion at the right strength
For the fictional table, a suitable conclusion is: “A larger share of Cohort B met the application criterion. The rates within the two prior-experience categories were unchanged, and the cohorts had different experience mixes.”
That conclusion explains the finding and suggests a next question. It does not declare one cohort better or one delivery approach more effective. An operational decision might be to review support for learners without prior experience, while retaining the limits of the evidence.
Your exercise
- Define two groups and the decision the comparison should inform.
- Check the outcome, unit, timing, instrument and coverage.
- Record one known difference and its evidence.
- Calculate the result with explicit denominators.
- Write what the comparison supports and what it does not.
- Identify the next evidence or analysis needed for a stronger question.
Frequently asked questions
How do I compare outcomes between two cohorts?
Start with compatible measures, explicit group definitions and observation windows. Show denominators and coverage, inspect relevant differences and interpret the result according to the evidence and design.
Can I normalize for cohort size and program length?
A rate can address a particular count denominator. There is no universal adjustment for program length or participant mix. Choose a method suited to the outcome and explain its assumptions.
Why is a raw difference not proof of improvement?
The groups may differ in composition, measurement, follow-up or other conditions. A raw difference describes what was observed; explaining why it occurred requires further evidence.