How do you combine results across different programs?
Combine only the measures whose definitions, populations, periods and calculation rules are sufficiently comparable. Keep the original program records, identify overlap between populations and show which programs are included or excluded. A shared outcome label is a starting point, not proof that the numbers can be added.
Use this reference when your course plan needs results across programs, sites or partners. Build a cross-program reporting table, calculate one justified combined rate and explain why another result remains separate. It does not require an earlier multi-rater exercise.
The method applies to an organization with several programs, a funder receiving partner reports or a business comparing delivery locations. The purpose is a useful shared view that preserves the meaning of each source, rather than a single total assembled from incompatible figures.
Start with the question the combined view must answer
A board asking “what did we achieve?” may need several answers: how many people participated, which outcomes were observed, where evidence is missing and what should change next. Those questions do not necessarily share one denominator.
Choose one reporting question to start. For example: “Among people with a documented six-month employment follow-up, how many met our agreed employment definition?” That is narrower and more testable than “What is our total impact?”
List the programs that can contribute evidence and the reason each belongs. Record the period, population and source. A program that does not measure employment is not a failed employment program; it may serve a different purpose.
A common dashboard can contain several outcome panels. A composite index can also be constructed for a defined purpose, but its normalization, weighting and interpretation need justification. Do not make unlike outcomes comparable merely by placing them on one chart.
A shared definition is necessary but may not be enough
Agree what the measure means: who qualifies, what event or response counts, when it is observed, how it is verified and how missing information is treated. Then check whether each program's collection actually supports that definition.
Two questions labeled “confidence” may measure different constructs or use response scales that cannot be directly pooled. Two programs can also ask the same words but apply different completion rules. Neither a shared label nor identical wording is sufficient on its own.
Some standardized measures benefit from consistent wording and administration. Other contexts need local questions with a documented mapping to a shared measure. Evaluate comparability rather than following an absolute rule to standardize all questionnaires or none of them.
The GIIN's Evaluating Impact Performance research examines comparisons within specified sectors. Its scope is a useful reminder that comparison needs context. A broad framework reference does not automatically make every program result interchangeable.
Use the data-dictionary lesson to document the local rule, source field, version, unit and owner. If you map to IRIS+, an SDG or another framework, preserve the exact definition and mapping rationale rather than just its label.
Classify each proposed measure before adding it
| Status | What it means | Reporting treatment |
|---|---|---|
| Comparable as collected | Definition, timing, population and source rules align for the stated question | Combine under a documented calculation |
| Comparable after a justified transformation | A documented conversion or mapping preserves the intended meaning | Retain original values and explain the transformation |
| Useful but not directly comparable | Related evidence uses a different horizon, construct or verification rule | Show separately with context |
| Insufficient evidence | Definition, denominator or source is missing | Flag for clarification; do not treat as zero |
Record who approved the classification and when. Some decisions require measurement expertise or consultation with program staff. A data owner can organize the process without pretending that every mapping is a simple text-matching task.
Keep a separate field for missing reports. A program excluded because its measure differs is not the same as a program that has not submitted its report. Both affect coverage, but they require different follow-up.
Worked example: two comparable programs and one separate result
All numbers below are fictional. Programs A and B use the same employment definition, the same six-month follow-up window and the same source rules. Their populations do not overlap in this exercise. Program C measures employment at exit, so it cannot join the six-month rate.
| Program | Eligible people | Known follow-up status | Met employment definition | Treatment |
|---|---|---|---|---|
| A | 100 | 80 | 48 | Include at six months |
| B | 50 | 40 | 32 | Include at six months |
| C | 60 | 60 at exit | 42 at exit | Separate: different observation time |
For A and B, the combined observed rate is (48 + 32) ÷ (80 + 40) = 80 ÷ 120 = 66.7%. Follow-up coverage is 120 ÷ 150 = 80%. Thirty eligible people have unknown status. These are two different percentages answering two different questions.
A simple average of the program rates would be (60% + 80%) ÷ 2 = 70%. That gives each program equal weight. The 66.7% figure weights each observed person equally. For a person-level combined rate, use the counts under the stated non-overlap assumption; if you report a program-level average, label it accordingly.
Program C's 42 of 60 at exit can be shown in a separate panel. Adding its counts to the six-month result would mix time horizons. The combined view should explain this exclusion instead of silently dropping C or presenting its result as a failure.
A clear report sentence is: “Across Programs A and B, 80 of 120 people with known six-month status met the employment definition; follow-up covered 120 of 150 eligible people. Program C reports at exit and is shown separately.” This describes observation, not the number of jobs caused by the programs.
Check overlap before reporting unique people
If someone participates in two programs, adding program counts may count them twice. That may be appropriate for a measure of program participations, but not for unique people served.
Use an authorized identity or linkage method where a unique-person total is required and feasible. Define the reporting window and how duplicate or uncertain matches are handled. Do not share personal information across organizations simply to obtain a cleaner headline.
If you cannot reliably deduplicate, label the total as reported participations or service contacts rather than unique people. Explain the limitation. An uncertain unique-person estimate is not improved by presenting it as an exact count.
The same issue applies to outcomes. Two partners can report the same person, event or enterprise. A total of reported outcomes and a deduplicated outcome count are distinct measures. Record which one the report uses.
Keep comments connected to the result they explain
Review qualitative evidence from the programs included in a finding, and from excluded programs where it adds relevant context. Preserve the source, question, language, period and program. A comment from last year's intake should not quietly explain this year's endline result.
Use shared themes where the questions and coding purpose support them, while allowing program-specific themes. One program's “access” may concern transport and another's digital authentication. A broad label can hide a useful difference.
Choose a review approach appropriate to the purpose. A documented sample can support some qualitative questions. Reviewing all available responses can improve coverage, but an automated full pass can still miss or misclassify important material. Neither approach guarantees that every consequential concern will be found.
Record which comments were available, which were analyzed and how quality was checked. If the task includes identifying concerns that require a response, define the separate review and escalation process rather than relying only on a thematic report.
Customer practice: the King Center
The King Center's published customer story describes bringing pre- and post-survey evidence and qualitative feedback together across seven programs. It illustrates why program teams need both the measured results and what participants say about their experience.

The lesson for this exercise is to keep evidence available across programs and reporting cycles. The customer story is not evidence that unlike measures can be combined without checking their definitions, or that the fictional employment calculation above represents the King Center's results.
Investigate a site difference before ranking teams
A lower observed rate at one site may reflect delivery, participant circumstances, timing, response coverage or another factor. Start by checking whether the definitions and usable populations are comparable.
Then examine the context and the qualitative evidence. Ask whether the site began later, serves a different group, uses another language or has a different follow-up process. Those are questions to investigate, not explanations to assume.
Report the size and base of a difference. Small sites can show large percentage swings from a few observations. Avoid turning a descriptive table into an unsupported performance league table.
If a comparison will drive consequential resource or staffing decisions, choose an analysis appropriate to that use. A simple dashboard can reveal a question; it does not necessarily answer the causal or fairness question behind the decision.
Keep the reporting definition current
Store the definition, version, owner and effective date with the reporting logic. Preserve the prior version so an issued report remains interpretable after a rule changes.
When a program revises its intake or follow-up, review whether the change affects eligibility, measurement or timing. Mark the break or approved mapping in the trend. Do not silently recalculate the past under a new rule without explaining the revision.
A spreadsheet can be a workable starting point if it is governed and maintained. The problem is not the file format itself; it is losing ownership, version history or the relationship between the rule and the reported number.
For an AI-assisted workflow, provide the approved definitions and source references before requesting a combined narrative. Test whether an incompatible measure is flagged. A plausible summary should not substitute for the inclusion rule.
Practice with three programs before expanding
- Inventory the measures: record population, unit, period, source, definition and missingness for each program.
- Choose one shared question: classify which evidence is comparable and document exclusions.
- Check overlap: establish whether counts represent people, participations or another unit.
- Calculate from the appropriate base: retain numerator, denominator and coverage alongside the rate.
- Review qualitative context: document the source pool and analysis approach.
- Trace the result: have a second reviewer recover the included records and rule version.
- Test a change: add an incompatible follow-up period and confirm it is not silently pooled.
In a Sopact evaluation, use that test with the proposed collection and analysis setup. Ask to see how a new partner submission, document or survey result is checked against the agreed context, and how a reviewer resolves a flagged inconsistency. Distinguish the configured workflow from a promise that AI automatically resolves every mapping.
Your deliverable is a small reporting table with a defensible combined result and visible boundaries. Once it works, expand to another shared question rather than forcing every outcome into the first calculation.
Turn the combined view into a readable report
Use How to Write an Impact Report for the report structure and the report examples for presentation ideas. Keep definitions, coverage and exclusions close enough to each chart that the reader can interpret it.
For the broader organizational reporting problem, see impact reporting. This Academy lesson is the practical aggregation exercise: decide what can be combined, show the calculation and explain the rest.
Watch: why a clean table still needs context
Watch the video · 6 minutes 7 seconds. A companion discussion of reporting context and usable evidence. Browse more videos in the video library.
Frequently asked questions
Can different programs contribute to one report?
Yes. Combine sufficiently comparable measures and show other evidence separately. A shared report does not require every program to contribute to every total.
Must every questionnaire be identical?
No, but shared labels are not enough either. Check the construct, wording, scale, population, timing and source rules. Use standardized measures or justified mappings where appropriate.
Should we average program percentages?
Only if you intend a program-level average and label the weighting. For a person-level pooled rate across non-overlapping comparable groups, combine the relevant numerators and denominators.
How do we prevent double counting?
Define the unit and use an appropriate authorized linkage process where needed. If deduplication is not possible, report participations or another accurate unit rather than claiming unique people.
Does a framework tag make data comparable?
No. Preserve the exact metric definition, version and calculation guidance, then check whether each source supports it. A tag describes an alignment; it does not perform that validation.
Must every comment be read to report a theme?
Not always. Choose and disclose a suitable qualitative approach. Sampling and full-dataset analysis have different limits, and neither automatically guarantees detection of every important concern.
What belongs beside the combined rate?
The population, period, numerator, denominator, follow-up coverage, definition and programs included or excluded. Include overlap and comparability limits where they affect interpretation.
For distributed member reporting
Continue to surveying a member network. That reference adds distributed collection and appropriate views for each participating organization.
Reviewed September 12, 2026. Programs A, B and C and all employment counts are fictional teaching examples.