Record shape C of the four in Connected Data Intelligence — many programs, one picture. It leans hardest on checks 05 Qualitative, 03 Volume and 08 Reliable.
A board asks a straightforward question: what did we achieve across everything we ran this year?
Seven programs go looking for the answer. Each was designed by different people at different times for different purposes, and each collects its own feedback in its own way. What comes back is seven spreadsheets that cannot be added together, and a fortnight of somebody trying anyway. The eventual answer to the board is a list of seven separate stories, which is not what was asked.
The instinctive fix is to make every program ask the same questions. That is the wrong fix, and it is worth understanding why before anyone rebuilds a single survey.
Can you combine programs that ask different questions?
Yes — because what has to match is not the wording of your questions but the meaning of what you are counting. Two programs can ask about confidence in completely different words and still roll up honestly, provided both are measuring against the same agreed definition of the outcome. And two programs asking a word-for-word identical question can be impossible to combine, if one counts someone as "completed" at eighty per cent attendance and the other at full attendance.
That is the whole thing in one line: standardise the definitions, not the questionnaires. It is a much smaller and much more achievable change, and it does not require flattening seven programs into one shape they were never designed for.
A youth program and an adult training program should not ask the same questions — their participants are different people in different situations. But if both report "gained employment," both need to be using the same rule about what counts: within what period, at what number of hours, self-reported or verified. Once that rule is shared, the two can be combined. Until it is, they cannot, no matter how similar the forms look.
Writing those rules down is its own discipline, and it has its own method — start with giving every number one definition, then keep them somewhere governed in a data dictionary.
Roll up to what the programs actually share — and name what you left out
The second mistake is aggregating to the highest level someone asked for rather than the highest level that is honest.
Leadership wants one number for the organisation. Frequently no such number exists, because the seven programs genuinely do not share an outcome — a food programme and a leadership academy are not two instances of one thing. Manufacturing a combined figure anyway produces a composite index that nobody can interpret and that falls apart the first time someone asks how it was calculated.
The honest move is to roll up by the dimension the programs really do share, and to say plainly which programs are not in it. Something like: five of our seven programs measure participant confidence; across those five, here is the picture; the remaining two do not measure it, and here is what they measure instead. That is a more useful sentence than any single index, and it survives scrutiny because its boundaries are visible.
This also tells you where to invest. If only two of seven programs measure anything about what happened to participants afterwards, that gap is the finding — and it is more actionable than an average.
The part that actually stalls: nobody reads the comments
Every program in a portfolio like this collects open-ended feedback, and in most organisations almost none of it gets read.
The King Center is a clear case, and a public one. Programs from Nonviolence365 to the Beloved Community Leadership Academy generated feedback across cities and platforms — more than ten thousand stakeholder voices — with no dedicated analysts on staff. Their own Chief Research, Education and Programs Officer described it exactly: gathering open-ended feedback was always part of the routine, yet it remained untouched. Not neglected out of carelessness. Untouched because reading ten thousand comments is a job nobody had.
What changed for them is worth being precise about, because it is not simply "faster reports." Once before-and-after surveys from different cities were brought into one analysed view, the analysis started happening during the training session itself — live, with the participants still in the room, feeding the discussion they were having that day. Work that had taken months compressed into minutes. Read the King Center story.
That is the actual prize in this shape. Not a tidier annual report. The ability to use what people told you while you are still with them.
Same program, different city — is the difference real?
Once you can compare, you will find differences between sites running the identical program, and there is a trap here.
A city showing weaker results might have a weaker delivery team. It might also serve a harder-hit population, run in a language the materials were not written for, or have started three months later. Those are completely different conclusions and they look identical in a bar chart.
So treat a site difference as a question rather than a finding. What was different about who enrolled, what was different about how it ran, and what did participants themselves say about it. The comments are usually where the answer is, which is another reason the qualitative side is not optional in this shape.
Doing this without any particular software
- List every outcome each program currently claims. One row per program per outcome. Expect the list to be longer and messier than anyone predicted.
- Find the overlaps and write one shared rule for each. For every outcome claimed by more than one program, agree a single definition — who counts, over what period, verified how — and record which programs have signed up to it.
- Let each program keep its own wording. They report against the shared rule; they do not have to ask the same question to do it.
- Build the roll-up by shared outcome, and next to each figure list the programs included and the programs excluded. The exclusions are part of the number, not a footnote.
- Read a sample of comments from every program — not just the ones with good numbers. Fifty per program will tell you more than a full read of one.
For three or four programs, this works and is worth doing by hand once, because the argument you have while agreeing the shared rules is the valuable part.
Where it breaks
It breaks on the comments first. Reading a sample from every program is manageable; reading everything is not, and a sample means you find only the themes that are common enough to appear in a sample. The specific, urgent thing one participant said in one city is exactly what a sample misses, and it is often the most useful thing in the whole dataset.
It breaks on drift second. A shared definition agreed in January quietly diverges by September because a program changed its intake and nobody updated the rule. Six months later two programs are reporting the same outcome name against different meanings, and nothing announces it.
And it breaks on people leaving. The shared definitions usually live in one person's head and one spreadsheet. When that person moves on, the next reporting cycle rebuilds the logic from scratch and gets slightly different answers.
What a system is for here is narrow: reading every comment as it arrives rather than a sample of them at the end, holding the shared definitions somewhere with an owner and a version rather than in a document, and letting a program lead ask a question about their own programme without waiting for someone at the centre to build a report. Your team still decides what the shared outcomes are, and whether a difference between cities means anything.
How to test this before you commit
Use: Three programs that all claim one outcome, where two define it the same way and the third does not. Include open-ended comments from all three, at least one in another language, and two sites running the same program.
Pass: The roll-up covers the two programs that share the definition and names the third as excluded, with the reason. Themes come from all comments, not a sample, and each cites the words it came from. The site difference is presented as a question, with the comments that bear on it.
Fail: A single combined figure appears with all three programs in it. Or the excluded program is silently dropped.
Frequently asked questions
Do all our programs have to ask the same questions?
No, and forcing it usually makes each program's data worse. What has to match is the definition of what you are counting — who qualifies, over what period, verified how. Once that is shared, differently worded questions can still be combined.
How do we give the board one number?
Often you should not. Roll up by the outcome your programs genuinely share, state which programs are included and which are not, and explain what the excluded ones measure instead. That is more defensible than a composite index and more useful to a board than an average of unlike things.
Is a sample of comments good enough?
For finding common themes, usually yes. For finding the urgent thing one person said, no — and that is frequently the most valuable item in the set. A sample systematically misses the rare and the specific.
One of our cities has worse results. What does that mean?
Treat it as a question, not a conclusion. It could be delivery, or a different population, a language mismatch, or a later start. Those lead to opposite decisions and look the same in a chart, so check who enrolled and what participants said before acting.
Our programs were designed years apart. Is it too late?
No. You are not rebuilding the programs, you are agreeing what shared words mean and recording it. That work is additive and can start with a single outcome that two programs both claim.
Who should own the shared definitions?
A named person, with the definitions written down somewhere versioned rather than in a spreadsheet on someone's drive. The failure mode is not disagreement — it is that the person who knew the rules leaves and the next cycle silently invents new ones.
The reason this shape is hard has very little to do with having many programs. It is that "what did we achieve" sounds like a reporting question and is actually a definitions question, asked years too late. Organisations that agree what a handful of shared words mean can answer the board in an afternoon. Organisations that skip it can spend a fortnight and still produce seven stories where one was wanted.
Next: Back to the eight checks