Seven programs, seven spreadsheets, and a board asking one question. The fix is not making every program ask the same questions — it is agreeing what a handful of shared words mean.
Record shape C of the four in Connected Data Intelligence — many programs, one picture. It leans hardest on checks 05 Qualitative, 03 Volume and 08 Reliable.
A board asks a straightforward question: what did we achieve across everything we ran this year?
Seven programs go looking for the answer. Each was designed by different people at different times for different purposes, and each collects its own feedback in its own way. What comes back is seven spreadsheets that cannot be added together, and a fortnight of somebody trying anyway. The eventual answer to the board is a list of seven separate stories, which is not what was asked.
The instinctive fix is to make every program ask the same questions. That is the wrong fix, and it is worth understanding why before anyone rebuilds a single survey.
Yes — because what has to match is not the wording of your questions but the meaning of what you are counting. Two programs can ask about confidence in completely different words and still roll up honestly, provided both are measuring against the same agreed definition of the outcome. And two programs asking a word-for-word identical question can be impossible to combine, if one counts someone as "completed" at eighty per cent attendance and the other at full attendance.
That is the whole thing in one line: standardise the definitions, not the questionnaires. It is a much smaller and much more achievable change, and it does not require flattening seven programs into one shape they were never designed for.
A youth program and an adult training program should not ask the same questions — their participants are different people in different situations. But if both report "gained employment," both need to be using the same rule about what counts: within what period, at what number of hours, self-reported or verified. Once that rule is shared, the two can be combined. Until it is, they cannot, no matter how similar the forms look.
Writing those rules down is its own discipline — start with giving every number one definition, then keep them somewhere governed in a data dictionary — but the sheet alone will not tell you whether two programs' fields are the same thing, and that takes four checks on each field: where the answer came from, at what stage it was collected, why it was asked, and what the definition says counts.
Watch (6:07): what a shared rule still leaves undecided — and why a portfolio total of 80 jobs created is fiction when two investees each counted a job their own way.
The second mistake is aggregating to the highest level someone asked for rather than the highest level that is honest.
Leadership wants one number for the organisation. Frequently no such number exists, because the seven programs genuinely do not share an outcome — a food programme and a leadership academy are not two instances of one thing. Manufacturing a combined figure anyway produces a composite index that nobody can interpret and that falls apart the first time someone asks how it was calculated.
The honest move is to roll up by the dimension the programs really do share, and to say plainly which programs are not in it. Something like: five of our seven programs measure participant confidence; across those five, here is the picture; the remaining two do not measure it, and here is what they measure instead. That is a more useful sentence than any single index, and it survives scrutiny because its boundaries are visible.
This also tells you where to invest. If only two of seven programs measure anything about what happened to participants afterwards, that gap is the finding — and it is more actionable than an average.
Every program in a portfolio like this collects open-ended feedback, and in most organisations almost none of it gets read.
The King Center is a clear case, and a public one. Programs from Nonviolence365 to the Beloved Community Leadership Academy generated feedback across cities and platforms — more than ten thousand stakeholder voices — with no dedicated analysts on staff. Their own Chief Research, Education and Programs Officer described it exactly: gathering open-ended feedback was always part of the routine, yet it remained untouched. Not neglected out of carelessness. Untouched because reading ten thousand comments is a job nobody had.
What changed for them is worth being precise about, because it is not simply "faster reports." Once before-and-after surveys from different cities were brought into one analysed view, the analysis started happening during the training session itself — live, with the participants still in the room, feeding the discussion they were having that day. Work that had taken months compressed into minutes. Read the King Center story.
That is the actual prize in this shape. Not a tidier annual report. The ability to use what people told you while you are still with them.
Once you can compare, you will find differences between sites running the identical program, and there is a trap here.
A city showing weaker results might have a weaker delivery team. It might also serve a harder-hit population, run in a language the materials were not written for, or have started three months later. Those are completely different conclusions and they look identical in a bar chart.
So treat a site difference as a question rather than a finding. What was different about who enrolled, what was different about how it ran, and what did participants themselves say about it. The comments are usually where the answer is, which is another reason the qualitative side is not optional in this shape.
For three or four programs, this works and is worth doing by hand once, because the argument you have while agreeing the shared rules is the valuable part.
It breaks on the comments first. Reading a sample from every program is manageable; reading everything is not, and a sample means you find only the themes that are common enough to appear in a sample. The specific, urgent thing one participant said in one city is exactly what a sample misses, and it is often the most useful thing in the whole dataset.
It breaks on drift second. A shared definition agreed in January quietly diverges by September because a program changed its intake and nobody updated the rule. Six months later two programs are reporting the same outcome name against different meanings, and nothing announces it.
And it breaks on people leaving. The shared definitions usually live in one person's head and one spreadsheet. When that person moves on, the next reporting cycle rebuilds the logic from scratch and gets slightly different answers.
What a system is for here is narrow: reading every comment as it arrives rather than a sample of them at the end, holding the shared definitions somewhere with an owner and a version rather than in a document, and letting a program lead ask a question about their own programme without waiting for someone at the centre to build a report. Your team still decides what the shared outcomes are, and whether a difference between cities means anything.
Use: Three programs that all claim one outcome, where two define it the same way and the third does not. Include open-ended comments from all three, at least one in another language, and two sites running the same program.
Pass: The roll-up covers the two programs that share the definition and names the third as excluded, with the reason. Themes come from all comments, not a sample, and each cites the words it came from. The site difference is presented as a question, with the comments that bear on it.
Fail: A single combined figure appears with all three programs in it. Or the excluded program is silently dropped.
No, and forcing it usually makes each program's data worse. What has to match is the definition of what you are counting — who qualifies, over what period, verified how. Once that is shared, differently worded questions can still be combined.
Often you should not. Roll up by the outcome your programs genuinely share, state which programs are included and which are not, and explain what the excluded ones measure instead. That is more defensible than a composite index and more useful to a board than an average of unlike things.
For finding common themes, usually yes. For finding the urgent thing one person said, no — and that is frequently the most valuable item in the set. A sample systematically misses the rare and the specific.
Treat it as a question, not a conclusion. It could be delivery, or a different population, a language mismatch, or a later start. Those lead to opposite decisions and look the same in a chart, so check who enrolled and what participants said before acting.
No. You are not rebuilding the programs, you are agreeing what shared words mean and recording it. That work is additive and can start with a single outcome that two programs both claim.
A named person, with the definitions written down somewhere versioned rather than in a spreadsheet on someone's drive. The failure mode is not disagreement — it is that the person who knew the rules leaves and the next cycle silently invents new ones.
The reason this shape is hard has very little to do with having many programs. It is that "what did we achieve" sounds like a reporting question and is actually a definitions question, asked years too late. Organisations that agree what a handful of shared words mean can answer the board in an afternoon. Organisations that skip it can spend a fortnight and still produce seven stories where one was wanted.
Start with data your teams struggle to bring together. Agree shared definitions, keep each source identifiable, and decide who can see what before asking AI for an answer.
Explore Connected Data Intelligence →