This chapter resolves check 05 Qualitative of the eight checks.
An after-school reading programme finishes its year. Two-thirds of the children gained at least a grade level; a third did not move, and a few slipped back. The team sits down with the parent and student comments — four hundred of them, pasted into one long document — and reads for an afternoon.
They come out with three themes: the tutors were caring, scheduling was hard, families wanted more sessions. All true. None explains why a third of the children did not improve, because all three came from a document where those children's comments sat mixed in with everyone else's, outnumbered two to one.
Buried in it, said by eleven families, is that the tutoring block ends at 4:30 and the bus they depend on leaves at 4:15. Ten of the eleven are in the group that did not improve.
Why doesn't our qualitative data explain our numbers?
Because you are reading everyone's comments together. Split people by what happened to their number first, then read each group's comments on their own. A merged read produces a blended narrative — a description of your programme in general, which is not an explanation of your result.
The test that shows a merged read explained nothing
Take the themes from your last merged read, cover up the outcome figure, and ask: if the numbers had gone the other way, would these themes look wrong?
Usually not. "Tutors were caring, scheduling was hard, families want more" fits a programme that worked and one that failed equally well. A finding that survives every possible result is not a finding — it is a portrait.
Two things cause it. The bigger group dominates by sheer volume, so if most participants improved you are mostly reading successful people's comments. And the longest, most fluent answers pull themes toward whoever writes best, who is rarely the person having the hardest time.
Split first, read second
The order matters more than the technique. Decide the groups from the numbers before you read a single comment, then read one group at a time.
Two or three groups is enough: improved and did not improve, or met the threshold, missed it, went backwards. The line comes from the programme's own standard where one exists — a grade level, a passing score, a placement — written down before the reading starts.
Illustrative example — the same themes, split by what happened
| What people said | Among those who improved | Among those who did not | What it is |
|---|
| Tutor was kind and patient | Most of them | Most of them | A feature of the programme. Worth keeping, explains nothing. |
| Wanted more sessions | Common | Common | A wish, not a cause. |
| Session ended after the bus left | Almost nobody | One in eight | The finding. One-sided, specific, fixable. |
Illustrative example — proportions show the shape of a result, not measured values.
A theme in both columns is a programme characteristic. A theme in one column is a lead. That is the whole logic, and it is why the merged read fails: merging collapses the two columns into one and throws away the only comparison that carried information.
The people who did not improve are the more useful half
Every incentive points the other way. The improvers are the larger group, their comments are warmer to read, and their words are what a funder report wants — so they get read first, quoted most, and shape the themes.
But improvers can only tell you what to keep. The people who did not improve are the only ones who can tell you what to change. There is also a distortion worth naming: participants who were helped write about gratitude, and gratitude crowds out mechanism. "Ms. Alvarez was wonderful" is sincere and says nothing about what she did. Someone who did not improve has no gratitude script to fall back on, so they describe the obstacle instead. Read that smaller, harder group first, while your attention is fresh.
Doing this without any particular software
- Get the number and the comment onto the same row. One row per person, outcome in one column, open response in the next. Everything below is impossible without it, and it is the step that usually fails.
- Write the split rule down before reading. "Improved means a gain of one grade level or more; no change means within one level; declined is below." Date it. A rule written afterwards is a rule shaped by the story you already formed.
- Sort by group and read one group in one sitting, writing that group's themes on its own sheet. Do not look at the other sheet yet — once you have seen it you will start matching.
- Read the smallest group first. It is usually the non-improvers, and it is where the unfamiliar material is.
- Put the two lists side by side and label each theme: in both, only among improvers, only among non-improvers.
- Count the one-sided ones as a share of their own group. Eleven of ninety non-improvers against four of two hundred and ten improvers is a real asymmetry. Raw counts hand it to the bigger group automatically.
- Pull two verbatim quotes per one-sided theme, exactly as typed. Those are what make the finding survive a sceptical reader.
Four hundred comments, two groups, an afternoon each. Not difficult work — just ordered differently from how everyone does it.
Where it breaks
The number and the comment live in different files. Test scores in one export, the feedback survey in another, matched on name or email. The match is never complete — nicknames, a second address, a form filled in by a parent — and the residue lands in an "unmatched" pile often bigger than the non-improver group itself. The split cannot be trusted then, because the people you failed to match are not a random selection either.
The split line moves to fit the story. Somebody reads first, forms a view, then picks a threshold that makes the view visible. Nobody does this dishonestly and it is invisible afterwards, which is exactly why the rule gets written and dated first.
You get to do it once. Having read four hundred comments split by outcome, you now want the split by site, by cohort, by language. Each means reading everything again, so the second cut never happens — and the interesting question, whether the bus problem is one site or all of them, goes unanswered.
The narrow claim: because each response stays attached to a person along with that person's numbers, Sopact Sense lets themes already applied be re-cut by any outcome group without re-reading — improvers against non-improvers today, by site tomorrow — and the original wording is kept, so each one-sided theme cites the participant's own words. It will not tell you where the split line belongs. That judgement is yours, and it belongs on paper before you look.
How to test this on your own data
Use: One outcome measure and one open question from the same cohort, where results were mixed. Write the split rule down and date it before opening any comments.
Pass: Every person's number and comment sit on one row with no unmatched pile. Two separate theme lists, each labelled with the share of its own group. At least one theme appears in only one group. Then re-cut by site without re-reading.
Fail: A single theme list covering everyone. Or an unmatched pile you cannot account for. Or a split rule that cannot be shown to predate the reading.
Frequently asked questions
What if the two groups are very different in size?
Compare shares within each group, never raw counts. Thirty comments out of ninety non-improvers is a third of that group; thirty out of three hundred improvers is a tenth. Reported as "thirty people said this" they look identical, which is how the larger group quietly wins.
Where should the line between improved and not improved go?
Use the standard your programme already commits to — a grade level, a passing score, a placement. Write it before reading. If no standard exists, pick one and then test a second line; if the themes hold under both, the finding is robust rather than an artefact of the cut.
We have no outcome number, only feedback. Can we still do this?
Yes. Split on something that happened rather than something said — completed or not, attended more than eight sessions or fewer, placed or not placed. The split needs to be a fact about the person recorded independently of their comment, not a score.
Isn't reading one group separately a form of cherry-picking?
Not when the rule is written first and both lists are reported. Cherry-picking is selecting quotes that support a conclusion. This is the reverse: every comment in each group is read, both lists are published, and shared themes are named as shared rather than offered as explanations.
What about people whose score improved but whose comments are negative?
That group is often the most informative in the whole exercise, so do not smooth the contradiction away. It usually means the measure is capturing something narrower than what the programme is meant to change — the score moved, the person's situation did not.
Do we still need an overall read of all the comments?
A quick one is fine for orientation, to learn the vocabulary people use. Just do not let it become the explanation section of the report. Written up as the finding, an overall read is the blended narrative again, wearing the clothes of analysis.
The received wisdom is that numbers tell you what happened and comments tell you why. That is close enough to true to be misleading. Comments do not explain the number; they explain the difference between the people the number treated differently — which is why the split comes first.
And the split is informative even when it fails. If the two lists come back genuinely identical, you have learned something specific: whatever separated the groups was not visible to the participants themselves, so nobody will find it in the feedback. That is the signal to go and look at attendance records, tutor assignments, session timing, room changes — the operational facts your programme holds and never asked anybody about.
Next: Analyze pre / mid / post data