This chapter resolves check 04 Longitudinal of the eight checks.
A workforce programme surveys its cohort at intake, again at week eight, and again at graduation. Average confidence moves from 3.1 to 3.6 on a five-point scale. The board sees plus half a point, agrees the programme works, and renews it. A year later someone opens the same file and counts people instead of averaging them: sixty-one participants rose, thirty-three did not move, twenty-six finished lower than they started.
Nothing in the original report was false. The average was arithmetically correct. It simply was not the shape of answer the board thought it was getting, because twenty-six people who got worse do not appear anywhere in a sentence about the mean.
How do you analyze pre, mid and post survey data?
Count how many people improved, how many stayed the same, and how many declined — and report those three counts instead of the change in the average. The average change is the one summary that cannot be acted on, because almost any pattern of individual movement can produce it. Three counts and a threshold take the same afternoon to produce and they answer a different, better question: not did the number go up, but for how many people, and at whose expense.
This chapter is about a fixed set of waves inside one programme — baseline, midpoint, endline, then the programme ends. If instead you are following the same individuals across many waves over several years, and what you need to know is the shape of each person's path over time, that is analyzing longitudinal survey data. Fixed waves, one cycle, is here.
The average change is compatible with almost any story
Take three programmes that each report a gain of half a point. They are not variations on one finding. They are three different situations, and the correct response to each is different.
| Same average gain | What actually happened to 100 people | What it asks you to do |
|---|
| Programme A | Almost everyone nudged up a little. 78 improved slightly, 20 flat, 2 declined. | Keep going, and ask whether a small gain for nearly everyone is worth what it costs. |
| Programme B | Half the cohort improved substantially, the other half got worse. 50 up two points, 50 down one. | Stop and find out who is in each half. This programme is doing harm to fifty people. |
| Programme C | Nothing moved for 85 people. 15 improved enormously. | Find out what the fifteen had in common. The programme currently works for a narrow group. |
Illustrative example. The point is not the specific counts. It is that the headline figure is identical in all three rows, and no reader of that figure can tell which row they are in.
What the midpoint is actually for
Baseline and endline tell you what happened after it is too late to change it. The midpoint is the only wave that exists to be acted on, and counting the distribution is what makes it actionable. At the midpoint, the people who have declined are not a percentage — they are a list of names, still enrolled, still reachable. A programme manager can do something with a list of eleven names in week eight. Nobody can do anything with a mean at graduation.
Which means the midpoint survey is only worth running if someone has agreed in advance to look at the decliners and respond. If that agreement does not exist, you are collecting a third wave to make a chart smoother.
How to do this without any particular software
This is a spreadsheet method. Build it by hand once and you will never again accept a bare average.
- One row per person, one column per wave. Baseline, midpoint, endline, side by side. This only works if the same person carries the same identifier through all three waves, which is a decision made before the first survey goes out, not during analysis.
- Analyse only the people who have all three readings, and write down how many that is. If 140 people gave a baseline and 96 gave all three, your counts describe 96 people and every sentence should say so.
- Decide what counts as a real change before you look. On a five-point scale, one full point is a defensible line. Write the rule down and date it. Choosing this after seeing the results is how a null finding becomes a positive one.
- Label each person. Endline minus baseline. Up by at least your threshold: improved. Down by at least your threshold: declined. Anything in between: unchanged. One word per person.
- Count the three labels. That is your finding. Do the same for baseline-to-midpoint, so you can see whether the decliners declined early or late.
- Repeat inside every subgroup you can support — site, cohort, age band, referral source. Do not report a count from a group of six.
- Publish five numbers together: people with all three waves, improved, unchanged, declined, and the threshold you used. Any reader can then recompute your rate, which is the whole point.
Run that and check 04 is satisfied for one cycle, with no software beyond the spreadsheet you already have.
Where the manual version breaks
It breaks in three places, and each one has the same effect: the distribution disappears and the average comes back.
The threshold moves. Somebody recomputes with a half-point threshold because the one-point version looked disappointing, and the new figure goes in the report without the old rule attached. Nobody is being dishonest; the rule was never written anywhere that a report could point at.
The three counts get computed on three different groups. A dozen people miss the midpoint. The baseline-to-midpoint count uses one set of people, the baseline-to-endline count another, and the two are then compared as if they described the same cohort.
The subgroup breakdown gets done once. It is the most laborious part by hand, so it happens in year one and never again — which is exactly the analysis that would have found the site where a third of participants decline every single cycle while the programme average stays healthy.
The narrow claim for a system is this: because each person's waves are already attached to one record, the matched set, the three counts and every subgroup breakdown are produced the same way each cycle rather than rebuilt, and the change threshold is stored with the result instead of living in someone's memory. It does not choose your threshold, decide which subgroups matter, or tell you what to do about the decliners. That is judgement, and it stays with your team.
How to test this on a real cycle
Use: One real cohort with all three waves collected, including at least a few people who scored lower at endline than at baseline, and at least one site or subgroup that performed worse than the rest.
Pass: The report states how many improved, stayed flat and declined; names the threshold used; names how many people had all three waves; and shows the same three counts for each subgroup large enough to report.
Fail: The headline is a change in the average, and the number of people who got worse appears nowhere.
Do it on a cycle you have already reported on. The interesting result is the gap between what you published and what the counts say.
Frequently asked questions
Should we stop reporting the average change entirely?
No — report it, but underneath the counts rather than as the headline. The average is a reasonable one-line summary for a funder skimming a page. It stops being reasonable the moment a decision rests on it, because the same average is produced by situations that call for opposite decisions.
What counts as a meaningful change on a five-point scale?
One full point is the usual defensible line, because a person moving from 3 to 4 has crossed a labelled category rather than drifted. Half-point rules make almost everyone a mover and the counts stop meaning anything. Whatever you choose, choose it before you look and write it down.
Do we really need a midpoint survey?
Only if someone will act on it. The midpoint's value is that the decliners are still enrolled and still reachable, so the count converts into a list of names and an intervention. If no one has agreed to look at that list, two waves will tell you the same story for less burden on participants.
What about people who answered only two of the three waves?
Report them, do not silently use them. Compute your counts on the people with all three, state that number, and state separately how many gave a baseline only. Partial responders quietly mixed into the same table make the counts look larger and more precise than they are.
How small is too small for a subgroup count?
Below about twenty people, one participant can swing the percentage by five points or more, so publish the raw counts rather than a rate — "four of nine" instead of "44 per cent". Small groups still deserve attention; they just cannot carry a percentage.
Why can two programmes with the same average need opposite decisions?
Because the average records net movement and decisions depend on who moved. A programme where everyone gains a little needs efficiency questions. A programme where half gain and half lose needs to know what distinguishes the halves before it enrols another cohort. The mean is identical and the responses share nothing.
Isn't a decline just measurement noise?
Some of it is, which is exactly why the threshold exists — a one-point rule filters out the drift. But a stable share of your cohort declining by a full point every cycle is not noise, it is a pattern, and a threshold you set in advance is what lets you tell the two apart without arguing about it afterwards.
The habit worth taking from this is a small change in what you ask for. Nobody in your cohort experienced the average. It is a property of the group, and the group is not the thing your programme serves. A programme can be working and failing at the same time — genuinely both — and the only report that can say both sentences at once is the one that counted people in three directions instead of moving one number.
Next: Track change across many years