This chapter resolves check 04 Longitudinal of the eight checks.
A youth employment programme has surveyed the same participants every six months for four years. One young woman scores 2 on a five-point confidence measure at intake. Six months in she is at 4. By month eighteen she is back to 2, and she is still at 2 at month forty-two. Her first and last readings are identical, so the analysis files her as no change — in the same row of the same table as a young man who answered 2 at every single wave and never moved at all.
The programme reached her and then lost her. It never reached him. Those are different failures with different fixes, and a table built from first and last readings cannot tell them apart. Both endpoints agree, and both are wrong about what happened.
How do you analyze survey data collected over many years?
Classify each person's whole sequence of readings into one of a small number of shapes, then count how many people are in each shape. Not the average at each wave, and not first-versus-last. The unit of analysis is one person's path, and the output is a handful of counts — improved and held, improved then fell back, late improver, never moved. Those counts can be acted on. A trend line cannot, because there is no intervention whose target is a slope.
If your programme has a fixed baseline, midpoint and endline and then ends, you do not need any of this — you need three counts and a threshold, which is analyzing pre, mid and post survey data. This chapter starts where that one stops: many waves, over years, following the same individuals.
A cohort average hides this twice over
The flat line in the annual report is flat for two independent reasons, and it is worth separating them because most teams only know about the first.
Across people. At month eighteen, some participants are climbing and others are falling back. Averaged together they cancel. The cohort mean sits still while a great deal is happening underneath it.
Within one person. Even if you look at a single participant, summarising her as first-versus-last throws away every reading in between. Her peak at month six — the moment the programme was working — has been deleted from the record by the summary, not by the data.
So a flat cohort line is consistent with total inertia and with a programme that produces real gains and then loses them. Those two possibilities have nothing in common except the picture.
The shapes worth naming
Four to six shapes is enough. More than that and staff disagree about the edges; fewer and you are back to up or down.
| Shape | The sequence looks like | What it asks of the programme |
|---|
| Improved and held | Rises early, stays at or near the peak through the last reading. | Nothing. This is the group your model is built for. Find out what they had in common. |
| Improved then fell back | Rises, peaks, returns towards the starting level and stays there. | A follow-up contact before the month the fall begins. This is the most actionable group you have. |
| Late improver | Flat through the programme, rises well after it ends. | A longer measurement horizon. If you had stopped measuring at exit you would have recorded a failure. |
| Never moved | Same value, within noise, at every reading. | Ask whether the programme reached them at all — attendance, language, timing — before questioning the outcome. |
| Declined throughout | Falls steadily from the first reading onward. | Individual review. Something outside the programme is usually driving this. |
| No settled direction | Up and down repeatedly with no trend. | Check the measure before the person. Volatility this high often means the question is being read differently each time. |
The counts across those rows are the finding. "Of 212 participants with four or more readings, 58 improved and held, 71 improved then fell back, 24 were late improvers, 44 never moved" is a paragraph a programme director can build next year's calendar around. The equivalent trend line is a decoration.
Time has to be measured from each person's own start
A wave number is not a date. In a rolling programme, one person's third survey lands fourteen months after her intake and another's lands twenty-two months after his. Line those up as "wave 3" and you have compared a fourteen-month reading to a twenty-two-month one, which is precisely the interval where the fall-back happens.
So every reading needs two things recorded next to it: the value, and the number of months since that person's intake. Shapes are read along that second column. This is the single most common reason a shape analysis produces counts that nobody can reproduce.
How to do this without any particular software
- Lay the data out long, not wide. One row per person per survey: identifier, value, date, and months since that person's intake. Wide layouts with a column per wave force the alignment error above.
- Sort by person, then by months, and read each person's values as a short sequence —
2, 4, 3, 2, 2. Do this for thirty people before you write any rules. The shapes you need will announce themselves.
- Write the rulebook before you classify. Each shape gets an explicit rule: "improved and held means the value rose at least one point above the first reading and the final reading is within one point of the peak." Date it and put a name on it.
- Give every person exactly one shape, and create an honest unclassified bucket for anyone who fits none. Keep that bucket visible in the output; a large one means the rulebook is wrong.
- Count the shapes. Then, for the fell-back group only, note the month at which each person's decline began and take the range. That month range is the most useful number the whole analysis produces.
- Read the open-ended answers one shape at a time. The fell-back group's own words at the wave after their peak will usually tell you what happened, and reading them mixed in with everyone else's guarantees you will miss it.
Where the manual version breaks
It breaks in three ways that are specific to many waves, and none of them is a discipline problem.
Two people apply the rules differently. One counts a dip of exactly one point as a fall-back, the other does not. The shape counts then move between reports for reasons that have nothing to do with participants.
Every new wave changes people's shapes. Someone classified as improved and held at wave five becomes improved then fell back at wave seven. That is not an error, it is the point — but it means the entire cohort has to be reclassified each wave, and by hand that means it is classified once, in the year someone had time.
People have different numbers of readings. Late joiners have three, the first cohort has nine. Count shapes across both and you are comparing paths of different lengths, where "held" means something much weaker for the short ones.
Here is the narrow claim for a system: because each person's readings are already joined to one record with their own elapsed time attached, the shape rules can be stated once and re-applied to everybody whenever a new wave lands, so the counts are recomputed rather than redone, and anyone can see which version of the rules produced a given number. It will not invent your shapes or tell you what the fall-back means. Reading the fell-back group's words and deciding what to do in month twelve is your work.
How to test this on your own cohort
Use: One cohort with at least four readings each, deliberately including one person who rose and fell back, one who never moved, one who improved only after the programme ended, and two people who joined a year apart.
Pass: The fall-back person and the never-moved person land in different shapes. The late improver is not recorded as a failure. The two staggered joiners are compared at the same months-since-intake, not the same wave number. Adding a new wave reclassifies everyone under the same dated rulebook.
Fail: The output is a line chart of cohort means, or the fall-back and never-moved participants share a row.
Frequently asked questions
How many waves do you need before shapes mean anything?
Three is the minimum — you cannot see a turn with two. Four is where "held" starts to mean something, because it takes one reading to establish the peak and another to show it stayed. Below three readings, use the fixed-wave method instead and say so plainly.
What if someone misses a wave in the middle?
Classify them on the readings you have and record how many they had. A person with readings at months 0, 6 and 30 can still be shaped, but the gap means you cannot say when the turn happened. Do not fill the hole with an interpolated value — that manufactures the smoothness the whole method is trying to avoid.
Isn't this just what a statistician would call growth curve modelling?
It is a plain-language cousin of it. The formal version fits a curve to each person and groups similar curves; this fits a labelled rule instead. The advantage of rules is that a programme manager can read them, argue with them, and check a specific participant against them, which is what makes the counts get used.
How many shapes should we use?
Four to six. Every extra shape creates a new boundary for two staff to disagree about, and boundaries are where reproducibility is lost. If a shape holds fewer than about five per cent of your cohort, fold it into a neighbour and note that you did.
Can the shape rules ever change?
Yes, but only as a versioned change applied to everybody at once. Change the rule, re-run the whole cohort, and report both the new counts and the fact that the rule moved. Changing a rule and republishing only the current wave produces a trend that is entirely an artefact of the definition.
Should we report the cohort trend line at all?
Report it if a funder expects it, underneath the shape counts, and let the counts explain it. A flat line above a table showing 71 people who improved and fell back is a useful pair. A flat line alone invites the conclusion that nothing worked, which in that cohort would be false.
Does this work when the measure is a written answer rather than a score?
Yes, and it is often better. A person's answers to the same open question across five waves form a sequence you can shape, as long as the wording never changed and the original text is kept. The shapes then come from what people say, not from a scale.
The idea worth leaving with is that long-term measurement does not produce a curve. It produces a month. The fell-back group has a point at which the decline starts, and once you know that point sits somewhere between month twelve and month eighteen, the finding stops being a research result and becomes a scheduling decision: put a contact in the calendar at month ten. Nothing in a trend line ever turns into a calendar entry.
Next: Measure how long outcomes last