This chapter resolves check 03 Volume of the eight checks.
A workforce programme surveys its participants three times: at intake, at week eight, and at completion. The confidence measure climbs at every wave, and the chart in the board pack goes up and to the right. Nobody asks how many people each point describes.
The intake point describes everyone. The completion point describes fewer than half of them — and the participants who stopped answering were, on average, the ones who had scored lowest at intake. The line went up. What it recorded was people leaving.
Our average is rising but fewer people are answering. Is the programme working?
You cannot tell from the average, and the fix is a comparison you can do in an afternoon: compare the people who stopped answering against the people who stayed, using what you knew about them at the start. If the leavers started lower, part of your improvement is not improvement. It is the departure of the people it was not working for.
Why the people who leave are never a random sample
Attrition gets treated as bad luck — busy lives, changed numbers, survey fatigue. Some of it is. But list the actual reasons someone stops answering and they lean one way:
- They left the programme, which is often why there is nothing good to report.
- They lost the job the programme placed them in, and the survey arrives as a reminder.
- They moved, changed number, or their circumstances became less stable.
- They have nothing positive to say, and no appetite for saying it to the organisation that helped them.
Somebody for whom the programme worked has an easy, mildly gratifying survey to fill in. Somebody for whom it did not has a harder one. That asymmetry is the whole mechanism: your remaining respondents get more favourable at every wave whether or not anything changes for anyone.
The arithmetic, on purpose
It helps to watch this with numbers small enough to check by hand.
Illustrative example — a rising average with nobody improving
| Group | How many | Score at intake | Score at endline |
|---|
| Everyone at intake | 100 | 5.0 | — |
| Answered at endline | 65 | 5.6 | 5.6 |
| Stopped answering | 35 | 3.9 | not collected |
The headline reads 5.0 at intake, 5.6 at endline — a gain of 0.6. But the 65 people who answered at endline had already averaged 5.6 at intake. Not one of them changed. The whole movement is the absence of the 35 who started at 3.9.
The error is comparing two different groups of people and calling it change over time. Compare the endline respondents to their own intake scores and the gain disappears. That single correction — same people, both ends — catches most of the damage, and needs no statistics beyond an average.
Two numbers, always together
From here on, every outcome you report carries two companions.
The matched change: the average change among people who answered at both points. Not wave one's average against wave two's — the same individuals, differenced.
The starting gap between leavers and stayers: the intake average of people who stopped answering, next to the intake average of people who stayed. If those two are close, your matched change is probably fair for the whole cohort. If the leavers started well below, say so in the same paragraph as the result. That gap is the honest size of your uncertainty, in the units your programme already uses.
Track who is missing as a list of names, not a percentage
"Endline response rate: 65%" is a fact you cannot act on. Nobody can call a percentage. The same information as 35 names, each with a last contact date and a staff member responsible, is a work queue — and it is also a finding, because the list has a shape. Twenty-eight of the 35 turn out to be from one site, one cohort, the evening group, or one referral partner. A percentage hides that completely.
Sort that list by intake score and the payoff is uncomfortable: it is a ranked list of the people the programme is losing track of, most-in-need first.
Doing this without any particular software
- Build one row per person and one column per wave. Wide, not stacked. The gaps become visible as blank cells instead of hiding as absent rows.
- Add a column that says whether they answered the latest wave. Yes or no. This column is the engine of everything below.
- Average the intake score twice, split by that yes/no column. Two numbers. Their difference is your attrition warning, and if it is large your headline gain is partly composition.
- Compute change only for people with both points, and write beside every figure how many people it describes. A number without that count is not reportable.
- Run the same split on the categories you care about — site, cohort, age band, language, referral source. Attrition concentrating in one of them is a programme finding, not a survey problem.
- Print the missing list with names, last contact and an owner, and work it before you finalise anything. One recovered response from that list is worth ten from people who were going to answer anyway.
- Repeat the split at every wave against the original intake, not the previous wave. Attrition compounds, and wave-on-wave comparison hides the accumulated tilt.
Where it breaks
Each wave arrives as its own export. Rebuilding the one-row-per-person sheet by hand is where identity slips — a name spelled differently, a new email address — and someone who did answer gets counted as missing, or worse, as a new participant. Your leaver list is then partly people who never left. This is the failure described in collecting feedback offline without losing who said what, and it makes everything here unreliable.
The missing list gets made once. Someone produces it in a burst of diligence at wave two, follow-up happens, results improve. Nobody produces it at wave three, because it was manual and that person is busy. Attrition work has to be routine to be worth anything, and a manual version is never routine for long.
The counts fall off in the last hundred metres. The analyst knows how many people each figure describes. The slide does not carry it, the annual report copies the slide, and the number gets quoted for two years detached from the 65 people out of 140 it was built from.
The narrow claim for a system: because responses stay attached to a person rather than to an export, Sopact Sense knows at any moment who has answered which wave and who has not — so the leaver-versus-stayer comparison and the named missing list are available without rebuilding anything, at every wave rather than the one wave somebody had time for. Outcome figures carry the number of people they describe. Your team still decides what counts as an acceptable gap.
How to test this on your own data
Use: A cohort with a baseline and at least two later waves, where some people stopped answering. Do not clean the missing people out first.
Pass: In one sitting you can produce the baseline average for people who answered the latest wave and for those who did not; the average change among people with both points; and a named missing list with last contact dates. Every figure shows how many people it covers.
Fail: The only available comparison is wave one's average against wave two's, across whoever happened to answer. Or the missing people exist as a percentage and cannot be named.
Frequently asked questions
How much attrition is too much?
There is no threshold worth trusting. Losing 40% at random damages your results less than losing 15% who all started at the bottom. The question is never how many left but whether they differed at the start, so the leaver-versus-stayer comparison replaces the rule of thumb rather than supplementing it.
Should we weight or fill in the missing responses?
Methods exist and can be appropriate with statistical support, but they are not the first move. Report the matched change and the starting gap first, because you can explain those to a board in two sentences. An adjustment nobody in the room can explain converts a visible problem into an invisible one.
We didn't collect a full baseline. Can we still do this?
Yes, with whatever you knew at intake — site, referral source, age band, attendance in the first two weeks, a single screening question. The comparison needs some characteristic recorded before the outcome, not the outcome itself. Even one such column tells you whether the leavers were a different sort of participant.
Can we just report results for people who completed?
You can, if you say that is what you are doing and publish how those people differed at intake. That is a defensible analysis of programme completers. What is not defensible is presenting it as the result for everyone who enrolled, which is what happens when the count is missing.
What do we say when the honest number turns out to be flat?
Lead with the leaver comparison, because it is the more useful finding. "Participants who stayed held steady; the people we lost started well below them, and most left in the first six weeks" tells a funder where the programme needs work. A rising average that dissolves under scrutiny tells them nothing they can fund.
Isn't chasing non-responders a way of biasing the sample?
It is the opposite. Extra effort aimed at the group most likely to differ from your respondents reduces the tilt rather than adding one. What biases things is chasing the easy contacts — people already inclined to answer — which raises the response rate while leaving the gap where it was.
The reframe worth carrying out of this chapter is that people stopping answering is not primarily a data problem. Non-response is the earliest outcome measurement you own. People go quiet before they tell you anything went wrong, and they go quiet in a pattern — one site, one shift, one intake month. So the missing list is not housekeeping to get through before the analysis starts. It is the first result of the wave, and it arrives weeks before the numbers do.
Next: Connect numbers to comments