This chapter resolves check 04 Longitudinal of the eight checks.
A training programme's annual report says its income gains are sustained over time. The sentence rests on one survey, sent nine months after that cohort finished. The survey ran three years ago. Nobody has measured those graduates since, and nobody has planned to. The claim is written in the present tense and the evidence stopped at month nine.
This is the most common overclaim in outcome reporting, and it is almost never deliberate. What happens is that a measurement loses its date. A finding travels from a spreadsheet into a slide into a grant application, and each time it moves it keeps the number and drops the horizon, until a nine-month reading becomes a claim about durability with nothing attached to it.
How do you measure how long an outcome lasts?
State the date you last measured, and the change between your final two measurements. Then stop. Everything after your last measurement is a forecast. It may be a reasonable forecast, and it still does not belong in a sentence that sounds like an observation. The honest version of the strongest claim most programmes can make is: we know it held at nine months; we do not know about year two.
That sentence is not a weaker claim. It is a claim with a boundary, which is the only kind a careful reader will accept — and it doubles as a budget line, because the way to extend it is to fund the eighteen-month survey.
The boundary, written out
Put your report's durability language next to what the data actually supports. The gap is usually a single missing clause.
What the report said, and what the data supports
Written in the report"Gains were sustained over time."Supported by the data"Of 84 graduates re-surveyed nine months after they finished, 61 were still in work. Between the six-month and nine-month surveys that figure fell by four."
Written in the report"Long-term impact."Supported by the data"Our longest measurement is nine months after exit. We have not measured beyond it."
Written in the report"Outcomes continue to hold."Supported by the data"They held as of March last year. Nothing has been measured since."
Illustrative example. Notice that the right-hand column is longer and more specific in every row, and that none of it is bad news. It is the same programme, described in a way a reader can check.
One follow-up cannot tell a plateau from a decline
Here is the part most measurement plans get wrong. A single survey after the programme ends gives you a level, and a level is not durability.
Suppose that nine-month reading is 61 people in work out of 84. That figure is equally consistent with two situations. In the first, the number was 61 at three months and 61 at nine — steady, and you can reasonably expect it to stay near there. In the second, it was 78 at three months and 61 at nine — falling hard, and by month eighteen it will be somewhere you would not want to publish. One reading cannot distinguish those. It is the same number in both.
So the rule is arithmetic rather than methodological: one point after exit gives you a level, two give you a direction, three tell you whether the decline is slowing or steady. Any claim with the word "sustained" in it needs at least the second. And the two points need real distance between them — two surveys a month apart give you a direction over one month, which is not what anybody means by durability.
How to do this without any particular software
- Record an exit date for every participant. Everything in this chapter is measured from that date, not from a calendar quarter. Without it, nothing below works.
- Choose your horizons before the programme ends. Three months, nine months, eighteen months after exit — pick two or three, and put the invitations in the calendar while the programme still has staff and attention. Durability you did not schedule is durability you will not have.
- Keep the question wording identical at every horizon. Not similar. Identical, word for word, including the response options.
- Store two things per person per follow-up: the value, and the exact number of days since their exit. Then "nine months" means nine months rather than an average of five to sixteen.
- Report the final two horizons side by side, with the number of people in each. The change between them is your durability finding. A single horizon reported alone is a level, and should be labelled as one.
- Write the claim as one sentence containing a horizon and a date, and put it exactly where the word "sustained" would have gone. "Held at nine months, last measured March 2026."
- Put the last-measured date on the chart itself, not in a footnote. This is the step that stops the finding from ageing invisibly once it starts travelling between documents.
Seven steps, all of them clerical. Nothing here requires a statistician; it requires that a date survive contact with a slide deck.
Where it breaks
Three failures, and the first is the one that produces the overclaim in the opening paragraph.
The date lives in the report, not in the data. A figure is computed once, written into a document, and copied forward. Next year's report inherits "gains sustained" without inheriting "measured nine months after a cohort that graduated in 2023". Nobody lied at any step; the claim simply outlived its evidence because the evidence and the date were stored in different places.
The horizons slip. The nine-month survey goes out when someone has capacity, so it reaches one person at five months and another at sixteen. The two readings you are comparing are now separated by an unknown amount of time, which means the change between them cannot be read as a rate of decline at all.
Somebody improves the question. Between the nine-month and eighteen-month waves, a colleague sharpens the wording or adds a response option. The eighteen-month figure comes back lower, and there is now no way to tell whether the outcome declined or the question did. This is the failure that is invisible in the output — the numbers look perfectly comparable.
What a system contributes here is narrow and mostly bookkeeping: each follow-up is timed from that person's own exit date rather than a shared calendar, the horizon and the last-measured date travel attached to the figure so a chart cannot be published without them, and the same question wording is reused at each horizon rather than retyped. It cannot decide how long you should measure for, and it cannot fund the eighteen-month survey. Those remain choices your organisation makes.
How to test this on your own reporting
Use: Your most recent published outcome claim that uses the words sustained, lasting, long-term or continues. Find the underlying measurement.
Pass: You can state, in under a minute and from the data rather than the document, the date of the last measurement, how many people it covered, how many months after their own exit it was taken, and the change from the previous horizon.
Fail: The claim rests on one post-programme reading, or the last-measured date can only be found by asking the person who wrote the report.
Most teams fail this on their first try, and the failure is recoverable in an afternoon: add the horizon and the date to the sentence and it becomes defensible without any new fieldwork.
Frequently asked questions
What is the minimum needed to claim an outcome lasted?
Two measurements after the programme ends, far enough apart that a change could show, using identical wording. One gives you a level at a point in time, which is worth reporting but is not durability. Two give you a direction, which is what the word "lasted" is actually asserting.
Can we use the word "sustained" at all?
Yes, with a horizon attached to it in the same sentence. "Sustained at nine months" is a fact. "Sustained" on its own is an implication about a period you never observed, and a careful funder will read it as one and discount everything near it.
How long after exit should the first follow-up be?
Long enough for the programme's immediate effect to have been tested by ordinary life, which for most training and support programmes is three to six months. Earlier than that and you are largely re-measuring the endline. The second follow-up matters more than the exact timing of the first.
What if the outcome improved after the programme ended?
Report it as what it is — an increase between two named horizons — and resist attributing all of it to the programme. A rise at month eighteen with nothing measured in between is compatible with the programme's effect emerging late and with something else entirely happening to that cohort.
Is a decline between follow-ups a failure?
Not by itself. Many outcomes are expected to soften once support stops, and a slow decline you measured is far more useful than a plateau you assumed. What makes a decline a failure is discovering it two years late, after the design that caused it has been repeated three times.
What if we can only afford one follow-up?
Then claim a level at a date and say plainly that you cannot speak to direction. "Sixty-one of 84 graduates were in work nine months after exit; this is our only post-programme measurement." That is an honest, publishable, useful finding. It is just not a durability finding, and calling it one is the failure this chapter exists to prevent.
Does a longer horizon always make a stronger claim?
No. A precise short claim beats a vague long one every time. Two well-timed readings at three and nine months, with identical wording and known dates, support more than a single loosely-timed survey somewhere in year three whose respondents were reached whenever they happened to answer.
What follows from all of this is slightly uncomfortable. Durability is not only a property of your programme. It is also a property of your measurement schedule, and the schedule is a choice you make while the programme is still running — not a finding you discover afterwards. An outcome you stopped measuring is not durable and not fading. It is unobserved, and the whole discipline here is refusing to let the third of those quietly get written up as the first.
Next: What the assistant may see