Record shape A of the four in Connected Data Intelligence — one person, followed for years. It leans hardest on checks 02 One record, 04 Longitudinal and 08 Reliable.
A patient association with more than eleven thousand members wanted to follow its community over time — three subtypes of a rare condition, plus carers and physicians, tracked properly for years rather than months.
The data already existed. An outside research firm ran the long-term study, on a fixed twelve-month window, and sold access to it. The association helped recruit for it. What the association could not do was extend the horizon past a year, add a question that mattered to them, or decide what got published — because it was not their dataset. As their research lead put it: they own the data.
That is the deepest thing about this shape, and it has nothing to do with software. Whoever owns the record decides how long the study can run. Ownership is a measurement decision dressed up as a legal footnote.
What has to be decided before the first survey goes out?
Almost everything expensive. The analysis work in a long-term study is well understood and there are good methods for all of it. What cannot be fixed later is the set of decisions made before wave one — ownership, permission to come back, and how you will recognise the same person in five years. Get those wrong and no amount of careful analysis later recovers the study.
This is why long-term studies rarely fail in year three. They fail in month one, invisibly, and nobody finds out until the second round cannot be joined to the first.
| Decide before wave one | Why it cannot wait |
|---|
| Who owns the data | Ownership sets your time horizon and your right to ask a new question. If it is someone else's dataset, their window is your window. |
| Permission to come back | Re-contact consent has to be collected up front, and it has to be worded to cover a round you have not designed yet. |
| How you recognise the person | A code you issue, not one the participant invents. This is the single most common point of failure. |
| Cadence and duration | Every six months or every year, and for how long. This determines what counts as a missing round later. |
| Which questions are fixed | The handful that must stay identical to make comparison possible, separated from the ones you expect to change. |
Never ask a participant to invent their own code
The most attractive-looking shortcut in this shape is also the one that reliably destroys it. To preserve anonymity, ask each person to build their own identifier from things only they know — first letters of a parent's name, a state abbreviation, a birthday.
The same patient association tried it across several surveys. People forgot the code they had invented, typed something different next time, and those answers could never be matched to their earlier ones. The anonymity worked perfectly. The study did not, which is the worst possible trade, because the code existed to make the follow-up possible in the first place.
The fix is not to abandon anonymity. It is to move the remembering from the participant to the system: a code you issue at first contact and carry in the invitation, so the participant is asked to remember nothing at all. Their answers can still be de-identified for analysis. The full version of this — including how a code survives changing devices, staff and addresses — is in collecting evidence offline and in the field.
Permission to come back is the thing that makes wave two possible
A long-term study is a series of returns, so the whole design rests on being allowed to return. Three things about that permission matter, and all three are decided at the start.
It has to be collected up front, because you cannot retrospectively acquire the right to contact someone. It has to be broad enough to cover a round you have not designed yet, without being so vague that it is not real consent — describing the purpose and the likely cadence usually threads that needle. And it has to be genuinely revocable: if someone declines to be contacted again, your system has to actually stop linking them, not merely stop emailing them.
That last point is where good intentions and real systems diverge. One global organisation running long-term follow-up was explicit that if a respondent said no, they must not be linked on return — and that this is exactly the kind of thing that has to be built rather than promised. A withdrawal that lives in a separate spreadsheet gets forgotten. A withdrawal attached to the record cannot be.
Expect the questions to change — plan for it rather than resisting it
Teams running long studies tend to fall into one of two failures. Either they freeze the questionnaire to protect comparability, and by year four are diligently collecting answers to questions that stopped mattering in year two. Or they revise freely and discover the trend line moved because the question moved.
Neither is necessary. The workable arrangement is to separate a small fixed spine — the handful of measures that must stay identical for comparison to mean anything — from everything else, which you should feel free to change as new issues emerge. The patient association wanted exactly this: a baseline with recurring questions, plus the ability to add new ones as new issues came up, without starting over.
The mechanics of changing a question without breaking the record are their own topic, covered in changing questions without breaking the record.
People will drop out, and that is data
Over years, some participants stop answering. The instinct is to treat them as a shrinking sample. That is dangerous, because the people who stop are rarely a random subset — they are frequently the ones whose circumstances got harder, which means a study that quietly loses them reports steadily improving outcomes for a steadily more comfortable group.
So who is missing has to be tracked as a named list and reported alongside every result. The methods for this — comparing the people who stayed against the people who left, on what you knew about them at baseline — are in who is missing survey waves.
Where to go for the analysis itself
Once the set-up above is right, the analysis work is well trodden and has its own chapters:
Doing this without any particular software
- Settle ownership in writing before any data is collected — who holds it, who may publish from it, and what happens if the relationship ends.
- Write the consent wording to cover re-contact for the cadence and duration you intend, and record each person's answer as a field on their record rather than in a separate log.
- Issue every participant a code at first contact and keep the code-to-person list in one access-controlled place, separate from their answers.
- Mark your fixed spine. Take your question list and flag the handful that must never change wording. Everything unflagged is free to evolve.
- Keep one row per person per round, with a round date — not one file per round. This single choice decides whether comparison is possible later.
- After every round, list who did not answer, by name, and note what you knew about them at baseline.
For one cohort over two or three rounds, that is entirely workable by hand, and doing it once is the best way to understand what matters.
Where it breaks
It breaks on time, in a way the other three shapes do not. Five years is longer than most staff tenures, most software contracts and most funding cycles. The person who designed the study leaves; the new person inherits a spreadsheet whose column names are not self-explanatory and a consent arrangement nobody can locate. The study does not collapse dramatically — it just gets quietly re-invented, and the new version is not comparable to the old one.
It also breaks on the question you did not ask. Three years in, something matters that nobody anticipated, and you can only ask it going forward. That cost is unavoidable, but it is much smaller if a fixed spine was marked at the start, because adding a question is then routine rather than a threat to the whole series.
What a system is for here is narrow: issuing and holding the identity so it survives staff turnover, keeping consent and withdrawal attached to the record so they are honoured rather than remembered, letting new questions extend the series instead of restarting it, and making a round-to-round comparison something you look up rather than rebuild. Your team still decides the cadence, the spine and what the findings mean.
How to test this before you commit
Use: Two rounds six months apart with the same twenty people. Between rounds, add two new questions, reword one existing question, and have one person withdraw consent to be contacted again. Have three people skip round two entirely.
Pass: Round two joins round one with nobody matching names by hand. The withdrawn person is not contacted and not linked. The two new questions show as not-asked in round one rather than as blanks. The reworded question is flagged so nobody reads its change as real change. The three non-responders appear as a named list with their baseline values.
Fail: Anyone has to reconcile two rounds by name, date of birth or email address.
Frequently asked questions
Should participants create their own anonymous code?
No. It sounds like it protects privacy at no cost, but people forget the code and type a different one next time, so their rounds cannot be matched. Issue the code yourself and carry it in the invitation — the participant then has to remember nothing, and the data can still be de-identified for analysis.
Why does data ownership matter for a long-term study?
Because it sets your time horizon. If an external firm owns the dataset, their collection window and their publication rules are yours too, and you cannot add a question that matters to you. Many organisations discover this only when they want to extend a study and find they cannot.
Can we add questions later without ruining comparability?
Yes, if you separated a small fixed spine at the start. New questions extend the series and show as not-asked for earlier rounds. What breaks comparability is quietly rewording one of the spine questions, which is a different action and needs to be visible in the output.
What if someone withdraws consent?
They must stop being contacted and stop being linked, which means the withdrawal has to live on their record rather than in a separate list somebody has to remember to check. Treat it as a field, not a note.
How do we handle people who stop responding?
Track them as a named list and report them alongside results. They are usually not a random subset — often the people whose circumstances worsened — so a study that loses them silently will show improvement that partly reflects who left.
How long should we plan to run it?
Decide a horizon and a cadence before wave one, because both determine what counts as a missing round. Changing cadence mid-study is possible but has to be disclosed, since a gap of twelve months and a gap of six are not the same measurement.
Is this only for research teams?
No. Any programme that wants to say whether change lasted is doing this shape, whether or not it calls itself research. The decisions are the same; the vocabulary is just less formal.
The strange thing about long-term studies is how little of the difficulty is in the analysis. The statistics are solvable and well documented. What actually determines whether you have a study in year five is a handful of unglamorous decisions taken before anyone collected anything — who owns it, whether you may come back, and whether you will still be able to tell who is who. None of those feel like measurement decisions at the time. All of them are.
Next: Back to the eight checks