What is a longitudinal study?
A longitudinal study observes the same units repeatedly over time to investigate change, continuity or the sequence of events. The units may be people, households, organizations or sites. The study needs a way to connect observations appropriately and a plan for interpreting differences across time.
Following learners from entry through later employment is one example. Following member organizations through annual returns is another. The value comes from the research question and design, not merely from sending the same survey several times.
This guide explains when a longitudinal study is useful, the main design choices, a worked study plan and the practical work of follow-up, measurement and governance.
When does the extra effort make sense?
A longitudinal approach is useful when the question concerns a trajectory, transition, duration or sequence. You might ask how a person's confidence develops, when an organization begins using a service, whether a reported outcome persists or which experiences precede a later change.
If the question is simply how a population views a service this year, a cross-sectional survey may be sufficient. Repeated cross-sections can monitor population trends without following the same individuals. They answer a different question and may suit the collection arrangements better.
Longitudinal work adds responsibilities: maintaining contact where needed, protecting linked records, keeping measures interpretable and understanding missing observations. Do not promise a multi-year study without a realistic plan for those tasks.
A well-known example is Understanding Society, which returns to participants over time. The UK Data Service describes the study's repeated collection. Your own study may be much smaller, but still needs explicit design and documentation.
Choose the design and the unit
Scroll horizontally to see all columns →
| Choice | Meaning | Practical implication |
|---|---|---|
| Prospective collection | Plan the study and follow units into future periods | You can design collection for the question, but need time and retention resources |
| Retrospective longitudinal analysis | Use existing records to reconstruct observations over time | Check historical coverage, linkage, definitions and why records were created |
| Panel | Repeatedly observe the same selected units | Maintain the sample relationship and account for missing waves |
| Cohort | Study units sharing a defined characteristic or starting event | Specify cohort membership and whether the same units are followed |
| Repeated cross-section | Observe a population at several times, often using different samples | Interpret population trends rather than individual matched change |
These labels describe different aspects of a design and can overlap. A prospective cohort study can follow a panel of people. State the actual sampling and observation plan instead of relying on a label alone.
Choose the unit carefully. A school panel follows schools even if the children change. An organizational survey may follow the same members while their representatives change. Those are meaningful studies, but their results should not be described as change in the same individuals.
A worked plan: follow learning into practice
The following plan is fictional. A training team enrolls 100 learners and wants to understand how skill and later use develop. The primary question is: “What changes in demonstrated skill and reported use occur during and after the program?” A separate evaluation design would be needed to estimate the program's causal contribution.
Scroll horizontally to see all columns →
| Collection point | Evidence | Purpose |
|---|---|---|
| Before training | A suitable skill assessment and relevant role context | Establish a starting observation |
| At program end | A comparable assessment and participant account | Describe measured change and experience |
| Later follow-up | Reported use, opportunity to apply the skill and barriers | Investigate transfer into practice |
| Optional later review | A defined sustained-use measure where justified | Assess whether observed practice continues |
The team chooses follow-up timing based on when learners are likely to have an opportunity to use the skill. A fixed “90 days” is not automatically appropriate for every role or program. Record actual dates and explain the window used.
Each observation is linked through an appropriate study key. Contact details are handled according to the agreed process, and people are told what follow-up involves. The team defines who is eligible for each wave and how nonparticipation will be recorded.
Before launch, it pilots the instrument, invitation, matching and reporting. The pilot includes a corrected record, a missing assessment and a returning participant with changed contact details. This tests the process rather than just the survey's appearance.
Plan measures that remain interpretable
A shared data dictionary should define each measure, its wording, scale, timing and calculation. It should distinguish a score from a category, a missing answer from zero and a reported experience from independently observed behavior.
Use measures suited to the purpose. A satisfaction rating is not a direct test of skill, and a self-reported skill rating is not interchangeable with an assessed demonstration. Keeping the labels clear prevents a later report from overstating what was measured.
Changes may be necessary. If wording is confusing or the service changes, record the version and decide how the change affects comparison. Do not quietly alter the instrument mid-wave and treat all responses as equivalent.
In a network or multi-site program, local teams may need different questions. Agree a limited shared core for the intended comparisons, while preserving local fields. Identical forms are not required when definitions and the scope of aggregation are clear.
Follow up without confusing missingness with failure
Track who was expected to contribute, who was invited, who responded and whose data can be used for the intended analysis. A missing follow-up is not proof that a participant failed, disengaged or experienced a poor outcome.
In the fictional study, 100 learners have a baseline, 80 provide an exit observation and 65 answer the later follow-up. Sixty have all three observations. The team reports these different coverage counts rather than describing one completion percentage as the whole story.
Attrition can be selective. People who remain may differ from those who leave, but the direction and size of bias are not known simply from the response count. Compare available characteristics and known reasons where appropriate, and acknowledge what remains unknown.
Retention measures might include clear expectations, proportionate reminders, accessible collection and keeping authorized contact information current. Respect a person's choice not to continue. Do not use repeated reminders as a substitute for an understandable and worthwhile request.
Record whether a missed wave reflects ineligibility, a declined invitation, a failed contact attempt or an import problem when known. These distinctions support analysis and practical improvement without manufacturing explanations.
Prepare the analysis before the final wave
Inspect a few complete histories and verify that the observations belong to the right units and periods. Check duplicate unit-wave keys, corrections, out-of-window responses and changes in instruments or group composition.
Then define the population behind each result. An entry-to-exit comparison may use 80 matched records, while a three-wave analysis uses 60. Report the denominator for each finding rather than hiding those differences behind a single sample size.
For illustration, suppose the 80 matched entry-to-exit records show a mean score increasing from 5 to 7 on the same ten-point assessment. That is a two-point mean increase in that matched group. It does not describe the missing 20 learners or establish that training caused the change.
Repeated observations from one unit are related. Select methods that account for that structure and any relevant clustering. Incomplete-data methods involve assumptions; they do not guarantee an unbiased result. Use suitable statistical expertise for models beyond the team's competence.
See longitudinal data structures for long and wide tables, and longitudinal data analysis for the analytical questions.
Use qualitative follow-up for context
Open responses or interviews can help explain experiences, barriers and unexpected patterns. Decide how they will be sampled and analyzed. A few accounts chosen to investigate a result are not a prevalence estimate for the entire cohort.
Keep the original account, its period and relevant context. A comment can provide a participant's explanation without proving a causal mechanism. Where ratings and accounts disagree, investigate rather than forcing them into one consistent story.
AI can assist with organizing text and suggesting themes. Review the source, preserve uncertainty and record significant coding changes. The same codebook does not guarantee identical model output, and a quoted passage does not automatically justify the conclusion attached to it.
Govern the study across staff and system changes
A study may outlast its first coordinator or software setup. Keep a clear record of the question, instruments, sample rules, identifiers, permissions and analysis decisions. Make ownership and handover part of the plan.
A reliable key reduces avoidable matching work, but imports and registrations can still contain mistakes. Review uncertain links and retain correction history. Historical linkage can be valid when it is appropriate and checked; it is not automatically inferior because it occurred after collection.
Sopact's relevant role is connecting collection, analysis and governance around recurring records. Test how it handles later waves, changed definitions, missing observations and restricted sources. Confirm the workflow rather than relying on a promise that all follow-ups connect automatically or that attrition disappears.
Evaluate setup, review effort and maintenance. For reporting, use the impact report guide and report examples. Explain the study design and coverage alongside the headline result.
Watch: connected evidence across time
This companion video introduces repeated evidence and connected records. Use the study plan above to decide what observations and safeguards your own question needs.
Frequently asked questions
How long must a longitudinal study last?
There is no universal minimum duration that fits every question. The study needs repeated observations over a period appropriate to the change or sequence being investigated.
Does a longitudinal study prove cause and effect?
No. It can establish the timing of observations and describe change, but causal interpretation requires a suitable design and consideration of other explanations, selection and missing data.
Can I use existing records for a longitudinal study?
Yes, if they provide suitable repeated observations and appropriate linkage. Check why the records were created, their coverage, changing definitions and the accuracy of the connections.
Does attrition always make the program look better?
No. Selective loss can bias findings in different directions. Examine what is known about those retained and missing, and report uncertainty rather than assuming a direction.
Can I run a small study without specialist software?
Yes, if the team can manage records, access, versions and analysis reliably. The need for additional software depends on scale, complexity and recurring work, not a rule that spreadsheets always invalidate a study.
What should I plan before the first wave?
Define the question, unit, sample, measures, timing, linkage, participant communication, governance and intended analysis. Pilot the complete process, including missing and corrected records.

