play icon for videos

Longitudinal Data: Definition, Examples and Long vs Wide Format

Understand longitudinal data with worked long and wide tables, missing-wave examples, record keys, comparability checks and practical analysis limits.

Updated
September 15, 2026
360 feedback training evaluation
Use Case
Case intelligence · Practical guide

Longitudinal Data: Definition, Examples and Long vs Wide Format

Understand longitudinal data with worked long and wide tables, missing-wave examples, record keys, comparability checks and practical analysis limits.

Read the guide ↓

What is longitudinal data?

Longitudinal data records observations of the same units over time. The units may be people, households, organizations, sites or other entities. Connecting observations to the correct unit makes it possible to examine change within that unit as well as patterns across the group.

A learner's assessments at entry and exit are one example. A member organization's annual returns are another. The data needs usable identifiers, time information and comparable measures, but having those fields does not automatically make the analysis complete or unbiased.

This guide explains the structure with small worked tables, the difference between long and wide formats, missing-wave problems and the checks needed before analysis. For designing the collection, see longitudinal studies.

Examples beyond individual participants

Scroll horizontally to see all columns →

UnitRepeated observationPossible question
LearnerAssessment and later practice checkHow does measured skill or reported use change?
Member organizationAnnual activity and service-use returnHow do participation and needs change across years?
Customer accountCheck-ins during onboarding and renewalHow does the account's reported experience develop?
Supplier siteRecurring delivery and evidence reviewsWhich operational measures improve or deteriorate?
School or program siteTermly measures under shared definitionsHow do site-level patterns develop?

Keep the level of the claim aligned with the unit. Tracking the same schools does not mean tracking the same children. Tracking a company does not mean the same representative answered each return. Changes in respondents or composition may affect interpretation.

A real example of repeated individual collection is Understanding Society, the UK Household Longitudinal Study. The UK Data Service describes its return to participants across years. The study's documentation is also a reminder that a longitudinal dataset needs more than a collection schedule.

Longitudinal data and repeated cross-sections

A cross-sectional dataset describes units at a particular time. Repeated cross-sectional surveys collect from a population at several times, often with different samples. They can show population-level trends when their designs support comparison, even though they cannot directly show each individual's change.

A panel follows the same units. If a workplace surveys different anonymous employee samples each year, a change in the average may reflect both changing experiences and changing respondents. If the same employees are appropriately linked across waves, a matched analysis can examine their individual changes.

Neither design is inherently superior. Repeated cross-sections may suit a population-monitoring question or an anonymity requirement. A panel may suit a question about trajectories, but creates retention, linkage and governance work. Select the structure for the question.

Long format: one row per unit and observation

In long format, repeated observations occupy separate rows. The example below is fictional. Two learners have scores on the same ten-point assessment at entry and follow-up. A third learner has an entry score but no follow-up observation.

Scroll horizontally to see all columns →

Learner keyWaveScoreObservation status
P-01Entry4Observed
P-01Follow-up7Observed
P-02Entry6Observed
P-02Follow-up8Observed
P-03Entry5Observed
P-03Follow-upNo follow-up observation

The explicit missing row is one way to document the expected wave; some systems instead retain expected-wave status in a separate table. Either way, make the missing observation visible. Do not store it as a score of zero.

Long format is useful for adding waves and working with uneven observation schedules. Include actual dates where timing matters. A label such as “follow-up” may hide a difference between two weeks and six months after an event.

Wide format: one row per unit

The same fictional data can be arranged with one row per learner and separate columns for each wave.

Scroll horizontally to see all columns →

Learner keyEntry scoreFollow-up scoreMatched change
P-0147+3
P-0268+2
P-035Not calculable from the observed pair

Wide format makes a simple pairwise calculation easy to inspect. For the two matched learners, the mean increase is (3 + 2) ÷ 2 = 2.5 points. That describes those two learners, not all three.

Changing format does not create missing information. Reshaping also requires checking the key: if the same learner has two follow-up records, you must understand whether they are corrections, repeated tests or different events before choosing a value.

The UK Data Service's data-skills glossary describes long and wide arrangements. Which you use depends on the analysis and software; neither format removes the need for sound definitions and quality checks.

Define the record key and the time key

A stable key should identify the intended unit without relying on a name that may be misspelled or changed. For a participant study, separate personal contact details from analytical data where appropriate. A pseudonymous identifier still requires responsible handling if it can be linked back to a person.

The key for an observation may include more than unit and wave. A learner might have several assessments, programs or attempts. A partner might submit separate site returns in the same quarter. Define the combination that makes a record unique.

A persistent identifier helps prevent avoidable matching work, but does not guarantee a correct match. Imports can contain wrong keys, duplicate registrations and mistaken assignments. Retain a review process for exceptions and corrections.

Historical linkage can also be legitimate when performed with suitable methods and verification. Data does not become invalid merely because a join occurs after collection; the issue is whether the connection is accurate, appropriate and documented.

Missing waves and uneven panels

A balanced panel has observations for the same units at the intended time points. An unbalanced panel has missing or differing observations. Unbalanced data can still be longitudinal, but the analysis must account for its structure and missingness.

In a fictional cohort of 100 people, 75 answer the next wave and 60 answer a later wave. These counts describe response coverage. They do not tell you whether the missing participants improved, deteriorated, left the program or simply did not answer.

Distinguish a person who was not eligible for a wave, one who declined, one who could not be reached and one whose record failed to import where that information is known. Do not invent reasons to fill the status column.

Analyzing only complete cases can change the group being described. More advanced methods may use incomplete observations under assumptions, but they do not make those assumptions disappear. Choose the approach with appropriate statistical expertise and explain its limits.

Keep measures comparable as the work changes

Track question wording, scales, calculation rules, collection modes and measurement dates. A change from a five-point to a ten-point response scale is not solved simply by doubling earlier scores. The response process and interpretation may have changed.

Some changes are necessary to improve a weak question or reflect a changed service. Record the change and decide whether to retain a separate series, bridge the measures through a justified design or report a break in comparison.

Federated organizations can maintain a limited shared core while local teams ask different questions. The shared dictionary defines which fields support comparison. It should not turn a local attendance count into a network count of unique people.

Also track changes in the unit itself. Organizations merge, households change membership and sites move. Decide how those events affect the identity and continuity of the record rather than treating every matching label as the same unchanged unit.

What longitudinal data can and cannot show

It can describe individual or organizational trajectories, transitions and the timing of observed events. It can support analysis of within-unit change and associations over time. These are valuable questions that a single snapshot cannot answer in the same way.

It does not automatically establish why change occurred. Other experiences, selection, measurement changes and missing data may influence the result. Observing a change after a program is not sufficient to attribute that change to the program.

Repeated observations from one unit are also related, rather than independent new participants. An analysis that treats every row as an unrelated person can give misleading uncertainty estimates. For detailed methods, continue with longitudinal data analysis.

Qualitative material can add context across waves, but it should be interpreted carefully. A later comment may explain how a participant sees their experience; it is not automatically a verified cause of a numerical change.

A readiness check before analysis

  1. Define the unit and the unique observation key.
  2. Inspect duplicate keys, corrections and unmatched records.
  3. Retain wave labels and actual timing where relevant.
  4. Check measure definitions, versions and collection conditions.
  5. Report expected and observed coverage by wave.
  6. Separate missing values from valid zeros and inapplicable items.
  7. Verify a few individual histories against their sources.
  8. Choose an analysis that fits the repeated structure and missing-data assumptions.

Sopact's relevant role is connecting recurring collection, analysis and governance so a team can maintain these records with their context. Test a real repeat-cycle workflow, including exceptions and changed definitions. A connected record supports the analysis; it does not replace the study design or guarantee that every observation is correct.

For collection planning, see longitudinal surveys. To present the findings clearly, use the impact report guide and report examples.

Watch: keeping observations connected

This companion video discusses connected evidence across time and methods. Apply the readiness checks above to the actual unit and measures in your study.

Frequently asked questions

What is an example of longitudinal data?

Repeated assessments from the same learners, annual returns from the same member organizations or recurring observations of the same supplier sites are examples. The observations must remain connected to the correct units and times.

Does longitudinal data have to be about people?

No. The units may be organizations, households, sites or other entities. Keep the interpretation at the same level as the unit being followed.

What is the difference between long and wide data?

Long format stores repeated observations in separate rows. Wide format stores one row per unit with separate columns for repeated measures. Reshaping requires a suitable unique key and explicit handling of duplicates.

Can longitudinal data have missing waves?

Yes. An unbalanced panel can still be analyzed, but missingness, eligibility and changing coverage must be considered. Missing observations should not be treated as zeros.

Are repeated anonymous surveys useless for change analysis?

No. They may support population-level trend analysis as repeated cross-sections when the designs are comparable. Without appropriate linkage, they do not directly establish each individual's change.

Does following the same people prove a program worked?

No. It can show observed change, but causal attribution requires a suitable design and consideration of other explanations, selection and missing data.