What is survey attrition in a longitudinal study?
Survey attrition is the loss of participants from later rounds of a study that follows the same people over time. It can reduce precision and distort findings when the people with missing outcomes differ in ways relevant to the question. A rising average among the remaining respondents does not, by itself, show improvement for the original cohort.
Use this reference when you need to understand who is missing from repeated collection. Build a wave-status register, compare available baseline characteristics and write a result that states who it describes. The aim is to notice missing evidence before turning it into an outcome claim.
In a fictional workforce program, confidence is measured at intake, week eight and completion. The average rises, but fewer people answer each time. The team needs to separate change among the same respondents from a change in which respondents appear in the chart.
Distinguish a missed wave from leaving the program
A participant can miss a survey and remain in the program. Someone can leave the program and still complete a follow-up. A person who misses week eight may answer at completion. Keep program participation, survey participation and individual missing answers as separate fields.
The Agency for Healthcare Research and Quality’s guidance on missing data distinguishes several reasons for missing observations, including circumstances related and unrelated to the study outcome. Although written for registries, that distinction is useful here: a missing answer does not identify its own cause.
| Status | Meaning | What to check |
|---|---|---|
| Wave not yet due | The scheduled follow-up has not arrived | Enrollment date and the allowed response window |
| Wave due, no response | No response recorded by the cutoff | Delivery, identity matching, contact permission and any known reason |
| Survey returned, score missing | Item nonresponse | Routing, optional question and usable analysis pair |
| Missed one wave, returned later | Intermittent nonresponse | Retain the later response; do not label permanent dropout |
| Withdrawn or no longer eligible | A documented change in study status | Protocol, recontact limits and the defined denominator |
Record “reason unknown” when that is the truth. Avoid labels such as “disengaged,” “unsuccessful” or “high risk” inferred solely from a missing survey. Administrative errors and changing contact details can also create apparent attrition.
Practice: a rising average with no observed change
Use this fictional example on a 0–10 confidence scale. At baseline, 100 people respond. Sixty later provide usable endline scores and 40 do not. For this exercise, assume each of the 60 respondents has exactly the same score at both waves.
| Group | People | Baseline mean | Endline mean |
|---|---|---|---|
| All baseline respondents | 100 | 5.0 | Unknown for the full cohort |
| Usable scores at both waves | 60 | 6.0 | 6.0 |
| No usable endline score | 40 | 3.5 | Missing |
The baseline calculation is (60 × 6.0 + 40 × 3.5) / 100 = 5.0. Comparing 5.0 at baseline with 6.0 at endline appears to show a one-point increase. But the same 60 respondents average 6.0 at both times: their observed mean change is zero.
The full cohort’s endline mean remains unknown. The missing 40 may have improved, declined or stayed the same. Their lower baseline score is a warning about selection, not a measurement of their later outcome.
In a real dataset, an unchanged group mean does not mean every individual stayed unchanged. Some may improve while others decline. The exercise explicitly assumes unchanged individual scores; do not infer that assumption from a flat average in your own report.
Report matched change and coverage together
Calculate each person’s change only where the relevant two scores are available and comparable, then summarize those changes. State the number of usable pairs and which waves they connect. A returned survey is not necessarily a usable pair if the outcome question is blank.
For the exercise, a defensible statement is: “Sixty of 100 baseline respondents supplied a usable endline score. Their mean confidence score was 6.0 at both waves. The 40 without an endline score averaged 3.5 at baseline; their subsequent outcomes are unknown.”
This is a description of paired respondents, not an automatic estimate for everyone enrolled. Matching prevents the simple comparison of different people, but does not remove selection bias or establish that the program caused a change.
AAPOR’s guidance emphasizes that response rates alone do not establish the amount—or presence—of nonresponse error. Report rates and case dispositions, then investigate what is missing. Use a longitudinal definition suited to your design; do not apply an unrelated response-rate calculator without checking its scope.
For a fixed cohort, define retention for a specified wave as the number meeting your response criterion divided by the eligible baseline cohort, with exclusions explained. In the exercise, 60/100 = 60% have usable paired scores and 40% do not. That is usable-outcome coverage, not necessarily the survey response rate: some of the 40 may have returned a survey without that score. Name the criterion before labeling a figure “attrition.”
Build a wave-status register before calculating rates
Start with the eligible cohort roster, not just the files containing responses. Otherwise a person absent from every later export disappears from the analysis. Define the due date, response window, cutoff and eligibility rules for each wave.
A wide view with one row per person and columns for waves makes gaps easy to inspect. A long table with one row per person-wave can work equally well if it includes expected waves, not only received responses. The layout is a tool; completeness of the expected-record list is the requirement.
- Identify the cohort. Keep a stable, permitted participant identifier and enrollment date. Resolve uncertain matches before classifying nonresponse.
- Generate expected waves. Distinguish not-yet-due, eligible-and-due and documented exclusions.
- Record response status. Keep received, partial, declined, undeliverable and unknown states separate where your process supports them.
- Identify usable outcome pairs. Check the measure, version, dates and missing items.
- Compare baseline characteristics. Review responders and nonresponders by relevant observed measures.
- Document follow-up. Record permitted attempts, outcomes and unresolved identity or delivery issues.
- Freeze a reporting snapshot. State the cutoff so late responses can be handled transparently in the next version.
Keep both cumulative retention from baseline and wave-specific participation where useful. Someone returning after a missed wave should remain visible. A changing eligibility rule should not silently make retention look better.
Compare responders and nonresponders without overclaiming
Compare available baseline outcomes, site, cohort, language, referral route and other relevant characteristics collected appropriately. Look at counts as well as averages; a small subgroup may be hidden in an overall result.
Differences can identify where further investigation is needed. Similar observed baseline averages do not prove the missing outcomes are comparable. The groups can differ on unmeasured characteristics or later events. A nonsignificant test also does not prove equivalence, particularly with small samples.
The baseline gap is not a confidence interval or the numerical size of attrition bias. It is a diagnostic. Likewise, the absence of a known reason does not prove missingness was random. Be precise about what the available records can and cannot show.
If baseline outcome data are incomplete, compare other available characteristics and disclose the gap. That still offers useful information, but it cannot recreate an unobserved baseline or establish how people changed.
When do you need statistical adjustment?
If you need estimates beyond the observed pairs, involve someone with the relevant methodological expertise. Complete-case analysis, weighting, multiple imputation and longitudinal models make different assumptions. None is a universal repair for missing outcomes.
A longitudinal simulation study of attrition examined how missing-data scenarios and analytical methods affected bias. The practical implication is to choose methods for the study’s missingness assumptions and target estimate, rather than assume deleting incomplete records is harmless.
Do not fill missing scores with zero, the respondent’s baseline score or the observed mean merely to complete a dashboard. Such replacements impose assumptions and can hide uncertainty. If an analytical method estimates missing values, preserve the distinction between observed and estimated data and document the method.
A simple sensitivity exercise can make the uncertainty visible without pretending to solve it. In the fictional cohort, if the missing 40 averaged 3.5 at endline, the full-cohort mean would be 5.0. If they averaged 6.0, it would be 6.0. These are scenarios, not estimates or confidence limits. Different plausible assumptions may change the conclusion.
Follow up respectfully and investigate delivery problems
A permitted contact list can turn missing-response checks into useful work. It might include participant ID, due wave, last attempt, preferred channel and assigned owner. Limit named records to staff who need them; an anonymous survey cannot be turned into a named follow-up list without changing its design and promises.
Make reminders understandable, accessible and consistent with consent and withdrawal preferences. Check whether timing, language, channel or survey burden creates avoidable barriers. Keep the invitation neutral; do not pressure someone to provide a positive account.
Targeted follow-up can improve coverage, but does not guarantee unbiased results. Record changes in collection mode or timing, and avoid pursuing only easy-to-reach people. A recovered response is valuable evidence, not automatically worth a fixed multiple of another person’s response.
When missingness clusters at one site, investigate delivery and operational context before assigning blame. A failed invitation batch and a service problem require different actions. The response register should help the team ask better questions, not supply an unsupported diagnosis.
What should a connected workflow make easier?
Repeated exports can lose the relationship between a person, a scheduled wave and an actual response. A changed email address may create a duplicate; an offline upload may arrive after the reporting cutoff. The offline collection lesson explains how to preserve identifiers and distinguish transmission issues from new observations.
In a configured Sopact workflow, test whether your roster, scheduled periods and incoming responses produce a clear status view without manually rebuilding the same spreadsheet. Ask to see how it handles late responses, partial surveys, uncertain identity matches and a participant who returns after missing a wave.
The outcome report should show its reporting base and source period. The system should support the team’s definitions and review; it cannot infer a missing person’s experience from silence. Keep these requirements in the implementation discussion rather than treating a dashboard as proof they have been met.
Test the register with six deliberate exceptions
Create a small test cohort containing someone not yet due, someone who declined, an undeliverable invitation, a returned survey with a blank score, a late offline response and someone who missed the middle wave but returned at the end. Use fictional records for setup, then validate the handling against authorized real records.
Pass when each has the correct status, usable pairs reconcile to the analysis, and no missing value is silently replaced by a score. Check that an update changes the appropriate reporting snapshot and that restricted follow-up details do not appear in public outputs.
Finish with three outputs: a status table by wave, a paired-result statement with coverage, and a short note on baseline differences and unresolved limitations. Those are more useful than a single upward trend line with no denominator.
Watch: keeping qualitative evidence connected
Watch the video · 2 minutes 34 seconds. This companion explains the broader connected-evidence approach. Use the exercise above for the attrition calculation. Browse more videos in the video library.
Frequently asked questions
How much attrition is too much?
No single percentage establishes bias. Consider the amount, reasons, group differences, outcome and intended inference. Report the rate and investigate the pattern rather than treating either as sufficient alone.
Does matching baseline and endline remove attrition bias?
No. It describes change among people with usable pairs. Those people may not represent the original cohort, and matching does not establish causation.
Are survey nonresponders program dropouts?
Not necessarily. Keep program status and survey status separate. A person can miss a wave while remaining active or respond after leaving.
Can I report only complete cases?
You can report that subgroup with a clear definition and limitations. Whether it supports broader inference depends on the design and missingness assumptions.
Should missing scores be replaced with zero?
Not simply to complete a table. Zero is an observed value only when the measure and evidence support it. Statistical treatment of missing data requires explicit assumptions.
What if someone returns after a missed wave?
Keep the later response and identify which comparisons have usable pairs. Do not automatically discard the person as a permanent dropout.
Does a similar baseline prove nonresponse is harmless?
No. Similar observed characteristics are reassuring on those measures, but cannot rule out unobserved differences or later events related to the outcome.
Use comments to investigate the observed pattern
Carry the wave-status register and reporting bases into connecting quantitative and qualitative survey data. Comments can help explain observed changes, while the missingness record keeps you clear about whose experience you have not heard.
Reviewed September 12, 2026. All cohort figures and program scenarios are fictional. They illustrate reporting choices, not a customer outcome.