How do you analyze pre, mid and post survey data?
Match each person's responses across comparable survey waves, calculate change for the relevant pairs, and report the distribution alongside any group summary. State who is included, who is missing and how you defined change. Use midpoint results to investigate emerging needs; use endline results to describe progress. A before-and-after difference alone does not establish that the program caused it.
Use this reference when your plan has baseline, midpoint and endline observations. Build a small matched dataset, compare two valid analysis sets and write a finding another person can reproduce. For the instrument and collection schedule, use the pre-and-post survey guide.
What do pre, mid and post mean?
Pre is the baseline before the activity or period being evaluated. Mid is an agreed checkpoint during it. Post is the endline after that period. Record actual dates as well as labels: a response collected before training and one collected after four sessions are not equivalent baselines.
A training team might measure job-search confidence at intake, week eight and completion. A customer onboarding team might use activation, day 30 and day 90. These are possible schedules, not universal standards. Choose timing around when change is plausible and when the team can use the findings.
Define the question first: “How did confidence change between intake and completion among people with both measurements?” is different from “When did confidence change across all three checkpoints?” The first needs a baseline/endline pair. The second needs information about the midpoint, or a suitable method for incomplete repeated measurements.
The CDC evaluation design guidance connects the evaluation question, intended use and design. Its distinction between observing outcomes and establishing causes is especially relevant to an uncontrolled pre/post survey.
Prepare comparable records before calculating change
Keep an identifier that consistently represents the same person, plus the program or cohort, wave, actual response date, instrument version and score. Store missingness separately from the score. A blank cell is not zero, and a skipped question is not necessarily a missed survey.
Check duplicate submissions before selecting a response. For example, use the latest valid response inside the agreed window only if that is your documented rule. Preserve earlier submissions and the reason for the selection. Do not choose whichever response creates the largest improvement.
Preserve the meaning, response scale and scoring direction of the core measures being compared. Local forms may include different additional questions. A crosswalk between different measures needs supporting evidence; simply giving them the same field name is insufficient. If an instrument changed, document whether the measures remain comparable. The lesson on changing survey questions explains how to retain versions without presenting a broken series as continuous.
A spreadsheet can place one person on each row and one wave in each column. A database may store one person-wave per row and produce the same analysis view. The shape of the storage is less important than preserving the relationship, date and scoring rule.
Worked example: six people, three checkpoints
This fictional exercise uses a single confidence item scored from 1 to 5, with higher values indicating greater confidence. It is not a validated scale or a customer result. Person F missed the midpoint but answered at baseline and endline.
| Person | Pre | Mid | Post | Post minus pre |
|---|---|---|---|---|
| A | 2 | 3 | 4 | +2 |
| B | 3 | 3 | 3 | 0 |
| C | 4 | 3 | 2 | −2 |
| D | 1 | 2 | 4 | +3 |
| E | 5 | 4 | 4 | −1 |
| F | 2 | Not collected | 3 | +1 |
For all six baseline/endline pairs, the baseline total is 17 and the endline total is 20. Their means are 2.83 and 3.33; the mean paired change is +0.50 points. The individual changes sum to three, and three divided by six is 0.50.
Those means are descriptive summaries of numeric category codes. On a single ordered response item, equal distances between categories are not guaranteed. Show category counts and movement as well; for a multi-item instrument, follow its scoring guidance rather than averaging items arbitrarily.
Using a one-point rule solely to practice classification, three people score higher, one stays the same and two score lower. That describes observed movement. It does not establish meaningful improvement, measurement reliability or program harm.
Choose the analysis set that answers the question
Do not automatically discard F from the baseline/endline comparison. Both required observations exist. If you want to describe complete three-wave paths, A through E form a separate set of five people. Keep those two sets visibly distinct.
| Question and set | Included | Result | Limit |
|---|---|---|---|
| Baseline to endline | A–F, six paired records | Mean 2.83 to 3.33; change +0.50 | Does not describe F's midpoint |
| Complete three-wave paths | A–E, five complete records | Means 3.00, 3.00, 3.40; pre/post change +0.40 | Excludes F; may differ from people with incomplete records |
The results differ because the included people differ. Neither calculation is automatically wrong. The error would be putting the six-person baseline mean beside the five-person midpoint mean and calling the line a within-person trajectory for one unchanged group.
In a larger dataset, report the number expected at each wave, usable responses, matched pairs and complete records. Describe known reasons for missingness without guessing them. Use the survey attrition lesson to examine whether missing follow-up could change the interpretation.
Separate observed change from meaningful change
A numerical difference, a statistically supported difference and a change that matters to participants are different judgments. There is no universal one-point threshold that resolves all three on every five-point scale.
Look for evidence about the instrument's reliability, measurement error, responsiveness and interpretation in a relevant population. COSMIN's measurement framework addresses these properties for health outcome instruments. It is a useful methodological reference, not a certification of the confidence question in this exercise.
If no defensible meaningful-change threshold exists, report the observed score movements and say that their practical significance is uncertain. You can adopt a provisional operational trigger for checking in, but label it as such and review whether it works. Do not describe it as a validated outcome threshold.
Record planned rules before examining the result when possible. If you explore alternative thresholds afterward, show that they are sensitivity analyses. A dated rule prevents silent redefinition; it does not, by itself, remove measurement noise.
Use the midpoint to decide what to investigate
In the fictional data, C's confidence falls at each checkpoint. E falls and then holds steady. A and D rise. These paths suggest different follow-up questions, even though the five-person mean is unchanged between baseline and midpoint.
Ask whether the question was understood consistently, whether training exposure differed and whether comments describe access barriers or changing expectations. A decline can reflect a real difficulty, measurement variation or a more realistic self-assessment after learning what a task involves. The score alone cannot distinguish them.
Midpoint collection can support a program adjustment, an aggregate learning question or an agreed individual follow-up. It does not always require identifying respondents. In an anonymous survey, examine group patterns without trying to reconstruct identities. Where named follow-up is appropriate, keep access and responsibilities clear.
Specify who reviews results, when they review them and what decision the checkpoint can inform. Endline evidence remains actionable too: it can shape the next cohort, continuing support or a later follow-up. Collecting another wave should have a purpose proportionate to the burden.
Keep the average, distribution and context together
A mean can summarize direction and magnitude, but it does not show how change is distributed. Pair it with counts or a suitable distribution chart, the analysis-set size and relevant uncertainty. The right headline depends on the question, not a rule that averages must always come last.
For this exercise, a transparent statement is: “Among six people with baseline and endline responses, the mean coded confidence score rose by 0.50 points. Three scored higher, one was unchanged and two scored lower. One person lacked a midpoint response. This small fictional example does not establish meaningful or causal change.”
For an actual cohort, add the collection dates, scale range, direction, threshold if used, missing-data handling and instrument version. Avoid percentage change on a scale without a meaningful zero: moving from two to four on a confidence item is not necessarily “100% more confident.”
Link comments to the relevant period before using them to interpret a path. A post-program comment can describe an earlier event, so distinguish the response date from the event date when known. Review contradictory accounts rather than selecting only quotes that support improvement.
Compare sites carefully and choose statistical methods deliberately
Site or cohort comparisons can reveal differences worth investigating. Show the reporting base for each group and check baseline differences, exposure, timing and missingness. A lower endline average does not automatically mean poorer delivery if the starting circumstances differ.
Small groups are not invisible, but their results can be unstable and identifiable. Use counts with context, apply disclosure rules and avoid league tables built from a handful of people. There is no single minimum group size appropriate for every analysis and privacy setting.
If your question requires statistical inference, choose a method suited to the measure, dependence between repeated observations and missingness. Do not treat repeated responses from the same person as independent samples. A paired analysis or repeated-measures model may be appropriate under its assumptions; the six-person exercise is a reconciliation exercise, not a basis for a significance claim.
Even a statistically supported pre/post difference does not rule out other causes such as external events, selection or changes in who responded. Seek evaluation expertise when the decision requires an attribution claim or a complex model.
Build and test the workflow once
- Define the question. Name the outcome, timing, included population and intended decision.
- Reconcile records. Check identifiers, duplicate rules, score direction, versions and missing waves.
- Calculate independently. Reproduce paired changes and summary counts on a small test set.
- Review the interpretation. Separate movement, meaningful change and causality; examine comments and alternative explanations.
- Retain the specification. Save the rule version, cutoff date, included records and reviewer decision with the report.
In a configured Sopact workflow, test whether incoming surveys, uploaded evidence and later check-ins stay connected to the appropriate contact and period. Ask for a demonstration of the matched analysis view, reviewable AI output and permissions. Software can make repeated work more consistent; it cannot choose a valid measurement rule for you.
Use the six records above as an acceptance exercise. A correct baseline/endline report includes F, shows six pairs and produces +0.50. A complete-three-wave report uses five people and produces +0.40 for baseline to endline. Neither should replace the missing midpoint with zero or silently merge the two sets.
Add a duplicate and a changed instrument version to the test. Confirm that the workflow follows the declared rules and flags comparability questions. A successful test can produce no improvement or an inconclusive result. It should not require discovering a declining subgroup.
Watch: following a training cohort through its stages
Watch the video · 4 minutes 13 seconds. This illustrative training workflow follows data across stages from application to placement. It provides context for connected records, not evidence that the fictional calculations above describe a customer. Browse more videos in the video library.
Frequently asked questions
Can I analyze pre and post data when the midpoint is missing?
Yes, if baseline and endline are valid and comparable for that person. State that the analysis uses paired baseline/endline records. A complete three-wave trajectory uses a different set or a method that explicitly handles missing observations.
Should I report averages or individual change counts?
Often both. Averages summarize numeric scores, while distributions and movement counts show variation. Choose summaries appropriate to the measurement scale and state the included population.
Is a one-point increase meaningful?
Not automatically. Meaning depends on the instrument, measurement properties, population and purpose. Label an unvalidated threshold as an operational or illustrative rule rather than proof of meaningful change.
Does a lower score show that the program caused harm?
No. It is a signal to investigate. A decline may have several explanations, and a before-and-after survey alone does not establish its cause.
How should missing scores be treated?
Keep them distinct from zero and report the relevant missingness. Do not automatically fill them with a mean or the previous score. Any estimation method needs appropriate assumptions and documentation.
Do we need all three waves?
Only when the question and intended use justify them. A midpoint can support learning or an adjustment during delivery. Baseline/endline pairs answer a narrower change question without describing the full path.
What is the next step after one program cycle?
Follow outcomes beyond completion using defined dates and periods. Preserve instrument versions and missingness so later comparisons do not lose the context of the original cycle.
When your question extends beyond one cycle
Carry your paired analysis view, measurement rules and limitations into the longitudinal survey analysis lesson. You will move from fixed checkpoints in one cycle to repeated observations across a longer period.
Reviewed September 12, 2026. All six-person scores and program scenarios in this lesson are fictional teaching examples. Methodological sources are linked where used.
Build this part of your plan in more detail
These lessons address the next practical questions in this module. Choose the detail your workflow needs, then return to complete the exercise.
- How to Analyze Pre, Mid and Post Survey Data
Calculate matched observations and distinguish endpoint pairs from complete-wave records.
- Survey Attrition in Longitudinal Studies: Track Missing Waves
Distinguish a change in respondents from a change in outcomes.
More practical questions in this module (4)
- How to Analyze Longitudinal Survey Data
Inspect learning and follow-up paths rather than only the first and last average.
- How to Measure Behavior Change After Training?
Plan workplace observations, opportunity to apply, coverage and support follow-up.
- How to Measure a Mentee's Growth Across Every Session
For mentoring or coaching, follow goal evidence, opportunity and actions across dated sessions.
- How Do I Evaluate an Accelerator or Startup Cohort?
For accelerator programs, connect selection with founder learning and company outcomes without mixing record levels.
Add this part to your plan
On workbook page 4, map a ten-person cohort with six usable pairs, one changed manager and one late response. State what you can compare.
Download workbook (fillable PDF)Compare with a suggested answer
The main paired result uses only the comparable observations and states its base. The changed manager and late response remain visible. Missing responses are not zeroes and are not described as successful outcomes.
Self-check: Would a reader understand exactly whose evidence supports the reported change?
Questions you may have
Does a Contact ID make scores comparable?
It links the record. Comparability also depends on the measure, timing, method and perspective.
Can late responses still be useful?
Yes. Preserve the actual date and decide whether they fit the intended window or need a separate interpretation.
Apply the method to your work
Use the five-part plan to assess the sources, analysis and permissions your team needs. The solution page shows where a connected platform can support that workflow.
Explore the training and programs solution →