This Foundations reading also supports Training & programs, module 3.
Academy / Foundations / Lesson 5
Course progress and additional readings
Foundations · Lesson 5 of 6
You leave with: A matched-pairs rule for your baseline and follow-up, and a result sentence another person could reproduce.
Where this fits: Lesson 5 asks whether you need a baseline to claim improvement. This deep dive answers the next question: once you have intake, mid-program and exit answers, how do you compare them without mixing up who answered what? You bring back a matched-pairs rule for your plan.
How do you analyze pre, mid and post survey data?
In short: Match each person’s answers across waves by ID, calculate change only where the same person answered both waves, and report the spread of changes beside any average. Say who is included and who is missing.
A before-and-after difference describes change. It does not show that the program caused it. For the survey design and collection schedule itself, see the pre-and-post survey guide.
What do pre, mid and post mean?
In short: Pre is the baseline before the program, mid is an agreed checkpoint during it, post is the endline after it. Record the actual date of every answer, not only the label.
The fictional training team asks confidence at intake, at a mid-program check-in and at exit. A response collected before the first session and one collected after four sessions are not the same baseline, even if both are labelled “pre”.
Write the question first. “How did confidence change between intake and exit among people with both answers?” needs a pair. “When did confidence change across all three checkpoints?” needs the midpoint too. The CDC evaluation design guidance ties the question to its intended use and draws the line between observing outcomes and establishing causes, which matters for any pre/post survey without a comparison group.
Why does the ID come before the calculation?
In short: Change is a difference within one person. Without a shared ID, you are comparing the intake group with the exit group and hoping they are the same people.
Training example · fictional
Maria, ID 0417, rated her confidence 2 out of 5 at intake and 4 out of 5 at exit. Both answers sit on the record created when she enrolled, so the pair is certain. In a survey-to-Excel workflow, the same comparison depends on matching two exports by name or email. A changed address or a second Maria in the cohort breaks the pair silently, and a model asked to “compare pre and post” will average whichever rows it received.
For each answer, keep the ID, cohort, wave, actual response date, form version and score. Store “not collected” separately from the score: a blank is not zero, and a skipped question is not a missed survey. If someone submitted twice, apply a written rule, such as the latest valid answer inside the window, and keep the earlier one. Never pick whichever answer shows the most improvement.
Keep the question wording, scale and scoring direction the same across waves. If a question changed between intake and exit, the pair may not compare; changing a question safely shows how to keep versions visible instead of presenting a broken series as continuous.
Worked example: six people, three checkpoints
In short: Six people have intake and exit answers; one missed the midpoint. The matched-pairs average rises by 0.50 points, and the spread matters as much as the average.
This fictional set uses one confidence question scored 1 to 5, higher meaning more confident. It is not a validated scale. Person F missed the mid-program check-in but answered at intake and exit.
| Person | Pre · Mid · Post | Post minus pre |
|---|---|---|
| A | 2 · 3 · 4 | +2 |
| B | 3 · 3 · 3 | 0 |
| C | 4 · 3 · 2 | −2 |
| D | 1 · 2 · 4 | +3 |
| E | 5 · 4 · 4 | −1 |
| F | 2 · not collected · 3 | +1 |
For all six intake/exit pairs, the intake total is 17 and the exit total is 20. The means are 2.83 and 3.33, and the mean paired change is +0.50: the individual changes sum to 3, and 3 divided by 6 is 0.50. Three people scored higher, one stayed the same and two scored lower.
These means summarize numbered categories, and the distance between 2 and 3 on a confidence question is not guaranteed to equal the distance between 4 and 5. Show the movement counts beside the average. And avoid percentage change on a scale with no true zero: Maria moving from 2 to 4 is not “100% more confident”.
Which people belong in the comparison?
In short: The matched-pairs rule: include a person in a comparison only if they answered both waves that comparison uses. Different questions use different sets, so name the set every time.
Do not drop F from the intake-to-exit comparison; both of F’s answers exist. If you want complete three-wave paths, A through E form a separate set of five. Keep the two sets visibly distinct.
| Question and set | Included | Result | Limit |
|---|---|---|---|
| Intake to exit | A–F, six pairs | Mean 2.83 to 3.33; change +0.50 | Says nothing about F’s midpoint |
| Complete three-wave paths | A–E, five records | Means 3.00, 3.00, 3.40; pre/post change +0.40 | Excludes F; may differ from people with gaps |
The results differ because the people differ. Neither is wrong. The error is plotting the six-person intake mean beside the five-person midpoint mean and calling the line one group’s path. In a real cohort, report the number expected at each wave, usable answers, matched pairs and complete records. When the gaps are large, who stops responding shows how to check whether missing answers could change the reading.
Is observed change the same as meaningful change?
In short: No. A number moving, a statistically supported difference and a change that matters to people are three separate judgments. No one-point rule settles all three.
Look for evidence about a measure’s reliability, measurement error and responsiveness in a similar population. COSMIN’s measurement framework sets out these properties for health outcome measures; it is a useful reference, not a certification of the confidence question here. If no defensible threshold exists, report the observed movement and say its practical meaning is uncertain. Write any threshold down before you look at results.
What is the midpoint for?
In short: Deciding what to look into while there is still time to change delivery. It is not a small version of the final result.
In the example, C falls at each checkpoint, E falls then holds, and A and D rise, even though the five-person mean is flat between intake and midpoint. Each path suggests a different question. Did C understand the question the same way each time? Did exposure differ? Does C’s comment describe a barrier, or a more realistic self-assessment after learning what the task involves? The score alone cannot tell. Read the comment from the same wave beside it, and keep contradictory accounts in view rather than quoting only the ones that fit.
Name who reviews the midpoint and what decision it can inform. If no one will act on it, the extra wave is burden without purpose.
When do you need statistics?
In short: When the decision needs an estimate beyond the people you observed. Then use a method built for repeated answers from the same people, and get help if the claim is causal.
Repeated answers from one person are not independent samples. A paired analysis or repeated-measures model may fit, under its assumptions; the six-person set is a reconciliation exercise, not a basis for a significance claim. Even a well-supported pre/post difference does not rule out outside events, selection, or a change in who responded.
What does governing the waves at collection change?
In short: The pairs exist before anyone calculates. Intake creates the ID; the mid-program and exit forms are added to that same record, so there is nothing to match by name.
In Sopact Sense, each person gets a persistent ID at the first form, and every later wave is another workflow on that record. The AI Assistant can answer “show confidence at intake and exit for people with both answers” and every line links to a record you can open, so you can check that F is in the pair set and the midpoint is not filled with a guess. Software keeps the pairs intact; it cannot choose a valid measurement rule for you.
ASK ANY TOOL, INCLUDING OURS
Load the six records above and ask for the intake-to-exit change. A correct answer includes F, uses six pairs and gives +0.50. Then ask for complete three-wave paths: five people, +0.40 from intake to exit. Neither answer should replace F’s missing midpoint with zero or merge the two sets. Add a duplicate submission and check the tool follows your written rule.
Try it on your own data
Open your working evidence plan ↗
- Write the change question in one sentence, naming the two waves it compares.
- Name the ID that links those waves and the form that creates it.
- Write your matched-pairs rule, your duplicate rule and how you will record “not collected”.
- Draft the result sentence with the number of pairs, the spread and one limitation.
Check your reasoning
A strong answer reads like this: “Among six people with intake and exit answers, the mean confidence score rose by 0.50 points. Three scored higher, one was unchanged and two scored lower. One person lacked a midpoint answer. This small fictional example does not establish meaningful or causal change.” If your sentence has no count of pairs, a reader cannot tell whose change it describes.
Questions teams ask
Can I analyze pre and post data when the midpoint is missing?
Yes, if the intake and exit answers are valid and comparable for that person. State that the analysis uses matched intake/exit pairs. A complete three-wave path uses a different, usually smaller, set of people, or a method that explicitly handles missing observations. Report which one you used, and never fill the missing midpoint with zero or with the average.
Should I report averages or individual change counts?
Often both. An average summarizes direction and size; counts of who rose, held and fell show how the change is spread. On a single 1–5 question, the counts are often the more honest headline. State how many matched pairs they come from.
Is a one-point increase meaningful?
Not automatically. Whether a change matters depends on the measure, its measurement error, the population and the decision. If you use a threshold without evidence behind it, label it as an operational rule for follow-up, not as proof of meaningful change, and write it down before you look at the results.
Does a lower score show that the program caused harm?
No. It is a signal to investigate. A decline can reflect a real difficulty, measurement noise, or a more realistic self-assessment once someone learns what a task involves. Read the person’s comment from the same wave, check attendance on the same record, and treat the result as a question for the next review.
Does an ID make scores comparable?
It makes them linkable. Comparability also depends on the same question wording, scale, timing and respondent. A confidence self-rating and a mentor’s rating of the same person can sit on one record, but they are different measures and should not be subtracted from each other.