Academy / Foundations / Lesson 5
Additional course links
You leave with: A follow-up schedule on one ID, with a comparability check for each wave and a way to report paths as well as averages.
Where this fits: Lesson 5 shows why waves only compare when they sit on the same ID. This deep dive follows people past the end of the program, from intake to exit, 30 days and 6 months, and answers how to read those waves without overclaiming. You bring back a follow-up schedule and a comparability check for your plan.
How do you analyze longitudinal survey data?
In short: Keep every wave on the same person’s ID, align answers to a clear time scale, and read individual paths beside the group average. Keep missing waves, question versions and the length of follow-up visible.
Longitudinal analysis answers whether a change seen at exit is still visible months later, and for whom. It cannot show what happened between checkpoints you did not observe, and naming the shape of a path does not explain it.
What does one ID make possible?
In short: A continuing record. Each wave is added to the person, not filed as a new spreadsheet, so a 6-month answer can be read beside the intake answer it is meant to compare with.
Training example · fictional
Maria, ID 0417, rated her confidence 2 out of 5 at intake. She attended 10 of 12 sessions, her mentor noted that she led a mock interview, and she rated herself 4 out of 5 at exit. At the 30-day follow-up she reported using the skill on the job. Her 6-month follow-up has not been sent yet. All of it sits on one record, in date order, so “has her use lasted?” is a question the team can answer when that wave arrives, not a matching project.
Compare that with the usual path. Each follow-up is a separate survey export. The 6-month file is matched to the exit file by email, some addresses have changed since the program ended, and the people who could not be matched drop out of the analysis without anyone deciding they should. A chat model given those files has no ID across surveys and compares whichever rows it receives.
Governed at collection, the order runs the other way. The first form creates the ID; the exit survey, the 30-day and 6-month follow-ups and even public data, such as a county employment rate loaded as its own survey, are added to the same record. The analysis starts from people, not from files.

Are you asking about people or about the group?
In short: Decide before you chart anything. “How did the average change?” and “Who rose and then fell back?” are different questions and need different summaries.
A panel follows the same people. Repeated cross-sectional surveys may reach different people each time; they can show a trend but not movement within a person. Count people and observations separately: if the 40 completers each answer four times, that is 40 people and up to 160 observations, not 160 participants.
A useful question reads like this: “For the spring cohort’s 40 completers, how does the same confidence question change from intake to 6 months, and how much follow-up is missing?”
What keeps waves comparable?
In short: The same question wording, scale, respondent and window at every wave, plus the actual date of each answer. The ID links the waves; these make them comparable.
| Check | What to keep the same | What breaks it |
|---|---|---|
| Question | Wording, scale and scoring direction | A reworded item or a 1–5 scale becoming 1–10 |
| Respondent | Self-report compared with self-report | A mentor rating at exit compared with a self-rating at 6 months |
| Window | A defined send and cutoff, such as day 25–35 | One person answering at day 20, another at day 90, both called “30 days” |
| Starting event | One clock: completion, enrollment or first session | Mixing time since enrollment with time since completion |
For rolling cohorts, elapsed time since completion compares similar stages; calendar time matters when a policy change or local hiring freeze affects everyone at once. Keep actual dates so both remain possible. State the horizon plainly: “higher at the available 30-day and 6-month checks” is more precise than “improved permanently”. If a question changes between waves, changing a question safely shows how to keep the break visible rather than smoothing it away.
Worked example: similar endpoints, different paths
In short: Two people can start and finish at the same score with very different histories. Keep the full sequence, and do not describe what you did not observe.
These fictional records use one confidence question scored 1 to 5 at intake, exit, 30 days and 6 months. They show observed sequences, not validated thresholds.
| Person | Intake · Exit · 30 days · 6 months | Observed description |
|---|---|---|
| A | 2 · 4 · 2 · 2 | Higher at exit, then back to intake level |
| B | 2 · 2 · 2 · 2 | Same score at each check |
| C | 2 · 3 · 4 · 4 | Higher at later checks |
| D | 2 · 2 · 2 · 4 | Higher only at 6 months |
| E | 4 · 3 · 2 · 1 | Lower at each check |
| F | 2 · 4 · missing · missing | Later path unknown |
A and B both start and finish at 2. Their middle records differ, which is a reason to keep the whole sequence, not proof the program reached A and missed B. Among A to E, who have all four answers, the means are 2.4, 2.8, 2.4 and 2.6. That summary is useful beside the paths; it does not mean anyone followed it.
F rose by exit and then stopped answering. Do not label F as having kept or lost the gain. If you plot every available answer, F counts at intake and exit but not later, so the base changes between points; show that on the chart. A line between checkpoints is a visual aid, not an observation.
How do you handle missing waves and unequal follow-up?
In short: Report how many people were eligible for each horizon and how many answered. Never turn a missing answer into zero or carry the last score forward as if it were observed.
In the training example, 25 of the 40 completers answered at 30 days. The 6-month wave will likely reach fewer. A later cohort has not yet been eligible for 6 months at all, so it belongs in no 6-month figure. If a score is higher at exit and lower at 6 months, the decline happened somewhere between those checks; you cannot name the month without more evidence. Who stops responding gives the full method for tracking gaps and reporting coverage.
If you label paths, such as “early gain, later lower”, write the rule first and keep “not enough follow-up” as its own label.
When does a statistical model add value?
In short: When you need to estimate change, compare groups or handle unequal schedules beyond what you observed. Models bring assumptions about timing, dependence and missing answers, so get qualified help.
UCLA’s longitudinal analysis seminar compares approaches and shows how multilevel models treat repeated answers as nested within people. A growth model estimates an average path and variation around it; it does not sort people into types. Rule-based labels, growth models and latent-group methods answer different questions, so do not present them as versions of one another.
How do comments fit into a path?
In short: Read each comment beside the wave it came from. A theme found after a decline is a lead to investigate, not its cause.
For A, a 30-day comment might mention a new role with no chance to use the skill, which is a different story from losing confidence. Note the date the comment describes as well as the date it was written. Put the number, the account, the limitation and the next question together, as numbers with their reasons explains, and keep a mentor’s view and the person’s own view labelled by who said them.
What does Sopact Sense do with the waves?
In short: It keeps them on one record from the first form. Each follow-up is another workflow on the same ID, and questions about paths come back with the records behind them.
In Sopact Sense, the ID assigned at the first form carries through exit, the 30-day and 6-month follow-ups and any public data you load as a survey. The Intelligence Row summarizes each person across their waves, and the AI Assistant answers questions such as “who was higher at exit but lower at 30 days?” with every line linked to a record you can open. It does not decide which model is valid or what a path means. Your team does.
A short explainer on keeping comments and responses attached to a continuing record, which is what lets you read a 30-day comment beside that person’s intake and exit answers. It is not a modeling tutorial. Watch on YouTube ↗
Try it on your own data
Open your working evidence plan ↗
- List your waves in order, from the first form to the last follow-up, with the send window for each.
- Name the form that creates the ID and confirm every later wave carries it.
- For one outcome question, fill in the four comparability checks: question, respondent, window and starting event.
- Write how you will report a person whose later waves are missing.
Check your reasoning
A strong schedule has one starting event, a defined window per wave and the ID created at the first form. It reports the 6-month result only for people eligible for 6 months, shows how many answered, and describes F as “later path unknown”, not as having kept the gain. If any wave is matched by name or email, that is the first thing to fix.
Questions teams ask
How many waves are required?
Two observations can describe a difference. More can reveal the shape of change, such as a gain that fades, but the number you need depends on the question, the timing and what you will do with the answer. Every extra wave costs respondents time, so each should have a decision attached.
Should I stop reporting the average trend?
No. A group trend is useful when its base and limits are clear. Pair it with the spread or with individual paths when the question is about who changed, and show how many people stand behind each point.
Can I use people with incomplete histories?
Often, with a method suited to the question and stated assumptions about the missing answers. Report their observed waves, keep “missing” distinct from any score, and do not compare groups with different lengths of follow-up as if they were the same.
Is time since completion always the right time scale?
No. Elapsed time compares people at similar stages after an event. Calendar time is better when something outside the program affects everyone at once, such as a local hiring freeze. Keep actual dates on every answer so you can choose the scale that fits each question later.
Does a later decline prove that an outcome was lost?
It shows a lower recorded value at that check. Whether the outcome was lost depends on the measure, the definition and other evidence, such as a comment about a new role. When the decline began, and why, may stay uncertain.