What is baseline data?
Baseline data is the first measurement of a group, taken before a program starts, so a later change can be measured against it. Its value is conditional: it only proves anything if the endline attaches to the same people. Sopact holds the baseline on the Outcome Thread, one participant record under a persistent Contact ID, so a midline and an endline join that same person instead of arriving as a fresh anonymous sheet.
Most teams collect a clean baseline and feel finished, then discover the problem a year later. The endline comes back as a separate export with new anonymous rows, and matching it to the baseline means joining on names, emails, or best guesses, which fails for the very people who moved, dropped out, or answered differently the second time. The baseline is fine; the join is what dies, and with it the whole comparison.
Key takeaways
- Baseline data only has value if the endline attaches to the same people, so the persistent ID matters more than the questionnaire.
- Sopact holds the baseline on the Outcome Thread: one participant record, under a persistent Contact ID, that a later wave rejoins instead of a fresh anonymous export.
- Matching two exports by name or email fails for exactly the people who changed, which is why post-hoc joins lose the movers you most need to see.
- Collect the baseline clean at the source and the later comparison is a query over one record, not a fuzzy merge across spreadsheets.
- Conventional collection treats each wave as a new sheet; the Outcome Thread keeps the baseline and endline on the same person.
The data-model gap: a baseline in a sheet, not on a record
A baseline stored as a spreadsheet is a snapshot of anonymous rows. When the endline arrives as its own sheet, nothing in the data says which endline row is the same person as which baseline row, so the analyst rebuilds the link by hand and accepts whatever match rate they can get.
Sopact is record-centric: the baseline lands on a persistent Contact ID at intake, so the endline attaches to the same person and change is a query over one Outcome Thread rather than a hand-matched join across two exports. See the collection layer on longitudinal data collection software, or the first-wave design on baseline survey.
The tools teams reach for, and the one test
Teams usually stand up a baseline in SurveyMonkey, Qualtrics, Google Forms, KoBoToolbox, SurveyCTO, or CommCare, then export to Excel and keep the file until the endline. Each of those tools captures a wave well, and each was built to capture a wave, so nothing carries the respondent’s identity forward to the next one.
The one test that sorts them: ask the tool to show one participant’s baseline answer and endline answer side by side, matched automatically, with no hand-keyed join. An export-per-wave workflow answers by making you merge two files. Sopact answers from the Outcome Thread, because both waves already sit on the same persistent Contact ID.
Matching after the fact vs collecting on a persistent ID from the start
The move that saves a study is deciding, before the baseline, that every wave will land on the same Contact ID. Collected that way, the baseline is not a file to protect; it is the first entry on a record the endline will extend, so attrition is visible as it happens and a comparison is defensible.
Kept on the Outcome Thread, the baseline stays useful for years: a participant’s answers across every wave read as one trajectory, each traceable to the person who gave it. Sopact collects clean at the source, so a finding rests on matched records rather than a merge no reviewer can re-check.
A baseline in a file vs a baseline on the Outcome Thread
A stored baseline file waits to be merged with a separate endline; the Outcome Thread holds the baseline on a persistent Contact ID that the endline rejoins. The difference is whether the comparison is a query or a hand-keyed join.
Two ways to hold a baseline
| The question | Baseline in a file | Outcome Thread |
|---|
| Match endline to baseline? | By hand, on name or email | Automatic, on one Contact ID |
| Keep the movers? | Often lost in the join | Yes: same record over time |
| See attrition early? | No: after the endline | Yes: as waves land |
| Defensible comparison? | Depends on match rate | Yes: matched records |
See the collection layer on longitudinal data collection software, or how the numbers are gathered on quantitative data collection methods.
A dataset tells you where a cohort ended. The Loop tells you who is drifting, in time to act.
A finished dataset is a snapshot of where a cohort landed by the time you cleaned the last wave. The value of a response is highest the moment it arrives, when a participant slipping between the baseline and the midline can still be reached, not in a report written after the endline closed. That is the premise of the Loop, Sopact’s method for continuous intelligence: collect clean at the source, so each wave is validated at intake on a persistent Contact ID with no post-hoc cleanup; analyze on arrival, so each wave is read as it lands and the open-text is themed rather than set aside; improve in time, so a participant drifting between waves surfaces mid-program instead of after it.
The Loop is also what keeps a longitudinal finding defensible: every trajectory traces back to the same person’s answers across waves on one persistent ID, the standard detailed in Loop traceability, so a conclusion rests on the Outcome Thread rather than a hand-matched merge of three spreadsheets no one can re-check.
One method, three moves that never stop
1 · CollectClean at the source; each wave validated at intake on a persistent Contact ID, so there is no anonymous sheet to clean and match to prior waves afterward.
2 · AnalyzeOn arrival; each wave read the moment it lands and the open-text themed, tied to the same person’s earlier answers on one Outcome Thread.
3 · ImproveIn time to act; a participant drifting between waves surfaces during the program, while you can still reach them, not at the end-of-program report.
Then the next wave reads a little sharper on the same record. Read the method: the Loop methodology →
Match a slice of your own baseline to a later wave
The fastest way to see the matching gap is to run it on your own data. Export a baseline and a later wave, each carrying a participant ID, then paste the prompts below into Sopact Sense’s Assistant, or reason through them with your team. The arrow above each links the Academy walkthrough with the expected output and tips.
Academy walkthrough → Analyze longitudinal survey data
Here are our baseline, midline, and endline responses, each row carrying the respondent’s persistent Contact ID: [ATTACH]. Match every wave to the same person by that ID, show each participant’s trajectory over time, quote the open-text behind any change, and keep it all on one Outcome Thread, so the change is a query over one record rather than a hand-matched join across three exports.
Academy walkthrough → Analyze pre, mid, and post data
Here are pre, mid, and post responses on the same participant IDs: [ATTACH]. For each person, line up the before, during, and after answers on their persistent Contact ID, compute the shift, quote the sentence that explains it, and keep every answer on the Outcome Thread, so a change is measured on one record instead of reconstructed from three anonymous sheets.
Academy walkthrough → Handle attrition across waves
Here are the responses to each wave with the respondent’s persistent Contact ID: [ATTACH]. Show me who answered the baseline but has not yet answered the latest wave, flag the drop-off by subgroup, and keep everyone on the Outcome Thread, so I can reach the people drifting away while the cohort is still reachable rather than discovering the gap after the study closes.
Academy walkthrough → Connect the number and the reason
Here is our quantitative data and the open-ended responses on the same participant IDs: [ATTACH]. For each rating, pull the open-text the same respondent wrote that explains it, quote the sentence, and show the number and the reason on one record, so a low score carries its reason on the Outcome Thread rather than sitting in a column with no explanation.
Learn the how-to in the Academy
Each walkthrough is short and practical: what to do, the prompt to run, the output to expect, and the tips that keep it reliable.
LongitudinalAnalyze longitudinal survey dataRead a baseline, midline, and endline as one trajectory on the Outcome Thread, so change is a query over one participant record instead of a fuzzy join across three separate exports.Pre / mid / postAnalyze pre, mid, and post dataCompare a person’s answers before, during, and after on the same persistent ID, so a shift is measured on one record rather than reconstructed from three anonymous sheets.AttritionHandle attrition across wavesSee who answered the baseline but not the endline while a cohort is still reachable, because every wave lands on the same Outcome Thread rather than in a pile of unmatched rows.ConnectConnect the number and the reasonPair each rating with the open-text explaining it on one record, so a score and its reason are read together instead of in two exports that never rejoin.
Watch: collecting clean at the source on a persistent Contact ID and reading each wave on arrival, so a baseline and an endline attach to the same person on one Outcome Thread.
Frequently asked questions
What is baseline data?
It is the first measurement of a group before a program starts, used to measure later change. Sopact holds it on the Outcome Thread under a persistent Contact ID, so a midline or endline attaches to the same person instead of arriving as an unmatched export.
Why does the baseline lose its value?
Because the endline comes back as a separate anonymous sheet and the join to the baseline fails for the people who moved or dropped out. Sopact keeps every wave on one Outcome Thread, so the baseline stays matched to the same person.
How do I match an endline to a baseline?
Ideally you do not match after the fact at all. Sopact collects each wave on a persistent Contact ID, so the endline is already attached to the baseline on the Outcome Thread rather than merged by hand on name or email.
Can I still see who dropped out?
Yes. Because everyone sits on the Outcome Thread, Sopact shows who answered the baseline but not the latest wave while the cohort is still reachable, rather than surfacing the gap after the study closes.
Do I need to clean the baseline first?
No. Sopact validates each response at intake on its persistent ID, so the baseline is analyzable on arrival on the Outcome Thread rather than after a round of post-hoc cleanup.
What if I already collected the baseline elsewhere?
You can bring it in and assign persistent Contact IDs, so future waves attach to it on the Outcome Thread. The sooner the record is persistent, the fewer movers you lose in later matching.
How is this different from keeping baseline and endline in Excel?
Two Excel files are two anonymous sheets you merge by hand. Sopact keeps both waves on one record, so a change is a query over the Outcome Thread rather than a fuzzy join you cannot fully trust.
Why does the persistent ID matter more than the questions?
Good questions on an unmatched wave still cannot show change. Sopact’s persistent Contact ID is what lets a baseline and endline sit on the same Outcome Thread, which is the precondition for measuring anything over time.
Next: design the first wave on baseline survey, or collect every wave on one record with longitudinal data collection software.
Baseline, kept
01CollectFirst wave on a persistent ID
02HoldOne record, not a stored file
03RejoinEndline attaches to the same person
04CompareChange as a query, not a join
A baseline is only worth collecting if the endline finds the same people.