Academy / Foundations / Lesson 3
Additional course links
You leave with: A change-log entry for one question and a written comparability decision you can add to your data dictionary entry.
Where this fits: Lesson 3 has you write a data dictionary entry for one measure. This page answers the question that entry raises as soon as the team learns something: how do we change the question without breaking the comparison with earlier waves? You bring back a change-log entry and a comparability decision for your entry.
After the spring cohort, the training team reads the 30-day follow-up and sees its limits. The question was “In the last 30 days, have you used [skill] at work?” with Yes or No. Fifteen of 25 respondents said Yes, 10 said No. But Yes covers someone who tried the skill once and someone who uses it daily, and the funder’s next question will be about how often. The team wants to ask about frequency for the summer cohort, and to add a question about practice, because it plans to test a practice session.
Someone remembers the last time a live form was edited: the numbers moved and it took two weeks to work out why. The safe-looking choice is to leave the question alone. That is how forms freeze: nothing breaks, and the team keeps collecting careful answers to a question that stopped doing its job.
How do you change survey questions without losing comparability?
In short: Do not write over the old question. Keep its wording with the answers already given to it, start the new wording as a separate question from a dated point, and decide in writing whether the two can be compared. Every report that spans the change says so.
Editing in place puts two different questions in one column, and any trend drawn across it partly measures your edit rather than your participants. Keeping both wordings protects the history. It does not make the two measures equivalent; that is a separate decision, and your team makes it.
Which kind of change is this?
In short: Adding a question, tidying wording and changing what is asked follow three different rules. The test between the last two: would someone plausibly answer differently because of this edit?
| The change | What it looks like | What to do |
|---|---|---|
| You add something new | A question this cohort is asked that earlier cohorts never saw | Add and test it, including where it sits and any skip logic. Earlier cohorts are “not asked,” never zero or blank |
| You tidy the wording | A typo; “Please describe” becomes “Describe” | Save the original wording, log the edit, and check that meaning, display and routing are unchanged |
| You change what you are asking | Yes/No becomes a frequency scale; a 1–5 scale becomes 1–7; the options change | Keep both. The old wording holds its answers, the new one starts its own run, and reports covering both say so |
The size of the edit is irrelevant. Adding the word “not” is three characters and reverses the question. If you are unsure whether an edit is a tidy-up or a change of meaning, treat it as a change of meaning.
What does this look like on the follow-up question?
In short: The spring answers stay under the spring wording, the summer question starts fresh, and any line drawn across the two is labelled as a derived measure with its rule written down.
The team has two changes. Changing Yes/No to “How often have you used [skill] at work in the last 30 days? Never / Once or twice / Weekly or more” is a change of meaning: the answer options differ, and someone who tried the skill once may have answered No to the old question and “Once or twice” to the new one. Adding “Did you practise the skill before using it at work?” is a new question.
CHANGE LOG ENTRY · EXAMPLE (ILLUSTRATIVE)
In the report, spring reads “15 of 25 respondents said they used the skill” and summer reads under its own wording and denominator. If the team shows an “any use” figure for both, it adds a sentence saying the question changed between cohorts and the summer figure is derived. The new practice question shows as “not asked” for spring, not as 40 blanks.
The same rule applies to a scale. If confidence moved from 1–5 to 1–7, Maria’s 4 of 5 at exit would not be a 4 of 7, and no percentage conversion makes it one. Report each scale on its own and mark the break.
Who is allowed to change what?
In short: Write the rules down before you need them. Teams freeze less for lack of permission than because nobody said which changes were safe.
| Kind of change | Who | What they need |
|---|---|---|
| Typos, help text, reminder wording | Anyone who edits forms | A log line with the old wording |
| New questions, changed options or scales, retiring a question | A named measurement owner | A change-log entry, a test, and a comparability decision |
| How people are identified across waves, consent wording, rules for hiding small groups | Nobody alone | A written, dated team decision, and specialist advice where needed |
How do you keep the change log by hand?
In short: A shared sheet with one row per change, written before the edit. It is a complete method, and the one to use whatever tool your forms live in.
01 · COLUMNS
Question, old and new wording in full, date, reason, change type, approver, report treatment
02 · BEFORE
Paste the old wording before you edit; afterwards it is gone
03 · SORT
New question, tidy-up or change of meaning, using the one test
04 · ADD
For a change of meaning, add a new question and retire the old one
05 · REPORT
Read the log before any report that spans the change date
A change log a team of two can keep in a shared sheet.
In Sopact Sense, each wave is its own survey on the person’s ID, so the spring follow-up answers stay where they were when the summer form changes. The change log itself is a team practice: do not assume any platform, ours included, records what changed and why on your behalf. Link the log from your data dictionary entry so the two are read together.
Where does it break?
In short: On repetition, in three predictable places: the log drifts from the form, its keeper leaves, or the one report that needed the note goes out without it.
A Friday fix never gets logged, and six months on the log describes a form that no longer exists. The keeper leaves, and nobody can tell which wording was live in June. Or the board pack, built in a hurry, draws one line across the change. A monthly check that the log matches the live form, and a named owner with a backup, cover most of this.
What else can change the answers?
In short: Question order, collection mode, reminders and who receives the form can all move a trend. Pretest a new wording with a few real respondents before it goes live.
AAPOR’s survey guidance advises keeping wording, framing and methods comparable when measuring change, and documenting and examining any change you have to make. Pew Research Center’s question-design guidance describes pretesting and the effect of wording, response choices and order. Ask a few people from the next cohort what they understand the new wording to mean.
Putting the new practice question before the frequency question may change what people have in mind; sending the follow-up by text instead of email may change who responds. These are reasons to check comparability, not proof the survey is spoiled.
If a trend must continue across a change of meaning, a survey specialist can design an overlap or split-sample study that separates the effect of the wording from changes in who answered. If your sample cannot support that, report two separate series. An interpretable history matters more than an unbroken line.
A short explainer (1:25) on writing each definition once and reusing it across reporting frameworks. When a definition changes, your log records when, and every report using it carries the note. Watch on YouTube ↗
ASK ANY TOOL, INCLUDING OURS
In a demo, change a question between two waves of test data and ask: “Show me the answers to the old wording and the new wording.” A good answer keeps the earlier responses readable with the wording people actually saw, shows the new question separately, and marks the new question “not asked” for earlier waves. Then ask for one figure across both waves; the answer should say the question changed rather than averaging across it.
Try it on your own data
Open your working evidence plan ↗
- Pick one recurring question you have wanted to change, or the measure from your lesson 3 dictionary entry.
- Write the old wording in full and your proposed new wording.
- Sort it: new question, tidy-up or change of meaning. Use the test: would anyone answer differently?
- Write the change-log entry: date, reason, approver, how existing answers stay, and your comparability decision.
- Write the one sentence a report spanning the change will carry, and add the entry to your data dictionary.
Check your reasoning
Moving from Yes/No to a frequency scale is a change of meaning, so the training team keeps spring’s 15 Yes, 10 No and 15 unknown under the old wording and starts the frequency question as its own question. A report spanning both cohorts reads: “The 30-day question changed from Yes/No to a frequency scale for the summer cohort. Spring and summer figures use different questions; the summer ‘any use’ figure is derived and not directly comparable.” The practice question is new, so spring shows “not asked.”
Questions teams ask
Can I edit a live form without ruining my data?
Yes, if you sort the change first and keep a record. A tidy-up needs a log line with the old wording. A change of meaning needs a new question beside the retired one, so earlier answers keep the wording people saw. Even adding a question can affect order or skip logic, so test the form before the next wave goes out.
What counts as a real change rather than a tidy-up?
Changing a scale’s range, changing the answer options, flipping a question so agreeing means the opposite, or rewording it to ask about something else. Moving from Yes/No to a frequency scale counts. The size of the edit is irrelevant: adding “not” is three characters and turns the question inside out.
How do I compare results either side of a scale change?
Start by reporting each version on its own. A top-two answer on a five-point scale is not automatically equivalent to one on a seven-point scale, and rescaling to a percentage changes the arithmetic without showing that people use the scales the same way. Combining them needs a justified method, evidence and a clear note in the report.
What should never change casually?
How you identify a person from one wave to the next, your consent wording, and any rule about hiding results for very small groups. Each can undo work downstream in ways nobody notices for months, so each belongs behind a written, dated team decision.