This chapter resolves check 05 Qualitative of the eight checks.
A programme asks its participants an open question: what barriers did you face? Hundreds answer. Some write paragraphs. A large number write a single word — none. During cleanup, those single-word answers look like noise, so they get emptied along with the stray whitespace and the "N/A"s. The programme then reports the barriers its participants faced, calculated on a denominator that has quietly excluded everyone who said they had none.
Nothing in that report is a lie. It is simply answering a different question than the one anybody thinks it is answering.
What does it mean to clean an open-ended response?
It means deciding what counts as an answer — not fixing spelling. Whitespace, casing and near-duplicates are the trivial part and any tool will do them. The consequential part is what you do with the responses that contain no prose: none, N/A, prefer not to say, and the empty cell. Those four look identical in a spreadsheet and mean four different things, and collapsing them into one is the single most common way a qualitative finding gets destroyed before anyone reads it.
Four kinds of non-answer, and what each one means
Sort every non-prose response into one of these before you clean anything. The distinction determines whether the response belongs in your denominator, and getting it wrong changes your headline figure.
| What they wrote | What it actually means | Where it belongs |
|---|
| "none", "nothing", "no problems" | A substantive answer. They engaged with the question and reported the absence of the thing. | In the denominator, as a finding. This is a zero, not a blank. |
| "N/A", "doesn't apply" | The question was not relevant to them. They are not refusing; the item was mis-targeted. | Out of the denominator, and counted — a high rate here means the question is wrong. |
| "prefer not to say", "no comment" | A refusal. They understood and declined, which is itself information about the question's sensitivity. | Out of the denominator, tracked separately from skips. |
| Empty | A skip. You do not know whether they had nothing to say, ran out of patience, or never saw it. | Out of the denominator, and reported as non-response. |
The first row is the one that costs programmes their findings. "None" is the answer you most want — it is the participant telling you the intervention removed a barrier — and it is the answer most likely to be discarded as junk because it is short.
The line between normalizing and rewriting
Once the non-answers are sorted, the remaining prose needs standardizing so theming works. There is a hard line in the middle of that job.
Normalizing is safe: trimming whitespace, lowercasing for comparison, treating "transport" and "Transport" as one thing, recognising that "bus fare" and "busfare" are the same barrier. None of it changes what the person said.
Rewriting is not safe: expanding abbreviations, fixing grammar, completing sentences, translating. Each of those substitutes your words for theirs. It usually improves readability and it always weakens the evidence, because the moment you quote that response back to a funder you are quoting yourself.
So the rule is procedural rather than stylistic: clean a copy and never the original. Every theme should be able to cite the response as the participant typed it, misspellings included. If your cleaning process overwrites the source, you have traded citability for tidiness — and citability is the entire reason qualitative evidence carries weight.
How to do this without any particular software
- Copy the raw column to a second column before touching anything. The first column is now evidence and is never edited again.
- Classify the non-answers in a third column using the four categories above. Do this before any find-and-replace, because a global replace of "none" to empty destroys the distinction irreversibly.
- Normalize the prose in the copy: trim, standardize case for comparison, unify obvious spelling variants of the same concept. Keep a note of each variant you merged.
- Read fifty responses before you build any categories. Themes that come from the data survive; themes you brought with you get imposed onto it.
- Record your denominator explicitly in the output: responses received, substantive answers, not-applicable, refusals, skips. Five numbers. Every rate you report is then interpretable, and a reader can recompute it.
That is a complete method. A careful person with a spreadsheet and an afternoon can do it well, and for a few hundred responses they should.
Where it breaks
It breaks on three things, and they compound.
Volume. The King Center collected more than ten thousand stakeholder voices across seven programmes with no dedicated analysts, and the honest result was that the open-ended feedback went unread — collected as a matter of routine and never analysed. That is not a discipline failure. It is arithmetic: hand-coding does not scale past the point where the reading itself takes longer than the programme cycle.
Arrival over time. Responses do not land all at once. You code four hundred, then two hundred more arrive, and some of them belong to a theme you had not invented yet. Now you either re-read everything or accept that your early and late responses were coded against different schemes.
More than one language. The normalizing rules are language-specific, and the temptation is to translate everything into one language first — which is rewriting, at scale, applied to every response at once. That is covered properly in analyzing multilingual feedback.
Where Sopact Sense helps is narrow and specific to those three: responses are classified and themed as they arrive rather than in a batch at the end, new themes back-fill across responses already processed, the original wording is retained so every theme cites the participant's own words, and the five denominator figures are produced rather than reconstructed. It does not decide what your categories mean. You still read the themes and judge whether they describe your programme.
Where the theming scheme belongs
Everything above produces a set of categories and a denominator rule, and those are definitions in exactly the sense a closed measure is. "Counts as a transport barrier" needs writing down as surely as "counts as enrolled" does, or the next person to code a wave will draw the boundary somewhere else and the trend will move for no reason.
So the theming scheme and the denominator rule go in the same place your closed measures live — the governed data dictionary, with an owner, a version, and an effective date. That is what makes a qualitative finding comparable across waves rather than merely repeatable by the person who invented it. The upstream discipline of writing a single measure so two people count it identically is covered in giving every number one definition; a theme category is the same problem with words instead of numbers.
How to test this on your own data
Use: One open-ended question with at least a few hundred responses, deliberately including some that say "none", some "N/A", some refusals, and some blanks. Then add fifty more responses after the first pass is complete.
Pass: The four non-answer types are counted separately; the reported rate states its denominator; every theme cites an unedited original; and the fifty late responses are themed against the same scheme, with any new theme applied back to the earlier ones.
Fail: "None" and blank end up in the same bucket. Or a quoted response does not match what the participant actually typed.
Frequently asked questions
Should "none" be treated as a missing response?
No. It is a substantive answer reporting the absence of the thing you asked about, and it belongs in your denominator as a zero. Treating it as missing removes your most favourable finding and inflates every rate calculated from the remainder.
Is it acceptable to fix spelling in open-ended responses?
In a working copy, yes. In the source, no. Correct spelling in the version you theme, and quote from the version the participant typed — otherwise the quote in your funder report is your sentence, not theirs.
How many responses should I read before creating categories?
Enough that you stop being surprised — around fifty is usually the point where new responses start fitting existing patterns. Building categories before reading imports your assumptions and then finds them.
What do I do when new responses arrive after coding?
Either re-read everything against the current scheme or state plainly that early and late responses were coded differently. The second is acceptable if disclosed and indefensible if not, which is why theming on arrival is worth the setup.
What is the difference between cleaning and analysis?
Cleaning decides what counts as an answer and makes comparable things comparable. Analysis decides what the answers mean. Doing them in one pass is how findings get smuggled in as data-tidying decisions that nobody reviews.
Why does the denominator matter so much for open text?
Because open questions have far more legitimate non-answers than closed ones, so the gap between "responses received" and "people who answered this question" is wide. Report a rate without naming which one you divided by and the number is uninterpretable even when it is correct.
The uncomfortable part of this chapter is that most of the damage happens in the step nobody documents. Theming gets scrutinised, sampling gets scrutinised, and the find-and-replace someone ran at four in the afternoon to tidy up a messy column does not — even though it is the step that decided who counted as having answered.
Next: Analyze multilingual feedback