This chapter resolves check 05 Qualitative of the eight checks.
A youth employment programme asks one open question at endline: what got in the way? Answers arrive in English, Spanish, Somali and Karen. Overnight, everything is machine-translated into English, and the next morning a coordinator reads the translated column and builds themes. The report says Spanish-speaking participants described transport and childcare in detail, and that Karen-speaking participants "gave brief, general feedback."
Two weeks later a caseworker who reads Karen opens the original column. The answers are not brief. One runs six lines and names a bus that stops running at 4:15 while the shift ends at 4:30. In English it had arrived as "transport was difficult sometimes."
Should we translate feedback before we analyze it?
No. Read and theme each response in the language it was written in, and translate only at the point where you publish a quote. Translating first feels like the tidy move — one column, one language, one afternoon of reading. But machine translation is not equally good at every language, so translating first does not just blur meaning evenly. It blurs some groups more than others, and then your themes report that blur as a fact about those people.
Why the loss is uneven rather than general
Machine translation learns from material that already exists in both languages. For Spanish or Mandarin there is an enormous amount of it. For Karen, Chuukese, Marshallese, Hmong, Mixtec or Haitian Creole there is far less — so the output comes back shorter, flatter and more generic. Specific nouns become categories. Six lines become one. That produces a chain nobody in the room intends:
How a translation step turns into a claim about people
Step oneThe machine is strong on widely-spoken languages and weak on the rest, so answers arrive with uneven richness.
Step twoThe flattened answers carry fewer specifics, so fewer themes attach to them. One vague line supports one vague theme.
Step threeCounted up, that group shows fewer themes per person and shorter answers than the Spanish or English speakers.
Step fourThe report says those participants were less engaged, less specific, less forthcoming — and a programme decision follows from it.
Nothing in step four is dishonest. It is a measurement of the translation, described as though it were a measurement of the participants. That makes it a fairness problem and not only a quality problem: the group whose language the machine handles worst is the group your report describes least favourably.
What you compare when you cannot compare the words
Teams translate because they want to compare groups, and comparison seems to need one shared language. It does not. What has to be shared is the rule, not the language.
Write down what counts as a theme in plain sentences — "counts as a transport barrier: the participant names getting to or from the site as something that made attending harder." Then apply that rule inside each language, by a reader or a tool that actually reads it. You now have counts of the same rule from four language groups, comparable even though no two responses share a word.
What you must never compare across languages is anything about the text itself: length, themes per person, richness, tone. Those are what translation quality contaminates most. Compare them within a language over time — Karen speakers this wave against Karen speakers last wave — and they mean something again.
Doing this without any particular software
- Group responses by the language the person answered in, not the language of the survey. People answer in whatever language they think in, including in a form labelled English.
- Write the theme rules on one page before anyone reads anything. For each theme give two examples of what counts and two of what does not. This page is what travels between languages.
- Assign a reader who reads the language. They tag against the written rules and change nothing in the response. Bilingual staff, a partner organisation, a community interpreter — the requirement is comprehension, not certification.
- Keep the original wording untouched in its own column, and put the tags in new columns beside it. Any translation lives in a third column, clearly labelled as a translation.
- Spot-check the translation, not the people. Take ten responses per language, have your reader read the original and the machine English side by side, and note what went missing. That list of losses is your evidence for why you did not translate first.
- Translate at the end and only what you publish. For every quote in the report, show the participant's own words with the translation underneath, and name who translated it.
- Report counts by rule and by language, and stop there. Never write a sentence comparing how much two language groups said.
Four languages and a few hundred responses is genuinely doable this way. It is slower than pasting a column into a translator, and it is the difference between a finding and an artefact.
Where it breaks
You will not have a reader for every language. The two big ones are covered by staff; the fifth language has eleven responses and belongs to the newest community you serve. Those eleven wait, then get handled by the machine anyway, then get described as thin. The smallest group is always the one that gets translated worst and read last, which is exactly backwards from where the new information is.
One reader per language means no cross-check. With a single Somali reader there is nobody to disagree with, so you cannot tell whether the rules were applied the same way in Somali as in English. When that person leaves, the Somali trend line moves and nobody can say whether the programme changed or the reader did.
This is the narrow place a system helps. Sopact Sense reads and themes each response in the language it was written in, against the theme rules you wrote, rather than translating first — and it keeps the original wording, so every quote can be shown as the participant typed it with a translation beside it. Detail measures are reported inside each language rather than across them. It does not decide whether your theme rules are the right ones, and it does not remove the need for someone who reads the language to check a sample of tags.
How to test this on your own data
Use: One open question answered in at least three languages, one of them a language with few speakers in your caseload. Inside that smallest group, include five answers you know are long and specific. Add two answers that switch between languages mid-sentence.
Pass: The five detailed answers carry as many specific themes as comparable English answers. Every quote in the output shows the original wording. The report gives theme counts by language and makes no claim about which group said more. The mixed-language answers are themed on what they say, not discarded.
Fail: The smallest language group shows systematically fewer themes per person. Or a quote appears in English only, with no original to check it against.
Run it on live responses. A test file translates cleanly, because someone wrote it in one language and then translated it — removing the only thing this check is about.
Frequently asked questions
Is machine translation ever the right tool here?
Yes — at the end, for quotes you are publishing, where someone who reads the language can check the result before it goes out. The problem is never the translator. It is putting the translator ahead of the analysis, where nobody sees what it dropped and every later step inherits the loss.
We only have staff who read two of our five languages. What do we do?
Cover what you can, and be explicit about the rest. Report themes for the languages you read properly, name the languages you could not, and give their response counts. An honest gap is usable. A translated-and-flattened group presented as equivalent is not, because it looks complete.
How do we know the translation caused the difference, and not the participants?
Read ten originals per language alongside the machine English and mark what disappeared. If specific details — a bus number, a supervisor's behaviour, a fee — survive in one language and vanish in another, you have your answer. This takes an hour and settles the argument better than any general claim about translation quality.
Wouldn't it be simpler to ask everyone to answer in one language?
It moves the same bias earlier. People writing in a second language write shorter and plainer, so you get the flattening without the translator — plus the people least comfortable in that language answer least. You have not removed the distortion, only made it invisible.
What about answers that mix two languages in one sentence?
Treat them as their own case and theme what is actually there. Mixed answers often carry the most precise word in whichever language has it — the borrowed term is usually the specific one. Forcing them into a single language bucket is where that detail gets lost.
Does this apply to interpreted interviews as well as written answers?
Yes, with a person in the seat instead of a machine. Summarising interpreters produce shorter accounts, and two interpreters produce different levels of detail from comparable interviews. Record who interpreted each session, so you can tell whether a pattern belongs to a community or to an interpreter.
How should we word this for a funder?
Report theme counts by language with response numbers, and say plainly that detail was not compared across languages because translation quality varies by language. Funders accept that readily. What they cannot use is a sentence claiming one community engaged less, resting on translation nobody checked.
The sentence worth striking from every report is "these participants gave less detailed feedback." It almost never describes participants. It describes the path their words took to reach you — how much translation, done by whom, checked by no one. Detail is not a property of a community. It is a property of your reading, and it can be improved by hiring rather than by asking people to say more.
Next: Find who is missing waves