Academy / Foundations / Lesson 4
Additional course links
You leave with: A language review sheet, a before-and-after check of what translation changed, a short software test and a methods note.
Where this fits: In the Lesson 4 review meeting, some of the training team’s 30-day answers arrived in Spanish, and the reviewer asks to see the original beside the translation for any answer that changed a code. This chapter shows how. You bring back a language review sheet and a methods note for your plan.
How do you analyze multilingual feedback?
In short: Keep every original answer on its record, store any translation in its own field beside it, and apply one set of theme definitions tested in each language. Report language coverage and unreviewed answers alongside the findings.
A language list does not show that a tool understands your respondents or local expressions. The test is whether a reviewer can put the original and the translation side by side and see where a code came from. The export-and-paste path breaks that at the first step: the English goes into a chat, the Spanish is gone from what the model reads, and a wrong theme has nothing to be checked against.
Training example · fictional
The spring cohort’s forms were offered in English and Spanish. Six of the 25 replies to the 30-day follow-up were written in Spanish, including three of the eight answers describing what stopped a learner from using the skill. Those three answers sit on the learners’ IDs beside their intake, attendance and exit, like every other answer.
Should you translate before coding?
In short: Sometimes. Coding in the original language, coding reviewed translations and a mix of both can each be defensible. Choose by the languages, the reviewers you have and the cost of an error.
| Approach | When it helps | What to check |
|---|---|---|
| Code in the original language | Competent readers can apply a shared codebook in context | Training, ambiguous terms, agreement between coders |
| Code reviewed translations | The team needs one working language | Lost detail, and access to the original when a code is questioned |
| Combine both | A small team or an AI step needs targeted language review | Which answers were reviewed, and whether errors cluster in one language |
Pew Research Center’s comparison of translation and coding approaches found that results depended on the language and topic. A later Pew methodology documents a mixed approach across many languages. The lesson for a small team is to describe how coding was checked, not to assume one route suits every language.
Why start with the question, not the answers?
In short: Differences can begin before anyone translates an answer. If the question means something different in each language, the answers will too.
“Use the skill at work” might read as “apply it in your job” in one version and “try it out at work” in another. Those are different bars. Pew’s explanation of questionnaire translation describes review, adjudication and pretesting for the questions people answer; keep that work separate from translating their answers afterward.
Record form language and answer language separately; people answer an English form in Spanish, or switch mid-sentence. Do not infer ethnicity, nationality or fluency from an answer’s language. If you reword either version between cohorts, see Change a question safely.
How do you keep the source language beside the translation?
In short: One row per answer: the original, the working translation, the codes, and who reviewed it. A reviewer should be able to go from a theme to the words in either language in one step.
LANGUAGE REVIEW SHEET · FIELDS
A spreadsheet can hold this for a small project. When data is governed at collection, most of it already sits on the record: the original answer, the learner’s ID, the form and its date. If you configure an Intelligence Cell prompt to produce a working translation or codes on arrival, keep its output in its own field so the original is never overwritten. Field selection keeps names and contact details out of what is sent to AI APIs; the answer text and the ID are what the analysis needs.
What did the first pass change?
In short: Compare the original, the working translation and the code for every answer where language could matter. Two of the team’s three Spanish answers needed a different code.
| Original | Working translation | First code | After review |
|---|---|---|---|
| “No me dieron chance de usarlo en el turno.” | “I didn’t get a chance to use it on my shift.” | No time to practise | No chance to use it at work yet: “they didn’t give me the chance”; the translation dropped who |
| “Me da pena hablar frente al grupo.” | “It pains me to speak in front of the group.” | Other | Not ready to use it for real: in much of Latin America “me da pena” usually means embarrassment |
| “Trabajo dos turnos, no hay tiempo para practicar.” | “I work two shifts, there’s no time to practise.” | No time to practise | No change |
Fictional answers written for this exercise.
Before review, the eight barrier answers counted six for no time, two for no chance at work, one for not ready and one “other”. After review: five, three and two, the counts in Clean open-ended answers. The headline barrier held, but the second theme grew by half, and that is the one that points at employers rather than learners. Both errors were in one language. That is the pattern to look for: not whether translation is “accurate” overall, but whether errors cluster in one group and move a finding.
How do you test shared definitions in each language?
In short: Write each code once, with examples and exclusions, then check it against real answers in every language before scaling.
“No chance to use it at work yet” means someone else (a manager, the employer, the tools) kept the learner from using the skill. “No time to practise” means their own workload did. The distinction lives in who is the subject of the sentence, which is exactly what the first translation lost. Add a borderline example in each language to the codebook. Watch negation too: “el supervisor no tuvo problema” (the supervisor had no problem with it) mentions the manager without reporting a barrier.
When reviewers disagree, decide whether the problem is the translation, the definition or the answer itself. Some answers should stay “awaiting review”. A visible open case is more useful than a confident invented reading.
How do you compare language groups without hiding gaps?
In short: Report counts beside any percentage, show coverage by language, and never treat unreviewed answers as “no theme”.
Six Spanish replies are too few for a rate comparison with 19 English ones. Say “three of the six Spanish replies described a barrier” rather than a percentage. Then check coverage: if learners who took the course in Spanish replied less often, their barriers are under-counted in the 25, and the 15 unknowns may not look like the respondents. Who stops responding covers that check.
Shared codes are necessary but not sufficient: question meaning, response rates and reviewer quality matter too. A difference between language groups is a question to investigate, not evidence that language or culture caused it.
What should you test in multilingual feedback software?
In short: Test it on your own answers, including your least-supported language, and check what a reviewer can see, not how fluent the summary sounds.
| Test | What a good result shows |
|---|---|
| Your language mix, including code-switching and local expressions | Errors reported per language, not only overall |
| Source access | Original and translation side by side; the translation never replaces the source |
| Corrections | A changed code updates the count, and the earlier reading is still visible |
| Exceptions | Unknown language, unreadable text and ambiguous answers stay flagged instead of forced into a theme |
| What reaches the model | You choose which fields go to AI; names and contact details can stay out |
| The next wave | A later answer from the same learner joins the same record, with its own language and date |
Count the whole cycle, review time and repeated corrections included, not the translation price alone. For a one-off, a reviewed translation may be enough; if feedback arrives every cohort, test the repeated cycle.
What if nobody on the team reads a language?
In short: Find a competent reviewer who also knows the subject, and until then mark those findings provisional with the number of answers affected.
A trusted language partner under a confidentiality agreement can review a targeted set: every answer that changed a code, every “other”, and a sample of the rest. A translation tool can help sort the pile, but its reading is not settled. Never drop the smallest language group from a report without saying so.
A short explainer on why open-ended answers get set aside when they are cut off from the records they belong to. Watch it with the Spanish answers in mind: the review above only works because each one stayed on its learner’s record. Watch on YouTube ↗
What goes in a methods note?
In short: The languages, how answers were translated and coded, what was reviewed by whom, and what is still open.
For the training team: “Forms were offered in English and Spanish; 6 of 25 replies were in Spanish. Answers were coded with one codebook. A Spanish-speaking reviewer checked every Spanish barrier answer against its working translation; two of three codes changed. Originals are kept beside translations on each record.” Do not claim a representative review if you checked a few convenient answers.
ASK ANY TOOL, INCLUDING OURS
Bring twenty of your own answers in your two most common languages, including one that switches language mid-sentence and one local expression. Ask for themes, then ask to see the original and the translation behind each coded answer. A good tool shows both side by side, keeps the mixed answer whole, and lets you correct a code and see the count change.
Try it on your own data
Open your working evidence plan ↗
- List the form languages and answer languages in your last wave, with counts for each.
- Pick every answer in your least-reviewed language that carries a theme code. Put original, translation and code side by side.
- Mark each code as confirmed, changed or awaiting review, and recount the themes.
- Write a three-sentence methods note and name the reviewer as a role in your Lesson 4 plan.
Check your reasoning
For the training team: 6 of 25 replies in Spanish; 3 Spanish barrier answers reviewed; 2 codes changed; counts move from 6/2/1/1 to 5/3/2. The finding “no time to practise” still leads, but “no chance to use it at work yet” is now a clearer second theme worth raising with partner employers. The six Spanish replies are reported as counts, and the note says whether Spanish-speaking completers replied at a similar rate.
Questions teams ask
Must feedback always be analyzed in its original language?
No. Original-language coding, reviewed translation and mixed approaches can all work. What matters is that the original stays available and a reviewer can check any code against it. Describe your route in the methods note.
What is the best platform for multilingual feedback analysis?
The one that performs acceptably on your own answers and supports review you can sustain. Test your language mix, source access, corrections and what is sent to the model, and check errors per language rather than overall.
Does a common codebook make language groups comparable?
It gives you shared definitions, which you need. Comparability also depends on what the question meant in each language, who was invited and replied, the collection mode and the quality of interpretation. Report counts beside percentages for small groups and show coverage by language.
How should mixed-language answers be handled?
Keep the whole answer, flag it as mixed and check that no part was dropped in translation. If the meaning cannot be established, leave it awaiting review rather than forcing a code.
What if some responses remain unreviewed?
Show how many, in which language, and mark the affected findings as provisional. Define the reporting base so unreviewed answers are not silently counted as “no theme”.