play icon for videos

NPS Verbatim Analysis: Turn Comments into Useful Decisions

Analyze NPS comments with clear themes, score context and source evidence. Compare groups, review coding, prioritize issues and track change without losing the original feedback.

Updated
September 15, 2026
360 feedback training evaluation
Use Case
Customer experience · Practical guide

NPS Verbatim Analysis: Turn Comments into Useful Decisions

Analyze NPS comments with clear themes, score context and source evidence. Compare groups, review coding, prioritize issues and track change without losing the original feedback.

Read the guide ↓

What is NPS verbatim analysis?

NPS verbatim analysis means analyzing the written comments that accompany a Net Promoter Score response. You group comments into defined themes, retain the supporting passages and compare those themes with respondents’ scores and relevant context. The useful result is an explanation the team can investigate: which experiences respondents describe, who reports them, and what deserves attention.

A score of 4 might accompany a complaint about slow setup, an unresolved support request or a missing feature. Those are different problems with different owners. Counting the score does not distinguish them. Reading the comment helps, but a comment still gives one person's account; it does not establish the cause of a score or predict that person's next action with certainty.

This guide follows the work from preparing comments to reviewing themes, comparing groups and checking whether a response helped. For the broader method, see survey analysis. For platform selection, use feedback analytics software.

Keep the score and the reason together

NPS uses a 0–10 recommendation rating. Promoters score 9–10, passives 7–8 and detractors 0–6. NPS is the percentage of promoters minus the percentage of detractors, using valid rating responses as the denominator. These are the standard scoring groups described by Bain.

The comment analysis needs its own denominator. If 500 people provide a valid rating and only 280 write a usable comment, the NPS calculation covers 500 responses. A theme percentage among comments covers 280, unless you explicitly choose another population. Do not describe the comments as the views of all respondents.

Keep a response identifier, rating, original comment, question, date and survey period together. Add account, location, product or relationship stage when they are relevant and permitted. An anonymous survey may support group analysis but cannot support named follow-up. Do not promise anonymity and later try to identify the author.

Prepare comments without losing the original evidence

Retain the original text before cleaning. Flag duplicates, empty comments, test records and responses that need review. Keep the reason for an exclusion. A short comment can be useful; an uncomfortable comment is not a quality failure.

Distinguish repeated submissions from separate experiences. A customer who responds at onboarding and renewal has contributed two observations. A duplicated import of the onboarding response should not count twice. Establish the response and period rules before computing theme totals.

For several languages, retain the source language and any translation used during review. Test whether the same definitions work across languages, including idioms and product terminology. A translated sentiment label should not replace examination of ambiguous wording.

If different branches ask different follow-up questions, choose a limited shared set of fields and comparable definitions. Local teams can ask additional questions. Compare only the portions that address the same concept in a compatible way, and make the local differences visible in a data dictionary.

Build a codebook that explains what a theme means

A codebook defines each category, what belongs in it, what does not, and examples. Start with a varied sample across scores, periods, languages and customer groups. This is a starting point for definitions, not permission to assume the sample contains every issue in the full dataset.

ThemeIncludeKeep separate
Delayed first responseThe customer waited too long for an initial replyA quick reply followed by a slow resolution
Slow resolutionThe issue remained unresolved longer than expectedDifficulty finding the contact channel
Setup difficultyInstructions, configuration or initial adoption prevented useA missing feature after successful setup
Helpful supportThe customer describes useful assistanceA positive rating with no supporting comment

One comment can receive several codes. “The team was helpful, but it took a week to fix access” can describe helpful support and slow resolution. Keep the issue and its evaluation distinct: the word “support” alone does not tell you whether the experience was good or bad.

Include a way to surface new or uncertain themes. Examine uncategorized comments and disagreements regularly. When you change a definition, assign a version and decide which earlier records need reprocessing. Comparing counts made under materially different definitions can make a coding change look like an experience change.

Apply the codes and review the difficult cases

Apply the approved definitions across the intended evidence set, preserving the supporting text for each classification. Automation can reduce repeated application work, but the review plan still matters. Inspect a varied sample, uncommon categories, contradictory comments and cases that could lead to consequential action.

Two reviewers can independently code an overlapping set, discuss differences and improve the definitions. The goal is not to force agreement by removing nuance. It is to understand where a category is unclear and whether the coding is consistent enough for the decision.

For automated coding, record the configuration and codebook version. Check both incorrect assignments and missed relevant comments. A system can appear accurate on common themes while missing a smaller group's recurring problem. Keep those cases visible rather than forcing every response into the nearest popular category.

Compare themes across scores and customer groups

Begin with detractors, passives and promoters, then compare relevant segments such as tenure, product or service stage. Do not assume all detractors plan to leave or that passives are always the easiest group to improve. Read what each group actually describes.

A useful comparison shows the count, eligible denominator, proportion and supporting comments. When respondents can mention multiple themes, theme percentages can add to more than 100%. State that explicitly. Avoid ranking groups with very small bases as though differences were stable estimates.

Fictional exampleDetractor commentsPromoter commentsWhat to examine
Delayed first response18 of 60 · 30%3 of 90 · about 3%Which service stages and channels appear in the comments?
Setup difficulty12 of 60 · 20%9 of 90 · 10%Is the problem concentrated among new customers?
Helpful support8 of 60 · about 13%36 of 90 · 40%What worked, including for customers whose overall score was low?

These invented counts illustrate a reporting structure, not a Sopact customer result. They identify associations worth investigating. They do not prove that support speed caused the score difference. Customer mix, issue complexity and response patterns may also differ.

Sentiment can help organize a large set, but it is not interchangeable with the recommendation score. A promoter may describe a serious problem; a detractor may praise an employee. Preserve that tension instead of adjusting the text interpretation to match the number.

Turn a theme table into a decision

Frequency is one input to prioritization. Add severity, affected groups, whether the issue is ongoing, the team's ability to respond and the evidence still needed. A rare accessibility or safety concern may need attention even when a common convenience complaint has more mentions.

For each priority, record the theme definition, evidence, responsible team, proposed response and review date. Separate a customer follow-up from a system improvement. Resolving one person's access issue is useful, but it does not establish that the onboarding process has improved for everyone.

Use quotations to make a pattern understandable, not to substitute for the pattern. Include contrary examples when they change the interpretation. Redact identifying details in wider reports and share individual records only with people authorized to see them.

Check what changes in the next survey period

Use compatible questions and theme definitions to compare periods. If the survey reaches different customers each time, describe a change in the responding population. If the same customers respond again, report matched change separately and show how many did not return.

A lower complaint share may reflect an improvement, different respondents, fewer comments or a changed codebook. Review those possibilities before attributing the movement to an intervention. Link to operational outcomes such as repeat contacts or cancellations only when the records and permissions support that comparison.

Keep the intervention date and the analysis period clear. For example, a setup change introduced halfway through a month cannot explain every response in that month's report. This is the practical reason to retain dates, relationship stage and source context alongside the text.

Where Sopact changes the work

Sopact's approach brings collection, record context, qualitative coding and quantitative analysis into the same workflow. The team defines what a theme means; configured analysis applies those definitions to eligible responses while retaining the original evidence. Ratings, comments and relevant history remain connected so a follow-up question does not require rebuilding the joins each time.

The useful difference is the work removed between definition and review. When a code changes, reprocess the affected scope and review the result rather than recoding every comment manually. When a manager asks what low-scoring customers said, inspect the linked comments and calculations rather than exporting two separate reports and reconciling them.

Other analysis products can also automate themes and connect contextual fields. Evaluate the complete workflow with your data: who maintains the definitions, how revisions are handled, how exceptions are reviewed, which records are included and how easily a result can be traced. A faster summary alone is not enough.

Measure total ownership effort across setup, coding, recoding, data joins, review, reporting and ongoing maintenance. The calculator below makes those assumptions visible. Its example can show hundreds of hours of difference, but it is an illustration; your savings depend on volume, existing automation and the review effort required for comparable quality.

Frequently asked questions

Is a word cloud enough for NPS comments?

A word cloud can suggest recurring vocabulary. It does not by itself distinguish praise from criticism or define a comparable issue category. Use coded themes, supporting passages and score context when deciding what to investigate.

Should every comment have only one theme?

No. A comment may describe several experiences. Allow multiple codes when the definitions justify them, and explain that theme percentages may exceed 100%. Count distinct respondents or responses consistently for each question.

How can we analyze thousands of comments?

Define the codebook, apply it with suitable automation, retain source passages and review a planned set of results and exceptions. Recheck coverage when new themes appear. Automation reduces repetitive work; it does not remove responsibility for the definitions or conclusions.

Can NPS comments predict churn?

Comments may reveal issues associated with later cancellation, but that relationship needs testing against actual outcomes and appropriate time periods. Do not label an individual as certain to leave based solely on their score, sentiment or one comment.

What should the final report include?

Include valid rating and comment counts, the NPS result, theme definitions and versions, group comparisons, missing evidence, representative passages, limitations, action owners and the next review date. Make the supporting records available at the appropriate permission level.

See how connected coding changes the workload

The expensive part is often what happens after the first analysis: a better definition, another collection cycle or a new question that requires the numbers and coded text to meet again.

A workflow with repeated manual work

  1. Define from an initial sampleRead material and agree on the codebook.
  2. Apply it across the datasetCode responses and check the result.
  3. Revise a definitionReturn to affected material and recode it.
  4. Reconnect the numbersReconcile coded results with ratings and context, then rebuild the view.

The Sopact workflow

  1. Your team owns the definitionsDecide what each code means and improve it as you learn.
  2. Apply coding across the eligible dataAutomate application; people review quality and exceptions.
  3. Reprocess after a definition changesReapply the revised definition across the configured scope instead of recoding each response by hand.
  4. Ask across coded text and numbersKeep the response, rating and relevant record context connected; inspect the evidence behind the result.

This compares workflow patterns, not a claim that every research tool requires manual coding or separate files. Some already automate parts of this work; compare the complete cycle.

The codebook can improve without another manual coding project

A sample helps a team develop its first definitions. It should not become an unspoken limit on what the final analysis considers. When thousands of later responses introduce something new, the team needs a practical way to improve the codebook and revisit earlier material.

Sopact's approach keeps that judgment with the team and automates application and reapplication across the configured data. The saving is the repetitive coding and reconnection work. Definition design, quality review, exceptions and interpretation still take time; processing and review are not free or instantaneous.

This applies to a codebook-based operational workflow. It is not a claim that every qualitative research method should use a fixed codebook or that all responses must identify a person. Use the appropriate response, account or participant relationship and respect access restrictions.

Reliable agentic analysis needs a route back to the data

An assistant should turn a question into a checkable operation on the data: apply the intended filters, calculate over the selected records, and return evidence that the reviewer can inspect. It should not invent a count from a generated summary.

1. Define the query“Show lower ratings with comments about delayed first replies.” Keep the scale, period and coding definition explicit.
2. Calculate from recordsFilter the appropriate dataset and count matching responses. Keep the denominator and review state visible.
3. Open the evidenceInspect the matching ratings and original comments, with only the identity and context the reviewer may access.

The reliability test is specific: can you reproduce the calculation on the same data and definitions, and inspect why a record was included? That does not mean an AI interpretation is infallible or that a codebook guarantees identical model output.

Try a definition change · fictional records

Same question. A sharper definition.

Query: ratings of 1–2, with a comment coded for a support delay. The denominator is six submitted responses.

V1 includes delays in either replying or resolving the issue.

3 of 6 responses match

Illustrative workflow, not a live Sopact session. These six synthetic records have predefined coding under each version. In real work, reprocessing and review must finish before the revised result is treated as ready.

The ownership cost is the recurring work

Include implementation and staff time, repeated coding, revision checks, data joins, reporting, platform and processing costs. A low license cost does not tell you how much capacity the workflow consumes.

Illustrative labor model—not a measured customer result or a guaranteed saving. The example below processes the same volume in both workflows. It does not claim that a manual team actually read only a sample, and it includes continuing human review in the Sopact scenario.

Estimate the annual staff hours

Adjust setup, review and reporting assumptions

Manual other work: 6 review + 6 join/reconciliation + 4 reporting hours. Sopact human work: 12 review/exception + 4 revision validation + 1 integration check + 4 reporting hours. Setup includes initial configuration and definition work. Replace these assumptions with observed effort. The reread share can exceed 100% if several revisions require repeat passes.

Scroll horizontally to see all columns →

Annual laborManual workflowSopact scenario
Setup16 h24 h
Initial manual application200 hAutomated; processing costs separate
Manual reapplication100 hAutomated; validation included below
Other human work64 h84 h
Total staff hours380 h108 h
272 fewer hours

In this illustrative annual scenario. Your result may be smaller, larger or negative.

See the calculation and excluded costs

Manual hours = setup + cycles × [(responses × minutes ÷ 60) × (1 + reread share ÷ 100) + other manual hours]. Sopact scenario = setup + cycles × human-work hours.

This estimates staff time only. Complete ownership cost also includes your actual platform, processing, storage, integration, procurement and training costs where not already counted. Translate hours into labor cost using your own rates. Measure processing delay separately. Do not add the same expense twice.

If another tool already automates coding, revisions or joins, reduce the manual baseline accordingly. Long interviews, complex codes or intensive review need different inputs. Comparable quality and coverage are conditions of a useful comparison.

Compare the complete cycle on your own data

Use a representative dataset, agree the coding definitions, introduce a meaningful revision and ask a question that combines a code with a rating or outcome measure. Record the staff hours required to get a reviewed answer, including corrections and rework. That test makes the ownership argument concrete.

The operational benefit is capacity: the team can revisit a better definition and ask another question without automatically starting another coding-and-joining project. Faster processing is valuable only if the evidence and review remain trustworthy.

Watch: why qualitative analysis stays small

This Sopact video explains the repeated work of applying and revising a codebook. Use it to evaluate your own analysis workflow.

Explore Connected Data Intelligence →