Qualitative data is the non-numerical evidence from words, images, and observation. Definition, four types, six characteristics, and eight worked examples.
Qualitative data is non-numerical evidence that describes qualities, experiences, and meaning — interview transcripts, open-ended survey answers, observation notes, documents, images, and recordings. It answers why and how something happened, where quantitative data answers how many and how much. It is the half of the evidence that explains the other half.
Qualitative data is rarely scarce. Most organizations hold far more of it than they use: recordings nobody transcribed, open-ended answers nobody coded, notes filed after a session and never opened. The problem is not collection — it is that the evidence arrives in fragments, in different tools, unattached to the person who gave it, so it never becomes something you can count on.
Key takeaways
The reason qualitative data underdelivers is structural, not analytical. The interview lives in a transcription tool, the open-ended answers in a survey platform, the case notes in a case system, and the demographics in a CRM. Each fragment is meaningful on its own and none of them is connected to the others, so a theme can be quoted in a report but never counted across a cohort or broken out by who said it.
Sopact calls that the Fragment Problem: qualitative evidence scattered across tools and detached from the participant who produced it. The fix is a persistent participant ID from first contact and a codebook fixed before collection, so every fragment lands on one record and is themed as it arrives. The seven methods that produce this evidence are on the qualitative data collection methods page, and the questions that elicit it on qualitative questions.
Once the fragments are joined, qualitative data stops being the part of the report that gets skimmed. A theme carries a count, a segment, and a verbatim line, which is what lets it sit beside a number rather than beneath it.
Qualitative data falls into four types: textual (interviews, open-ended answers, case notes), observational (field notes, behavior records), audio-visual (recordings, photographs, video), and documentary (reports, policies, existing records). The type decides how you capture it; it does not change what has to happen next.
Whatever the type, the same two things determine whether it becomes evidence: is it attached to a person, and is it read against a consistent codebook. The table pairs each type with an example and what it takes to make it analyzable. Running it at scale in a survey is covered on qualitative survey.
Qualitative data becomes evidence at the moment it is attached to a participant and themed against a fixed codebook — not when it is collected, and not when someone finally reads it. The stage below shows the same workflow run the usual way and run as a loop.
The comparison that matters is not manual versus automated; it is whether the reading keeps pace with the collection. Choosing the software that does this is on qualitative data analysis software, and the difference from numeric evidence on qualitative vs quantitative.
Qualitative data comes in four types — textual, observational, audio-visual, and documentary — and each needs the same two things to become evidence: a person attached to it, and a codebook applied to it. Read the last column.
| Type | Examples | What makes it analyzable |
|---|---|---|
| Textual | Interviews, open-ended answers, case notes | Themed against a fixed codebook on arrival |
| Observational | Field notes, behavior records | A protocol, so two observers record comparably |
| Audio-visual | Recordings, photos, video | Transcribed and bound to the participant |
| Documentary | Reports, policies, existing records | Coded against a rubric, with the source cited |
Read the last column and the types converge on one requirement: attached to a person, read against a consistent rule. That is what turns four kinds of raw material into evidence, and its absence is the Fragment Problem.
Fifty transcripts collected in March and read in November explain a cohort that has already left. Reading each response as it lands means the theme is available while the program can still act on it, and the codebook holds instead of drifting. That is the premise of the Loop, Sopact's method for continuous impact intelligence: collect clean at the source, analyze the moment data arrives, improve while you can still act.
The Loop is also what makes a qualitative claim defensible. Every theme count traces back to the response it came from, so a percentage resolves to the people who said it. That standard has its own chapter in traceability and transparency.
One method, three moves that never stop
Then the cycle runs again, a little sharper each wave. Read the method: the Loop methodology →
The fastest way to solve the Fragment Problem is to code one batch against a fixed codebook. Each prompt below pastes into Sopact Sense's Assistant, or reasons through with your team; the arrow above each links the Academy walkthrough that shows the expected output and the tips.
Academy walkthrough → Fix the codebook before you code
Draft a codebook for this qualitative data from our framework and a sample: [PASTE FRAMEWORK + 10-15 RESPONSES]. For each code give a short name, a one-line definition, an include-when rule, an exclude-when rule, and an example quote. Keep it to 6-10 codes and flag overlaps where two codes would catch the same sentence.
Academy walkthrough → Theme it on arrival
Theme this batch against the codebook, one row per respondent: [PASTE CODEBOOK + RESPONSES with respondent_id]. Return respondent_id, assigned theme(s), sentiment, and the percentage distribution across the batch. Keep the codebook fixed; only add NEW_THEME if more than 5% fit nothing.
Academy walkthrough → Attach every fragment to a person
Review how we collect qualitative data today: [PASTE SOURCES AND TOOLS]. For each source, say whether a response can be traced to a specific participant and to their quantitative record. Flag every fragment that arrives detached, and give the fix that would attach it at collection rather than afterward.
Academy walkthrough → Give the theme a segment
Using this themed dataset with demographics on each record: [PASTE], show the theme distribution by [SITE / GENDER / COHORT], report where a theme appears in one subgroup but not another, and cite the strongest verbatim line for each difference.
Each walkthrough is short and practical: what to do, the prompt to run, the output to expect, and the tips that keep it reliable.
Watch: Unified Qualitative Analysis | What Changes Everything.
Qualitative data is non-numerical evidence describing qualities, experiences, and meaning — interview transcripts, open-ended survey answers, observation notes, documents, images, and recordings. It answers why and how, where quantitative data answers how many. In Sopact's framing, it only becomes evidence once it escapes the Fragment Problem: attached to a participant and read against a fixed codebook.
Four types cover most of it: textual (interviews, open-ended answers, case notes), observational (field notes, behavior records), audio-visual (recordings, photos, video), and documentary (reports, policies, existing records). The type decides how you capture it; all four still need to be attached to a person and coded consistently to be analyzable.
Examples include an interview transcript about why a participant left a program, an open-ended survey answer explaining a low confidence rating, a facilitator's field notes on group dynamics, a photograph documenting a site condition, and a policy document coded against a rubric. Sopact keeps each example bound to the participant it came from so it can be counted, not just quoted.
Quantitative data is numerical — counts, ratings, measurements — and answers how many and how much; qualitative data is non-numerical and answers why and how. Most real evidence needs both, ideally captured together so a rating and its reason sit on one record. The fuller comparison is on the qualitative vs quantitative page.
Define a codebook before collection, theme each response against it as the response arrives, keep sentiment and the verbatim line, and disaggregate the themes by segment. Hand-coding does not scale past a few dozen responses, and an unconstrained model re-guesses each run. Sopact themes on arrival against a locked codebook, which keeps the analysis reproducible.
Through seven established methods — interviews, focus groups, open-ended surveys, document analysis, observation, case studies, and ethnography — covered in depth on the qualitative data collection methods page. The decision that matters more than the method is whether each response binds to a participant at collection, which is what Sopact's persistent Contact ID does.
Because it arrives in fragments across different tools, detached from the person who gave it, so reading it is a separate project nobody has time for. Fifty transcripts nobody coded contribute nothing to a finding. Sopact calls this the Fragment Problem and solves it by binding every fragment to one participant record and theming it as it lands.
It is as reliable as the method applied to it. A fixed codebook applied the same way to the same data returns the same distribution twice, which is what makes a qualitative finding defensible; ad-hoc reading does not. Sopact fixes the codebook and keeps every theme traceable to its source sentence, so a qualitative claim can be checked rather than trusted.
Next: see the seven methods on the qualitative data collection methods page, or compare the tools on the qualitative data analysis software page.