Before asking AI to write the explanation, write down what the number means. You will leave this exercise with one calculation another person can repeat.
Give the denominator a name.
In the fictional training example, 40 people complete the course, 25 answer the follow-up and 15 report using the skill after 30 days.
15 ÷ 25 = 60% of respondents report using the skill.
25 ÷ 40 = 62.5% of completers responded.
15 people have no follow-up answer. Do not call them non-users.
The team can also report that 37.5% of all completers are confirmed users. That is a different statement, not a replacement for the respondent rate. Keep the definition beside the figure so the reader does not have to guess.
Save enough to repeat the check.
Keep the dated records used, the definition of completion and skill use, the follow-up period, the duplicate rule and the formula. Ask a colleague to calculate the result from that material. Repeating the same answer is useful only if the rules and source records are also correct.
AI can draft a query or suggest a calculation. Inspect what it counts and test it on records with known answers. A number written fluently in a generated paragraph is not a substitute for that check.
Test the awkward cases too.
Try a duplicate response, a missing answer and a corrected record. Write the expected treatment before testing the tool. If a late response arrives, preserve the earlier result and issue a new version with a short explanation of what changed.
For written feedback, keep the theme definition and reviewed source passages. Test clear and ambiguous examples separately from arithmetic. Counting stored, reviewed codes is a different task from asking AI to interpret the comments again.
Before you move on
Give a colleague your figure, population, period, formula and source snapshot. Can they reproduce it—and explain who is missing? Resolve any difference before putting the figure in a report.
For recurring releases, multiple definitions or model changes, the full reliability guide covers versioning and change checks. This shorter revision keeps the first exercise focused.
Practice: reconcile a changed number
This separate fictional exercise uses its own dataset; do not combine its counts with the earlier training example.
Fictional continuation of the training course. The approved first release has 80 starters, 60 people with known employment status 90 days after exit and 36 employed. The known-status rate is 60%, with 75% coverage. Keep that release as a fixed record.
| Release or test | Known status | Employed | Rate among known | Explanation |
|---|---|---|---|---|
| Release A | 60 of 80 | 36 | 60% | Approved original snapshot; 20 unknown |
| Rerun A unchanged | 60 of 80 | 36 | 60% | Same records, reviewed statuses and rules must reproduce the count |
| Release B: four late confirmations | 64 of 80 | 39 | 60.94% | Three employed and one not employed; coverage now 80%; 16 unknown |
| Wrong denominator test | 64 of 80 | 39 | 48.75% of all starters | A different valid view if labeled; not the same measure as 60.94% among known statuses |
The increase from 36 to 39 confirmed employed is new evidence about the same follow-up date. It is not evidence that three people became employed between releases. Confirm the date each late response describes. The employment rate alone also hides the change in response coverage.
Keep a short release record
- Scope: cohort, period, follow-up reference date and inclusion rules.
- Inputs: snapshot ID, source versions and approved status or theme decisions.
- Calculation: saved formula or query, denominator, rounding and exclusions.
- Exceptions: missing, conflicting and late records, with the owner’s decision.
- Change: previous and new values, the records responsible and whether the measures remain comparable.
- Approval: reviewer, release date and where the earlier release is retained.
This is the practical record called a reviewed result package in this lesson; it is not an external certification. Reproducing the same answer does not establish that the underlying data or method is valid. A consistently wrong denominator produces a consistently wrong answer.
Test qualitative coding separately from arithmetic
Keep a small reviewed test set with clear examples, ambiguous passages, multiple themes and missing context. Re-run a proposed coding change against that set. Compare source-level decisions and investigate disagreements before replacing released codes. Set the acceptance criteria for the task and its consequences; there is no universal agreement percentage that makes every qualitative analysis valid.
Once codes are approved, count the stored decisions using the same unit and rule. Re-generating themes from raw text on every request creates a new analysis, even if the underlying interviews have not changed.
When should a stable result legitimately change?
A result may change when evidence, identity resolution, definitions, codebooks, or approved methods change. The system should never overwrite the prior release silently. It should produce a new version and a reconciliation that explains the difference.
| Change | Correct response | What must remain visible |
|---|---|---|
| A late attendance record arrives | Create a new evidence snapshot and rerun the same approved query | Previous value, new value, late record, and release date |
| Two participant records are confirmed as duplicates | Apply the approved identity correction and issue a reconciliation | Merge decision, owner, affected results, and reason |
| A qualitative codebook is improved | Create a new coding version; review material changes before release | Old and new codes, affected excerpts, overrides, and reviewer |
| The definition or denominator changes | Create a new metric version or restatement under an approved policy | Effective date, comparability warning, and authorized approval |
| The underlying AI model changes | Test against reviewed examples; do not silently replace released qualitative decisions | Model/rule version, test results, changed proposals, and approval |
Add this to your plan
Record what you decided in this chapter in the same working evidence plan you started in Foundations, so definitions, sources and checks stay in one place.
