This chapter resolves check 06 Documents of the eight checks.
A programme director is asked a simple question in a board meeting: which of our partner organisations told us that staff turnover was their binding constraint? She knows the answer exists. She read it herself, eight months ago, in a partner's strategic plan — around page twelve, a paragraph explaining that they could not expand because they could not keep case managers. She also knows exactly where the document is: a shared drive folder called "Partner Docs 2024", alongside forty-one other PDFs.
She cannot answer the question. Not because the evidence is missing, but because it is in a folder, and a folder can tell you a file's name and nothing about what is inside it.
Where do documents belong in an impact measurement system?
Attached to the record they describe — the partner, the participant, the site — and citable down to the passage, not just the file. A folder organises documents by name, which makes them findable if you already know which one you want. That is the opposite of what analysis needs. The value in a forty-page strategic plan is one paragraph you did not know was there, and reaching it requires the document to be part of the record it belongs to rather than a separate object sitting next to it.
Documents are not just long surveys
It is tempting to treat a document as a big open-ended answer. It behaves differently in four ways, and each one changes how you handle it.
| A survey you designed | A document someone sent you |
|---|
| Asks your questions | Answers questions you did not think to ask — which is its whole value, and why you cannot plan for what you will find |
| Arrives when you scheduled it | Arrives on the author's schedule: at a grant deadline, after a board meeting, when a crisis has already passed |
| In your format, one row per person | In their format: a scanned annual report, a slide deck, a transcript with no headings, a spreadsheet inside a PDF |
| A direct answer from the person | A prepared statement. It says what its author chose to say, to the audience they had in mind |
That last row is the honest limit of the whole method, and it is worth stating before anything else. A document is evidence that a claim was made. It is not proof that the claim is true. A partner's annual report saying they served 1,400 people is a record of what they reported, and treating it as a headcount is a decision you should make deliberately rather than by accident.
A finding from page 12 has to be able to cite page 12
The test for whether a document is really in your evidence base is whether a claim drawn from it can point back to the exact passage it came from. Not the file. The passage.
This matters for a practical reason more than a purist one. Someone will eventually challenge the finding — a board member, an auditor, a partner who disagrees with your characterisation of their programme. If your answer is "it's in their strategic plan somewhere", the finding does not survive the challenge, and the effort you spent reading forty documents is written off. If your answer is "page 12, third paragraph, and here is the sentence", the conversation is about the substance instead.
Citing at passage level also protects you from a subtler failure. When people cite whole files, they paraphrase from memory, and the paraphrase drifts each time it is repeated. "Staff turnover is their binding constraint" is a summary. What the partner wrote was that they had lost four case managers in eighteen months and had paused two new sites as a result — a sharper and more useful statement than the summary that replaced it.
How to do this without any particular software
You can build a working document evidence base with two spreadsheet tabs and a filing rule. It is worth doing by hand once, because the manual version makes the requirements obvious.
- Give every document a home record before you file it. Decide which entity it describes — this partner, this participant, this site — and write that record's identifier down with it. A document that belongs to no record is a document you will not find again.
- Keep a document register. One row per document: record identifier, document title, who wrote it, date, type (strategic plan, transcript, annual report, monitoring visit note), and how you received it. This is the tab that replaces the folder.
- Keep a separate excerpt table. One row per passage that matters: record identifier, document, page or timestamp, and the passage quoted exactly. Never paraphrase into this column. The quote column is your evidence; a summary column can sit beside it.
- Mark what kind of statement each excerpt is. Two values are enough: the author's own claim, or something you have corroborated elsewhere. A number from a partner's report is the first kind until you have checked it against something independent.
- Note what the document does not say. A one-line absence note per document — "no mention of waiting lists", "no disaggregation by site" — is often the most valuable row you will write, because absences are invisible when you search by keyword.
- Read for the register, not for the report. Read each document once, on arrival, and record excerpts then. Reading forty documents in the week before a report is how the length-and-recency bias gets in.
Run that on twenty documents and you have something a keyword search of a drive cannot give you: a set of passages, each attached to a record, each traceable to a page.
Where it breaks
Three failure modes, and they arrive in this order.
The register drifts behind the inbox. Documents arrive faster than anyone reads them, because they arrive when the sender is ready rather than when you have capacity. Four weeks of grant reporting deadlines produce a backlog, the backlog gets skimmed, and the excerpt table quietly becomes a record of the documents one person had time for.
The format defeats the reading. Page 12 of a scanned annual report has no searchable text at all. A number inside a table inside a PDF does not come out as a number. A ninety-minute interview recording has the explanation at minute thirty-eight and no headings to get you there. These are not edge cases; in partner and grantee reporting they are the normal condition.
Two documents disagree and nobody notices. A partner's annual report says 1,400 people served; the grant report for the same year says 1,120. Both statements are in your possession. Neither is wrong on its face — different periods, different definitions of "served" — but the discrepancy is only visible if the two documents are attached to the same record and someone is looking. In a folder they are two filenames eleven rows apart.
What a system does about this is specific and limited: in Sopact Sense a document is attached to the record it describes rather than to a folder, its contents are read so that a finding cites the passage and page it came from, and the same themes that run across your survey responses run across the document text, so a partner's written explanation sits next to their reported numbers instead of in a different place entirely. It does not decide which claims to believe. A document is still its author's account, and judging it is still work for a person who knows the programme.
How to test this on your own documents
Use: Five real documents of different kinds — one long strategic plan, one interview transcript, one scanned report, one slide deck, one document that contradicts a number you already hold. Then ask a question whose answer is buried in the middle of one of them.
Pass: The answer comes back with the document, the page or timestamp, and the passage as written. Each document is visibly attached to the partner or participant it describes. The contradiction between the two numbers is surfaced rather than averaged.
Fail: The answer names a file but not a place in it. Or the scanned report returns nothing and no one is told it returned nothing.
Include the scan deliberately. A document set with no scans, no slides and no transcripts is not your document set.
Frequently asked questions
Do we have to read every document?
No, but you should know which ones you have not read. An unread document is not neutral — if it is in your evidence base and you cite the partner it belongs to, you are implicitly accepting whatever is in it. A register with a blank excerpt column at least tells you where the unexamined claims are.
What if a document contradicts our survey data?
Keep both and record the disagreement rather than resolving it silently. Most contradictions turn out to be definition differences — a different period, a different meaning of "completed" — which is useful information about your measures. Averaging the two numbers, or quietly preferring the one that suits the report, is the failure to avoid.
Can a document be the source for a reported metric?
It can, if you say so. Record that the figure is as-reported by its author and has not been independently verified. That single label is the difference between a defensible number and one that collapses the first time someone asks where it came from.
How do we handle transcripts of people who were promised confidentiality?
Attach the confidentiality status to the record, not to a note in a separate file, and decide before analysis what may be quoted and at what level of detail. A transcript is the single easiest way to identify someone accidentally, because the identifying detail is in the story rather than in a name field. This is covered properly in what the assistant may see.
What about scanned or handwritten documents?
Scans need their text recognised before anything can cite a passage inside them, and recognition on a poor scan or handwriting is imperfect. The rule that matters is that a failure has to be visible: a document that could not be read must be marked unread, not returned as empty.
Isn't a keyword search of the shared drive good enough?
It finds documents containing a word, which is a different task from answering a question. It cannot tell you how many partners raised an issue, it cannot connect what a partner wrote to what they reported numerically, and it cannot show you the thing you did not know to search for — which is the reason the document was worth having.
How long should an excerpt be?
Long enough to stand on its own when read cold, months later, by someone who has not read the document. Usually two to four sentences. Anything shorter tends to lose the condition attached to the claim, which is normally where the meaning lives.
The thing worth noticing at the end of this is where the documents are kept. Survey data usually sits in something with rules — required fields, defined measures, someone who owns it. Documents sit in a folder, and folders were designed by people who wanted to retrieve a file by its name. Your richest evidence has ended up in the part of your organisation with the weakest rules, and not because anyone decided it should. It is just where files go.
Next: Clean open-ended responses at the source