By Sopact · Updated September 12, 2026
SOPACT ACADEMY · PORTFOLIO INTELLIGENCE · ANALYZE
To compare investees whose metrics differ, compare the definitions before you compare the numbers. Pull each investee's field as it is actually written, run four checks on it — source, stage, period, boundary — and sort the comparison into one of three kinds: Identical, Translatable, or Incommensurable. Identical definitions permit a direct comparison, but ranking or aggregation still requires checks on overlap, scale, uncertainty and the decision being made. Translatable ones can be compared only with the conversion rule written down and shown. Incommensurable ones are reported side by side with definitions visible, never averaged and never ranked.
That choice helps determine which comparisons the evidence supports. A portfolio team may already hold a data dictionary that names each field, its type and its allowed values — and then find at roll-up time that the dictionary defines the field but cannot tell them how to read the answer, because the context that makes two answers comparable was never captured next to the field.
Watch the portfolio reporting demonstration · 2 minutes 16 seconds
A Sopact demonstration of connected portfolio records and reporting. The fictional comparisons below are exercises in this article, not claimed results from the video.
How do you run a comparability check across your portfolio?
Take one metric at a time, read each investee's definition as written, run four checks on it, and sort the comparison into one of three kinds before any number is added, averaged, or ranked.
- Pick one metric — not the whole report. Comparability is decided for a stated question, group of observations and reporting period. The same field can support one comparison and fail another.
- Pull each investee's definition as written — from their reporting form, data dictionary, or the sentence in their narrative report. Not the label, not your own restatement.
- Run the four checks — source, stage, period, boundary — and record the value of each for every investee, with the sentence you took it from.
- Sort the comparison — Identical, Translatable, or Incommensurable. Record the kind for this comparison, including its intended use and reporting cycle.
- Write the conversion rule for anything Translatable, naming the common basis and the field each investee must supply.
- Decide what you may publish — a comparison table, a translated table with the rule shown, or a side-by-side view with definitions visible. Review overlap, scale and uncertainty before ranking.
- Record the decision next to the field in the data dictionary, and set the fix-forward ask for the next reporting cycle.
What you will produce
One line per metric: the four check values for every investee, the comparison kind, the conversion rule if there is one, the publishable form, and the specific field each investee must start supplying next cycle. That line lives on the field record — not in a deck.
What four checks decide whether two investee numbers can be compared?
Source, stage, period, and boundary. These four are the context that sits underneath the number; when any of them differs and is not disclosed, the comparison is broken whether or not it looks clean.
| Check | The question it answers | What breaks when it differs silently |
|---|---|---|
| Source | Who produced this number and to what standard — self-reported, management-reviewed, or externally audited? | An externally verified figure and a self-reported figure have different evidence histories. Check the verification scope and disclose uncertainty before comparing them. |
| Stage | At what moment in the program or the investment relationship is the event counted? | Counting at hire and counting at 90 days retained measure different things and produce different numbers from the same events. |
| Period | What window does the value cover — calendar year, fiscal year, trailing twelve months, since inception? | A three-year cumulative total placed next to a one-year total looks like outperformance and is arithmetic. |
| Boundary | Who and what is inside the number, and who is deliberately outside it? | Full-time only against all employment, direct against direct-plus-supply-chain — the boundary usually moves the number more than performance does. |
Stage, period and boundary need to match or have a defensible reconciliation. Source quality also matters: disclose verification differences and consider whether uncertainty prevents the intended comparison. Add population, denominator and overlap checks before ranking or aggregation.
Identical, Translatable, or Incommensurable — which kind of comparison is this?
This working method sorts a comparison into three kinds, with unresolved cases held for clarification. The classification informs the form of reporting; it is not permission to ignore other limitations. Record the reasoning and unresolved information; the review may require clarification from the investee.
Identical
Same substantive definition, period and boundary. Source types may differ; assess and disclose verification differences. Compare directly only after checking population, overlap and uncertainty; matching definitions do not automatically justify a ranking or total.
Translatable
Definitions differ but a stated conversion rule exists and both sides can supply what it needs. Comparable only with the rule written down and shown.
Incommensurable
The underlying question differs, or the missing context cannot be recovered. Report side by side with definitions visible. Never average, never rank.
| Kind | Condition | What you may publish | What you must not do |
|---|---|---|---|
| Identical | All four checks match, or differ only on source and the difference is disclosed. | A direct comparison; totals and rankings require additional overlap, scale and uncertainty checks. | Assume it stays Identical next cycle without re-checking. |
| Translatable | Definitions sit on one shared underlying question, and every investee can supply the common basis. | Converted table with the rule printed beside it and the original values retained. | Apply the rule silently, or drop the original values. |
| Incommensurable | The underlying question differs, or a required check is unrecoverable from what was reported. | Side-by-side values, each with its definition and window printed. | Sum, average, index, rank, or draw one line through both. |
Worked example 1 — two investees each report 40 jobs created
Two identical-looking forties fail three of the four checks, so the portfolio total of 80 is fiction. As reported the comparison is Incommensurable; it may become Translatable if suitable underlying records can be recovered. Otherwise, collect the required detail in future cycles.
Illustrative example. The investees, figures and definitions below are constructed to demonstrate the method; they are not measurements from a client portfolio.
| Check | Investee A — 40 jobs created | Investee B — 40 jobs created |
|---|---|---|
| Source | Self-reported by the founder in the quarterly template. | Externally verified for this specific employment measure in the exercise; a statutory financial audit alone would not establish this scope. |
| Stage | Counted at hire — the job exists on the start date. | Counted at 90 days retained — the job exists once it survives probation. |
| Period | Calendar year 2025. | Since inception — three years of trading. |
| Boundary | Full-time permanent roles only. Part-time, seasonal and contract excluded. | All employment including part-time and seasonal, counted as headcount not FTE. |
| Verdict | Incommensurable as reported. Stage, period and boundary all differ, and none of the three can be recovered from the numbers that were submitted. | |
Read the four checks in sequence and the total falls apart. A's forty covers one year; B's covers three. A's forty counts a job the day someone starts; B's counts only jobs that lasted ninety days — a stricter test that suppresses B's count. A's forty is full-time permanent roles; B's includes part-time and seasonal headcount, which inflates B's count relative to A's boundary. Two of these differences push in one direction and one pushes the other way, so you cannot even say which investee is ahead. Adding them produces 80, a number that describes no population over no window under no definition.
Note also what is not the problem. The source difference — self-reported against audited — is real and matters, but it requires an explicit quality assessment. Disclose it beside the value and consider whether uncertainty limits the intended comparison.
What would move this to Translatable. Additional source records are needed; the reported totals alone are insufficient. Each investee would have to supply:
- Investee B: a calendar-year 2025 cut of the same figure, and a full-time-permanent subtotal alongside the all-employment headcount.
- Investee A: a 90-days-retained subtotal alongside the at-hire count, so both sides can meet at the stricter stage.
- Both: the confirmation moment stated explicitly on the field, so it is not inferred next cycle.
With those three additions the common basis exists: full-time permanent jobs, confirmed at 90 days retained, in calendar year 2025. The conversion rule is written down, the original values stay visible beside the converted ones, and the source difference is printed rather than erased. That is a Translatable comparison — provided the supporting records establish the common basis. If historical records permit a valid restatement, label and retain that restatement; otherwise apply the change prospectively.
Worked example 2 — twelve investees reporting people trained
The same label covering three different cut points on one funnel is Translatable, not Incommensurable — because the underlying question is shared and the exercise assumes each investee can supply the earlier counts. Verify that availability before classifying a real comparison.
Illustrative example, constructed to contrast with the case above.
Twelve investees in a workforce portfolio all report a field labelled people trained. Reading the definitions rather than the labels: four count everyone who enrolled, five count everyone who completed the full course, and three count everyone who attended at least one session. Boundary, source and period match across all twelve. Only stage differs — and it differs along a single ordered funnel: enrolled, then attended, then completed.
That is what makes this Translatable rather than Incommensurable. All twelve are answering one question — how many people did this program reach and carry through? — and for this exercise, all twelve providers have confirmed that complete enrollment records for the same population and period are available. This cannot be assumed for every provider. The conversion rule is therefore available: report all twelve at the widest common cut point, enrollment, and publish completion rate as a separate field for the investees that can produce it. The five who count completions are not penalised; their completion rate becomes a second, more informative column rather than a hidden discount on their headline number.
Compare that with the jobs case. There the underlying question genuinely differed — jobs created in a year against jobs sustained since inception — and the missing cut could not be produced from what had been submitted. One check differing along a funnel both sides can walk is Translatable. Several checks differing on questions that were never the same is not.
Should you benchmark against an external standard or against your own prior period?
An external benchmark tells you where an investee sits in a market; your own prior period tells you whether it is moving. Both references need a comparability check; neither is valid merely because it is external or historical.
Against an external standard. A sector median, an IRIS+-aligned reference series, or a development finance institution's published figures can only be used if the reference publishes its own definition and denominator — and if your investee's field passes the four checks against that published definition, not against its label. If the reference publishes a number but not a definition, you can cite it as context in a narrative; you cannot score an investee against it. Mapping your fields onto published frameworks is its own piece of work, covered in how to map portfolio data to IRIS+, GRI, and ESRS.
Against your own prior period. Here Identical is actually achievable, because the definition is one you control. The condition is that the definition did not change between the two periods — and definitions change more often than teams expect, usually because a form was edited, a field was renamed, or a new cohort type was folded in. Check whether the change affects meaning or only presentation. A renamed label need not break a trend. A changed population, period or definition requires a documented compatibility check: retain the series only where the source evidence supports a valid mapping; otherwise show a series break. Preserve original observations and any restated comparison values separately.
Why is a benchmark only as good as its denominator?
Every rate needs a denominator, and the denominator is where the boundary check usually fails. A rate comparison is more dangerous than a count comparison because the count is visible and the denominator is not.
Take one investee with 140 people placed into work, and four defensible denominators someone might have used to build a placement rate:
| Denominator | Count | Reported placement rate |
|---|---|---|
| Everyone enrolled | 400 | 35% |
| Everyone who completed | 250 | 56% |
| Everyone job-seeking at exit | 200 | 70% |
| Everyone reached at follow-up | 175 | 80% |
Illustrative figures. One numerator, four honest denominators, a 45-point spread — and every one of these rates is defensible on its own terms.
Each rate is meaningful only if the same 140 observed placements belong within its stated denominator; that condition is assumed in this fictional example. The last rate is highest in this example. Nonrespondents may differ from respondents, but the direction of that difference is not established here. Report response coverage and investigate nonresponse before interpreting the rate as representative. The practical rule: never accept a rate without its denominator definition, and when denominators differ across investees, compare the counts and the denominators separately rather than the rates. A benchmark expressed as a rate inherits every weakness of the denominator behind it.
What do you do when an investee cannot supply the missing context retroactively?
Improve collection at the next reporting cycle and mark the historical series as a break. Do not back-fill assumptions into the periods you already have.
- Write down what the historical number does and does not include — in the field's own note, in the investee's own words where you have them, and marked as unrecoverable where you do not.
- Mark the series break at the boundary date. Any chart drawn across that date carries the break rather than a continuous line, and no total combines incompatible observations without a justified reconciliation.
- Change the ask, not the analysis. The fix belongs in the collection form the investee fills in next cycle — the field, the help text, the required subtotal — not in a prompt you run afterwards. Defining the ask so both sides can live with it is covered in how to define measures the organization and funder can both use.
- Report the two sides of the break next to each other, each with its own definition printed, until enough cycles accumulate on the new basis to carry a trend on their own.
The reason not to back-fill is simple and unglamorous: a back-filled assumption becomes indistinguishable from reported data the moment it sits in the column. It survives the analyst who made it, it gets rolled up by someone who never saw the reasoning, and it is defended in an LP meeting by a person who believes it was collected. A break is visible and asks a question. An assumption is invisible and answers one it shouldn't.
Where does the comparability decision belong?
On the field record in the data dictionary, next to the definition — not in a reporting prompt, a slide footnote, or an analyst's head.
A prompt is written per question and per person. Two analysts asking the same portfolio the same question next quarter will phrase it differently, supply different context, and get two different answers — and neither will know the other exists. The decision has to sit where the field sits, so it is inherited rather than re-derived. Add these to the field record:
| Added to the field record | What it holds |
|---|---|
| Comparison kind | Identical, Translatable, Incommensurable, or awaiting clarification, for the specified comparison and reporting cycle. |
| Four check values | Source, stage, period, boundary as stated by each investee, with the sentence each was taken from. |
| Conversion rule | If Translatable: the common basis, the rule, and the field each investee must supply. |
| Decision provenance | Who decided, on what date, against which version of each investee's definition. |
| Review trigger | What forces a re-check — a definition edit, a new investee on the metric, a standards version change, a series break. |
If you do not yet have a field record to attach this to, start with how to build a governed data dictionary and how to define an impact metric so everyone counts it the same way. A shared worksheet is enough to begin: record the comparison purpose, definitions, sources, decision and owner. Link it to the dictionary as that structure develops.
Prompt: run a comparability check on one metric across the portfolio
Use this on one metric at a time, with each investee's definition text pasted in verbatim. It produces the four check values and a proposed kind — a draft for a person to confirm, not a decision.
What safeguards belong on a ranking that moves capital?
A documented comparability review can be easier to inspect and repeat. AI assistance does not guarantee greater accuracy or consistency, and it is not deterministic and must not be the final word when the output changes who gets follow-on capital.
- Cited evidence for every check value. Each of the four values traces to a verbatim sentence and a named document or field. If direct source evidence is unavailable, label the value as an inference or unresolved rather than presenting it as verified.
- Human review and final authority. A person confirms every comparison kind and every conversion rule before anything is ranked. The model proposes; the investment or impact committee decides.
- Documented overrides. When a reviewer changes a proposed kind, record what changed, why, and who — on the field record, in the same place as the original decision.
- Test across the portfolio's real variety. Run the check against investees of different sizes, in different languages, and reporting in different formats — narrative PDF, spreadsheet, structured form. Extraction quality varies by format, and a systematic extraction weakness becomes a systematic ranking bias.
- Show the kind wherever the ranking appears. If a table is Translatable, the conversion rule travels with it. A ranked table that has lost its rule is indistinguishable from an Identical one to the person reading it.
- Access and retention. Definition text is investee material. Record who can read it, how long it is kept, and what happens to it when the investment exits.
Where does Sopact Sense help — and where do people decide?
The whole method above works in a spreadsheet. Add a definition column next to every reported value, four columns for the checks, one for the kind, one for the rule. For a portfolio of five investees and a handful of metrics, that is enough if the team can maintain the evidence and review history reliably.
It stops being enough at scale and across cycles. Thirty investees, twenty-five metrics, four checks each is three thousand judgements — and a poorly maintained process may repeat them each cycle if ownership and the previous reasoning are lost. That is the failure mode: analysis whose reasoning cannot be carried into the next review.
Test the workflow in Sopact Sense with one metric and two authorized investee records. Check that the definition stays linked to the number, that configured analysis can prepare source-linked comparison questions, and that a reviewer can retain the confirmed rule for the next cycle. Verify change detection, permissions and question behavior in your setup rather than assuming every step is automatic. People confirm every kind and every conversion rule. What the system contributes is repeatability and traceability, not the decision.
How this fits the wider portfolio workflow is covered on portfolio data management and impact measurement and management.
Frequently asked questions
How do you compare investees when each one defines its metrics differently?
Compare one metric at a time. Read each investee's definition as written, run four checks on it — source, stage, period, and boundary — then sort the comparison as Identical, Translatable, or Incommensurable. Matching definitions support comparison; ranking also requires attention to population, scale and uncertainty. Translatable can be compared once the conversion rule is written down and shown. Incommensurable is reported side by side with definitions visible, and never averaged.
What makes two investee metrics comparable?
Matching definitions, not matching labels. Two fields are comparable when they count the same event, at the same point in the program, over the same window, within the same boundary of who and what is included. A difference in verification — self-reported against audited — must be disclosed and assessed; in some cases the uncertainty will limit the comparison.
Can you add up an impact metric across a portfolio?
Only when the fields share a compatible definition and unit, their populations and periods are appropriate, and overlap is addressed. Summing values counted over different windows or different boundaries produces a number that describes no population. If the fields are Translatable, convert them to the common basis first and show the rule. If they are Incommensurable, publish the components and their definitions, not the total.
What is the difference between a benchmark and a comparison?
A comparison examines similarities and differences; it need not rank investees. A benchmark places them against an outside reference or against their own prior period. For a benchmark, the four checks must also pass against the reference's published definition and denominator. If the reference publishes neither, you can cite it as narrative context but cannot score an investee against it.
How do you handle an investee reporting since inception when the others report calendar year?
Ask for the calendar-year cut rather than converting it yourself. If they can produce it, the comparison becomes Translatable and you state the rule beside the table. If they cannot, report the since-inception figure separately with its window labelled, change the ask on next cycle's form, and mark the point where the basis changes as a series break.
Should you back-fill missing context for past reporting periods?
Do not fill gaps with unlabelled assumptions. You can reconstruct or correct a past value when source records support it: preserve the original, record the method and evidence, and label the revision. If that evidence is unavailable, improve collection for the next cycle and show the historical gap or series break rather than inventing continuity.
Can AI decide which investees are comparable?
It can extract the four checks from each definition, cite the sentence each came from, and propose a classification — which can support a documented review when extraction and interpretation are checked. It is not deterministic and should not be treated as final. A person confirms every classification and every conversion rule before a ranking is published or capital moves.
What is a series break, and how do you mark it?
A series break marks a change that prevents a valid comparison across the boundary. A change in definition, population or verification requires review; it does not automatically invalidate the series when a documented reconciliation remains possible. Mark it on the field record with the boundary date, what changed, and why. Charts drawn across that date should show the break rather than a continuous line, and totals should not combine incompatible observations without a justified reconciliation.
Where should the comparability decision be stored?
On the field record in the data dictionary, alongside the definition — with the comparison kind, the conversion rule if there is one, the four check values, who decided and when, and what triggers a review. Stored in a reporting prompt or a slide footnote, it disappears; two analysts asking the same question next quarter may get different answers.
Sources and scope
- All investees, figures, and definitions in this chapter are illustrative and labelled as such. They are constructed to demonstrate the method and are not measurements from a client portfolio.
- The four checks (source, stage, period, boundary) and the three-kind classification are Sopact practice drawn from portfolio reporting work, not a published standard. Treat them as a working method to adapt, not a compliance requirement.
- Verification language follows the ordinary distinction between self-reported, management-reviewed, and externally audited figures used in investor reporting. Which one applies to a given investee is governed by that investee's own reporting agreement.
- Where published frameworks are in scope — IRIS+, GRI, ESRS — mapping a field onto them is separate work with its own version-control problem, covered in the standards mapping chapter.
For published guidance on aggregation limitations, see Operating Principles monitoring guidance. This article’s classifications remain a practical review aid, not an official standard.
For the reporting step, use portfolio dashboards and reporting. Carry the comparison rules and unresolved gaps into the display.