Sopact is a technology based social enterprise committed to helping organizations measure impact by directly involving their stakeholders.
Copyright 2015-2026 © sopact. All rights reserved.
Most training surveys only measure Level 1 — did people enjoy it. Here is how to map every question to Kirkpatrick's four levels, flag what you don't cover, and fill the gaps with Sopact Sense.
In short: All four Kirkpatrick levels run on ONE persistent participant ID, so reaction, learning, behavior, and results stay on the same record instead of scattering across four disconnected surveys. You set up four instruments — Level 1 reaction at session end, Level 2 pre/post around the training, Level 3 behavior at 60–90 days, Level 4 results against an organizational metric — and map each to the same ID. Design all four before the cohort starts, and Sopact Sense chains them into one evaluation you can trace from a satisfaction score all the way to a business result.
The whole model only works if the same person's reaction, learning gain, on-the-job behavior, and downstream result live on one record. In Sopact Sense every participant gets a persistent contact ID at intake, and each of the four instruments below writes back to that ID. That is what lets you later ask "did the people who reacted well also learn, apply it, and move the metric?" — a question you cannot answer if Level 1 sits in one form tool and Level 4 sits in a spreadsheet no one joined.
You are the Sopact Sense Assistant working over the Training / Kirkpatrick dataset (clean data + persistent contact IDs). Load my Decision Brief (decision, audience, outcomes, indicators, evidence standard) first, then wait for my task.
Level 1 measures how participants reacted to the training itself, collected immediately at the end of the session while it is fresh. Frame these as instrument design, not a satisfaction dump: each question should map to a driver you can act on — content, facilitator, relevance, pace, and would-recommend. Sample questions to build into the reaction form: "How satisfied were you with this session overall?", "How relevant was the content to your day-to-day work?", "How would you rate the facilitator?", and one open field, "What is the one thing that would have made this session more useful?" Every response lands on the participant's persistent ID so the reaction score can later be correlated with whether they actually learned and applied anything.
Once the post-session feedback is in, run the reaction analysis:
Analyze Level 1 (Reaction) for [COHORT]: code the post-session feedback into themes (content, facilitator, relevance, pace, would-recommend), give a reaction score per participant on their persistent ID, rank the top two drivers of low satisfaction with a representative quote each, and flag any participant whose reaction predicts drop-off.
Expected output. A reaction score per participant tied to their ID, the post-session comments coded into the five themes, the top two drivers of low satisfaction each with a representative verbatim quote, and a flag list of participants whose reaction predicts they will disengage before Level 2 or Level 3. Input: the post-session feedback form on the persistent ID. Output: a scored, themed reaction table plus a drop-off watch list.
Level 2 measures what participants actually learned, captured as a pre/post pair around the training so the gain is the change on the same ID rather than a single after-the-fact score. Build a short knowledge or skills assessment and administer it once before the training and again immediately after. Sample questions to build into it: a scored knowledge check on the core competency, a confidence-to-apply self-rating, and a short scenario item that tests judgement rather than recall. Because both sittings write to the same persistent ID, the pre-to-post delta per participant is the learning gain — the number Level 3 and Level 4 later lean on. This is exactly the measurement covered in how to analyze pre / mid / post survey data, the Level 2 how-to in this chain.
Level 3 measures whether the learning changed behavior back on the job, and it cannot be asked at session end — behavior needs time to show up, so this instrument fires 60–90 days later against the same ID. Design it as more than a self-report: pair a self-rated application scale with a manager or peer verification and one open field on what blocked application. Sample questions: "How often have you used [skill] in your work since the training?", a manager-rated "How consistently is this person applying [skill]?", and "What has made it hard to apply what you learned?" Treat self-report alone as amber and aim for the observed check. The full method — application rate, transfer barriers, and correlating behavior back to the Level 2 gain — is in how to measure behavior change after training.
Level 4 measures whether the behavior change produced an organizational result, and you plan it first even though it fires last. Pick ONE metric the organization already records — retention, promotion rate, error rate, productivity, customer satisfaction — set a baseline, and where possible hold a comparison group. Sample questions are mostly not survey questions at all: they are the metric join. On the participant's persistent ID you attach the recorded outcome (for example a 6-month promotion/retention flag) so the result sits on the same record as their reaction, learning gain, and behavior. Building the board-ready summary that traces the result back through behavior, learning, and reaction — while stating the attribution limits — is covered in how to connect training to organizational results.
Because all four instruments write to the same ID, Sense can read up the chain: a reaction score, a pre/post learning gain, a 60–90 day application rate, and one organizational result, per participant. Grade the coverage so you can see at a glance where the evaluation is solid and where it is thin.
GRADE: green | L1/L2 | reaction + learning on one ID; amber | L3 | behavior self-report at 60–90 days; red | L4 | result not yet joined to the record
Green is Levels 1 and 2 covered — reaction and learning are measured on the same ID. Amber is Level 3 resting on self-report rather than an observed or manager-verified check. Red is Level 4: the organizational result exists somewhere but is not yet joined to the participant record.
Design all four instruments before the cohort starts. Level 3 fires at 60–90 days and Level 4 against a metric that needs a baseline — if you wait until after the training to design them, you have already lost the pre-measure and the baseline. Draft all four forms and the metric join up front, even though they fire at different times.
Reuse the same IDs and codebook across waves. Run cohort after cohort on the same persistent contact IDs and the same theme codebook (content, facilitator, relevance, pace, would-recommend) so waves stack into a trend instead of four unrelated one-off reports. A shared codebook is what keeps "relevance" meaning the same thing in wave 1 and wave 4.
Plan Level 4 first. The organizational result is the hardest to retrofit, so pick the ONE metric and its baseline before the first session. Everything upstream — which behaviors you watch, which knowledge you test, which reaction drivers you code — gets sharper once you know the business result you are ultimately trying to move.
Treat self-report behavior as amber, not green. "Do you use this on the job?" is Level 3 in name only. Grade it amber and add a manager or on-the-job check so behavior is observed, not merely claimed.
Applying the Kirkpatrick model to a survey means running all four levels on one persistent participant ID: a Level 1 reaction form at session end, a Level 2 pre/post assessment around the training, a Level 3 behavior follow-up at 60–90 days, and a Level 4 join to an organizational metric. Because every instrument writes to the same ID in Sopact Sense, you can trace one participant from their satisfaction score through their learning gain and on-the-job behavior to a business result — rather than collecting four disconnected surveys that never join.
Level 1 (Reaction) is collected at session end while the experience is fresh. Level 2 (Learning) is a pre/post pair administered right before and right after the training so the gain is a change on the same ID. Level 3 (Behavior) fires 60–90 days later, once there has been time to apply the learning on the job. Level 4 (Results) is measured against an organizational metric on whatever cadence that metric already reports, with a baseline set before the training begins.
Without a shared persistent ID, reaction, learning, behavior, and results sit in separate tools and can never be joined, so you can never answer whether the people who reacted well also learned, applied it, and moved the metric. Anchoring all four levels to one ID in Sopact Sense turns four isolated surveys into a single chained evaluation you can trace end to end.
Open Sopact Sense, paste your program description, and put it to work.
Try in Sopact