Use the Kirkpatrick Model to decide what evidence you need about reaction, learning, behavior and results. A survey can contribute to that evidence, but four survey forms do not constitute a complete evaluation. Start with the intended result, identify the behavior that could support it and choose suitable collection methods for each question.
This lesson is for training and program teams planning an evaluation. Bring one learning objective and one intended operational result. You will create an evidence plan showing the question, source, record level, timing and interpretation for each relevant level.
Start from the result and work backward
The Kirkpatrick Model organizes evaluation around reaction, learning, behavior and results. Planning from the intended result helps identify the critical behaviors and learning that matter. The levels are not four mandatory forms or a rigid timetable.
For a fictional service team, the goal is fewer incorrectly routed requests. Staff need to classify a request, use the routing guide and confirm ambiguous cases. Training can support that work, while system design, workload and supervisor support also affect the result.
Choose the evidence before writing the questions
| Level | Question in the example | Possible evidence | Limit |
|---|---|---|---|
| Reaction | Was the practice relevant and usable? | Brief feedback and comments | Positive feedback is not proof of learning. |
| Learning | Can staff classify the sample requests? | Reviewed scenario task | A task result is not necessarily workplace performance. |
| Behavior | Do staff use the guide on real requests? | Appropriate observation or work review | Opportunity and working conditions affect application. |
| Results | Did the team’s routing errors change? | Defined operational records | Before/after change alone does not establish cause. |
Keep reaction feedback useful
Ask about the aspects the team can act on: relevance, practice, accessibility or what remains unclear. Choose timing that lets people comment usefully. A short session-end check can help, but it need not cover every concern.
Decide whether identification is necessary. Anonymous feedback may support a delivery decision even though it cannot be linked to an individual’s later performance. Explain the privacy arrangement rather than requiring identity simply to complete an analytical chain.
Do not infer disengagement risk from a low satisfaction score. Review the comment and context, and offer an appropriate route for follow-up where that was agreed.
Measure learning with an appropriate task
Use a method that fits the learning objective. A knowledge check, demonstration or scored scenario can provide evidence of different capabilities. A confidence self-rating is a reported perception; label it separately from demonstrated skill.
If you want to examine individual change, collect comparable observations at suitable points and match the intended records. Check whether repeated exposure to the same assessment or a changed test affects interpretation. Without a baseline you can still describe observed performance, but should not invent the amount learned.
Observe application when there is an opportunity
Plan behavior collection around when people can use the skill. Some behaviors can be observed quickly; others require a later task. There is no universal 60–90-day waiting period.
Record the behavior definition, source, date and opportunity to apply. A manager observation can add evidence but is not automatically accurate or independent. Self-report can be useful when its purpose and limitations are clear. Use more than one suitable perspective when it improves the decision.
Keep “no opportunity yet,” “not observed” and “observed not applying” distinct. They imply different next actions.
Keep organizational results at the right level
A team’s error rate belongs to the team and period. Do not copy the total saving or error reduction to every participant and sum it. Link person-level participation and observations to the relevant team or site context where appropriate.
Choose the result and supporting measures needed for the decision; there is no rule that exactly one metric is sufficient. Check definitions, reporting coverage, workload and other changes before attributing a result to training.
Work through the evidence gaps
A fictional team trains 24 staff. Eighteen complete a later scenario task, and 15 meet the defined criterion: 15/18, or 83.3%, among those assessed, with 18/24, or 75%, coverage. Without a comparable earlier task, this describes current performance rather than a measured gain.
In a separate workplace review, 12 of 16 staff observed using the process meet the behavior criterion. Eight staff were not observed. Keep that result distinct from the assessment. Meanwhile the team’s recorded routing errors fall from 12 in 100 requests to eight in 100 requests. That is a four-percentage-point change in those samples, not proof that the training caused it.
The next decision may be to check observation coverage, review difficult request types and examine whether another process change contributed.
Allow local evidence with a shared comparison core
Across sites, use the few compatible definitions needed for comparison. Local teams can retain different examples, support questions and observation methods where appropriate. Record method differences, and do not combine results unless the interpretation remains defensible.
In Sopact, connect authorized participant observations, learning tasks, documents and team results using their proper identities and dates. Configure analysis around the approved dictionary. Test the joins and calculations before using a combined view. Connected records support review; they do not make every level a person-level survey or establish causality.
Practice: build the evaluation plan
- Write the intended result and critical behavior.
- Choose relevant evidence for each level.
- Specify the unit, source, collection moment and access rule.
- Identify what can be compared and what remains separate.
- Write one limitation and one next action for the fictional example.
Frequently asked questions
Do we need four surveys?
No. Use suitable evidence for the question. Assessments, observations and operational records may answer questions a survey cannot.
Must every level use the same participant ID?
No. Use person identifiers for appropriate individual continuity, but retain team or organizational results at their own level. Anonymous collection may be appropriate for some purposes.
When should behavior be measured?
When there is a meaningful opportunity to apply and observe the behavior. Choose timing from the work rather than a universal delay.