To measure behavior change after training, define the action people should use in their work, collect appropriate evidence when they have an opportunity to use it, and distinguish reported application from observed application. Keep the dates, source and relevant starting evidence so you can assess change without confusing missing follow-up with failure.
In the Kirkpatrick Model, Level 3 concerns critical behaviors in the performance environment, including the support that enables people to carry them out. It is not simply a survey sent on a fixed day after a course. See Kirkpatrick Partners’ explanation of the model.
Define the behavior in observable terms
“Communicates better” leaves too much room for interpretation. A fictional service program might instead ask whether a handoff includes the next action, the responsible team and the expected response date. A reviewer can look for those elements without deciding whether someone is generally a good communicator.
Define what counts as application and what does not. Keep the unit clear: a person, a handoff or an observation period may each support a different question. Decide whether you are assessing one successful attempt, repeated use or an improvement relative to a starting point.
Plan support and observation together
People need a realistic opportunity to use the behavior. The timing of follow-up should reflect that opportunity and the decision the team needs to make. A rare annual task calls for a different plan from a task used every shift.
Kirkpatrick Partners emphasizes early monitoring and support rather than waiting until 90 days to begin examining application. Plan check-ins around actual work so the team can respond while the information is useful. Source: the official Level 3 description.
For the service example, the team might review an early sample of handoffs, ask what made the method difficult and revisit the evidence after a practical adjustment. This is a hypothetical collection plan, not a universal schedule.
Choose evidence and keep perspectives distinct
| Source | What it can contribute | Limit to retain |
|---|---|---|
| Learner account | When the learner says they used the method and what happened | A self-report is not the same as direct observation |
| Work sample | Whether specified elements appear in a handoff | The sample may not represent all work |
| Manager or peer observation | Another perspective on a defined action | The observer may not see every opportunity and may apply criteria differently |
| Operational record | Dates or events relevant to use | A recorded event may not establish the quality of the behavior |
Use the sources suited to the task. A manager rating is not mandatory in every design and is not automatically more accurate than a learner’s account. Record whose evidence it is and review disagreements rather than averaging them into apparent agreement.
Connect the record without changing the privacy promise
Where identified follow-up is appropriate, link observations to the continuing learner record, course episode and relevant dates. Reuse stable registration information rather than asking the person to re-enter it. Keep the observer role and measure version with each observation.
Anonymous feedback can still help explain group-level barriers, but it cannot support a claimed individual learning-to-application history. Choose the design deliberately and explain its limits. Do not retrospectively identify respondents because matching would be convenient.
In Sopact, a continuing record can keep task evidence, learner comments and later observations together where permitted. The data dictionary gives reviewers shared criteria. The collection and review configuration still needs to be tested; a stored identifier does not establish that a measure is valid or that training caused a change.
Calculate an application rate with the right base
Suppose 40 people completed a course. At the review date, 30 are known to have had an opportunity to perform the task. Twenty of those 30 have usable follow-up evidence, and 14 of the 20 meet the agreed application criterion.
The observed application rate among those with usable evidence is 14/20, or 70%. Follow-up coverage among those known to have an opportunity is 20/30, or about 67%. Show both. Do not describe the missing ten observations as non-application, and do not present 70% as a measured result for all 40 completers.
If the opportunity status of other learners is unknown, show that too. If some were not expected to use the task yet, retain that distinction. The reporting base should follow the question and the evidence, not whichever denominator produces the most favorable result.
Investigate barriers to use
Ask what made application difficult. Possible explanations include unclear guidance, no opportunity, unavailable tools, conflicting priorities or uncertainty about the method. Treat them as reported barriers until the relevant evidence or review supports more.
Code comments with definitions and keep the source passages. One person can report several barriers. Frequency does not establish severity, and a lack of comments does not establish that no barriers exist.
Assign an owner to the response. A request for a job aid, approval of the request, delivery of the aid and later use are separate observations. Keep that chain visible so reporting does not mistake an assigned action for a resolved problem.
Read learning and application together
A linked record lets you examine whether people with particular learning evidence later demonstrate the target behavior. First confirm that the starting and later measures mean what you say they mean. A confidence increase is not automatically a learning gain, and an unobserved behavior is not automatically a failure to transfer.
For a descriptive comparison, report the matched group and its coverage. Differences in opportunity, task assignment, support and observation can influence what you see. Do not describe association as proof that training caused the application or that a learner’s lack of application was caused by the course.
Your exercise
- Define one observable behavior and the criterion for application.
- Choose when people are likely to have a real opportunity to use it.
- Select appropriate sources and record their limitations.
- Define the eligible group, usable follow-up and missing states.
- Calculate a hypothetical rate and show coverage separately.
- Assign an owner to investigate a reported barrier.
- Describe the next observation that would tell you whether the response helped.
Frequently asked questions
What is Kirkpatrick Level 3?
It examines critical behavior in the performance environment and the support around it. Use the official model guidance when applying the framework; the practical example here illustrates a collection and review plan.
When should behavior follow-up happen?
Plan it around meaningful opportunities to use the behavior and the need for timely support. Do not treat a fixed 60–90-day interval as a requirement for every program.
Does a manager rating confirm behavior change?
It supplies another observation or judgment. Its usefulness depends on what the manager saw, the criteria used and the context. Preserve differences between sources and inspect the evidence before concluding that change occurred.