← The library

Training evaluation / L&D ROI

Kirkpatrick Four Levels

The standard model for evaluating training: Reaction, Learning, Behavior, Results.

Each level is harder to measure and more valuable to know. Most L&D functions stop at level 1 or 2, which is exactly why training budgets get cut.

Problem
Training evaluation / L&D ROI
Altitude
Program
Effort to run
Moderate
Evidence base
Strong

Theory & origin

Donald Kirkpatrick set out the four levels in his 1950s University of Wisconsin dissertation, and they became the default language of training evaluation. The levels, reaction, learning, behavior, results, rise in value and in difficulty to measure. That is exactly why most L&D stops at the 'smile sheet' and then loses its budget the moment finance asks about impact. The discipline the model imposes is designing evaluation backward from the business result, before the program is even built.

Key components

The parts at a glance. Click any term for the full definition, a field example, and the common failure, in the model below.

Explore the model

How a consultant runs it

  1. 01 Start at level 4: agree with the sponsor on the business metric the program is meant to move, and how you will argue attribution.
  2. 02 Design the level 3 behavior measure and the environment that has to support it. Manager reinforcement is usually the make-or-break factor.
  3. 03 Build level 1 and 2 instruments (reaction, pre/post knowledge) as leading indicators, not the endpoint.
  4. 04 Run the chain after launch and find where it breaks, often at behavior, because managers never reinforced the training.
  5. 05 Report the full chain to the sponsor, so the fix targets the real gap, not just course quality.

When to use

  1. 01 Designing evaluation before a training program launches (define level 4 first, then work backward)
  2. 02 Justifying or defending an L&D budget with evidence
  3. 03 Auditing an existing curriculum for programs that never get past the smile sheet

When not to use

  1. 01 When the real problem is not training. Role design or incentive problems cannot be trained away.
  2. 02 For informal or social learning, where isolating a level is artificial
  3. 03 When no one will act on the results. Measurement without decisions is theater.

Worked example

A retailer invests in service training. Level 1: 4.5 out of 5. Level 2: mystery-shop role-play scores up 25%. Level 3: 60 days later, only 30% of stores show the behaviors, and investigation reveals store managers never reinforced them. Level 4, NPS, is unchanged.

The Kirkpatrick chain locates the break: the problem is manager reinforcement, not course quality.

Common pitfalls

  1. 01 Measuring only levels 1 and 2 and calling the program a success
  2. 02 Deciding the evaluation strategy after the program is built instead of before
  3. 03 Claiming level 4 results with no control group or baseline comparison

Sample deliverable

One real engagement, start to finish. Watch the numbers travel from raw input, onto the chart, into the finished artifact.

Break-point analysis: service training

Input

  • L1 Reaction4.5 / 5
  • L2 Learning74% pass
  • L3 Behaviour30% adopt
  • L4 ResultsNPS +0

Process

Each level's measure is charted, and the chain breaks where the bars collapse

OutputDeliverable

Break-point analysis: service training

  • Chain breaks at L3behavior
  • Causemanagers never reinforced it
  • Fixmanager toolkit, not more course

Sources

Next in the library Job Profiling (Job Analysis)