Kirkpatrick Model of Training Evaluation: 4 Levels
The four levels of the Kirkpatrick model, how to run them backwards on a small team without software, the New World update, and the real criticisms.
Kirkpatrick Model of Training Evaluation
The four levels and the questions each one actually answers, why planning backwards from results is the whole trick, how to run all four on a team of ten in under three hours with no software, the federal regulation that made this framework the basis of government training evaluation, and the criticisms most guides leave out
Almost every explanation of this framework draws the pyramid, walks up it from the bottom, and stops. Which is exactly the wrong direction, and it is the reason most organizations run Levels 1 and 2 forever and never reach the two that matter.
The model is four questions. Did they like it. Did they learn it. Are they doing it. Did anything change. The first two can be answered before people leave the room. The last two require you to go back thirty and ninety days later, which is why they get skipped, and which is why training evaluation in most businesses consists of an average satisfaction score reported as if it were evidence.
This guide covers the four levels and what each one genuinely tells you, the backwards planning sequence that is the single most useful idea attached to the framework, the New World update and what it added, the criticisms that most explanations omit, and the version that works for a business with no training function. That last part is the gap: every ranking explanation of this model assumes a learning platform, a training manager, and cooperative middle management. I build the people tooling for businesses that have none of those at FirstHR.
What the Kirkpatrick Model Is
The Kirkpatrick model is a training evaluation framework that measures the impact of a learning program across four sequential levels: Reaction, Learning, Behavior, and Results. Each level asks a different question, at a different time, at a different cost, and the answers get more valuable and more difficult as you move up.
The four questions are also the reason the framework survived sixty years while most management models of its era did not. They are the questions a sceptical owner asks anyway, in roughly that order, whether or not they have heard of the model. Naming them and numbering them is most of what the framework contributes.
The origin is worth knowing because it explains the shape. Kirkpatrick developed the levels out of doctoral research at the University of Wisconsin and published them one per article, in sequence, in the journal of what was then the American Society of Training Directors. He did not originally call them a model or draw them as a pyramid. They were four techniques for evaluating training programs, and the hierarchy that later got attached to them is an interpretation rather than the original claim.
That distinction matters more than it sounds, because the pyramid implies a causal ladder that the four questions do not require. Read as four independent questions, the framework is durable and useful. Read as a proof that reaction causes learning causes behaviour causes results, it makes a claim that has been contested for decades.
The Four Levels in Detail
Each level answers a different question about the same training, and the sensible way to hold them is by what each one can and cannot tell you rather than by their number.
| Level | The question | How it is measured | When | Difficulty |
|---|---|---|---|---|
| 1. Reaction | Did they find it engaging and relevant? | Short survey, or asked out loud | Immediately | Trivial |
| 2. Learning | Did knowledge, skill, or confidence increase? | Before-and-after check, demonstration, exercise | Same day | Low |
| 3. Behavior | Are they doing it differently at work? | Manager check-in, observation, work sample | 30 to 90 days | High |
| 4. Results | Did a business outcome move? | One tracked measure against a baseline | 3 to 12 months | High, plus attribution problems |
Level 1 is the easiest to collect and the easiest to over-trust. A four point six out of five average feels like success and is compatible with nobody changing anything. It is genuinely useful for a narrow purpose: finding out whether the session was clear, correctly pitched, and worth the time. Treat it as feedback on the delivery, not as evidence of effect.
Level 2 is where a small change in method pays disproportionately. A quiz measures recall on the day. A demonstration measures capability. If you can watch the person do the thing, do that instead, because the difference between someone who can describe the procedure and someone who can perform it is exactly the difference you are trying to detect.
Level 3 is where the model earns its keep and where almost everyone stops. It is also where the constraint is organizational rather than technical: measuring behaviour requires someone to go back thirty days later and involve the person's manager. That is a scheduling problem disguised as a measurement problem.
The Level 3 constraint deserves one more sentence, because it is the reason the whole model stalls. Behaviour change depends on the person's manager noticing and reinforcing it, which means measuring Level 3 requires cooperation from someone who did not attend the training and did not ask for it. In a large organization that is a political problem. In a small one it is usually the owner, which makes it simpler to arrange and easier to keep postponing.
Level 4 is the level executives actually asked about, and the one with the honest caveat attached. A number moving after training does not mean the training moved it, and the framework provides no mechanism for separating the two. Report it as a correlation with the confounders named, and you keep your credibility. Report it as proof, and you lose it the first time someone asks about seasonality.
Run It Backwards
The single most useful idea attached to this framework is that you plan in the reverse of the order you measure in. Start at Level 4 and work down. Most training programs are designed by picking content and then wondering afterwards how to evaluate it, which produces an evaluation plan bolted onto a decision already made.
The discipline this imposes is uncomfortable and valuable. If you cannot name a business number at step one, you have learned something important before spending anything: either the training is a compliance requirement, which is a legitimate reason to run it without pretending to measure impact, or it is a solution looking for a problem.
Backwards planning also fixes a scoping problem. Starting from content produces a session that covers everything anyone might want to know, because there is no principle for excluding anything. Starting from one behaviour gives you a sharp test for every candidate topic: does knowing this help someone perform the behaviour. Most of what would have gone in fails that test, and the session gets shorter, which is separately valuable given that every hour of it is paid.
The second thing it exposes is when training is the wrong intervention entirely. If the target behaviour is not happening because people lack a tool, or because the old behaviour is still what gets rewarded, no amount of training changes it. Working backwards surfaces that before the session rather than at the ninety-day review.
The Four Levels Without an L and D Team
Every explanation of this model assumes a learning platform, a training manager, and managers who will cooperate with a behavioural observation protocol. A business with no training function has none of those and can still run all four levels properly.
The Level 3 item is the one to fight for. It is a single calendar reminder and a fifteen-minute conversation, and it is the difference between knowing whether training worked and guessing. The question that does most of the work is show me a recent example, because it cannot be answered from memory of the session. Either an example exists or it does not.
The Level 4 item is easier than it sounds at small scale, for a slightly counterintuitive reason. A small business has fewer confounders. When twelve people work somewhere and one process changed, attribution is considerably more tractable than at a company where forty things changed the same quarter. The disadvantage of small numbers, noise, is real, but the advantage of a simple system is real too.
One cost line belongs in the calculation and is almost never in it. Required job-related training is hours worked, so the denominator of any return calculation includes the wages of everyone in the room plus whoever delivered it. For non-exempt staff those hours also count toward overtime under the Fair Labor Standards Act. A half-day session for twelve people is a real number, and knowing it changes how seriously you take the day-ninety check.
| A | B | C | D | E | F | G | H | I | J | |
|---|---|---|---|---|---|---|---|---|---|---|
| 1 | Planning order | Level | What we are asking | How we will measure it | Baseline value | Target | Measure on | Owner | Result | What we changed |
| 2 | 1st | Level 4 Results | Which business number should move | |||||||
| 3 | 2nd | Level 3 Behavior | What people must do differently | |||||||
| 4 | 3rd | Level 2 Learning | What they must know or be able to do | |||||||
| 5 | 4th | Level 1 Reaction | Was it clear and worth attending | |||||||
| 6 | Required drivers | What reinforces the behavior afterward | ||||||||
| 7 | ||||||||||
| 8 |
The first sheet is the evaluation plan laid out in planning order rather than level order, so the first row you fill in is Level 4 and the last is Level 1, with a separate row for the required drivers that will reinforce the behaviour afterward. The second is a question bank you can lift directly, with the level, the format, and when to ask each one. The third is the results log, which includes a column for confounders noted, because writing down what else changed is the step that keeps a Level 4 claim honest.
The New World Kirkpatrick Model
In 2016, James D. Kirkpatrick and Wendy Kayser Kirkpatrick published a substantial update to the original framework. It keeps the four levels and adds the machinery that the original left implicit.
| Element | Original model | New World model |
|---|---|---|
| Planning direction | Implicitly bottom-up, level 1 first | Explicitly plan from Level 4 backwards |
| Level 2 scope | Knowledge and skill | Knowledge, skill, attitude, confidence, and commitment |
| Level 3 scope | Behavior change | Critical behaviors plus required drivers that reinforce them |
| What happens after training | Largely unaddressed | Required drivers treated as part of the program, not an afterthought |
| Level 4 definition | Organizational results | Results plus leading indicators that arrive sooner |
| Underlying stance | Evaluate the training | Evaluate and support the whole chain to the result |
Two additions do most of the work. Required drivers is the recognition that behaviour change after training is not automatic and needs reinforcement, coaching, and accountability built into the program rather than hoped for. Leading indicators solve a practical problem: Level 4 results arrive months later, and having a nearer-term signal keeps the evaluation alive in the meantime.
The required drivers idea is the one small businesses should take most seriously, because it names the thing that actually determines whether training sticks. It is rarely the content. It is whether the person's manager asks about it, whether the new way is easier than the old way, and whether anyone notices when they do it. In a small business the manager is usually the owner, which makes this both simpler and harder to delegate.
Written down, required drivers are unglamorous: a question in the next one to one, a change to the checklist, a note in the responsibilities that makes the new behaviour someone's actual job. None of that is training. All of it determines whether the training survives contact with a busy week, which is the point the original framework left for the reader to work out.
What the Model Gets Wrong
The framework has been criticised steadily for decades, and the criticisms are worth knowing because they tell you what the model cannot do rather than whether to use it.
The causal criticism is the substantive one. Alliger and Janak, writing in Personnel Psychology in 1989 under the title Kirkpatrick's Levels of Training Criteria: Thirty Years Later, examined the assumptions the levels carry and challenged the notion that they are causally linked or positively intercorrelated. Subsequent work has broadly supported the concern: how much someone enjoyed a session tells you very little about whether they learned anything, and less about whether they changed.
The second criticism, about intervening variables, is the one small businesses feel most directly. Whether someone applies training depends on their workload that week, whether the tooling supports the new way, whether their colleagues are doing it, and whether the old way still gets results faster. A review of training evaluation models made exactly this point, noting that the framework is too simple and does not account for the variables affecting learning and transfer. The New World revision is largely an attempt to answer that criticism from inside the model.
The practical consequence is a rule rather than a rejection. Do not treat a good Level 1 as a reason to skip Level 3. The whole appeal of the pyramid is that it lets you infer upward from the cheap measurement to the expensive one, and that inference is exactly what does not hold.
The Federal Version, and the Caveat Attached to It
The largest employer in the United States is required by regulation to evaluate its training, and the guidance it publishes for doing so is built on this framework. That is a stronger adoption signal than any vendor case study, and it comes with a caveat that vendor pages never mention.
Federal agencies are required to evaluate their training programs annually to determine how well those plans and programs contribute to mission accomplishment and meet organizational performance goals, under 5 CFR 410.202 (Office of Personnel Management). The underlying training regulations sit in 5 CFR Part 410, and the evaluation field guide OPM publishes to help agencies meet that obligation uses the Kirkpatrick four levels as its basis (OPM reference materials).
The distinction OPM draws is worth sitting with. Evaluating a training program and evaluating a performance problem are different tasks. The Kirkpatrick levels are built for the first. If your actual question is why is this not working, a framework that begins by assuming training was the correct intervention will not tell you that it was not.
Kirkpatrick and the Alternatives
Most alternatives extend the model rather than replacing it, and each was built to fix a specific weakness in it.
| Framework | What it adds | Best for | Cost to run |
|---|---|---|---|
| Kirkpatrick four levels | The common vocabulary and four clear questions | Almost any training, as a default | Low to moderate |
| Phillips ROI Methodology | A fifth level converting results to money, plus isolation techniques | When you must defend spend in financial terms | High |
| CIRO | Evaluates context and inputs before training, not only outcomes | Deciding whether to train at all | Moderate |
| Brinkerhoff Success Case Method | Studies the best and worst cases in depth instead of averaging | Understanding why it worked for some and not others | Moderate |
| Modified versions with Level 0 and 5 | Participation at the bottom, return at the top | Academic and clinical evaluation | High |
The Phillips extension is the most commonly cited because it addresses the attribution gap directly, adding a fifth level that converts Level 4 results into monetary terms and, critically, includes techniques for isolating the training's contribution from other factors. That isolation step is the part the original framework lacks, and it is also the part that makes the method expensive to run properly.
Extensions of this kind are not only commercial. Clinical and academic evaluations have adapted the levels in both directions, adding a Level 0 for participation and activity accounting and a Level 5 for return on investment or expectation, as in published evaluations of simulation training in emergency medicine (National Library of Medicine). The same source base shows the framework in routine use for the first three levels in clinical training evaluation (National Library of Medicine).
For a small business the practical recommendation is short: use the four levels, plan backwards, and borrow the isolation question from the ROI methodology without building the whole apparatus. Asking what else could explain this before you claim a result gets you most of the credibility at none of the cost.
There is also a case for borrowing the CIRO instinct rather than the CIRO method. Before designing anything, ask whether the gap is knowledge, tools, incentives, or clarity about who owns what. That question routes a meaningful share of apparent training problems into process and structure instead, and it costs one conversation.
Where This Goes Wrong
The failure patterns are consistent and almost all of them come from measuring what is easy rather than what was asked.
Reporting Level 1 as the result is first and by far the most common. An average satisfaction score presented as evidence that training worked is not a small overstatement, it is a different claim entirely, and the people receiving it usually know that.
Designing forwards is second. Content chosen first, evaluation appended afterwards, business outcome never named. The plan that results can only ever measure whether the session went smoothly.
Skipping the thirty-day check is third, and it is a calendar failure rather than a resource one. Fifteen minutes per person, once, is the entire Level 3 apparatus at small scale.
Claiming Level 4 causation is fourth. A number moved, training happened, therefore training moved the number. Name the confounders yourself before somebody else does, and the finding survives scrutiny.
Treating the levels as a ladder you can climb by inference is fifth. Good reaction does not imply learning, and learning does not imply transfer. This is the substance of the academic criticism and it is also just observably true to anyone who has run training.
And evaluating training when the problem was never a training problem is last. If people know what to do and are not doing it, the constraint is a tool, an incentive, or a manager, and no framework aimed at learning programs will surface that. Working backwards from the result is what protects you, which is the same reason it belongs in the design of the development plan rather than only in the evaluation of a session.
Working backwards is also how you avoid running the same training twice for the same reason. If day ninety shows the number did not move and the day-thirty conversation showed the behaviour did not appear, the diagnosis is transfer rather than content, and rerunning the session will produce the same result. That distinction is worth more to a small business than any refinement of the measurement, and it belongs alongside the workforce plan that decides what gets trained next.
Whatever you conclude, keep the record. What was delivered, to whom, when, and what happened at day thirty and day ninety is the institutional memory that stops you rerunning a program that did not work. Storing it with the rest of your employee records rather than in a deck is the low-effort version, and keeping the completion record beside the roster in a single employee directory is what makes it retrievable a year later.
Frequently Asked Questions
What is the Kirkpatrick model of training evaluation?
The Kirkpatrick model is a framework for evaluating training across four sequential levels. Level 1 Reaction asks whether participants found the training engaging and relevant. Level 2 Learning asks whether knowledge, skill, confidence, or commitment increased. Level 3 Behavior asks whether people apply the training on the job. Level 4 Results asks whether an outcome the organization cares about actually moved. Donald Kirkpatrick published the four levels as a series of journal articles at the end of the nineteen fifties, and they remain the most widely used vocabulary in training evaluation.
What are the 4 levels of the Kirkpatrick model?
Level 1 is Reaction, measured with a short survey immediately after the session. Level 2 is Learning, measured with a check before and after, a demonstration, or a practical exercise. Level 3 is Behavior, measured thirty to ninety days later through observation, a manager conversation, a work sample, or system data. Level 4 is Results, measured over three to twelve months against a business number tracked before the training. Difficulty and cost rise sharply from Level 1 to Level 4, which is why most organizations stop after Level 2.
Who created the Kirkpatrick model and when?
Donald L. Kirkpatrick developed the framework from his doctoral research at the University of Wisconsin in the mid nineteen fifties and published it as a four-part series in the Journal of the American Society of Training Directors across 1959 and 1960, with one article per level. He consolidated the work in a book in 1994. In 2016, James D. Kirkpatrick and Wendy Kayser Kirkpatrick published a substantial update known as the New World Kirkpatrick Model, which added required drivers, confidence and commitment at Level 2, and the practice of planning backwards from Level 4.
What is the difference between the Kirkpatrick model and the New World Kirkpatrick model?
The original model describes four levels of evaluation. The New World version keeps the same four levels and adds practical machinery around them. It introduces required drivers, meaning the processes and reinforcement that make behaviour stick after training. It splits Level 3 into critical behaviours and the systems that support them. It adds confidence and commitment as Level 2 measures alongside knowledge and skill. And it establishes the principle of planning backwards from Level 4, starting with the business outcome rather than with the training content.
Why do most organizations stop at Level 2?
Because Levels 3 and 4 require time, cooperation, and a delay. Level 3 means going back to people thirty days later and involving their managers, which requires someone to own the follow-up. Level 4 means agreeing on a business number before the training and waiting months to look at it, which requires a decision up front that most programs never make. Levels 1 and 2 can be completed in the room on the day. The irony is that Levels 3 and 4 are the only ones anybody outside the training function cares about.
What are the criticisms of the Kirkpatrick model?
The main criticism is that the model implies a causal chain it does not establish. A frequently cited review challenged whether the four levels are causally linked or even positively intercorrelated, and satisfaction has repeatedly proven a poor predictor of learning. A second criticism is that the model ignores the intervening variables between training and behaviour, such as the manager, workload, and whether the old way is still rewarded. A third is that Level 4 provides no method for isolating training's contribution from everything else that moved the number.
How do you use the Kirkpatrick model in a small business?
Run it backwards and keep it cheap. Start by naming one business number you already track and want to move. Define one observable behaviour that would move it. Derive the training content from that behaviour. Then measure: three questions after the session, a demonstration rather than a quiz, one calendar reminder at thirty days asking people to show a recent example, and a look at the baseline number at ninety days. For a team of ten this is under three hours in total across three months and requires no software.
What is the alternative to the Kirkpatrick model?
The most common alternatives extend rather than replace it. The Phillips ROI Methodology adds a fifth level that converts results to money and isolates the training's contribution, which addresses the attribution gap. The CIRO model evaluates context and inputs before the training rather than only outcomes after it. Brinkerhoff's Success Case Method takes a different approach entirely, examining the most and least successful cases in depth rather than averaging across everyone. Some medical and academic adaptations add a Level 0 for participation and a Level 5 for return.
Is the Kirkpatrick model still relevant?
It remains the common vocabulary of training evaluation and is used widely enough that the four level numbers function as shorthand across the field. It is also the basis of the training evaluation field guide published by the US Office of Personnel Management for federal agencies, which are required by regulation to evaluate their training programmes annually. The fair assessment is that it is an excellent set of four questions and a weak method for proving causation, and it works best when treated as the former.