Every year, in almost every organisation that invests in developing its managers, the same quiet ritual plays out. A programme runs. Managers attend. At the end, they fill in a form. The scores come back warm — four-point-something out of five, a page of encouraging comments — and that number becomes the evidence that the money was well spent. The training "worked."
Except nobody in the room actually believes that a feedback form tells them whether anything changed. The managers who rated the course highly go back to their teams and, in the next genuinely difficult conversation, do more or less exactly what they did before. This is the uncomfortable truth sitting underneath the L&D budget: the thing we measure and the thing we care about are not the same thing, and everybody knows it.
The good news is that this is not an unsolvable measurement problem. It is a solved framework applied at the wrong level, for a very human reason — the level that matters has, until recently, been genuinely hard to see.
What does it actually mean for training to "work"?
The most durable answer is over sixty years old. In 1959, Donald Kirkpatrick set out four levels at which any training can be evaluated, and the model — now stewarded by Kirkpatrick Partners — has outlasted almost everything that came after it. Level 1 is reaction: did people like it. Level 2 is learning: can they recall the concepts. Level 3 is behaviour: are they doing anything differently on the job. Level 4 is results: did the business outcome move.
Read those back and the problem announces itself. A manager can love a course (Level 1), pass the quiz (Level 2), and still handle the next underperformance conversation exactly as badly as before. Reaction and learning are necessary conditions and terrible proxies. The level that carries the meaning of the word "worked" — the level a Head of L&D is really being asked about when a CFO queries the spend — is Level 3. Behaviour. What the manager actually does when it counts.
Why do most organisations only measure the first level?
Because Level 1 is almost free and Level 3 has always been expensive. A reaction score is a form handed out while everyone is still in the room. Behaviour change happens later, in private, in conversations no one is watching — and capturing it has traditionally meant slow, costly, subjective work: manager self-reports, delayed 360 cycles, line-manager observation that rarely happens.
So the industry does the affordable thing and stops. By the Association for Talent Development's own reporting, only around a third of organisations evaluate beyond participant reaction and learning — which means the majority never formally look at whether behaviour changed at all. The picture in the UK is starker still. In the CIPD's Learning at Work 2023 survey of 1,108 organisations, just 7% strongly agreed that they had a process in place to support the transfer of learning back into the job. Not measure it. Support it. The measurement gap and the transfer gap are the same gap, and they open up precisely where the value is supposed to land.
This is why so many L&D leaders feel they are defending real work with unreal evidence. The number they can produce on demand — the satisfaction score — has close to zero relationship with the number that would actually justify the programme.
Doesn't a follow-up survey ninety days later solve this?
It feels like it should, and it is better than nothing. But the common Level 3 instruments share a flaw that a survey cannot fix: they measure perception, not behaviour. A manager reporting that they "now give feedback more directly" is telling you what they believe, or what they think you want to hear, filtered through three months of recall. A 360 tells you how a manager's conduct is perceived by people with their own incentives and blind spots. Both are gameable, both are slow, and neither shows you the actual conversation.
The behaviour-change literature has been consistent about the consequence for decades: what gets learned in a room mostly does not survive contact with the job. The most cited meta-analytic review of transfer of training, Blume and colleagues (2010), found transfer to be modest and highly conditional on the work environment — echoing the long-standing field estimate that only a low share of trained skill, often put at 10–30%, reliably shows up in on-the-job behaviour. For behavioural skills specifically — staying composed while someone is angry, naming a problem without damaging the relationship — the transfer is at the lower end, because those skills are motor skills. You cannot survey your way to them, and you cannot survey your way to knowing whether someone has them.
What would measuring behaviour actually look like?
The shift is smaller than it sounds. Instead of measuring behaviour retrospectively — guessing, months later, whether something changed — you measure it at the point of practice, where the behaviour is visible in real time.
Give a manager a realistic, emotionally loaded conversation to work through — the kind they avoid in real life because the stakes are too high to experiment on a real employee. Let them rehearse it privately, as many times as they need. Now you can see, in the practice itself, what they actually do: whether they name the issue or dance around it, whether they hold their composure when the other person pushes back, whether they check understanding or steamroll it. Run the same scenario before a development programme and again after, and the change is no longer a matter of opinion. It is a movement you can point to.
The distinction that matters here — and the one it is easy to get wrong — is that this is not a test of the manager. It is a picture of how their practice is developing. The scenario is a rehearsal space, not an exam and not surveillance; the value is that the manager gets to fail safely and try again, and that the organisation gets an honest, consistent read of whether the development is landing. That is Level 3 turned from a retrospective survey into a live signal, and it is what "training made measurable" actually means: not a new layer of scrutiny on people, but a way of seeing the practice that the old instruments could never catch.
What does "good" look like in the data?
Not a high score. Movement. A single number out of a hundred tells a Head of L&D almost nothing; the change between a manager's first attempt and their tenth tells them almost everything. Directness rising. Composure holding for longer under pressure. More listening, fewer premature solutions. Those are the leading indicators — the behaviours that, if they shift, plausibly move the lagging ones the business already tracks.
And those lagging measures are exactly where most organisations are trying, and struggling, to prove impact. In LinkedIn's 2024 Workplace Learning Report, the top metrics L&D functions use to track business impact were performance reviews (36%), employee productivity (34%) and employee retention (31%) — all real, all valuable, and all downstream of manager behaviour by months or years. Behaviour at the point of practice is the missing middle term: the evidence that connects the development you paid for to the retention number you are eventually judged on. Without it, you are asking a board to take the whole chain on faith.
The instrument, not the intent
None of this asks L&D to care about different things. Everyone already knows behaviour is what matters; the Kirkpatrick model has said so since 1959. What has been missing is not the will to measure Level 3 but a practical instrument for it — one that is fast enough to use at scale, consistent enough to compare, and honest enough to show what a manager did rather than what they say they did. That instrument now exists, and it changes the answer L&D can give when someone asks whether the training worked. The honest answer stops being a shrug dressed up as a satisfaction score, and becomes: here is what our managers could do before, here is what they can do now, and here is the difference.
This is the problem HumanVantage was built to solve. Managers rehearse difficult conversations in a private, safe practice environment — the same conversations they would otherwise be avoiding on real people — and our conversational judgment model reads what actually happens in each rehearsal: how directly the issue was named, how composure held, how well the manager listened. Because the same scenario can be practised before and after development, the platform turns Level 3 from a retrospective guess into an observable change, giving L&D a straight line from practice to behaviour to the business outcomes they are ultimately asked to defend. It measures existing managers as they build confidence, not candidates being screened — because the point was never to grade people. It was to make the development visible.
Management training has spent decades unable to prove its own worth, not because behaviour change is unmeasurable, but because we kept measuring the wrong level at the wrong moment. Measure practice, and the question that has haunted every L&D budget — did it actually work — finally has an answer you can put in front of the board.