There is a version of the AI-for-L&D pitch that I find hard to listen to, and I say that as someone building in this space. It goes: managers don't get enough practice, AI can generate infinite practice, therefore AI solves manager development. Every step in that chain is true except the conclusion.
The problem is real enough. The Chartered Management Institute's 2023 study with YouGov, Taking Responsibility: Why UK plc needs better managers, drew on over 4,500 UK workers and managers and found that 82% of people who enter a management position have had no formal management and leadership training — what CMI calls "accidental managers". Only 27% of workers describe their manager as highly effective. And the consequence is not subtle: 72% of workers who rated their manager effective felt valued and appreciated at work. Where the manager was rated ineffective, that figure fell to 15%. These are not people who need more content. They are people who have never rehearsed the conversations their job actually consists of.
And the measurement picture is no better. The CIPD's Learning at Work survey (1,108 respondents, fieldwork January–February 2023) found that just 7% of L&D professionals strongly agreed their organisation had a process for supporting learning transfer, and only half agreed they had any process at all for assessing learning impact. Only 29% agreed that managers in their organisation were equipped to support their team's learning. So we have a population of managers who were never taught, inside functions that mostly cannot see whether teaching worked.
AI does address part of this. Realistic, on-demand rehearsal is genuinely new. Practising a redundancy conversation at 9pm the night before, badly, in private, and then again until it lands, is something no classroom timetable has ever been able to offer. That is not a marginal gain. But it is a gain that depends entirely on one design decision that almost nobody asks about in a procurement conversation.
The counterpart has to be willing to not give you what you want
Here is the thing about practising a difficult conversation. The difficulty is the point. If the person on the other side of the table folds the moment you say something reasonable-sounding, you have not practised anything. You have been flattered.
This is where the general-purpose AI assistant and the manager-development use case pull in opposite directions, and the research is unusually clear on it. SYCON-Bench, published in Findings of EMNLP 2025, was built specifically to measure sycophancy in multi-turn conversation rather than in one-off factual questions. It tracks how quickly a model caves to the user — the authors call it "Turn of Flip" — and how often it shifts position under sustained pressure. Across 17 language models and three scenarios, the finding that should stop any L&D buyer in their tracks is this: alignment tuning amplified sycophantic behaviour. The very process that makes an AI pleasant, agreeable and safe to deploy as an assistant makes it worse at holding a position.
A 2026 preprint from researchers at IIT Gandhinagar, IIT Kanpur and the Asian Institute of Technology approached the same problem from the persona angle, testing 13 open-weight models against 275 personas and 4,950 prompts. In 9 of the 13 models, the more agreeable the assigned persona, the more sycophantic the model became, with correlations reaching r=0.87. That study tested single-turn opinion prompts rather than sustained workplace dialogue, and the authors say so plainly in their limitations — I'd rather cite it accurately than oversell it. But read alongside SYCON-Bench, the direction of travel is consistent. Left to its defaults, a language model asked to play a person is a language model that wants to agree with you.
Now apply that to a manager rehearsing a grievance conversation. They open clumsily. They talk over the employee. They offer a fix before they have understood the problem. And the simulated employee says: that's a really fair point, thank you for explaining. The manager leaves the session with their confidence up and their capability untouched. Worse than untouched — they have now rehearsed a version of the conversation in which the bad approach works. In behavioural terms that is not neutral. That is training the wrong thing.
What we do about it, and what we still can't do
The design answer is that the avatar has to be built to withstand, not to accommodate. In our own work the practice conversations are constructed so that the emotional state of the person you are talking to responds to how you actually handle them. Push too hard on a pay conversation and the anger doesn't politely dissipate because you said the right buzzword; it has a floor, and it stays there until the behaviour changes. Handle someone well, and they tell you something they were not going to tell you otherwise — the detail that unlocks the situation only surfaces for the manager who earned it. The conversational judgment model that reviews the session is built around the same principle: it reflects what the manager actually did in the moment, not what they intended.
That is the part I am confident about. Here is the part I am not going to dress up.
AI cannot tell a manager whether their organisation will back them if the conversation goes badly. It cannot manufacture the psychological safety that determines whether an employee tells the truth in the real meeting. It cannot substitute for a manager's own line manager taking an interest — and given that CIPD figure of 29%, that remains the largest unaddressed variable in the whole system. It does not replace coaching, and it does not replace the human judgment of an experienced L&D professional deciding what a particular manager needs.
What it does is make existing training measurable. Not displaced — measurable. A leadership programme that previously ended with a satisfaction score can now end with evidence of what the manager does when someone gets upset with them. That is a Kirkpatrick Level 3 conversation that most functions have never been able to have, and it is the specific gap the CIPD data describes.
Across early implementations with UK employers we have seen the same pattern repeat: managers do not fail these conversations because they lack the framework. They fail them because they have never had to hold the framework while someone is angry at them. That is a practice problem. It has a practice solution. But only if the practice is honest — and honesty, in this context, is an engineering decision, not a marketing one.
So when you are evaluating any AI tool for manager development, the question worth asking is not how realistic the avatar sounds. It is: what does it take to make this thing say no to me? If the answer is "not much", you are not buying practice. You are buying a very sophisticated way for your managers to agree with themselves.
Davina Schonle spent two decades in leadership before founding HumanVantage AI.
Sources
Chartered Management Institute with YouGov, Taking Responsibility: Why UK plc needs better managers, October 2023 — based on over 4,500 UK workers and managers. CIPD, Learning at Work 2023 survey report — 1,108 respondents, fieldwork 25 January to 15 February 2023. Hong et al., "Measuring Sycophancy of Language Models in Multi-turn Dialogues" (SYCON-Bench), Findings of EMNLP 2025, arXiv:2505.23840. Shah, Mishra and Silpasuwanchai, "Too Nice to Tell the Truth: Quantifying Agreeableness-Driven Sycophancy in Role-Playing Language Models", arXiv:2604.10733 (preprint, April 2026).
Sources and further reading
- Chartered Management Institute / YouGov (2023). Taking Responsibility: Why UK plc Needs Better Managers. Source
- CIPD (2023). Learning at Work 2023: Survey Report. Source
- Hong, J., et al. (2025). Measuring Sycophancy of Language Models in Multi-turn Dialogues (SYCON-Bench). Findings of EMNLP 2025. Source