Definition: AI roleplay for manager development is the use of conversational AI — often with realistic video avatars — to let managers rehearse difficult workplace conversations in private, receive feedback on how they actually behaved, and repeat the scenario until they improve. It turns management training from content consumption into observable practice.

How does AI roleplay actually work?

A manager chooses or is assigned a scenario — say, addressing persistent underperformance in someone they personally like. Instead of reading about the conversation, they have it. The AI counterpart plays the employee: it responds to what the manager actually says, reacts emotionally, pushes back, goes quiet, gets defensive. In well-designed systems the counterpart’s emotional state tracks the manager’s handling — push too hard and the resistance stays until the behaviour changes; handle the person well and the conversation opens up.

After the session, the manager gets structured feedback on the behaviour itself: whether they named the issue or danced around it, acknowledged emotion or steamrolled it, listened before proposing fixes, held their message under pressure. Then — and this is the point — they can try again. Immediately, privately, as many times as it takes. The rehearsal loop that a classroom rarely offers — at anything like this frequency or privacy — is the product.

What is the evidence that practice-based development works?

The underlying science is older and stronger than the technology. Ericsson’s work on deliberate practice established that expertise comes from structured, effortful practice with feedback — not from experience or exposure alone. Taylor and colleagues’ meta-analysis of behaviour modelling training found that programmes built on observing, practising and getting feedback on behaviour reliably improve skills, with the strongest results where trainees actively rehearse. Lacerenza and colleagues’ meta-analysis of leadership training reached a parallel conclusion: programmes that include practice and feedback, spaced over multiple sessions, substantially outperform information-only delivery.

The forgetting curve completes the argument — applied honestly. Murre and Dros replicated the characteristic shape of Ebbinghaus’ forgetting curve under controlled conditions: retention generally declines when learning is not retrieved or reinforced, though the rate varies substantially by learner and material. The practical implication is not that every learner forgets at an identical rate; it is that one-off exposure should not be treated as durable learning. A two-day workshop, however brilliant, needs reinforcement to survive — and practice that is available on demand, the night before the real conversation if that is when it is needed, is how it gets it.

One boundary worth stating plainly: the evidence base is strongest for the learning mechanisms underlying AI roleplay — active practice, feedback, repetition and spacing. Direct evidence comparing AI roleplay with other forms of manager development, particularly evidence of sustained workplace behaviour change, is still emerging. Buyers should therefore examine both a platform’s design and the quality of its outcome evaluation.

What does a manager actually practise?

The scenarios worth rehearsing are the ones managers avoid in real life because the stakes are too high to experiment on a real person: raising a performance concern for the first time; delivering feedback that may land badly; a wellbeing check-in where something personal is clearly going on; telling a valued team member they are not getting the promotion or pay rise; mediating conflict between two colleagues; a return-to-work conversation after long absence; announcing a restructuring decision they did not make but must own. Each type demands different behaviour, which is why practising one does not automatically transfer to the others — and why a categorised scenario library matters more than a big one.

How does AI roleplay compare with the alternatives?

ApproachRealismPrivacyRepetitionConsistency & measurabilityCost to scale
Peer role-play in workshopsLow — colleagues break character, soften, and everyone knows it’s a colleagueLimited — performed in front of peersOne or two attempts, then the agenda moves onInconsistent; nothing capturedHigh per rehearsal (facilitator + everyone’s time)
Professional actorsHighLow — usually observedVery limited — actor time is the constraintSomewhat consistent; feedback is the actor’s impressionVery high; does not scale beyond small cohorts
E-learning / video contentNone — describes conversations rather than simulating themFullUnlimited replays of content, zero rehearsalMeasures completion and recall onlyLow — which is why it dominates, despite weak transfer
1:1 coachingVariable — good coaches may use live rehearsal, but it is coach-dependentHighLimited by sessionsDepends entirely on the coach; rarely comparable across managersVery high per manager
AI roleplayHigh when the counterpart is built to resist, not pleasePotentially high — no audience by design; depends on data governance and reportingHigh — subject to platform accessSame scenario for every manager; behaviour observed and comparable over timeLow per manager once deployed

None of these are mutually exclusive. The strongest programmes use AI roleplay as the practice layer inside a broader development architecture — the rehearsal room attached to the classroom, not a replacement for human coaching or leadership judgement.

What are the limitations?

Honest ones first, because vendors in this category — ours included — are better at listing strengths. The most important technical risk is sycophancy: general-purpose AI models are tuned to be agreeable, and research such as SYCON-Bench shows that assistant-style tuning makes models cave to users faster, not slower. A counterpart that folds the moment the manager says something reasonable-sounding is not practice; it is flattery, and it trains the wrong behaviour. Ask any vendor how their counterpart is built to withstand pressure. We have written more about what AI can and cannot do for manager development.

Beyond the technology: AI roleplay cannot manufacture the organisational backing a manager needs when the real conversation goes hard; it cannot create the psychological safety that determines whether an employee tells the truth in the real meeting; and rehearsed behaviour still has to survive contact with the job — the transfer research is clear that the work environment conditions how much sticks. And a practice platform used as covert assessment will fail: managers rehearse honestly only in a space that is genuinely theirs.

Frequently asked questions

Is AI roleplay just e-learning with a chatbot attached?

No. E-learning delivers content and checks recall; AI roleplay requires the manager to conduct the conversation, in their own words, against a counterpart that reacts. The output is observed behaviour, not a completion record. If a product cannot show a manager being pushed back on and having to respond, it is a course, not practice.

Do managers take practising with an AI seriously?

They do when two conditions hold: the scenario is emotionally realistic enough to be uncomfortable, and the space is genuinely private. Managers disengage from role-play when a colleague is watching or the counterpart is a pushover. Remove the audience and the flattery and the rehearsal instinct takes over quickly.

Is this assessment or surveillance of managers?

It should never be either, and buyers should walk away from designs where it is. Practice works because managers can fail safely and try again. Reporting to the organisation should default to aggregate, cohort-level patterns, not individual verdicts — the session is the manager's rehearsal space, not an exam.

How often should a manager practise?

Little and often beats a single binge. Memory research shows unreinforced learning decays within days, so short rehearsals spaced over weeks — and a fresh rehearsal just before a real high-stakes conversation — produce more durable change than one long session.

How do you know whether it is working?

Look for behavioural movement across several attempts, not a single before-and-after on an identical scenario — repeating the same scenario can measure familiarity rather than transfer. Use comparable but different scenarios, a delayed follow-up, and appropriate on-the-job indicators, reported in aggregate rather than as isolated individual scores.

Sources and further reading

  • Ericsson, K. A., Krampe, R. T., & Tesch-Römer, C. (1993). The Role of Deliberate Practice in the Acquisition of Expert Performance. Psychological Review, 100(3), 363–406. Source
  • Taylor, P. J., Russ-Eft, D. F., & Chan, D. W. L. (2005). A Meta-Analytic Review of Behavior Modeling Training. Journal of Applied Psychology, 90(4), 692–709. Source
  • Lacerenza, C. N., Reyes, D. L., Marlow, S. L., Joseph, D. L., & Salas, E. (2017). Leadership Training Design, Delivery, and Implementation: A Meta-Analysis. Journal of Applied Psychology, 102(12), 1686–1718. Source
  • Murre, J. M. J., & Dros, J. (2015). Replication and Analysis of Ebbinghaus’ Forgetting Curve. PLOS ONE, 10(7), e0120644. Source
  • Hong, J., et al. (2025). Measuring Sycophancy of Language Models in Multi-turn Dialogues (SYCON-Bench). Findings of EMNLP 2025. Source
  • Blume, B. D., Ford, J. K., Baldwin, T. T., & Huang, J. L. (2010). Transfer of Training: A Meta-Analytic Review. Journal of Management, 36(4), 1065–1105. Source

About the author

Davina Schonle is the founder and CEO of HumanVantage AI. She spent two decades in leadership — leading teams, turning around a private-equity-backed division and closing multi-million-pound global deals — before building HumanVantage to make manager capability practisable and measurable. Connect with her on LinkedIn.