In brief: To practise difficult conversations at scale, give every manager a private, realistic rehearsal environment with an emotionally responsive counterpart, structured feedback and unlimited retries — then track behavioural change, not attendance. Consistency, privacy and measurement are what separate genuinely scaled practice from the one-off role-play workshop.

Why is scale the hard part?

The need is not niche. In the Chartered Management Institute’s 2023 study with YouGov, 82% of people who step into management had no formal training for it — a population of accidental managers now expected to handle underperformance, conflict, wellbeing and bad news as a routine part of the job. And the conversations they duck or fumble are expensive: research for Acas by Saundry and Urwin puts the cost of workplace conflict to UK employers at roughly £28.5 billion a year, most of it incurred where issues escalate into formal processes that an earlier, better conversation might have prevented.

Every organisation already knows the answer is practice. The problem is arithmetic. The traditional practice formats — facilitated role-play, actor-based simulation, observed coaching — are handcrafted. They reach one cohort at a time, deliver one or two rehearsals per manager, vary with whoever is in the room, and leave nothing behind that can be measured. Meanwhile memory research is unsentimental about what happens next: retention generally declines when learning is not retrieved or reinforced — one-off exposure should not be treated as durable learning. Scale is not a nice-to-have; it is the difference between a programme that touches managers once and a capability managers actually keep.

What must practice include to change behaviour?

Four conditions, all evidence-backed, and all four have to hold at once.

Realism. The rehearsal has to recruit the same pressure the real conversation will. Ericsson’s deliberate-practice research is explicit that improvement comes from effortful practice at the edge of current ability — a counterpart that folds when the manager says something reasonable-sounding is not practice, it is reassurance. The emotional resistance is the training load.

Safety. Managers only experiment where failure is survivable. Edmondson’s work on psychological safety shows that people take learning risks — admitting error, trying the unfamiliar — only where it is safe to do so. In practical terms: the rehearsal must be private by design. No audience, no line manager watching, no individual verdict flowing upward. The moment practice feels like assessment, managers perform instead of learning.

Repetition with feedback. Taylor and colleagues’ meta-analysis of behaviour modelling training shows skills build through cycles of attempt, feedback and re-attempt — not exposure. One rehearsal is a demonstration; five rehearsals with honest feedback between them is development.

Spacing. Lacerenza and colleagues’ meta-analysis of leadership training found spaced, multi-session designs substantially outperform single events, and the forgetting-curve replication by Murre and Dros explains why. Practice has to be available when it is needed — including the night before the real conversation — not when the training calendar says so.

The same evidence boundary we state in our guide to AI roleplay applies here: these studies establish the learning mechanisms — practice, feedback, repetition, spacing — rather than direct comparisons between delivery formats, so examine any platform’s own outcome evaluation as well as its design.

How do you roll it out across a management population?

Map the conversations first. Not all difficult conversations are the same skill. Categorise what your managers actually face — underperformance, wellbeing, conflict, pay and promotion disappointment, restructuring — and identify where the risk concentrates. Grievance data, exit interviews and employee-relations caseloads will tell you.

Start with one cohort and a real scenario. New managers are usually the highest-leverage starting point: the transition is where avoidance habits form, and the five conversations they consistently get wrong are well documented. Give the cohort the same scenario so results are comparable.

Make privacy the visible promise. Tell managers explicitly what is and is not reported. Aggregate patterns to the organisation; the rehearsal itself belongs to the manager. Adoption follows trust, and trust follows the design being said out loud.

Attach practice to the programmes you already run. Rehearsal before a workshop surfaces where each manager actually struggles; rehearsal after converts the frameworks into behaviour while they are still fresh. The practice layer makes the existing spend work harder rather than competing with it.

Measure movement, then report patterns. Run comparable — but not identical — scenarios before and after a development period (an identical rerun can measure familiarity rather than skill) and compare behaviour: directness, composure under pushback, listening before fixing. Across hundreds of managers, the aggregate picture — where capability gaps cluster, which conversation types are weakest, whether investment is shifting behaviour — is the organisational intelligence a people function can take to the board. We cover the measurement design in detail in How to Measure Whether Management Training Actually Worked.

What are the options for scaling practice?

ApproachConsistencyPrivacyAvailabilityMeasurabilityCost per manager at scale
Peer role-play workshopsLow — depends on the roomLimitedScheduled, rareNoneHigh
Actor-based simulationMediumLow — observedScheduled, very limitedFacilitator notes at bestVery high
External 1:1 coachingVaries by coachHighSession-limitedAnecdotalVery high
E-learning modulesPerfect — but consistent content, not practiceFullOn demandCompletion onlyLow
AI roleplay platformHigh — same scenario for every managerPotentially high — depends on data governanceOn demand, repeated retriesBehavioural, comparable over timeLow

The pattern is blunt: the formats with the realism have never had the scale, and the formats with the scale have never had the practice. AI roleplay is the first format that offers both — provided the counterpart is engineered to withstand pressure rather than agree, which is the single technical property to verify before buying. Our guide to what AI roleplay is and how it works covers that in depth.

What gets in the way?

Three failure modes account for most disappointments. First, practice deployed as covert assessment: if managers suspect the rehearsal space reports on them individually, they stop rehearsing honestly and the data becomes theatre. Second, scenario quality: a generic scenario produces generic practice, and managers spot the difference immediately — the library must reflect the conversations your managers actually have. Third, missing cultural reinforcement: practice builds capability, but managers use it only where leaders model direct conversations and HR processes do not punish early honesty. A rehearsal platform inside an avoidant culture will produce well-rehearsed avoiders.

And one honest limitation: rehearsed behaviour still has to transfer. The transfer-of-training literature is consistent that the work environment — manager support, opportunity to use the skill — conditions how much survives. Practice at scale narrows that gap; it does not abolish it.

Frequently asked questions

How many managers do we need before scale becomes the constraint?

Earlier than most people think. Peer role-play and actor-based simulation strain past a single cohort of 12–20; consistency collapses first, then cost. If you have 50+ managers who each face difficult conversations weekly, ad-hoc practice is already failing quietly — you just cannot see it, because nothing is measured.

Can we scale by training internal facilitators instead?

Train-the-trainer scales delivery of content, not practice. Each facilitator still runs live role-play one room at a time, quality varies with the facilitator, and managers still will not rehearse honestly in front of colleagues. It improves workshops; it does not create a rehearsal space every manager can use privately, repeatedly, on demand.

How long before we see behaviour change?

In our experience, managers stop making their most obvious error in a specific scenario (burying the message, filling silences, proposing fixes before listening) within a handful of attempts — though this varies by manager and scenario. Durable change across real conversations takes longer and depends on reinforcement and the environment managers return to. Measure movement between attempts, not calendar time.

Do we need video avatars, or is text enough?

Text rehearsal is better than nothing, but difficult conversations are carried by tone, pace and the discomfort of facing a person. A realistic spoken counterpart with a face recruits the same nerves the real conversation will — and practising under those conditions is what desensitises the panic response.

How does this fit the leadership programme we already run?

As the practice layer, not a replacement. Keep the frameworks and the cohort experience; attach rehearsal before and after, so concepts are immediately practised and the programme's impact becomes observable. The workshops supply the what; the practice supplies the can.

Sources and further reading

  • Chartered Management Institute / YouGov (2023). Taking Responsibility: Why UK plc Needs Better Managers. Source
  • Saundry, R., & Urwin, P. (2021). Estimating the Costs of Workplace Conflict. Acas. Source
  • Ericsson, K. A., Krampe, R. T., & Tesch-Römer, C. (1993). The Role of Deliberate Practice in the Acquisition of Expert Performance. Psychological Review, 100(3), 363–406. Source
  • Edmondson, A. (1999). Psychological Safety and Learning Behavior in Work Teams. Administrative Science Quarterly, 44(2), 350–383. Source
  • Taylor, P. J., Russ-Eft, D. F., & Chan, D. W. L. (2005). A Meta-Analytic Review of Behavior Modeling Training. Journal of Applied Psychology, 90(4), 692–709. Source
  • Lacerenza, C. N., et al. (2017). Leadership Training Design, Delivery, and Implementation: A Meta-Analysis. Journal of Applied Psychology, 102(12), 1686–1718. Source
  • Murre, J. M. J., & Dros, J. (2015). Replication and Analysis of Ebbinghaus’ Forgetting Curve. PLOS ONE, 10(7), e0120644. Source
  • Blume, B. D., et al. (2010). Transfer of Training: A Meta-Analytic Review. Journal of Management, 36(4), 1065–1105. Source

About the author

Davina Schonle is the founder and CEO of HumanVantage AI. She spent two decades in leadership — leading teams, turning around a private-equity-backed division and closing multi-million-pound global deals — before building HumanVantage to make manager capability practisable and measurable. Connect with her on LinkedIn.