Guide7 min read
Does practising with an AI client improve coaching skills?
The research reviewed here does not establish whether AI-client practice improves professional coaches' work with human clients.
The research reviewed here does not establish whether AI-client practice improves professional coaches' work with human clients.
The closest evidence comes from counsellor training, where one randomised study found that practice with a simulated client changed novices' behaviour in both directions depending on whether they got feedback. None of the studies discussed here follows professional coaches into sessions with human clients to see whether AI rehearsal made any difference to their work. That does not prove there is no benefit; a claim of professional-coach improvement needs evidence beyond the studies reviewed here.
Decide which "improve" you mean
Three different things get compressed into that word:
- Feeling readier. Confidence before a session with a real person.
- Behaving differently. More reflections, fewer premature suggestions, questions that follow the client rather than your plan.
- Making a difference to clients. Whether the people you coach get somewhere they wouldn't otherwise have got.
In the CARE study described below, confidence and measured behaviour moved in opposite directions in the practice-only group.
The CARE study: what practice alone did
Researchers at Stanford and Georgia Tech built a training system called CARE and ran a randomised experiment with 94 novice counsellors recruited through Prolific in the US and UK, with recruitment criteria specifying an educational background in psychology, counselling, social work or nursing, no completed graduate degree, and less than a year of counselling-related experience (CARE, CHI 2026).
The session lasted 75 minutes over Zoom, including surveys and an interview. Its practice activities were a five-minute tutorial, a ten-minute conversation with a simulated patient to establish a baseline, a 20-minute practice block, then a ten-minute conversation with a new simulated patient. Everyone practised. One group could also review structured feedback on their responses, labelling strengths, weak points and alternative phrasings. The simulated patients were GPT-4o prompted with behaviour principles written by experienced counsellors, so that they would resist, stay vague or hold back rather than co-operate.
The practice-only group used less empathy after practice. They also used fewer inappropriate suggestions, with no statistically significant improvement in reflections or questions.
The practice-plus-feedback group improved on reflections and questions within that group. Its empathy scores did not change significantly. But when the researchers compared the groups' changes directly, only empathy showed a statistically significant difference. The within-group improvements in questions and reflections do not establish that feedback outperformed practice alone on those skills.
These are the directions reported in CARE's March 2026 manuscript. Its table and prose differ slightly on some numerical results, so the account here does not rely on their exact magnitudes.
The authors offer two likely explanations. Participants may have adapted to patients who resisted suggestions but did not respond differently to empathic statements. Also, the post-test patient was less emotionally forthcoming, prompting more information-gathering at the expense of empathic reflections. These are proposed mechanisms from one immediate simulated counselling experiment, not an established general effect of unsupervised rehearsal.
Confidence went up regardless
Both groups reported modest gains in self-efficacy, and those gains did not track their measured behaviour. The practice-only group's confidence in exploration skills rose significantly while their empathy fell. Comparing participants' self-ratings against their actual classified performance, the researchers found that people in the lowest performance quartile overestimated themselves most, while the strongest performers slightly underestimated themselves.
Feeling more prepared may help you approach a session. It does not measure your coaching or establish that you are ready to work with paying clients.
What CARE cannot tell you
Performance was scored by fine-tuned language classifiers, not by human supervisors watching whole sessions. Three domain experts annotated a sample described as 10% of participants’ transcripts to support classifier validation. Classifier performance varied by skill; the main text and appendix give different ranges for the selected classifiers, so their scores should not be treated as precise human judgements. The authors flag that fine-tuning on a subset of the study's own data raises "potential concerns about data leakage and overfitting", and that utterance-level microskills miss session-level qualities.
The design measured immediate change within a single 75-minute lab session, "rather than long-term retention or transfer to real-world clinical encounters with actual patients". There was no non-AI comparison group at all: no classroom training arm, no human roleplay arm. As the paper puts it, without direct comparisons to established training methods or evaluation in real educational settings, "we cannot determine the relative effectiveness, acceptability, or practical integration challenges of AI-enhanced training".
And it is counselling. Empathy, reflections and open questions matter in coaching too, but a coaching session is not a counselling session, and transfer between the two professions is an assumption rather than a finding.
The qualitative picture
A smaller study asks a different question. Counsellor trainees in a Texas programme ran a 30 to 40 minute text session with ChatGPT playing a fictional client, then discussed it (Behavioral Sciences, 2025). Nine submitted transcripts; four joined the focus group.
They described a low-pressure place to rehearse before working with real clients, and they were candid about the limits: an overly agreeable client, thin emotional nuance, a persona that stayed culturally neutral unless prompted otherwise, and sessions that tidied themselves up too quickly. They came out calling it a supplement to live practice, not a substitute for it.
These accounts describe how practice felt; the study did not measure whether anyone coached better afterwards.
The wider review
A 2026 systematic review's abstract reports searches of eight databases up to March 2024 and 16 included studies covering 2,312 participants (Passmore, Olafsson and Tee, Journal of Work-Applied Management). It says AI coaches can be useful and match human coaches on specific tasks.
That is a claim about AI doing coaching. The abstract does not establish that a human learns to coach better by practising with a fictional AI client. This account is limited to the abstract; it cannot assess the individual studies or their quality assessments.
Watching for change in your own practice
An informal before-and-after record cannot isolate what caused a change. You can still make the behaviour you're watching specific enough to notice whether it changes.
Pick one behaviour, in observable terms: "I check what the client wants from this session before I start exploring." Then, over a few weeks:
- Note where you currently are, using a session with a consenting practice partner rather than a rehearsal. One or two lines is enough: did you do it, when in the session, what happened next.
- Rehearse that one behaviour with the AI client. Don't change three things at once.
- Go back to a human session and note the same thing again.
- Ask your partner what they noticed, and ask about the behaviour rather than about you. "Did I check what you wanted from the session, and did it land?" gets you further than "how did I do?"
- Check whether it survives a month later, when it isn't the thing you're working on any more.
That gives you a record of what changed for one person in one setting. Other things changed at the same time, including the fact that you were paying attention. Treat it as a reason to keep going or to stop, not as a finding.
What would actually answer the question
If a study lands claiming AI practice improves coaching, these are the things to look for:
- Coaches, not counselling students, and not undergraduates recruited as coachees.
- A comparison group doing something real, such as peer practice or supervised roleplay, rather than nothing.
- Behaviour rated by qualified people who don't know which group the coach was in.
- Sessions with actual clients, some weeks after the training ended.
- Some measure of what the clients got, and a look at whether anything got worse.
Sources
- Can LLM-Simulated Practice and Feedback Upskill Human Counselors? A Randomized Study with 90+ Novice Counselors, Proceedings of CHI 2026, DOI 10.1145/3772318.3791821. 94 novice counsellors, single 75-minute lab session, automated utterance classifiers, no non-AI control group. This account uses arXiv v2, 9 March 2026; numerical inconsistencies limit precision.
- Learning Through Simulation: Counselor Trainees' Interactions with ChatGPT as a Client, Behavioral Sciences, 2025. Nine transcripts, four focus-group participants, one fictional case; qualitative.
- Passmore, Olafsson and Tee, systematic literature review of AI in coaching, Journal of Work-Applied Management, 2026, DOI 10.1108/JWAM-11-2024-0164. Publisher-deposited abstract: eight databases searched to March 2024; 16 studies, n=2,312. AI-delivered coaching; full review not assessed here.