Research

7 MIN READ

What VR Public-Speaking Coaching Should Measure Beyond Confidence

Adaptive VR coaching can improve how practice feels without proving better real-world performance. Use this five-level evidence ladder to plan the next test.

Adaptive VR coaching can improve how practice feels without proving better real-world performance. Use this five-level evidence ladder to plan the next test.

A blue clay speaker considers a blank cue card offered by a lavender coach beside a lectern and a paper audience.

SNIPPETS BLOG / ARTICLE

SNIPPETS BLOG / ARTICLE

Feeling calmer in a virtual rehearsal is a useful result. It is not the same result as speaking better in front of a real audience.

That distinction matters as training teams add adaptive feedback, large language models and embodied coaches to VR practice. A system can make practice feel safer, more personal and more actionable while leaving its effect on workplace performance unknown. The right response is not to dismiss the early result. It is to name precisely what has been learned and design the next test.

A practical evaluation should climb five levels: practice response, self-report, in-environment behavior, retention and real-world transfer. Each level supports a different claim. Skipping a level does not make a product more advanced; it makes the evidence harder to interpret.

What one randomized VR coaching study actually found

A 2026 Frontiers in Psychology paper provides a useful example. The researchers randomized 120 participants into two equal groups. Both delivered two speeches in a VR classroom and received a feedback segment of similar duration. One group received adaptive, performance-contingent coaching from embodied virtual judges. The control group received generic feedback.

The adaptive system combined automatic speech recognition, structured LLM prompting, text-to-speech, gaze-related cues and supportive, non-shaming feedback. It did not interrupt the speech. Coaching arrived between rounds so participants could apply a short, specific suggestion in the next attempt.

The adaptive group reported larger reductions in state anxiety and public-speaking anxiety appraisals across the repeated speaking rounds. Lower self-reported anxiety remained observable when participants returned after one week and one month.

Those follow-ups happened in the same VR classroom with the same equipment and comparable tasks. The primary outcomes were self-reported. The study did not include blinded external ratings of speaking performance, physiological outcomes or a test before a real audience. Its authors therefore described the result as maintenance within the VR practice setting, not proof of stable trait change, objective performance improvement or real-world transfer.

That boundary is the most useful part of the result for a training team. It shows how an encouraging outcome can be both meaningful and incomplete.

The Coaching Evidence Ladder

Use the ladder before deciding what to instrument, what to promise or what to test next.

1. Practice response

First ask whether the learner can and will use the experience. Completion, repeated attempts, dropout, clarification requests, technical failures and willingness to practise again belong here.

This level tells you whether the practice loop is operable and acceptable. It does not tell you that the learner improved. A polished session with high satisfaction can still teach the wrong behavior or avoid the pressure that matters outside VR.

2. Self-report

Next measure what the learner says changed: anxiety, confidence, perceived usefulness, cognitive load, trust or sense of safety.

Self-report is especially relevant when the experience is designed to make a difficult task feel manageable enough to repeat. It can reveal a real product benefit. But the wording of the claim must stay close to the measure. “Participants reported lower anxiety in the practice setting” is defensible. “Participants became better public speakers” is not the same claim.

3. In-environment behavior

Now ask whether observable behavior inside the simulation changed. Depending on the task, that might include speaking duration, pause distribution, gaze coverage, response timing, task completion or an independently rated performance recorded in VR.

Behavior needs a defined rubric and a reason to matter. A gaze pattern is not automatically a performance score. More words are not automatically a clearer speech. Choose measures that connect to the instructional objective and keep exploratory signals separate from confirmatory outcomes.

A clay speaker climbs a five-step staircase from a small VR practice stage toward a separate live audience platform, with a visible gap before the final step.
The final step is deliberately separate: improvement inside the practice environment does not establish transfer by itself.

4. Retention and variation

Test whether the change persists after time has passed and when part of the practice context changes. Use a new room, new audience behavior, a new topic, different equipment or a delayed session.

Repeating the same context can still answer a useful maintenance question, but familiarity and habituation remain plausible explanations. Variation makes the retention claim stronger because it asks whether the learner can carry the behavior beyond one rehearsed configuration.

5. Real-world transfer

Finally, test the performance the organization actually cares about: a live presentation, interview, safety briefing, customer conversation or assessed workplace scenario.

Define success before the intervention. Use an appropriate comparison, blinded raters where practical and a rubric tied to the real task. Measure both benefit and failure: clearer delivery, stronger audience understanding, fewer critical errors, better recovery after interruption or lower anxiety under authentic pressure.

This is the level required for a real-world effectiveness claim. It is also the most expensive and context-sensitive level, which is why it should be planned before production rather than added as an afterthought.

Turn the ladder into an evaluation plan

Start with the claim the team eventually wants to make, then work backwards.

If the goal is a lower-pressure rehearsal tool, practice response and self-report may be sufficient for the first release decision. If the goal is better eye contact in a virtual assessment, add a behavior measure with a pre-defined rubric. If the commercial promise is better workplace presentations, the plan needs delayed, varied practice and a real-audience transfer test.

For each level, record five things:

  1. Claim: What exact sentence could this result support?

  2. Outcome: What will be measured and by whom?

  3. Context: Where, when and under what pressure will it be measured?

  4. Comparison: What would have happened with generic feedback, repeated exposure or no coach?

  5. Failure condition: What result would make the team narrow or reject the claim?

This forces a useful distinction between an experience metric and an effectiveness metric. Satisfaction can justify improving the experience. It cannot substitute for the performance outcome promised to a buyer.

Keep the coaching system and the performance layer separate

An adaptive VR coach is a system, not a single feature. Speech recognition, LLM instructions, scenario state, safety rules, feedback timing, assessment and analytics all shape the result. The character's voice, gaze, facial performance and body language shape how that feedback is delivered and received.

Snippets is a controllable production layer for creating and delivering synchronized 3D character performances in Unity. It can support authored coach, judge or audience performances that combine voice, lip-sync, animation, gaze, text and runtime events. The surrounding application owns real-time conversational AI, speech recognition, branching, learner scoring, analytics, LMS integration and domain validation.

That separation improves evaluation. A team can test whether the wording is adaptive, whether the character delivery feels supportive and whether the learner behavior changes without treating the whole experience as one opaque “AI coach” effect. It also makes failure easier to diagnose: a relevant suggestion can arrive with poor timing, a credible performance can carry generic advice or an enjoyable session can miss the transfer task.

Decide the next rung before building the next feature

The 2026 study gives credible preliminary evidence for a specific result: a combined adaptive coaching package was associated with larger self-reported anxiety reductions than attention-matched generic feedback during repeated VR practice, with the lower self-reports still observable in the same VR context at follow-up.

It does not answer whether participants became better speakers before a live audience. That is not a flaw to hide. It is the next research question.

For a training team, the practical move is to name the highest rung already earned and define the next rung before expanding the feature set. A coach that makes practice feel manageable may be worth building. A coach sold on workplace effectiveness needs evidence that leaves the headset.

For the wider instructor-design decision, read Designing Digital Instructors for Better Learning.

Sources

Revolutionising how you create and manage 3D character content.

©

2026

Snippets3D. All rights reserved.

Revolutionising how you create and manage 3D character content.

©

2026

Snippets3D. All rights reserved.

Revolutionising how you create and manage 3D character content.

©

2026

Snippets3D. All rights reserved.