Skip to main content
This guide explains how Real Talk Studio scores a session. Share it with client admins, L&D leads, and programme managers when someone asks “How is my score calculated?”
This page covers 1:1 roleplay — the most common session type. Team roleplay and presentation practice follow similar principles but may differ in detail.

The short answer

Scoring has two ingredients: 1. What you set up in the scenario
When authoring a scenario, you define a learning objective and up to three skills — each with a name and description. Optionally you can add good/bad examples and specific achievements to look for. This is the foundation: the session is scored against your criteria, not a generic checklist.
2. What happened in the conversation
When the session ends, the transcript is reviewed for evidence that the learner demonstrated those skills and met the learning objective.
From that review, learners receive:
  • A rating for each skill — from Novice to Expert, with examples and tips
  • An overall score out of 100 — how well they met the learning objective as a whole
  • Highlighted moments from the conversation — where they did well and where they could improve
Scoring rewards meaningful engagement and demonstrated competence, not length alone. Very short sessions or conversations that never get going will score low or may not receive full feedback at all.

Ingredient 1 — Skills and the learning objective

Everything starts with how the scenario is configured.
Write skill descriptions that describe observable behaviour — what someone would actually say or do in the conversation — rather than abstract traits like “good communicator.” The description is the main criterion for the rating. Good/bad examples help the system recognise those patterns in the transcript and write sharper comments on specific moments — one clear pair per skill is enough. Extra variants rarely change the rating.
See Scenario creation for the full authoring workflow.

Ingredient 2 — Reviewing the transcript

When the session finishes, the conversation transcript is analysed for evidence of the skills in action and progress toward the learning objective. The review looks at what the learner said and did — not what they intended. Each skill is judged on the whole conversation: how often, how clearly, and how consistently that skill showed up. The system is looking for evidence of the skill you described, not a keyword match or a single required phrase. One good moment on its own is usually not enough for a high rating. Learners also see key moments pulled from their conversation — where a skill was demonstrated well, or where there was a missed opportunity. The full transcript is always available for admins and reviewers.
When a learner asks “How did I get this rating?”, open their session — not this page. Point them to the score explanation on Snapshot, the quotes, strengths, and tips under each skill on Dig Deeper, and the highlighted turns on Review. Those are the evidence for that attempt. Share this page with admins and stakeholders who want the method.

What learners receive

Per-skill rating (Novice → Expert)

Each skill gets one of five levels: Each rating comes with specific examples from the conversation, strengths, and practical tips for next time.

Overall score (0–100)

This is a separate assessment of how well the learner met the learning objective across the whole conversation. It is not a simple average of the skill ratings. In general:
  • Low scores (roughly below 50) — limited engagement, missed the objective, or significant gaps in execution
  • Mid scores (roughly 50–70) — partial success, some skills shown but inconsistent
  • Strong scores (roughly 70–89) — clear progress toward the objective, skills applied well
  • Exceptional scores (90+) — outstanding performance; intentionally hard to achieve
Each score includes a short explanation of why it was given.

What affects the score — and what does not

What matters

  • Engaging meaningfully with the scenario topic
  • Demonstrating the defined skills through what they actually said
  • Consistency across the conversation, not just one strong reply
  • Appropriate tone and professionalism for the situation

What does not automatically help

  • Longer conversations do not automatically score higher. A long but unfocused session can still score poorly. What counts is the quality of engagement, not the word count.
  • The overall score is not an average of skill ratings. A learner could do well on one skill but still miss the broader objective.
  • Very short sessions may not receive full feedback. Scenarios can require a minimum session length, and sessions with almost no content will score at or near zero.

Common questions

No. Length gives more to assess, but quality matters more than quantity. A focused session that meets the minimum can score well; a long, unfocused one can score poorly.
No. The 0–100 score and the per-skill ratings are assessed independently. They usually align, but not always — a learner might demonstrate one skill well yet miss the overall objective.
Usually because the session was too short, contained only greetings, or the learner did not engage with the topic. These sessions do not provide enough evidence to assess fairly.
Scores in the 90s are reserved for truly exceptional performance — full mastery of the objective and consistent skill demonstration throughout. Most strong sessions land in the 70s or 80s.
Scores reflect a specific conversation. Like human coaching, two sessions of similar quality may receive slightly different scores depending on the evidence in each transcript.
Their own completed session. The rating is explained there: Snapshot has the overall score and why it was given; Dig Deeper has each skill rating with quotes, strengths, and tips; Review highlights the turns that counted. This page is for admins and stakeholders who want the method — learners should start from the evidence in their conversation.
Both, in different ways. The rating is holistic — how well the skill showed up across the whole conversation, not a tick-box for one phrase. The comments then cite particular turns as evidence. Optional achievements are the binary checks: did this specific outcome happen, or not.
Prefer one strong pair per skill over a long list. The skill description is what the rating is judged against. Examples help illustrate what strong or weak looks like in this scenario, which sharpens the moment-level comments. After one clear “do this / don’t do this” pair, more examples rarely change the rating. Keep them short, scenario-specific, and observable — a phrase someone would actually say, not a paragraph of theory.
That skill is still rated. With little or no evidence, it typically lands Novice or Beginner, and the comments will say the skill was not demonstrated. It is not skipped or marked N/A. If a skill often does not come up, remove it or rewrite the scenario so the conversation naturally reaches it. Use an achievement if you need a yes/no check for a specific moment.
Yes, the cap of three is deliberate. Feedback stays focused when each skill is distinct and likely to appear. Two well-chosen skills often beat three overlapping ones. Three is a maximum, not a target. If a skill is a stretch for the scenario, leave it out — unused skills pull ratings down and dilute the debrief.
They should land in the same band (for example both in the 70s, both Advanced), not as identical numbers. The avatar will not say the same thing twice, so the evidence is never a carbon copy. Treat the score as a coach’s judgement of that attempt, not a lab measurement. If two attempts feel the same to you but the band jumps (say 58 vs 81), that is worth a review — a point or two, or one skill level, is normal.

Worked example

A coaching scenario with this setup: A six-minute session that asked good questions but never booked the follow-up might receive: The learner sees those quotes and tips on Dig Deeper, and the same turns marked on Review. That is the explanation of the rating — not a hidden formula.

Suggested talking points

When a learner or line manager asks about their score:
  1. Open their session first — the score explanation, skill quotes, and highlighted turns are the rationale.
  2. The score reflects the learning objective — not attendance, effort alone, or how long they talked.
  3. Skills are defined upfront — the scenario author decides what good looks like for that conversation.
  4. Feedback is evidence-based — ratings and highlights refer to specific moments from the conversation.
  5. A missed skill still gets a rating — no evidence usually means Novice or Beginner, not “not applicable.”
  6. Low scores are actionable — learners get concrete next steps and can practise again immediately.
  7. Practise again is the point — a score measures one attempt; improvement comes from repeated practice.