Taizen
Scoring

Score real calls

Apply your scorecards to actual customer calls, and compare practice against the job.

Scoring → Real calls applies your scorecards to actual customer conversations from your connected call source. Same rubrics, same evidence requirement, same dimensions as practice.

Prerequisites

  • A connected call source — Gong, Genesys Cloud, or Fireflies — with transcripts syncing
  • At least one scorecard that matches the call type you're scoring

Settle the rubric on practice first

A scorecard that produces fair, evidence-backed scores on roleplays is one you can trust on real calls. Tune it against practice sessions where the stakes are low, then point it at live conversations — not the other way round. See the scorecard builder.

Scoring at volume

Scoring every call by hand doesn't scale. Attach it to an agent instead: an event-triggered agent that fires when a call ends, scores it against the right scorecard for that call type, and delivers the result to the rep and their manager. See triggers and the cookbook.

Two things to get right when you automate it:

  • Match the scorecard to the call type. A demo scored on a cold-call rubric produces nonsense. Route by call type, or run only on the call type you care about most.
  • Exclude internal calls. Team syncs scored as customer conversations pollute every average.

Reading real-call scores

The same rules as practice: dimensions over totals, evidence over numbers (how scoring works). Two differences worth knowing:

  • Real calls are messier. Multi-party calls, partial recordings, calls that get cut short. Expect more Not scored results than in practice sessions.
  • The stakes make the score more useful, not more accurate. A low score on a real call is a coaching conversation, not a verdict on the deal — the deal outcome is its own signal.

The comparison that matters

This is why both surfaces share one rubric. Put a rep's practice scores and real-call scores side by side on the same dimensions:

PracticeRealDiagnosisAction
UpUpDrills are transferringKeep the plan, raise the bar
UpFlatLearning the bot, not the buyerChange the scenarios: build bots from real lost deals (build a bot)
FlatUpPlan isn't testing the real jobRebuild the plan around what real calls demand
FlatFlatNothing has changed yetCheck reps are completing stages (performance)

Feeding it back

Real-call scores are the best source of material for the next training plan. If the whole team is weak on one dimension across real calls, that dimension becomes the next stage — with a bot built from the deals where it cost you.

The full loop, in one line: real calls reveal the weakness → a training stage drills it → practice scores show progress → real calls confirm it transferred.

A note on trust

Reps will ask whether they're being graded by a machine. The honest answer: the rubric is written by your enablement team, every rating cites the moment in the transcript it came from, and the score exists to tell a manager where to listen. If you can't say that truthfully about your setup — because nobody reads the evidence, or because scores feed a ranking — fix that before you roll real-call scoring out to the team.

On this page