How scoring works
The AI judge, evidence citations, why it's strict, and what to do about a score you disagree with.
The judge
An AI judge scores each call against the rubric you control. It reads the transcript, applies the dimensions and weights from the scorecard, and produces a score per dimension plus a total.
You control the rubric. You do not have to trust the judge's taste, because of the next part.
Every rating cites its evidence
Each rating must cite the exact moment in the transcript it came from. That's the design constraint that makes the output checkable: you can read the moment and decide whether the rating is fair, instead of accepting or rejecting a number on faith.
Practically, this means the right way to read a score is:
- Look at the dimension breakdown, not the total.
- Open the evidence for the lowest dimension.
- Replay that moment (scores and feedback).
If you find yourself arguing with a number without having read its citation, you're arguing with the wrong thing.
It's built to be strict
Sessions fail. That's intentional: a generous score tells a manager nothing, and a failing score with evidence tells them exactly what to coach. Expect first attempts against a hostile buyer to score low, and expect the distribution to look harsher than a human coach's would.
A rep scoring 19 out of 100 on their first hostile-buyer roleplay is a normal starting point, not a performance problem. What matters is the second and third attempt on the same scenario.
Not scored
A call can come back Not scored:
- it was too short to evaluate — a few seconds of audio isn't a conversation
- scoring hasn't finished yet
- the transcript didn't contain enough of the rep's speech to judge
"Not scored" is not a zero. It means there's nothing to score, and the fix is to run the session properly rather than to argue with the result.
Consistency between calls
Two things make scores comparable:
- Structured templates. The rubric has explicit dimensions and rules, so the same behaviour gets rated the same way across calls and across weeks.
- Not editing the rubric mid-cycle. If you change dimensions or weights, scores before and after aren't comparable. Create a new template instead of mutating a live one.
Disagreeing with a score
The process that works:
- Read the citation. Most disagreements dissolve here — the judge is usually pointing at something real.
- If the citation is right but the rating is wrong, your rubric is underspecified. Say explicitly in the scorecard what should have scored higher and why.
- If the citation is wrong — it's quoting the buyer as the rep, or citing a moment that doesn't support the claim — that's a scoring error worth reporting, with the session and the dimension.
- Re-score after fixing the rubric. Don't hand-adjust individual scores; that breaks comparability, which is the only reason the scores are useful.
What scoring is not for
- Not for compensation. Tie pay to roleplay scores and reps optimise for the bot. You lose the diagnostic value of comparing practice to real calls.
- Not for stack-ranking on one call. A single call is noisy. Trends per rep and per dimension are the signal.
- Not a substitute for listening. The score tells a manager where to listen. It doesn't replace the listening.