Ask Spotlight: How to Interrogate Your Customer Analytics
Most Ask Spotlight conversations run to about two messages. The longest one we found ran to forty-eight. Both can be
A scorecard is the difference between a subjective impression and a repeatable quality standard, but most teams never see what a well-built one actually looks like. This guide breaks down real scorecard anatomy, what categories teams typically measure, how weighting and auto-fail criteria should work, and the specific ways scorecards fall short when they're the…

A quality monitoring scorecard is the standard every reviewed conversation gets measured against: the difference between “that call felt fine” and an actual, repeatable answer to what “good” looks like. This article covers what scorecards look like in practice, what teams typically choose to measure, how scoring and weighting actually work, and where scorecards fall short when used on their own.
This is part of our series on contact centre quality assurance: read the full pillar guide for the broader picture.
A quality monitoring scorecard is a structured set of criteria used to evaluate a customer conversation, a call, chat, or email, consistently, against a defined standard. Instead of a reviewer judging a conversation on gut feel, a scorecard breaks quality down into specific, answerable questions: was the required disclosure given, did the agent verify identity, did they show empathy, was the issue resolved. Each answer contributes to an overall score, which is what actually gets tracked, reported, and coached against over time.
The point of a scorecard is that two different reviewers, scoring the same conversation, should land on roughly the same result. Without a shared, specific standard, that consistency is impossible.
Most scorecards are organised into a handful of sections, each containing a small number of specific questions. A simplified example for a support call might look like this:
| Section | Example criterion | Scoring type |
| Opening | Did the agent greet the customer and confirm the reason for the call? | Yes / No |
| Compliance | Did the agent verify the customer’s identity before discussing account details? | Yes / No — auto-fail |
| Accuracy | Was the information or policy the agent gave correct? | Yes / No |
| Soft skills | Did the agent show empathy appropriate to the situation? | 1–5 scale |
| Resolution | Was the customer’s issue resolved, or was a clear next step agreed? | Yes / No |
| Closing | Did the agent summarise the outcome and close professionally? | Yes / No |
A real scorecard usually has more granularity than this; most teams land somewhere between 10 and 15 individual criteria, enough to be meaningful without turning every review into a lengthy audit. Scorecards are also commonly customised by role and by channel: a sales team’s scorecard weights persuasion and upselling; a technical support scorecard weights troubleshooting accuracy; a chat scorecard drops criteria that only make sense for voice, like tone of voice or hold-time handling.
Scorecard criteria tend to fall into a handful of recurring categories, regardless of industry:
Many scorecards also include non-scoring questions; criteria tracked for reporting purposes without affecting the agent’s overall score, such as flagging that an issue was caused by a system outage or policy gap rather than agent error. These don’t penalise the agent, but they give QA teams visibility into problems that sit outside any individual agent’s control.
Not every criterion on a scorecard should carry the same weight, and how a team handles that distinction matters more than the specific questions chosen.
Weighted vs. equal scoring: Treating every criterion as equally important is a common starting point for a new scorecard, but it’s rarely accurate; a missed compliance disclosure and a slightly flat greeting are not the same severity of issue. As teams accumulate data, the more rigorous approach is to weight criteria based on how strongly they actually correlate with outcomes like CSAT, complaints, or repeat contact, rather than by instinct alone.
Auto-fail criteria: Some failures are severe enough that they should void the entire score, regardless of how well the rest of the conversation went; a compliance breach, a missed identity check, sharing information the agent shouldn’t have. The most common design mistake is treating a critical criterion as just a heavily weighted item rather than a true auto-fail; a well-designed scorecard keeps a small number of clear, binary auto-fail criteria — usually three to five — separate from the weighted scoring entirely, so a serious violation can’t be averaged out by an otherwise good call.
Keeping criteria answerable: A good scorecard question can be answered consistently by two reviewers who’ve never spoken to each other. “Was the agent rude?” is too subjective to score reliably. “Did the agent interrupt the customer more than twice?” is answerable. This distinction is most important for auto-fail criteria specifically, since the consequence of getting one wrong is severe.
Ordering the scorecard: Most teams find it easier for reviewers to score consistently when scorecard sections follow the natural flow of a conversation – opening, then process and compliance, then resolution, then closing – rather than grouping by category in a way that has the reviewer jumping around the transcript.
A scorecard is a necessary tool, but it isn’t a complete quality programme by itself, and treating it as one creates its own problems:
None of this is an argument against scorecards, it’s an argument for treating them as one part of a QA programme, alongside calibration, coverage, and a willingness to retire criteria that have stopped earning their place.
A quality monitoring scorecard is a structured set of criteria used to evaluate a customer conversation consistently against a defined standard, covering things like compliance, accuracy, tone, and resolution. It’s what turns a subjective impression of a conversation into a specific, comparable score.
Most scorecards cover five recurring areas: compliance and process adherence, accuracy of information given, soft skills and tone, resolution effectiveness, and general professionalism. Most teams keep the total list to roughly 10–15 criteria to stay specific without turning every review into a lengthy audit.
An auto-fail is a scorecard criterion serious enough that failing it voids the entire score, regardless of how well the rest of the conversation went — typically compliance or data-handling violations. Well-designed scorecards keep auto-fail criteria few, clearly answerable yes or no, and separate from weighted scoring entirely.
No. Equal weighting is a reasonable starting point for a brand-new scorecard, but it treats every behaviour as equally important, which usually isn’t accurate. More mature programmes weight criteria based on how strongly they actually correlate with outcomes like CSAT or repeat contact.
No. A scorecard defines the standard, but it only measures whatever share of conversations actually gets reviewed against it. A well-designed scorecard applied to a 2–3% manual sample still leaves the vast majority of interactions unmeasured — coverage and calibration matter as much as the scorecard itself.
This article is part of a broader series on contact centre quality assurance:
EdgeTier Coach keeps scorecards fully customisable, so teams can build criteria around their own compliance requirements, channels, and role types rather than working within a rigid template and adjust them as products or policies change, without needing to coordinate a manual rollout across every reviewer.
Because Coach’s AI completes scorecards automatically across 100% of conversations, the coverage limitation described above largely disappears: the scorecard gets applied consistently to every call, chat, and email, not just a manually selected sample. Heat-map analytics then make it easy to spot exactly which scorecard questions have stopped providing useful signal, criteria everyone passes every time, or questions where scoring drifts noticeably by reviewer or team, so scorecards can be refined based on real data rather than guesswork.
Most Ask Spotlight conversations run to about two messages. The longest one we found ran to forty-eight. Both can be
Learn how we evaluated EventBridge, Kinesis, SNS, and SQS FIFO to build a secure, ordered, multi-tenant events API architecture on
Contact centres don't lack data, they lack fast answers. This guide compares the 10 best call center analytics software platforms
"It has reduced the time for the quality assurance process as it provides clear data and a very robust direction on where to look and what matters the most."
"EdgeTier is really shining when it comes to responsible gambling. We can proactively track critical issues and take actions, reducing human error."
"You’ve got an issue, but you don’t know how many people are affected. You don’t know the scale. You don’t even know if it’s real."



Let us help your company go from reactive to proactive customer support.
Unlock AI Insights