Ask Spotlight: How to Interrogate Your Customer Analytics
Most Ask Spotlight conversations run to about two messages. The longest one we found ran to forty-eight. Both can be
Call center quality monitoring is the day-to-day act of observing and scoring customer conversations; the operational engine underneath a wider quality assurance programme. This article covers how monitoring actually runs week to week, how teams decide what gets sampled, where coverage gaps form, and the operational challenges that show up once a monitoring programme is…

Call center quality monitoring is the day-to-day act of observing and scoring customer conversations; the operational engine underneath a wider quality assurance programme.
This article covers how monitoring actually runs week to week, how teams decide what gets sampled, where coverage gaps form, and the operational challenges that show up once a monitoring programme is running for real.
This is part of our series on contact centre quality assurance: read the full pillar guide for the broader picture.
Call center quality monitoring is the process of observing, listening to, or reading customer conversations and scoring them against a quality standard. It’s the operational activity inside a broader quality assurance programme: QA is the discipline (standards, coaching, calibration); monitoring is the specific act of watching or listening to a conversation and recording how it measures up.
Monitoring happens in two distinct forms, and most contact centres run some mix of both:
Most of what follows in this article is about post-interaction monitoring, since that’s the process most contact centres actually run their QA scoring through.
Stripped down, day-to-day monitoring follows the same basic cycle regardless of team size:
In most manually run programmes, this cycle repeats on a fixed schedule (weekly or monthly reviews per agent) rather than continuously. That fixed cadence is itself a limitation: it means monitoring reflects a snapshot from whenever the sample was pulled, not what’s happening in the contact centre right now.
Take a QA analyst on a 60-agent travel support team, covering both calls and chat. On a given Monday, last week’s sample lands in her queue: eight calls and four chats, spread across the dozen agents she’s responsible for. One call is from an agent handling a cancelled-flight refund request. Scoring it against the scorecard, she catches something specific: the agent has quoted a 14-day refund window, when policy changed to 21 days the previous month. That’s a fail on the “accuracy” line of the scorecard, not on tone or effort (the agent handled the customer well), they just had outdated information.
She logs the score with a note explaining exactly where the policy reference went wrong, and flags it for a coaching conversation later in the week. Pulling up her notes ahead of that 1:1, she notices this isn’t the first time this month she’s seen the same 14-day figure quoted: two other agents made the same error in her sample. That’s no longer an individual coaching issue; it’s a training gap serious enough to raise at the team’s next calibration session, where the fix is to reissue guidance to the whole team rather than just correct one agent.
That one flagged sentence, buried in a few minutes of a single call, is exactly the kind of thing a monitoring programme is supposed to catch, and exactly the kind of thing that’s easy to miss if the week’s sample happens not to include it.
Since most teams can’t monitor everything, the sampling method, or how conversations get chosen for review, matters as much as the scorecard itself. In practice, teams tend to use some combination of the following:
Most manually run QA programmes lean on a mix of random and targeted sampling: a baseline random sample to catch the average conversation, plus targeted review for anything flagged as high-risk. The trade-off is unavoidable with a manual process: broaden the sample and reviewer time runs out; narrow it to targeted conversations only and the average conversation goes unmonitored entirely.
In practice, the numbers here are smaller than most people expect. Industry estimates consistently put manual QA coverage at somewhere between 1% and 5% of total interactions; commonly cited guidance suggests reviewing a minimum of 5 to 10 calls per agent per month as a manual baseline.
That gap matters more than it might sound. One estimate from Clarity suggests that for an agent handling around 1,400 conversations a month, something in the region of 65 monitoring sessions would be needed to draw statistically reliable conclusions about their performance at a 90% confidence level, far more than most manual programmes review even for new agents, let alone established ones. In other words: most QA scores aren’t just incomplete, they’re often not statistically meaningful representations of an agent’s actual performance.
The coverage gap also isn’t evenly distributed. Random sampling misses rare-but-serious issues by design. Targeted sampling misses the ordinary, unremarkable conversations that make up the bulk of an agent’s day, which are exactly the conversations most likely to reveal a quietly forming habit before it becomes a bigger problem.
Beyond the coverage gap itself, a handful of problems show up consistently once a monitoring programme has been running for a while:
Call center quality monitoring is the process of observing or reviewing customer conversations, live or after the fact, and scoring them against a defined quality standard. It’s the operational activity that sits underneath a broader quality assurance programme.
Quality monitoring is the act of scoring individual conversations. Quality assurance is the wider discipline that includes monitoring, but also covers setting standards, calibrating reviewers, and turning scores into coaching and process improvement.
Manual programmes commonly aim for somewhere around 5 to 10 calls per agent per month as a baseline, though this is well below what’s needed for statistically reliable conclusions about an agent’s overall performance. AI-assisted monitoring can score every conversation rather than a fixed monthly sample.
Live monitoring involves listening to a call as it happens, sometimes escalating to real-time coaching or intervention. Recorded monitoring involves scoring a call, chat, or email after the fact against a scorecard; this is how most routine QA scoring is actually done, since it doesn’t depend on a reviewer being available in the moment.
Sampling is the method used to decide which conversations get reviewed, since most teams can’t review every interaction manually. Common approaches include random sampling, targeted sampling based on risk signals, and stratified sampling across agents or queues, each with different trade-offs, covered in detail above.
Yes. AI-assisted monitoring can score every conversation automatically against a scorecard, removing the need to choose a sample at all. Human reviewers then focus on calibration, coaching, and the conversations flagged as highest priority. See Automated Quality Management in Contact Centres for more detail.
This article is part of a broader series on contact centre quality assurance:
The sampling trade-offs covered throughout this article — random versus targeted, coverage versus reviewer time — exist because manual monitoring has a hard ceiling on how much a human can realistically review. EdgeTier Coach removes that ceiling by scoring 100% of chat, call, and email conversations automatically, so the sampling question largely disappears: every conversation gets scored, and reviewers spend their time on calibration, coaching, and the conversations AI flags as highest priority.
That shift also solves the cross-channel fragmentation problem directly: Coach applies the same scorecard logic across voice, chat, and email, so an agent’s quality is measured consistently regardless of which channel a conversation happened on. Heat-map analytics then surface exactly where scoring patterns differ across agents, teams, or scorecard questions, which is normally the hardest thing to see in a manually run, sample-based programme.
Most Ask Spotlight conversations run to about two messages. The longest one we found ran to forty-eight. Both can be
Learn how we evaluated EventBridge, Kinesis, SNS, and SQS FIFO to build a secure, ordered, multi-tenant events API architecture on
Contact centres don't lack data, they lack fast answers. This guide compares the 10 best call center analytics software platforms
"EdgeTier is really shining when it comes to responsible gambling. We can proactively track critical issues and take actions, reducing human error."
"We now have highly detailed understanding of agent performance, not just on key agent metrics, but also on how customers react to our agents and the emotions of our customers feel when talking to our team."
"We’re a big business, so getting the right people to agree and fix something hasn’t always been easy. Now we’ve got one version of the truth—it’s much easier to align and act"



Let us help your company go from reactive to proactive customer support.
Unlock AI Insights