Half the Calls Your QA Team Reviews Never Happened

We analyzed over 400,000 recorded sales calls. Nearly half ended in under two minutes. What non-connects do to your QA sample and your quality scores, and the boring fix that works.
Shishir Agarwal
August 2026

We analyzed over 400,000 recorded sales calls, more than 30,000 hours of audio. Nearly half ended in under two minutes. One in five never crossed the first minute.

Those numbers surprise nobody who has run an outbound floor. Every sales leader knows most dials do not connect: voicemails, instant hangups, wrong numbers, "call me later" brush-offs that last eleven seconds. The connect rate is a fact of life, and teams manage it as one.

Here is the part that gets missed. Modern dialers and telephony stacks record everything, and everything means everything. The voicemail gets a recording. The four-second hangup gets a recording. Each of those recordings lands in the same pool as your real conversations, and that pool is what your quality program draws from. The non-connects do not just waste airtime. They contaminate two of the most important numbers your QA function produces.

Histogram of recorded sales calls by duration: 47% end under two minutes. Gistly call corpus, August 2026.
How long recorded sales calls actually run. Share of the corpus by duration bucket; exact percentages and the snapshot date are in the methodology note below.

Problem one: your QA sample

Most manual QA programs review a small slice of calls, typically 1 to 2 percent, sampled more or less at random from the recording pool. Random sampling has an underrated property: the sample inherits the composition of the pool. If half the pool is voicemails and instant hangups, half of your sample is voicemails and instant hangups.

Play that forward. Your auditors sit down to review their weekly allocation, and every second recording contains no conversation at all. There is nothing to score, nothing to coach, nothing to learn. Half of your audit effort, half of the salary cost, half of the calendar time, goes to calls where nothing happened.

The arithmetic on the other side is worse. A roughly 2 percent sample, applied to a pool where only about half the recordings are real conversations, means a real conversation gets reviewed roughly once in a hundred. The calls that carry all of the risk, the mis-selling, the missed discovery, the compliance slips, are precisely the calls your program almost never sees.

Problem two: your quality score

The second distortion is quieter and harder to catch, because it hides inside an average.

Run a ninety-second call against a full quality rubric and it fails everything. No opening, because the prospect hung up. No discovery, no objection handling, no closing, because there was no conversation to have. The rep did nothing wrong, and the scorecard reads like a disaster.

Leave those calls in the denominator and your "quality score" stops measuring quality. It becomes a blend of two unrelated things: how well your reps sell, and how often phones get answered. When list quality dips or a carrier has a bad week, your connect rate falls, your short-call share rises, and your quality score sinks, with not a single rep selling any differently. Teams then chase the phantom: retraining on openings that were never given a chance to happen, coaching against a number that was never about coaching.

We watched this happen in production

This is not a hypothetical. In a real deployment, a scoring template showed entire parameters failing for every single call. Every call, zero percent pass. Read at face value, the floor had a catastrophic skills problem.

It did not. The cause was non-connected calls being scored as if they were conversations. Short recordings with no dialogue were being run through parameters that assume a dialogue exists, and they failed them all, dragging the aggregate to the floor. The fix was not a coaching program. It was a classification gate.

The fix is boring, and it works

  1. Classify every call first. Before any scoring happens, decide: was this a conversation or a non-connect? Duration is a strong first signal; a transcript check settles the ambiguous middle.
  2. Exclude non-connects from quality metrics automatically. They belong in your connect-rate and dialing metrics, which are worth tracking, as their own numbers. They do not belong in a quality denominator.
  3. Audit conversations, not recordings. Sample from the pool of real conversations. The same auditor headcount roughly doubles its effective coverage, and every reviewed call is one a rep can actually be coached on.

None of this is sophisticated. That is rather the point. The distortion comes from a default nobody chose, and it disappears the moment someone chooses otherwise.

Try this on your own data this week

Pull the duration histogram from your dialer. It is a five-minute export. If a large share of short calls is sitting inside your QA averages, your quality number is not measuring what you think it measures, and your auditors are spending real hours on calls that never happened.

Methodology

All figures in this article trace to a read-only snapshot of the Gistly production corpus taken on August 13, 2026. The snapshot queries are archived and reproducible. Exact counts appear once, here; the article body uses rounded figures by house style.

  • Corpus: 422,834 recorded sales calls analyzed, totalling 30,597 hours of audio, collected from January 2023 through August 2026. Duration shares are computed over the 422,806 calls with a recorded duration.
  • Composition: the corpus is weighted toward India-based ed-tech sales and BPO operations. Treat the percentages as descriptive of that mix, not of every sales team everywhere. The under-two-minute share also varies several-fold across operation types in the corpus: counselor-led sales operations sit well below the blended 47%, while support and collections operations sit well above it. Check your own histogram rather than assuming the blend.
  • Duration mix: under 1 minute: 21.8%; 1 to 2 minutes: 25.2%; 2 to 5 minutes: 26.8%; 5 to 10 minutes: 15.7%; 10 to 20 minutes: 7.9%; over 20 minutes: 2.6%. "Nearly half ended in under two minutes" is the sum of the first two buckets: 47.0%.
  • The "half of your audit effort" arithmetic assumes manual QA sampling is random with respect to call duration. If your team already filters short calls before sampling, your exposure is smaller than described.
  • The 1 to 2 percent manual review rate is an industry framing for audit-depth review programs, not a measurement from this corpus.
  • The deployment example is a real customer engagement reported in anonymized, aggregate form. We do not name clients, identify industries at the account level, or quote call content.

See what your calls look like without the noise

Gistly classifies every call automatically, gates non-connects out of your quality metrics, and scores 100 percent of real conversations. Bring your own duration histogram and we will walk through it together. Book a walkthrough.

Get a live walkthrough from the founder.

30 minutes. No SDR, no script. Book directly with Ashit, founder of Gistly.

Book 30 min with the founder →

Explore other blog posts

see all