Voice quality, turn taking, and reliability
The reliability metrics (relevancy, consistency, gibberish, and infrastructure) were near perfect across almost every agent, so the separation came from voice tone, humanness, interruption recovery, and latency. The widest spread by far was stop time after an interruption: from 0.62s (Bland) to 7.55s (Cognigy), a more than 12× difference in how quickly agents stopped talking when the caller cut in.
Ranked by voice quality composite
Scores marked /5 are qualitative (higher is better); % values are pass rates. Latency and stop time are in seconds, lower is better.
| Agent | Composite | Humanness /5 | Voice tone /5 | Interruption /5 | Latency ↓ | Stop time ↓ |
|---|---|---|---|---|---|---|
| 1Bland | 92.8 | 4.13 | 4.30 | 4.61 | 2.62s | 0.62s |
| 2Retell | 92.2 | 3.57 | 4.15 | 4.90 | 2.38s | 2.74s |
| 3Cresta | 90.1 | 3.60 | 4.06 | 4.97 | 2.78s | 1.56s |
| 4ElevenLabs | 90.0 | 3.93 | 4.08 | 5.00 | 3.08s | 1.23s |
| 5Decagon | 87.4 | 3.60 | 4.23 | 4.81 | 3.13s | 1.76s |
| 6Sierra | 85.5 | 3.63 | 3.81 | 4.67 | 3.42s | 1.06s |
| 7PolyAI | 82.6 | 3.73 | 4.22 | 3.68 | 2.99s | 3.97s |
| 8Vapi | 78.3 | 3.47 | 3.38 | 4.52 | 3.34s | 5.34s |
| 9Cognigy | 73.7 | 3.60 | 3.92 | 4.65 | 3.76s | 7.55s |
Every metric, side by side
One chart per metric, each agent bar for bar. The reliability metrics (relevancy, consistency, gibberish, infrastructure) sit near the ceiling for almost every agent; the real spread shows up in voice tone, humanness, latency, and stop time after an interruption.
Every agent, every metric
Open each agent for the full ten metric snapshot and the run review observation.
01BlandProduction voice agentComposite92.8+
Metric snapshot · three run average
Highest composite. Led humanness (4.13) and the quickest barge in recovery in the field (0.62s, averaged across every interruption) with fast turns. A couple of relevancy and gibberish edge cases were the only blemishes.
02RetellProduction voice agentComposite92.2+
Metric snapshot · three run average
The cleanest reliability sheet in the field: perfect relevancy, consistency, gibberish and infrastructure, zero unnecessary repetition, the fastest turns, and a solid 2.74s barge in recovery once first message overlaps are excluded.
03CrestaProduction voice agentComposite90.1+
Metric snapshot · three run average
Second-best interruption handling (4.97) with fast, clean turns and a quick 1.56s barge in recovery. A few gibberish edge cases were the only mark against it.
04ElevenLabsProduction voice agentComposite90.0+
Metric snapshot · three run average
Perfect interruption score (5.0), near-perfect no repetition, a fully clean reliability sheet, and a quick 1.23s barge in recovery. Mid-pack latency is the main thing left to improve.
05DecagonProduction voice agentComposite87.4+
Metric snapshot · three run average
A clean sheet on every reliability metric with strong voice tone (4.23) and a quick 1.76s barge in recovery. Mid-pack latency keeps it just outside the top group.
06SierraProduction voice agentComposite85.5+
Metric snapshot · three run average
A clean reliability sheet and the second-quickest barge in recovery (1.06s), but the lowest voice tone clarity in the upper group and the slowest turns (3.42s) hold it mid pack.
07PolyAIProduction voice agentComposite82.6+
Metric snapshot · three run average
Strong voice tone and a near clean reliability sheet, but the weakest interruption handling in the field (3.68) and a slow ~4s barge in recovery pull the composite down.
08VapiProduction voice agentComposite78.3+
Metric snapshot · three run average
Low unnecessary repetition (4.92), but the lowest voice tone clarity in the field (3.38) and a slow 5.34s barge in recovery place it near the bottom of this quality slice.
09CognigyProduction voice agentComposite73.7+
Metric snapshot · three run average
Solid interruption score and low repetition, but the slowest latency (3.76s) and by far the slowest barge in recovery in the field (7.55s across 21 interruptions), plus an infrastructure edge case, place it last.
Methodology
Each agent was reached on the phone number its business publicly lists for inbound customer calls, the same public support line any customer would dial, and driven by the same simulated caller through ten conditional action scenarios: barge in, silence, hold, hallucination bait, a vague caller, escalation, out of scope requests, a compound request, interrupt and switch, and a frustrated caller. We attributed each agent to a platform based on that platform vendor's own public claim of the customer, in its published case studies and customer references, and we retain the source for each attribution. Each row is a single production deployment; a business that has since changed vendors could still be attributed to its former one. Every agent ran the identical scenarios, caller personality, and metric set three times; only the number differed. The ten metrics shown are averaged over the three runs. Stop time after an interruption is averaged across every qualifying interruption rather than per call, so each barge in is weighted equally; and we ignore any interruption in the first ten seconds of a call, since a caller's opening words routinely overlap the agent's greeting and that is not a barge in the agent should be scored on.
Score your own production agent the same way
This is exactly what Cekura does for its customers: probe a live agent, score it on these ten metrics from real calls, and rerun the suite after every prompt or config change to catch regressions before they ship. Point Cekura at your own number and get this same scorecard.
Frequently asked questions
What exactly was tested?+
Nine real production voice agents, each running on a different platform (Bland, Retell, Cresta, Decagon, ElevenLabs, PolyAI, Sierra, Cognigy, Vapi), reached on their live public phone lines. Every agent got the same ten conditional action scenarios and the same simulated caller personality, three times each, 270 live calls in total.
How is the composite computed?+
Each of the ten metrics is normalized direction aware (higher is better metrics against their scale; latency and stop time inverted because lower is better), then averaged into a 0 to 100 composite. Boolean reliability metrics are counted as their pass rate.
Is this a fair, apples-to-apples comparison?+
The test harness is identical for every agent: same scenarios, same caller personality, same metrics, same three repetitions; only the phone number differs. Because each agent operates in a different business domain, residual domain effects can't be fully removed, so read the ranking as indicative rather than a definitive verdict on any platform.
Can I run this on my own production agent?+
Yes. This is exactly what Cekura does for its customers: probe a live agent, score it on these metrics from real calls, and retest after each prompt or config change so quality moves in the right direction. You can point Cekura at your own number and get the same scorecard.