Experiment 03 · September 2026

Every major voice platform, tested live on the phone

Nine real customer agents already live in production, one on each major platform, put through one identical ten scenario stress test on their public phone lines, three times each (270 live calls). This view scores ten qualities measured in the sampled agents: how each one sounds, takes turns, recovers from a barge in, and holds up under a difficult caller. Mean stop time after an interruption ranged from 0.62s to 7.55s.

One deployment per platform, 30 calls each. These are live production agents from different businesses, so each result reflects that specific agent, not a definitive verdict on its platform. The harness is identical for every agent; treat small composite gaps as ties, not rankings.
Result summary

Voice quality, turn taking, and reliability

The reliability metrics (relevancy, consistency, gibberish, and infrastructure) were near perfect across almost every agent, so the separation came from voice tone, humanness, interruption recovery, and latency. The widest spread by far was stop time after an interruption: from 0.62s (Bland) to 7.55s (Cognigy), a more than 12× difference in how quickly agents stopped talking when the caller cut in.

Composite leader
Bland · 92.8
0 to 100 quality composite
Fastest turns
Retell · 2.38s
Median response latency
Fastest to stop on barge in
Bland · 0.62s
Stop time after interruption
Total volume
270 calls
9 agents · 10 scenarios × 3
Want this scorecard for your own agent?Sign up to test your own agent →

Ranked by voice quality composite

Scores marked /5 are qualitative (higher is better); % values are pass rates. Latency and stop time are in seconds, lower is better.

How the composite is computed. Each of the ten metrics is put on a common 0 to 100 scale in the direction where higher is better: the qualitative /5 scores are scaled against 5, the pass rates use their own percentage, and latency and stop time are ranked across the field so the fastest agent scores 100 and the slowest 0. The ten normalized scores are then averaged with equal weight into a single 0 to 100 composite. Because it is a relative, equal weight blend on a small sample, treat gaps of a point or two as ties rather than a ranking.
AgentCompositeHumanness /5Voice tone /5Interruption /5Latency ↓Stop time ↓
1Bland92.84.134.304.612.62s0.62s
2Retell92.23.574.154.902.38s2.74s
3Cresta90.13.604.064.972.78s1.56s
4ElevenLabs90.03.934.085.003.08s1.23s
5Decagon87.43.604.234.813.13s1.76s
6Sierra85.53.633.814.673.42s1.06s
7PolyAI82.63.734.223.682.99s3.97s
8Vapi78.33.473.384.523.34s5.34s
9Cognigy73.73.603.924.653.76s7.55s
Metric by metric

Every metric, side by side

One chart per metric, each agent bar for bar. The reliability metrics (relevancy, consistency, gibberish, infrastructure) sit near the ceiling for almost every agent; the real spread shows up in voice tone, humanness, latency, and stop time after an interruption.

Humanness
ProviderHigher is better
Bland4.13
ElevenLabs3.93
PolyAI3.73
Sierra3.63
Cresta3.60
Decagon3.60
Cognigy3.60
Retell3.57
Vapi3.47
Voice tone + clarity
ProviderHigher is better
Bland4.30
Decagon4.23
PolyAI4.22
Retell4.15
ElevenLabs4.08
Cresta4.06
Cognigy3.92
Sierra3.81
Vapi3.38
Interruption score
ProviderHigher is better
ElevenLabs5.00
Cresta4.97
Retell4.90
Decagon4.81
Sierra4.67
Cognigy4.65
Bland4.61
Vapi4.52
PolyAI3.68
No unnecessary repetition
ProviderHigher is better
Retell5.00
ElevenLabs4.97
Vapi4.92
Cognigy4.91
Bland4.74
Cresta4.72
Sierra4.70
Decagon4.59
PolyAI4.53
Relevancy
ProviderHigher is better
Retell100%
Cresta100%
ElevenLabs100%
Decagon100%
Sierra100%
PolyAI100%
Vapi100%
Cognigy100%
Bland95%
Response consistency
ProviderHigher is better
Bland100%
Retell100%
Cresta100%
ElevenLabs100%
Decagon100%
Sierra100%
Vapi100%
Cognigy100%
PolyAI95%
Gibberish free
ProviderHigher is better
Retell100%
ElevenLabs100%
Decagon100%
Sierra100%
PolyAI100%
Vapi100%
Cognigy100%
Cresta97%
Bland95%
Infrastructure clean
ProviderHigher is better
Bland100%
Retell100%
Cresta100%
ElevenLabs100%
Decagon100%
Sierra100%
PolyAI100%
Vapi95%
Cognigy95%
Median turn latency
ProviderLower is better
Retell2.38s
Bland2.62s
Cresta2.78s
PolyAI2.99s
ElevenLabs3.08s
Decagon3.13s
Vapi3.34s
Sierra3.42s
Cognigy3.76s
Stop time after interruption
ProviderLower is better
Bland0.62s
Sierra1.06s
ElevenLabs1.23s
Cresta1.56s
Decagon1.76s
Retell2.74s
PolyAI3.97s
Vapi5.34s
Cognigy7.55s
Per agent detail

Every agent, every metric

Open each agent for the full ten metric snapshot and the run review observation.

01BlandProduction voice agent+

Metric snapshot · three run average

Composite
92.8
Humanness
4.13
Voice tone + clarity
4.30
Interruption score
4.61
No repetition score
4.74
Relevancy
95%
Response consistency
100%
Gibberish clean
95%
Infrastructure clean
100%
Median latency
2.62s
Stop time after barge in
0.62s
Observed in run review

Highest composite. Led humanness (4.13) and the quickest barge in recovery in the field (0.62s, averaged across every interruption) with fast turns. A couple of relevancy and gibberish edge cases were the only blemishes.

02RetellProduction voice agent+

Metric snapshot · three run average

Composite
92.2
Humanness
3.57
Voice tone + clarity
4.15
Interruption score
4.90
No repetition score
5.00
Relevancy
100%
Response consistency
100%
Gibberish clean
100%
Infrastructure clean
100%
Median latency
2.38s
Stop time after barge in
2.74s
Observed in run review

The cleanest reliability sheet in the field: perfect relevancy, consistency, gibberish and infrastructure, zero unnecessary repetition, the fastest turns, and a solid 2.74s barge in recovery once first message overlaps are excluded.

03CrestaProduction voice agent+

Metric snapshot · three run average

Composite
90.1
Humanness
3.60
Voice tone + clarity
4.06
Interruption score
4.97
No repetition score
4.72
Relevancy
100%
Response consistency
100%
Gibberish clean
97%
Infrastructure clean
100%
Median latency
2.78s
Stop time after barge in
1.56s
Observed in run review

Second-best interruption handling (4.97) with fast, clean turns and a quick 1.56s barge in recovery. A few gibberish edge cases were the only mark against it.

04ElevenLabsProduction voice agent+

Metric snapshot · three run average

Composite
90.0
Humanness
3.93
Voice tone + clarity
4.08
Interruption score
5.00
No repetition score
4.97
Relevancy
100%
Response consistency
100%
Gibberish clean
100%
Infrastructure clean
100%
Median latency
3.08s
Stop time after barge in
1.23s
Observed in run review

Perfect interruption score (5.0), near-perfect no repetition, a fully clean reliability sheet, and a quick 1.23s barge in recovery. Mid-pack latency is the main thing left to improve.

05DecagonProduction voice agent+

Metric snapshot · three run average

Composite
87.4
Humanness
3.60
Voice tone + clarity
4.23
Interruption score
4.81
No repetition score
4.59
Relevancy
100%
Response consistency
100%
Gibberish clean
100%
Infrastructure clean
100%
Median latency
3.13s
Stop time after barge in
1.76s
Observed in run review

A clean sheet on every reliability metric with strong voice tone (4.23) and a quick 1.76s barge in recovery. Mid-pack latency keeps it just outside the top group.

06SierraProduction voice agent+

Metric snapshot · three run average

Composite
85.5
Humanness
3.63
Voice tone + clarity
3.81
Interruption score
4.67
No repetition score
4.70
Relevancy
100%
Response consistency
100%
Gibberish clean
100%
Infrastructure clean
100%
Median latency
3.42s
Stop time after barge in
1.06s
Observed in run review

A clean reliability sheet and the second-quickest barge in recovery (1.06s), but the lowest voice tone clarity in the upper group and the slowest turns (3.42s) hold it mid pack.

07PolyAIProduction voice agent+

Metric snapshot · three run average

Composite
82.6
Humanness
3.73
Voice tone + clarity
4.22
Interruption score
3.68
No repetition score
4.53
Relevancy
100%
Response consistency
95%
Gibberish clean
100%
Infrastructure clean
100%
Median latency
2.99s
Stop time after barge in
3.97s
Observed in run review

Strong voice tone and a near clean reliability sheet, but the weakest interruption handling in the field (3.68) and a slow ~4s barge in recovery pull the composite down.

08VapiProduction voice agent+

Metric snapshot · three run average

Composite
78.3
Humanness
3.47
Voice tone + clarity
3.38
Interruption score
4.52
No repetition score
4.92
Relevancy
100%
Response consistency
100%
Gibberish clean
100%
Infrastructure clean
95%
Median latency
3.34s
Stop time after barge in
5.34s
Observed in run review

Low unnecessary repetition (4.92), but the lowest voice tone clarity in the field (3.38) and a slow 5.34s barge in recovery place it near the bottom of this quality slice.

09CognigyProduction voice agent+

Metric snapshot · three run average

Composite
73.7
Humanness
3.60
Voice tone + clarity
3.92
Interruption score
4.65
No repetition score
4.91
Relevancy
100%
Response consistency
100%
Gibberish clean
100%
Infrastructure clean
95%
Median latency
3.76s
Stop time after barge in
7.55s
Observed in run review

Solid interruption score and low repetition, but the slowest latency (3.76s) and by far the slowest barge in recovery in the field (7.55s across 21 interruptions), plus an infrastructure edge case, place it last.

Methodology

Each agent was reached on the phone number its business publicly lists for inbound customer calls, the same public support line any customer would dial, and driven by the same simulated caller through ten conditional action scenarios: barge in, silence, hold, hallucination bait, a vague caller, escalation, out of scope requests, a compound request, interrupt and switch, and a frustrated caller. We attributed each agent to a platform based on that platform vendor's own public claim of the customer, in its published case studies and customer references, and we retain the source for each attribution. Each row is a single production deployment; a business that has since changed vendors could still be attributed to its former one. Every agent ran the identical scenarios, caller personality, and metric set three times; only the number differed. The ten metrics shown are averaged over the three runs. Stop time after an interruption is averaged across every qualifying interruption rather than per call, so each barge in is weighted equally; and we ignore any interruption in the first ten seconds of a call, since a caller's opening words routinely overlap the agent's greeting and that is not a barge in the agent should be scored on.

Score your own production agent the same way

This is exactly what Cekura does for its customers: probe a live agent, score it on these ten metrics from real calls, and rerun the suite after every prompt or config change to catch regressions before they ship. Point Cekura at your own number and get this same scorecard.

Frequently asked questions

What exactly was tested?+

Nine real production voice agents, each running on a different platform (Bland, Retell, Cresta, Decagon, ElevenLabs, PolyAI, Sierra, Cognigy, Vapi), reached on their live public phone lines. Every agent got the same ten conditional action scenarios and the same simulated caller personality, three times each, 270 live calls in total.

How is the composite computed?+

Each of the ten metrics is normalized direction aware (higher is better metrics against their scale; latency and stop time inverted because lower is better), then averaged into a 0 to 100 composite. Boolean reliability metrics are counted as their pass rate.

Is this a fair, apples-to-apples comparison?+

The test harness is identical for every agent: same scenarios, same caller personality, same metrics, same three repetitions; only the phone number differs. Because each agent operates in a different business domain, residual domain effects can't be fully removed, so read the ranking as indicative rather than a definitive verdict on any platform.

Can I run this on my own production agent?+

Yes. This is exactly what Cekura does for its customers: probe a live agent, score it on these metrics from real calls, and retest after each prompt or config change so quality moves in the right direction. You can point Cekura at your own number and get the same scorecard.