Voice agent workflow benchmark

Compare complete voice agent stacks on reliability, task completion, and response time.

8
configurations
82
scenarios
3
repeats
The tradeoff

Reliability vs response time

Most attractive quadrantPareto line
Repeatable reliability
◤ better
Mean response time
ElevenLabs is fastest, Retell leads reliability, and LiveKit balances both. Gemini Live reflects the completed hybrid-VAD rerun.
Key results

Compare the metrics that matter

Completed the caller’s task

Task completion

The share of scored calls that fully reached the expected outcome.

Telnyx reaches 97.56% across all 246 calls. Vapi matches that score across 205 calls with outcome evidence; its 41 no-connect calls are reflected separately in infrastructure reliability.

How this is measured

A call counts as complete only when Expected Outcome receives the full 5/5 score. Coverage varies when a call has no outcome evidence.

ProviderHigher is better
Telnyx97.56%
Vapi97.56%
LiveKit95.12%
Pipecat94.21%
Retell93.88%
GPT Realtime92.68%
ElevenLabs91.46%
Gemini Live87.80%
Calls connected and completed cleanly

Infrastructure reliability

This is the share of calls that completed without a provider-side or connection issue. Failed connections remain visible instead of being removed from the results.

ElevenLabs and Telnyx are infrastructure-clean on every retained call; LiveKit and Retell are above 98%.

How this is measured

All 246 retained calls per configuration are included. A call counts as clean only when Infrastructure Issues receives 5/5.

ProviderHigher is better
ElevenLabs100.00%
Telnyx100.00%
LiveKit99.19%
Retell98.37%
Pipecat97.15%
GPT Realtime95.53%
Vapi82.93%
Gemini Live72.36%
Handled natural turn changes

Interruption handling

This score measures how naturally the agent handles barge-in and conversational turn changes. Higher scores mean smoother turn taking.

Retell scores 5.00/5. The next five configurations are between 4.96 and 4.98; Vapi and Telnyx score 4.73 and 4.63 respectively.

How this is measured

Mean Interruption Score out of 5 across calls where the interruption evaluator applied.

ProviderHigher is better
Retell5.00/5
GPT Realtime4.98/5
LiveKit4.97/5
Pipecat4.97/5
Gemini Live4.97/5
ElevenLabs4.96/5
Vapi4.73/5
Telnyx4.63/5
Sounded clear and natural

Voice naturalness

The benchmark’s Voice Tone + Clarity score rates the voice heard during each scored call.

ElevenLabs leads at 4.47/5. Retell and LiveKit follow at 4.36/5.

How this is measured

Mean Voice Tone + Clarity score out of 5. Gemini Live is not shown because this evaluator was not included in its hybrid-VAD rerun.

ProviderHigher is better
ElevenLabs4.47/5
Retell4.36/5
LiveKit4.36/5
GPT Realtime4.25/5
Telnyx4.21/5
Vapi4.08/5
Pipecat3.74/5
Frozen matched study · 8 configurations · 82 scenarios · 3 repeats

Leaderboard

Retell leads the frozen cohort: 62 of 82 scenarios passed on all three retained runs. Telnyx ranks fourth with 56 of 82.

Primary rank uses repeatable reliability, pass^3, only. Task completion, infrastructure, interruption, voice naturalness, and response time provide supporting evidence. Select any column to explore another ordering.

RankEvidence
01Retell
75.61%
93.88%
98.37%
5.00
4.36
2.21s
View runs ↗
02LiveKit
70.73%
95.12%
99.19%
4.97
4.36
2.59s
View runs ↗
03ElevenLabs
69.51%
91.46%
100.00%
4.96
4.47
1.27s
View runs ↗
04Telnyx
68.29%
97.56%
100.00%
4.63
4.21
1.82s
View runs ↗
05GPT Realtime
64.63%
92.68%
95.53%
4.98
4.25
1.58s
View runs ↗
06Pipecat
63.41%
94.21%
97.15%
4.97
3.74
1.97s
View runs ↗
07Vapi
59.76%
97.56%
82.93%
4.73
4.08
3.08s
View runs ↗
08Gemini Live
30.49%
87.80%
72.36%
4.97
3.05s
View runs ↗

pass³ is the share of 82 scenarios where all three retained runs passed. Task completion is the full-score success rate among calls with Expected Outcome evidence; coverage varies by configuration. Interruption and Voice Tone + Clarity are mean scores out of 5. Response time is measured by Cekura at the main-agent layer, not from provider-native component timing.

Provider notes

What stands out

Retell

What’s good

Leads repeatable reliability at 75.61% pass³.

What can be improved

The transcript captured a phone number correctly, but a different number was sent to the tool.

LiveKit

What’s good

Ranks second on repeatable reliability with 99.19% infrastructure-clean calls.

What can be improved

Consent was collected, but consent_id was omitted from the handoff tool.

ElevenLabs

What’s good

Fastest response at 1.27s, highest Voice Tone + Clarity at 4.47/5, and 100% infrastructure-clean calls.

What can be improved

The agent narrated a tool call and continued with an invented result.

GPT Realtime

What’s good

Second-fastest response at 1.58s with 95.53% infrastructure-clean calls.

What can be improved

In a noisy-audio run, a long pause was followed by lost digits and a skipped tool action.

Pipecat

What’s good

A 1.97s response time with 94.21% task completion across 242 scored calls.

What can be improved

Routing completed, but the returned route ID was omitted from the handoff.

Vapi

What’s good

Tied for first in task completion at 97.56% across 205 calls with outcome evidence.

What can be improved

The remaining 41 of 246 calls did not connect and remain visible in infrastructure reliability.

Gemini Live

What’s good

Second-highest repetition score at 4.78/5 across 174 scored calls.

What can be improved

One rerun had no latency evidence and scored 0/5 for both task outcome and tool accuracy.

Telnyx

What’s good

Tied for first in task completion at 97.56% across all 246 calls.

What can be improved

An opening interruption led the agent into plan discussion before completing the recording notice and TPMO disclosure.

Scope and interpretation
  1. 01

    Providers chose what to test

    Cekura invited voice agent platforms to submit the configuration they wanted benchmarked. Providers chose their models, speech components, and settings.

  2. 02

    Everyone received the same brief

    Cekura shared the system prompt, tool definitions, test-case summaries, and test data. Transfers to other agents were the only configuration restriction.

  3. 03

    Each scenario ran three times

    The same 82 caller situations and evaluator suite were used for every configuration. A scenario earns pass³ only when all three retained runs pass.

  4. 04

    Failures remain in the results

    Calls that did not connect or produced no transcript stay in the denominator. Missing provider evidence is not removed from the release.

Common questions

Frequently asked questions

What does pass^3 mean?

Share of scenarios where all three runs passed the applicable configured rubric gates.

Which configuration ranked first?

Retell ranked first at 75.61% pass^3.

Can the Cekura Bench leaderboard be sorted by a different metric?

Yes. The published table can be reordered by provider, repeatable reliability, task completion, infrastructure reliability, interruption handling, conversational naturalness, or mean main-agent response time.

How was Gemini Live configured for its rerun?

The published Gemini Live result uses high reasoning and hybrid VAD. Its completed rerun is included in the leaderboard and reliability versus response-time chart.

How many calls were retained?

246 calls per configuration, 1968 across the eight-configuration cohort.

Submit your agent

See where your voice agent stands

Send us your production agent. Cekura will run the standard benchmark and share how it compares with the published cohort.

Submit your agent ↗