Cekura Bench

Benchmark for Voice Agents

See how leading voice agent platforms perform on reliability, response time, task completion, and real-world call handling.

The tradeoff

Reliability vs response time

More reliable ↑Faster response ←
Mean main-agent response time
ElevenLabs is fastest, Retell leads reliability, and LiveKit balances both. Gemini Live reflects the completed hybrid-VAD rerun.
Key results

What the benchmark shows

Four practical views of task completion, infrastructure, interruption handling, and voice naturalness.

Completed the caller’s task

Task completion

The share of scored calls that fully reached the expected outcome.

Vapi reaches 97.56% across 205 scored calls. LiveKit reaches 95.12% across all 246 calls. Vapi’s 41 no-connect calls are reflected separately in infrastructure reliability.

How this is measured

A call counts as complete only when Expected Outcome receives the full 5/5 score. Coverage varies when a call has no outcome evidence.

ProviderHigher is better
Vapi97.56%
LiveKit95.12%
Pipecat94.21%
Retell93.88%
GPT Realtime92.68%
ElevenLabs91.46%
Gemini Live87.80%
Calls connected and completed cleanly

Infrastructure reliability

This is the share of calls that completed without a provider-side or connection issue. Failed connections remain visible instead of being removed from the results.

ElevenLabs is infrastructure-clean on every retained call; LiveKit and Retell are above 98%.

How this is measured

All 246 retained calls per configuration are included. A call counts as clean only when Infrastructure Issues receives 5/5.

ProviderHigher is better
ElevenLabs100.00%
LiveKit99.19%
Retell98.37%
Pipecat97.15%
GPT Realtime95.53%
Vapi82.93%
Gemini Live72.36%
Handled natural turn changes

Interruption handling

This score measures how naturally the agent handles barge-in and conversational turn changes. Higher scores mean smoother turn taking.

Retell scores 5.00/5. Five other configurations are between 4.96 and 4.98; Vapi scores 4.73.

How this is measured

Mean Interruption Score out of 5 across calls where the interruption evaluator applied.

ProviderHigher is better
Retell5.00/5
GPT Realtime4.98/5
LiveKit4.97/5
Pipecat4.97/5
Gemini Live4.97/5
ElevenLabs4.96/5
Vapi4.73/5
Sounded clear and natural

Voice naturalness

The benchmark’s Voice Tone + Clarity score rates the voice heard during each scored call.

ElevenLabs leads at 4.47/5. Retell and LiveKit follow at 4.36/5.

How this is measured

Mean Voice Tone + Clarity score out of 5. Gemini Live is not shown because this evaluator was not included in its hybrid-VAD rerun.

ProviderHigher is better
ElevenLabs4.47/5
Retell4.36/5
LiveKit4.36/5
GPT Realtime4.25/5
Vapi4.08/5
Pipecat3.74/5
Frozen matched study · 7 configurations · 82 scenarios · 3 repeats

Leaderboard

Retell leads the frozen cohort: 62 of 82 scenarios passed on all three retained runs.

Primary rank uses repeatable reliability, pass^3, only. Task completion, infrastructure, interruption, voice naturalness, and response time provide supporting evidence. Select any column to explore another ordering.

RankEvidence
01Retell75.61%93.88%98.37%5.004.362.21sView runs ↗
02LiveKit70.73%95.12%99.19%4.974.362.59sView runs ↗
03ElevenLabs69.51%91.46%100.00%4.964.471.27sView runs ↗
04GPT Realtime64.63%92.68%95.53%4.984.251.58sView runs ↗
05Pipecat63.41%94.21%97.15%4.973.741.97sView runs ↗
06Vapi59.76%97.56%82.93%4.734.083.08sView runs ↗
07Gemini Live30.49%87.80%72.36%4.973.05sView runs ↗

pass³ is the share of 82 scenarios where all three retained runs passed. Task completion is the full-score success rate among calls with Expected Outcome evidence; coverage varies by configuration. Interruption and Voice Tone + Clarity are mean scores out of 5. Response time is measured by Cekura at the main-agent layer, not from provider-native component timing.

Provider notes

What stands out

The clearest strength from each configuration, paired with a separate example issue and its source.

Retell

Leads repeatable reliability at 75.61% pass³.

Example issue observed

The transcript captured a phone number correctly, but a different number was sent to the tool.

LiveKit

Ranks second on repeatable reliability with 99.19% infrastructure-clean calls.

Example issue observed

Consent was collected, but consent_id was omitted from the handoff tool.

ElevenLabs

Fastest response at 1.27s, highest Voice Tone + Clarity at 4.47/5, and 100% infrastructure-clean calls.

Example issue observed

The agent narrated a tool call and continued with an invented result.

GPT Realtime

Second-fastest response at 1.58s with 95.53% infrastructure-clean calls.

Example issue observed

In a noisy-audio run, a long pause was followed by lost digits and a skipped tool action.

Pipecat

A 1.97s response time with 94.21% task completion across 242 scored calls.

Example issue observed

Routing completed, but the returned route ID was omitted from the handoff.

Vapi

Highest task-completion rate at 97.56% across 205 calls with outcome evidence.

Example issue observed

The remaining 41 of 246 calls did not connect and remain visible in infrastructure reliability.

Gemini Live

Second-highest repetition score at 4.78/5 across 174 scored calls.

Example issue observed

One rerun had no latency evidence and scored 0/5 for both task outcome and tool accuracy.

Scope and interpretation

Methodology

01

Providers chose what to test

Cekura invited voice agent platforms to submit the configuration they wanted benchmarked. Providers chose their models, speech components, and settings.

02

Everyone received the same brief

Cekura shared the system prompt, tool definitions, test-case summaries, and test data. Transfers to other agents were the only configuration restriction.

03

Each scenario ran three times

The same 82 caller situations and evaluator suite were used for every configuration. A scenario earns pass³ only when all three retained runs pass.

04

Failures remain in the results

Calls that did not connect or produced no transcript stay in the denominator. Missing provider evidence is not removed from the release.

Tested configurations

Models and speech components

Providers selected these configurations. The OpenAI row is the exception: Cekura tested gpt-realtime-2.1 directly, without a configuration submitted by OpenAI.

ElevenLabs

STT · LLM · TTS

scribe_realtime → qwen35-397b-a17b → eleven_flash_v2

Vapi

STT · LLM · TTS

stt-rt-v5 → gpt-4.1-2025-04-14 → vapi-v2 (Clara)

Retell

STT · LLM · TTS

accurate (Retell-managed STT) → gpt-4.1 → eleven_flash_v2

LiveKit

STT · LLM · TTS

nova-3 → openai/gpt-4.1 → sonic-3

Pipecat

STT · LLM · TTS

nova-3-general → gpt-4.1 → sonic-3.5

GPT Realtime (OpenAI)

Speech to speech

gpt-realtime-2.1

Gemini Live

Speech to speech

reasoning: high → hybrid VAD

Gemini Live reflects the completed rerun using high reasoning and hybrid VAD.

Submit your agent

See where your voice agent stands

Send us your production agent. Cekura will run the standard benchmark and share how it compares with the published cohort.

Submit your agent ↗