Archive · August 2026

Four-provider configuration study

An earlier four-provider release. Each team received the same agent and test suite, then selected the LLM, speech-to-text, text-to-speech, and voice configuration it believed would perform best. These results remain separate from the current seven-configuration leaderboard.

Archived release. This study covers an earlier four-provider cohort and should not be compared directly with the current Benchmark v1 leaderboard.
Result summary

Reliability, response time, and tool accuracy

Retell delivered the strongest repeatable reliability. ElevenLabs was fastest and led transcription accuracy. LiveKit tied ElevenLabs on pass³ while recording the stronger pass¹ and tool score.

Reliability leader
Retell · 75.6%
pass³ · 62 of 82 scenarios
Response-time leader
ElevenLabs · 1.26s
Mean evaluated response time
Test coverage
82 scenarios
Three repetitions per platform
Total volume
984 calls
246 calls per provider

Four-provider study results

pass¹ is the share of individual calls that passed. pass³ is stricter: all three attempts for a scenario must pass. Evaluator scores use a 0–5 scale.

Platformpass¹pass³Mean responseOutcomeTool accuracyReport
1Retell88.2%75.6%2.23s4.694.84View runs ↗
2ElevenLabs81.3%69.5%1.26s4.634.49View runs ↗
3LiveKit84.2%69.5%2.58s4.744.67View runs ↗
4Vapi78.5%59.8%3.08s4.834.12View runs ↗
Provider-selected stacks

What each platform ran

Open each provider for its exact model and speech configuration, complete metric snapshot, coverage, run-review observation, and public report.

01Retellgpt-5.5+

Selected configuration

LLM
gpt-5.5
STT
Retell managed STT · accurate mode
TTS
Cartesia · Cimo
Voice / config ID
cartesia-Cimo

Metric snapshot

pass¹
88.2%
pass³
75.6%
Mean response
2.23s
Expected outcome
4.69
Tool accuracy
4.84
Exact tool inputs
93.7%
Transcription
4.72
Interruption
5.00
Repetition
4.86
Clean end call
100.0%
Infrastructure clean
98.8%
Observed in run review

Strongest repeatable reliability and tool-input accuracy. A small number of runs showed premature tool calls or hallucinated inputs.

Coverage · core 245/246 · voice 176/177 · repetition 173/177

Open public report ↗
02ElevenLabsQwen 3.5 397B A17B+

Selected configuration

LLM
Qwen 3.5 397B A17B
STT
ElevenLabs Scribe Realtime · high quality
TTS
ElevenLabs Flash v2 · Bella
Voice / config ID
EXAVITQu4vr4xnSDxMaL

Metric snapshot

pass¹
81.3%
pass³
69.5%
Mean response
1.26s
Expected outcome
4.63
Tool accuracy
4.49
Exact tool inputs
82.7%
Transcription
4.81
Interruption
4.97
Repetition
4.63
Clean end call
96.0%
Infrastructure clean
100.0%
Observed in run review

Lowest mean response time and strongest transcription score. Some runs hallucinated tool inputs, including a Medicare consent identifier.

Coverage · core 246/246 · voice 177/177 · repetition 174/177

Open public report ↗
03LiveKitOpenAI gpt-4.1 · temperature 0+

Selected configuration

LLM
OpenAI gpt-4.1 · temperature 0
STT
Deepgram Nova-3
TTS
Cartesia Sonic-3
Voice / config ID
9626c31c-bec5-4cca-baa8-f8ba9e84c8bc

Metric snapshot

pass¹
84.2%
pass³
69.5%
Mean response
2.58s
Expected outcome
4.74
Tool accuracy
4.67
Exact tool inputs
84.4%
Transcription
4.68
Interruption
4.97
Repetition
4.37
Clean end call
96.6%
Infrastructure clean
99.2%
Observed in run review

Tied ElevenLabs on pass^3 while posting the stronger pass^1 and tool score. One observed Medicare flow omitted a consent ID returned by the prior tool.

Coverage · core 246/246 · voice 177/177 · repetition 174/177

Open public report ↗
04VapiOpenAI gpt-4.1-2025-04-14 · temperature 0+

Selected configuration

LLM
OpenAI gpt-4.1-2025-04-14 · temperature 0
STT
Soniox STT-RT v5 · Deepgram Nova-3 fallback
TTS
Vapi Clara v2
Voice / config ID
Vapi Clara v2

Metric snapshot

pass¹
78.5%
pass³
59.8%
Mean response
3.08s
Expected outcome
4.83
Tool accuracy
4.12
Exact tool inputs
79.8%
Transcription
4.47
Interruption
4.73
Repetition
4.50
Clean end call
100.0%
Infrastructure clean
99.0%
Observed in run review

Highest expected-outcome mean among the four, but intermittent answering and connectivity reduced metric coverage and repeatable pass rates.

Coverage · core 211/246 · voice 164/177 · repetition 161/177

Open public report ↗

Controlled methodology

Every provider received the same agent definition, system prompt, tools, mock data, first message, metrics, and scenario coverage. Providers selected their preferred production stack. Cekura then ran all 82 scenarios three times on every platform. pass³ rewards repeatability; it only counts a scenario when all three runs pass.

Frequently asked questions

Which configuration was most reliable in this study?+

Retell led this four-provider study with 75.6% pass³ (62 of 82 scenarios passing all three attempts) and 88.2% pass¹ (217 of 246 calls).

Which configuration had the lowest mean response time?+

ElevenLabs recorded the lowest mean response time at 1.26 seconds, ahead of Retell at 2.23 seconds, LiveKit at 2.58 seconds, and Vapi at 3.08 seconds.

Did every provider run the same LLM, STT, and TTS stack?+

No. Every provider received the same agent definition, prompt, tools, mock data, first message, metrics, and scenario coverage, but selected its own LLM, speech-to-text, text-to-speech, and voice configuration.

What does pass³ mean in this benchmark?+

Pass³ is the share of scenarios where all three repeated calls passed the configured rubric. It is the strict repeatability measure; pass¹ reports the share of individual calls that passed.