Reliability, response time, and tool accuracy
Retell delivered the strongest repeatable reliability. ElevenLabs was fastest and led transcription accuracy. LiveKit tied ElevenLabs on pass³ while recording the stronger pass¹ and tool score.
Four-provider study results
pass¹ is the share of individual calls that passed. pass³ is stricter: all three attempts for a scenario must pass. Evaluator scores use a 0–5 scale.
| Platform | pass¹ | pass³ | Mean response | Outcome | Tool accuracy | Report |
|---|---|---|---|---|---|---|
| 1Retell | 88.2% | 75.6% | 2.23s | 4.69 | 4.84 | View runs ↗ |
| 2ElevenLabs | 81.3% | 69.5% | 1.26s | 4.63 | 4.49 | View runs ↗ |
| 3LiveKit | 84.2% | 69.5% | 2.58s | 4.74 | 4.67 | View runs ↗ |
| 4Vapi | 78.5% | 59.8% | 3.08s | 4.83 | 4.12 | View runs ↗ |
What each platform ran
Open each provider for its exact model and speech configuration, complete metric snapshot, coverage, run-review observation, and public report.
01Retellgpt-5.5pass³75.6%+
Selected configuration
- LLM
- gpt-5.5
- STT
- Retell managed STT · accurate mode
- TTS
- Cartesia · Cimo
- Voice / config ID
- cartesia-Cimo
Metric snapshot
Strongest repeatable reliability and tool-input accuracy. A small number of runs showed premature tool calls or hallucinated inputs.
Coverage · core 245/246 · voice 176/177 · repetition 173/177
Open public report ↗02ElevenLabsQwen 3.5 397B A17Bpass³69.5%+
Selected configuration
- LLM
- Qwen 3.5 397B A17B
- STT
- ElevenLabs Scribe Realtime · high quality
- TTS
- ElevenLabs Flash v2 · Bella
- Voice / config ID
- EXAVITQu4vr4xnSDxMaL
Metric snapshot
Lowest mean response time and strongest transcription score. Some runs hallucinated tool inputs, including a Medicare consent identifier.
Coverage · core 246/246 · voice 177/177 · repetition 174/177
Open public report ↗03LiveKitOpenAI gpt-4.1 · temperature 0pass³69.5%+
Selected configuration
- LLM
- OpenAI gpt-4.1 · temperature 0
- STT
- Deepgram Nova-3
- TTS
- Cartesia Sonic-3
- Voice / config ID
- 9626c31c-bec5-4cca-baa8-f8ba9e84c8bc
Metric snapshot
Tied ElevenLabs on pass^3 while posting the stronger pass^1 and tool score. One observed Medicare flow omitted a consent ID returned by the prior tool.
Coverage · core 246/246 · voice 177/177 · repetition 174/177
Open public report ↗04VapiOpenAI gpt-4.1-2025-04-14 · temperature 0pass³59.8%+
Selected configuration
- LLM
- OpenAI gpt-4.1-2025-04-14 · temperature 0
- STT
- Soniox STT-RT v5 · Deepgram Nova-3 fallback
- TTS
- Vapi Clara v2
- Voice / config ID
- Vapi Clara v2
Metric snapshot
Highest expected-outcome mean among the four, but intermittent answering and connectivity reduced metric coverage and repeatable pass rates.
Coverage · core 211/246 · voice 164/177 · repetition 161/177
Open public report ↗Controlled methodology
Every provider received the same agent definition, system prompt, tools, mock data, first message, metrics, and scenario coverage. Providers selected their preferred production stack. Cekura then ran all 82 scenarios three times on every platform. pass³ rewards repeatability; it only counts a scenario when all three runs pass.
Frequently asked questions
Which configuration was most reliable in this study?+
Retell led this four-provider study with 75.6% pass³ (62 of 82 scenarios passing all three attempts) and 88.2% pass¹ (217 of 246 calls).
Which configuration had the lowest mean response time?+
ElevenLabs recorded the lowest mean response time at 1.26 seconds, ahead of Retell at 2.23 seconds, LiveKit at 2.58 seconds, and Vapi at 3.08 seconds.
Did every provider run the same LLM, STT, and TTS stack?+
No. Every provider received the same agent definition, prompt, tools, mock data, first message, metrics, and scenario coverage, but selected its own LLM, speech-to-text, text-to-speech, and voice configuration.
What does pass³ mean in this benchmark?+
Pass³ is the share of scenarios where all three repeated calls passed the configured rubric. It is the strict repeatability measure; pass¹ reports the share of individual calls that passed.