Retell
- Strength
- Leads repeatable reliability at 75.61% pass³.
- What can be improved
- The transcript captured a phone number correctly, but a different number was sent to the tool.
- accurate (Retell-managed STT)
- gpt-5.5
- eleven_flash_v2
Retell is ahead on repeatable reliability, infrastructure reliability, tool call accuracy, voice tone and clarity, interruption handling and response time, and Vapi on task completion.
Last updated Methodology by Luis Ojeda
Pick Retell for reliability across repeated runs (15.85 points higher), calls that connect and stay up (15.44 points higher), accurate tool calls (0.82 higher), how the agent sounds (0.28 higher), handling interruptions (0.27 higher) and fast replies (0.87s faster). Pick Vapi for completing the caller's task (3.68 points higher).
Bars share one scale across all 8 platforms. The place next to each figure is its rank in the field.
Share of the 82 scenarios that passed on all three runs.
Retell, 15.85 points higher
Calls where the expected outcome was fully reached.
Vapi, 3.68 points higher
Calls with no connection, audio or platform failure.
Retell, 15.44 points higher
Mean score for calling the right tool with the right arguments.
Retell, 0.82 higher
Mean score for how clear and natural the agent sounds.
Retell, 0.28 higher
Mean score for yielding and recovering when the caller cuts in.
Retell, 0.27 higher
Mean time for the agent to start replying after the caller stops.
Retell, 0.87s faster
246 calls per platform: 82 scenarios, 3 runs each.
Retell is more reliable, by 15.85 points. Retell passed 75.61% of the 82 scenarios on all 3 runs and Vapi passed 59.76%. A scenario only counts when every run of it passed.
Retell starts replying sooner, 0.87s faster. Mean response time is 2.21s for Retell and 3.08s for Vapi.
Retell reached the expected outcome on 93.88% of calls and Vapi on 97.56%. Vapi is ahead on task completion.
Both ran the same Appointment and Medicare agents through the same 82 caller scenarios, 3 times each, with the same evaluators and mock tools. Only the platform and its speech and model components changed.