LiveKit
- Strength
- Ranks second on repeatable reliability with 99.19% infrastructure-clean calls.
- What can be improved
- Consent was collected, but consent_id was omitted from the handoff tool.
- nova-3
- openai/gpt-4.1
- sonic-3
LiveKit is ahead on repeatable reliability, task completion, infrastructure reliability, tool call accuracy and voice tone and clarity, and Pipecat on response time. The two tie on interruption handling.
Last updated Methodology by Luis Ojeda
Pick LiveKit for reliability across repeated runs (7.32 points higher), completing the caller's task (0.91 points higher), calls that connect and stay up (2.04 points higher), accurate tool calls (0.09 higher) and how the agent sounds (0.62 higher). Pick Pipecat for fast replies (0.62s faster). They are level on handling interruptions.
Bars share one scale across all 8 platforms. The place next to each figure is its rank in the field.
Share of the 82 scenarios that passed on all three runs.
LiveKit, 7.32 points higher
Calls where the expected outcome was fully reached.
LiveKit, 0.91 points higher
Calls with no connection, audio or platform failure.
LiveKit, 2.04 points higher
Mean score for calling the right tool with the right arguments.
LiveKit, 0.09 higher
Mean score for how clear and natural the agent sounds.
LiveKit, 0.62 higher
Mean score for yielding and recovering when the caller cuts in.
Level
Mean time for the agent to start replying after the caller stops.
Pipecat, 0.62s faster
246 calls per platform: 82 scenarios, 3 runs each.
LiveKit is more reliable, by 7.32 points. LiveKit passed 70.73% of the 82 scenarios on all 3 runs and Pipecat passed 63.41%. A scenario only counts when every run of it passed.
Pipecat starts replying sooner, 0.62s faster. Mean response time is 2.59s for LiveKit and 1.97s for Pipecat.
LiveKit reached the expected outcome on 95.12% of calls and Pipecat on 94.21%. LiveKit is ahead on task completion.
Both ran the same Appointment and Medicare agents through the same 82 caller scenarios, 3 times each, with the same evaluators and mock tools. Only the platform and its speech and model components changed.