GPT Realtime
- Strength
- Second-fastest response at 1.58s with 95.53% infrastructure-clean calls.
- What can be improved
- In a noisy-audio run, a long pause was followed by lost digits and a skipped tool action.
- gpt-realtime-2.1
GPT Realtime is ahead on tool call accuracy, voice tone and clarity, interruption handling and response time, and Telnyx on repeatable reliability, task completion and infrastructure reliability.
Last updated Methodology by Luis Ojeda
Pick GPT Realtime for accurate tool calls (0.08 higher), how the agent sounds (0.04 higher), handling interruptions (0.35 higher) and fast replies (0.24s faster). Pick Telnyx for reliability across repeated runs (3.66 points higher), completing the caller's task (4.88 points higher) and calls that connect and stay up (4.47 points higher).
Bars share one scale across all 8 platforms. The place next to each figure is its rank in the field.
Share of the 82 scenarios that passed on all three runs.
Telnyx, 3.66 points higher
Calls where the expected outcome was fully reached.
Telnyx, 4.88 points higher
Calls with no connection, audio or platform failure.
Telnyx, 4.47 points higher
Mean score for calling the right tool with the right arguments.
GPT Realtime, 0.08 higher
Mean score for how clear and natural the agent sounds.
GPT Realtime, 0.04 higher
Mean score for yielding and recovering when the caller cuts in.
GPT Realtime, 0.35 higher
Mean time for the agent to start replying after the caller stops.
GPT Realtime, 0.24s faster
246 calls per platform: 82 scenarios, 3 runs each.
Telnyx is more reliable, by 3.66 points. GPT Realtime passed 64.63% of the 82 scenarios on all 3 runs and Telnyx passed 68.29%. A scenario only counts when every run of it passed.
GPT Realtime starts replying sooner, 0.24s faster. Mean response time is 1.58s for GPT Realtime and 1.82s for Telnyx.
GPT Realtime reached the expected outcome on 92.68% of calls and Telnyx on 97.56%. Telnyx is ahead on task completion.
Both ran the same Appointment and Medicare agents through the same 82 caller scenarios, 3 times each, with the same evaluators and mock tools. Only the platform and its speech and model components changed.