Gemini Live
- Strength
- Second-highest repetition score at 4.78/5 across 174 scored calls.
- What can be improved
- One rerun had no latency evidence and scored 0/5 for both task outcome and tool accuracy.
- reasoning: high
- hybrid VAD
GPT Realtime is ahead of Gemini Live on repeatable reliability, task completion, infrastructure reliability, tool call accuracy, interruption handling and response time.
Last updated Methodology by Luis Ojeda
GPT Realtime is the better pick over Gemini Live on every measure that separates them: reliability across repeated runs (34.14 points higher), completing the caller's task (4.88 points higher), calls that connect and stay up (23.17 points higher), accurate tool calls (0.39 higher), handling interruptions (0.01 higher) and fast replies (1.47s faster).
Bars share one scale across all 8 platforms. The place next to each figure is its rank in the field.
Share of the 82 scenarios that passed on all three runs.
GPT Realtime, 34.14 points higher
Calls where the expected outcome was fully reached.
GPT Realtime, 4.88 points higher
Calls with no connection, audio or platform failure.
GPT Realtime, 23.17 points higher
Mean score for calling the right tool with the right arguments.
GPT Realtime, 0.39 higher
Mean score for how clear and natural the agent sounds.
Not comparable
Mean score for yielding and recovering when the caller cuts in.
GPT Realtime, 0.01 higher
Mean time for the agent to start replying after the caller stops.
GPT Realtime, 1.47s faster
246 calls per platform: 82 scenarios, 3 runs each.
GPT Realtime is more reliable, by 34.14 points. Gemini Live passed 30.49% of the 82 scenarios on all 3 runs and GPT Realtime passed 64.63%. A scenario only counts when every run of it passed.
GPT Realtime starts replying sooner, 1.47s faster. Mean response time is 3.05s for Gemini Live and 1.58s for GPT Realtime.
Gemini Live reached the expected outcome on 87.80% of calls and GPT Realtime on 92.68%. GPT Realtime is ahead on task completion.
Both ran the same Appointment and Medicare agents through the same 82 caller scenarios, 3 times each, with the same evaluators and mock tools. Only the platform and its speech and model components changed.