Is Azure Realtime or Gemini 3.8 Live more reliable?+
Azure Realtime is more reliable, by 29.3 points. Azure Realtime passed 58 of 82 scenarios on all three runs, and Gemini 3.8 Live passed 34. A scenario only counts when every run of it passed.
Which should I pick, Azure Realtime or Gemini 3.8 Live?+
For a phone agent, Azure Realtime, since it gets more scenarios right on every run. Azure Realtime is ahead on reliability, success rate, agent responses, data accuracy and response time, and Gemini 3.8 Live on stalled calls and interruption.
Which saves caller data more accurately, Azure Realtime or Gemini 3.8 Live?+
Azure Realtime, by 15.6 points. Azure Realtime saved everything exactly right on 86.1% of the calls that save data, and Gemini 3.8 Live on 70.5%. On Appointments, 94.2% against 78.4%. On Medicare, 65.2% against 50.0%. A call counts only when nothing is missing, wrong, extra or saved twice.
Is Azure Realtime faster than Gemini 3.8 Live?+
Yes. Azure Realtime starts replying 0.25s sooner at the median. Azure Realtime takes 1.66s at the median and 2.22s at p90; Gemini 3.8 Live takes 1.91s and 2.23s. Both are measured on the call audio, from the caller finishing to the agent starting to reply.
Which handles interruptions better, Azure Realtime or Gemini 3.8 Live?+
Gemini 3.8 Live, by 0.02 on a 0 to 5 scale. Azure Realtime scores 4.96 and Gemini 3.8 Live 4.98 for how well the agent handled being talked over.
Which goes silent on callers less often, Azure Realtime or Gemini 3.8 Live?+
Gemini 3.8 Live, by 2.0 points. The agent went silent for 10 seconds or more on 4.1% of Azure Realtime calls and 2.0% of Gemini 3.8 Live calls.
Which is cheaper, Azure Realtime or Gemini 3.8 Live?+
There is no verified per-minute rate for Gemini 3.8 Live, so cost cannot be compared yet.
Where does the Azure Realtime vs Gemini 3.8 Live data come from?+
Cekura ran both as the whole agent in the same open-source Pipecat pipeline, with the same prompt, tools and connection. A simulated caller worked through all 82 scenarios on live calls, three times each, 2,952 calls across the benchmark. A call passes only when the agent said the right things and saved the right data.