Is Azure Realtime or Gemini 3.8 Live Extended Thinking more reliable?+
Azure Realtime is more reliable, by 18.3 points. Azure Realtime passed 58 of 82 scenarios on all three runs, and Gemini 3.8 Live Extended Thinking passed 43. A scenario only counts when every run of it passed.
Which should I pick, Azure Realtime or Gemini 3.8 Live Extended Thinking?+
For a phone agent, Azure Realtime, since it gets more scenarios right on every run. Azure Realtime is ahead of Gemini 3.8 Live Extended Thinking on reliability, success rate, agent responses, data accuracy, stalled calls, interruption and response time.
Which saves caller data more accurately, Azure Realtime or Gemini 3.8 Live Extended Thinking?+
Azure Realtime, by 1.7 points. Azure Realtime saved everything exactly right on 86.1% of the calls that save data, and Gemini 3.8 Live Extended Thinking on 84.4%. On Appointments, 94.2% against 88.9%. On Medicare, 65.2% against 72.7%. A call counts only when nothing is missing, wrong, extra or saved twice.
Is Azure Realtime faster than Gemini 3.8 Live Extended Thinking?+
Yes. Azure Realtime starts replying 0.41s sooner at the median. Azure Realtime takes 1.66s at the median and 2.22s at p90; Gemini 3.8 Live Extended Thinking takes 2.07s and 2.74s. Both are measured on the call audio, from the caller finishing to the agent starting to reply.
Which handles interruptions better, Azure Realtime or Gemini 3.8 Live Extended Thinking?+
Azure Realtime, by 0.01 on a 0 to 5 scale. Azure Realtime scores 4.96 and Gemini 3.8 Live Extended Thinking 4.95 for how well the agent handled being talked over.
Which goes silent on callers less often, Azure Realtime or Gemini 3.8 Live Extended Thinking?+
Azure Realtime, by 1.2 points. The agent went silent for 10 seconds or more on 4.1% of Azure Realtime calls and 5.3% of Gemini 3.8 Live Extended Thinking calls.
Which is cheaper, Azure Realtime or Gemini 3.8 Live Extended Thinking?+
There is no verified per-minute rate for Gemini 3.8 Live Extended Thinking, so cost cannot be compared yet.
Where does the Azure Realtime vs Gemini 3.8 Live Extended Thinking data come from?+
Cekura ran both as the whole agent in the same open-source Pipecat pipeline, with the same prompt, tools and connection. A simulated caller worked through all 82 scenarios on live calls, three times each, 2,952 calls across the benchmark. A call passes only when the agent said the right things and saved the right data.