Is Azure Realtime or Grok Voice Think Fast 2.0 more reliable?+
Azure Realtime is more reliable, by 3.7 points. Azure Realtime passed 58 of 82 scenarios on all three runs, and Grok Voice Think Fast 2.0 passed 55. A scenario only counts when every run of it passed.
Which should I pick, Azure Realtime or Grok Voice Think Fast 2.0?+
For a phone agent, Azure Realtime, since it gets more scenarios right on every run. Azure Realtime is ahead on reliability, success rate, agent responses, data accuracy, stalled calls and interruption, and Grok Voice Think Fast 2.0 on response time. The two tie on cost.
Which saves caller data more accurately, Azure Realtime or Grok Voice Think Fast 2.0?+
Azure Realtime, by 7.2 points. Azure Realtime saved everything exactly right on 86.1% of the calls that save data, and Grok Voice Think Fast 2.0 on 78.9%. On Appointments, 94.2% against 97.1%. On Medicare, 65.2% against 31.8%. A call counts only when nothing is missing, wrong, extra or saved twice.
Is Azure Realtime faster than Grok Voice Think Fast 2.0?+
No. Grok Voice Think Fast 2.0 starts replying 0.03s sooner at the median. Azure Realtime takes 1.66s at the median and 2.22s at p90; Grok Voice Think Fast 2.0 takes 1.62s and 1.94s. Both are measured on the call audio, from the caller finishing to the agent starting to reply.
Which handles interruptions better, Azure Realtime or Grok Voice Think Fast 2.0?+
Azure Realtime, by 0.04 on a 0 to 5 scale. Azure Realtime scores 4.96 and Grok Voice Think Fast 2.0 4.92 for how well the agent handled being talked over.
Which goes silent on callers less often, Azure Realtime or Grok Voice Think Fast 2.0?+
Azure Realtime, by 0.4 points. The agent went silent for 10 seconds or more on 4.1% of Azure Realtime calls and 4.5% of Grok Voice Think Fast 2.0 calls.
Which is cheaper, Azure Realtime or Grok Voice Think Fast 2.0?+
They cost the same. Azure Realtime costs $0.080 a minute of call and Grok Voice Think Fast 2.0 $0.080, at each vendor's published rates.
Where does the Azure Realtime vs Grok Voice Think Fast 2.0 data come from?+
Cekura ran both as the whole agent in the same open-source Pipecat pipeline, with the same prompt, tools and connection. A simulated caller worked through all 82 scenarios on live calls, three times each, 2,952 calls across the benchmark. A call passes only when the agent said the right things and saved the right data.