Is Azure Realtime or GPT Realtime 2.1 more reliable?+
GPT Realtime 2.1 is more reliable, by 8.5 points. Azure Realtime passed 58 of 82 scenarios on all three runs, and GPT Realtime 2.1 passed 65. A scenario only counts when every run of it passed.
Which should I pick, Azure Realtime or GPT Realtime 2.1?+
For a phone agent, GPT Realtime 2.1, since it gets more scenarios right on every run. GPT Realtime 2.1 is ahead on reliability, success rate, agent responses, data accuracy, stalled calls and interruption, and Azure Realtime on response time and cost.
Which saves caller data more accurately, Azure Realtime or GPT Realtime 2.1?+
GPT Realtime 2.1, by 5.9 points. Azure Realtime saved everything exactly right on 86.1% of the calls that save data, and GPT Realtime 2.1 on 92.0%. On Appointments, 94.2% against 95.9%. On Medicare, 65.2% against 81.8%. A call counts only when nothing is missing, wrong, extra or saved twice.
Is Azure Realtime faster than GPT Realtime 2.1?+
Yes. Azure Realtime starts replying 0.28s sooner at the median. Azure Realtime takes 1.66s at the median and 2.22s at p90; GPT Realtime 2.1 takes 1.94s and 2.46s. Both are measured on the call audio, from the caller finishing to the agent starting to reply.
Which handles interruptions better, Azure Realtime or GPT Realtime 2.1?+
GPT Realtime 2.1, by 0.01 on a 0 to 5 scale. Azure Realtime scores 4.96 and GPT Realtime 2.1 4.98 for how well the agent handled being talked over.
Which goes silent on callers less often, Azure Realtime or GPT Realtime 2.1?+
GPT Realtime 2.1, by 0.8 points. The agent went silent for 10 seconds or more on 4.1% of Azure Realtime calls and 3.3% of GPT Realtime 2.1 calls.
Which is cheaper, Azure Realtime or GPT Realtime 2.1?+
Azure Realtime is $0.006 cheaper a minute. Azure Realtime costs $0.080 a minute of call and GPT Realtime 2.1 $0.085, at each vendor's published rates.
Where does the Azure Realtime vs GPT Realtime 2.1 data come from?+
Cekura ran both as the whole agent in the same open-source Pipecat pipeline, with the same prompt, tools and connection. A simulated caller worked through all 82 scenarios on live calls, three times each, 2,952 calls across the benchmark. A call passes only when the agent said the right things and saved the right data.