Is Azure Realtime or GPT Realtime 2.1 Mini more reliable?+
Azure Realtime is more reliable, by 20.7 points. Azure Realtime passed 58 of 82 scenarios on all three runs, and GPT Realtime 2.1 Mini passed 41. A scenario only counts when every run of it passed.
Which should I pick, Azure Realtime or GPT Realtime 2.1 Mini?+
For a phone agent, Azure Realtime, since it gets more scenarios right on every run. Azure Realtime is ahead on reliability, success rate, agent responses, data accuracy and response time, and GPT Realtime 2.1 Mini on stalled calls, interruption and cost.
Which saves caller data more accurately, Azure Realtime or GPT Realtime 2.1 Mini?+
Azure Realtime, by 21.5 points. Azure Realtime saved everything exactly right on 86.1% of the calls that save data, and GPT Realtime 2.1 Mini on 64.5%. On Appointments, 94.2% against 88.8%. On Medicare, 65.2% against 0.0%. A call counts only when nothing is missing, wrong, extra or saved twice.
Is Azure Realtime faster than GPT Realtime 2.1 Mini?+
Yes. Azure Realtime starts replying 0.43s sooner at the median. Azure Realtime takes 1.66s at the median and 2.22s at p90; GPT Realtime 2.1 Mini takes 2.08s and 2.75s. Both are measured on the call audio, from the caller finishing to the agent starting to reply.
Which handles interruptions better, Azure Realtime or GPT Realtime 2.1 Mini?+
GPT Realtime 2.1 Mini, by 0.02 on a 0 to 5 scale. Azure Realtime scores 4.96 and GPT Realtime 2.1 Mini 4.99 for how well the agent handled being talked over.
Which goes silent on callers less often, Azure Realtime or GPT Realtime 2.1 Mini?+
GPT Realtime 2.1 Mini, by 1.6 points. The agent went silent for 10 seconds or more on 4.1% of Azure Realtime calls and 2.5% of GPT Realtime 2.1 Mini calls.
Which is cheaper, Azure Realtime or GPT Realtime 2.1 Mini?+
GPT Realtime 2.1 Mini is $0.060 cheaper a minute. Azure Realtime costs $0.080 a minute of call and GPT Realtime 2.1 Mini $0.020, at each vendor's published rates.
Where does the Azure Realtime vs GPT Realtime 2.1 Mini data come from?+
Cekura ran both as the whole agent in the same open-source Pipecat pipeline, with the same prompt, tools and connection. A simulated caller worked through all 82 scenarios on live calls, three times each, 2,952 calls across the benchmark. A call passes only when the agent said the right things and saved the right data.