All comparisons
MicrosoftOpenAI

Azure Realtime vs GPT-Live 1 + GPT-6 Sol (low thinking)

Azure Realtime is ahead on success rate, stalled calls, interruption and response time, and GPT-Live 1 + GPT-6 Sol (low thinking) on reliability, agent responses, data accuracy and cost.

Last updated Methodology by Dileep Chagam

Compare

Success rate by suite

Higher is better
Appointments59 scenariosMedicare23 scenarios
  • Azure RealtimeMicrosoft
    91%
    67%
  • GPT-Live 1 + Sol lowOpenAI
    97%
    49%
Each model's share of calls that passed every check, on each suite.
Verdict

Which to pick

Pick Azure Realtime for calls that pass every check (0.4 points higher), fewer stalled calls (8.1 points lower), handling interruptions (0.08 higher) and fast replies (0.36s sooner). Pick GPT-Live 1 + GPT-6 Sol (low thinking) for reliability across repeated calls (4.9 points higher), getting the agent's replies right (1.2 points higher), saving the caller's data correctly (3.0 points higher) and lower cost ($0.024 cheaper).

Every metric, side by side

Reliability

Scenarios that passed on all three runs.

Azure Realtime
70.7%3rd
GPT-Live 1 + Sol low
75.6%2nd

GPT-Live 1 + Sol low, 4.9 points higher

Success rate

Calls that passed every check.

Azure Realtime
84.1%2nd
GPT-Live 1 + Sol low
83.7%3rd

Azure Realtime, 0.4 points higher

Agent responses

Calls where the agent's replies reached the expected outcome.

Azure Realtime
98.8%3rd
GPT-Live 1 + Sol low
100.0%1st

GPT-Live 1 + Sol low, 1.2 points higher

Data accuracy

Calls where everything the agent saved matched exactly.

Azure Realtime
86.1%3rd
GPT-Live 1 + Sol low
89.0%2nd

GPT-Live 1 + Sol low, 3.0 points higher

Stalled calls

Calls where the agent went silent for 10 seconds or more.

Azure Realtime
4.1%7th
GPT-Live 1 + Sol low
12.2%11th

Azure Realtime, 8.1 points lower

Interruption

How well it handled being talked over, 0 to 5.

Azure Realtime
4.966th
GPT-Live 1 + Sol low
4.8910th

Azure Realtime, 0.08 higher

Response time

Median time to start replying after the caller finishes.

Azure Realtime
1.66s3rd
GPT-Live 1 + Sol low
2.02s6th

Azure Realtime, 0.36s sooner

Cost / min

One minute of call at the vendor's published rates.

Azure Realtime
$0.0804th
GPT-Live 1 + Sol low
$0.0552nd

GPT-Live 1 + Sol low, $0.024 cheaper

The small figure is the place among ranked models. The cascade is not ranked.

Azure Realtime answers first

Azure Realtime 1.66s, GPT-Live 1 + Sol low 2.02s at the median.

Response time, median to p90

Lower is betterMedianTo p90
Medianp90
  • Azure RealtimeMicrosoft
    Azure Realtime: median 1.66s, p90 2.22s
    1.66s2.22s
  • GPT-Live 1 + Sol lowOpenAI
    GPT-Live 1 + Sol low: median 2.02s, p90 2.42s
    2.02s2.42s
1.5s2.5s
Seconds from the caller finishing to the agent starting to reply, per model.

What we saw on the calls

Azure Realtime

  • Handles a caller correcting themselves better than any other model.
All 4 notes on Azure Realtime
Model
azure-realtime
Setting
No reasoning setting
Voice
ava
Turn taking
Service turn detection
Audio
24 kHz
Cost
Each call's metered tokens at Voice Live's Pro rates, read from Azure's price list on 7 Oct 2026. Excludes caller transcription, which Azure bills separately as speech input.

GPT-Live 1 + GPT-6 Sol (low thinking)

  • Near perfect on bookings, run after run.
All 4 notes on GPT-Live 1 + Sol low
Model
gpt-live-1
Setting
Delegates reasoning and tools to Sol
Voice
marin
Turn taking
Service turn detection
Audio
24 kHz
Cost
Includes the delegated backend model.

Frequently asked questions

Is Azure Realtime or GPT-Live 1 + GPT-6 Sol (low thinking) more reliable?

GPT-Live 1 + GPT-6 Sol (low thinking) is more reliable, by 4.9 points. Azure Realtime passed 58 of 82 scenarios on all three runs, and GPT-Live 1 + GPT-6 Sol (low thinking) passed 62. A scenario only counts when every run of it passed.

Which should I pick, Azure Realtime or GPT-Live 1 + GPT-6 Sol (low thinking)?

For a phone agent, GPT-Live 1 + GPT-6 Sol (low thinking), since it gets more scenarios right on every run. Azure Realtime is ahead on success rate, stalled calls, interruption and response time, and GPT-Live 1 + GPT-6 Sol (low thinking) on reliability, agent responses, data accuracy and cost.

Which saves caller data more accurately, Azure Realtime or GPT-Live 1 + GPT-6 Sol (low thinking)?

GPT-Live 1 + GPT-6 Sol (low thinking), by 3.0 points. Azure Realtime saved everything exactly right on 86.1% of the calls that save data, and GPT-Live 1 + GPT-6 Sol (low thinking) on 89.0%. On Appointments, 94.2% against 100.0%. On Medicare, 65.2% against 60.6%. A call counts only when nothing is missing, wrong, extra or saved twice.

Is Azure Realtime faster than GPT-Live 1 + GPT-6 Sol (low thinking)?

Yes. Azure Realtime starts replying 0.36s sooner at the median. Azure Realtime takes 1.66s at the median and 2.22s at p90; GPT-Live 1 + GPT-6 Sol (low thinking) takes 2.02s and 2.42s. Both are measured on the call audio, from the caller finishing to the agent starting to reply.

Which handles interruptions better, Azure Realtime or GPT-Live 1 + GPT-6 Sol (low thinking)?

Azure Realtime, by 0.08 on a 0 to 5 scale. Azure Realtime scores 4.96 and GPT-Live 1 + GPT-6 Sol (low thinking) 4.89 for how well the agent handled being talked over.

Which goes silent on callers less often, Azure Realtime or GPT-Live 1 + GPT-6 Sol (low thinking)?

Azure Realtime, by 8.1 points. The agent went silent for 10 seconds or more on 4.1% of Azure Realtime calls and 12.2% of GPT-Live 1 + GPT-6 Sol (low thinking) calls.

Which is cheaper, Azure Realtime or GPT-Live 1 + GPT-6 Sol (low thinking)?

GPT-Live 1 + GPT-6 Sol (low thinking) is $0.024 cheaper a minute. Azure Realtime costs $0.080 a minute of call and GPT-Live 1 + GPT-6 Sol (low thinking) $0.055, at each vendor's published rates.

Where does the Azure Realtime vs GPT-Live 1 + GPT-6 Sol (low thinking) data come from?

Cekura ran both as the whole agent in the same open-source Pipecat pipeline, with the same prompt, tools and connection. A simulated caller worked through all 82 scenarios on live calls, three times each, 2,952 calls across the benchmark. A call passes only when the agent said the right things and saved the right data.