All comparisons
OpenAIAmazon

GPT-Live 1 + GPT-6 Sol (low thinking) vs Amazon Nova 2 Sonic

GPT-Live 1 + GPT-6 Sol (low thinking) is ahead on reliability, success rate, agent responses, data accuracy and response time, and Amazon Nova 2 Sonic on stalled calls and interruption.

Compare

Success rate by suite

Appointments59 scenariosMedicare23 scenarios
  • GPT Realtime 2.1OpenAI
    92%
    81%
  • GPT-Live 1 + Sol lowOpenAI
    97%
    49%
  • Gemini 3.8 Live ThinkingGoogle
    82%
    61%
  • Grok Think Fast 2.0xAI
    96%
    25%
  • Gemini 3.1 Flash LiveGoogle
    93%
    32%
  • Phonic v1Phonic
    97%
    14%
  • Realtime 2.1 MiniOpenAI
    84%
    4%
  • Nova 2 SonicAmazon
    80%
    6%
  • Cascade, GPT-4.1Not ranked
    97%
    83%
Each model's share of calls that passed every check, on each suite.

Every metric, side by side

Reliability

Scenarios that passed on all three runs.

GPT-Live 1 + Sol low
75.6%2nd
Nova 2 Sonic
50.0%7th

GPT-Live 1 + Sol low, 25.6 points higher

Success rate

Calls that passed every check.

GPT-Live 1 + Sol low
83.7%2nd
Nova 2 Sonic
59.3%8th

GPT-Live 1 + Sol low, 24.4 points higher

Agent responses

Calls where the agent's replies reached the expected outcome.

GPT-Live 1 + Sol low
100.0%1st
Nova 2 Sonic
93.9%7th

GPT-Live 1 + Sol low, 6.1 points higher

Data accuracy

Calls where everything the agent saved matched exactly.

GPT-Live 1 + Sol low
89.0%2nd
Nova 2 Sonic
59.9%8th

GPT-Live 1 + Sol low, 29.1 points higher

Stalled calls

Calls where the agent went silent for 10 seconds or more.

GPT-Live 1 + Sol low
12.2%8th
Nova 2 Sonic
1.2%2nd

Nova 2 Sonic, 11.0 points lower

Interruption

How well it handled being talked over, 0 to 5.

GPT-Live 1 + Sol low
4.897th
Nova 2 Sonic
4.991st

Nova 2 Sonic, 0.10 higher

Response time

Median time to start replying after the caller finishes.

GPT-Live 1 + Sol low
2.02s4th
Nova 2 Sonic
2.18s6th

GPT-Live 1 + Sol low, 0.16s sooner

Cost / min

One minute of call at the vendor's published rates.

GPT-Live 1 + Sol low
$0.0552nd
Nova 2 Sonic
No verified rate

Not comparable

The small figure is the place among ranked models. The cascade is not ranked.

GPT-Live 1 + Sol low answers first

GPT-Live 1 + Sol low 2.02s, Nova 2 Sonic 2.18s at the median.

Response time, median to p90

MedianTo p90
Medianp90
  • Phonic v1Phonic
    Phonic v1: median 1.59s, p90 1.86s
    1.59s1.86s
  • Grok Think Fast 2.0xAI
    Grok Think Fast 2.0: median 1.62s, p90 1.94s
    1.62s1.94s
  • GPT Realtime 2.1OpenAI
    GPT Realtime 2.1: median 1.94s, p90 2.46s
    1.94s2.46s
  • GPT-Live 1 + Sol lowOpenAI
    GPT-Live 1 + Sol low: median 2.02s, p90 2.42s
    2.02s2.42s
  • Realtime 2.1 MiniOpenAI
    Realtime 2.1 Mini: median 2.08s, p90 2.75s
    2.08s2.75s
  • Nova 2 SonicAmazon
    Nova 2 Sonic: median 2.18s, p90 2.76s
    2.18s2.76s
  • Gemini 3.8 Live ThinkingGoogle
    Gemini 3.8 Live Thinking: median 2.32s, p90 3.06s
    2.32s3.06s
  • Gemini 3.1 Flash LiveGoogle
    Gemini 3.1 Flash Live: median 2.88s, p90 4.00s
    2.88s4.00s
  • Cascade, GPT-4.1Not ranked
    Cascade, GPT-4.1: median 2.17s, p90 2.51s
    2.17s2.51s
2.0s3.0s4.0s
Seconds from the caller finishing to the agent starting to reply, per model.

What we saw on the calls

GPT-Live 1 + GPT-6 Sol (low thinking)

  • Near perfect on the short form, run after run.
  • Every failure is in the saved record; the conversation itself always reaches the outcome.
  • On the long form it routes the caller wrong or skips a step.
  • Says goodbye without hanging up, and stalls more than any other model.
  • Lowest voice score of the field.
Model
gpt-live-1
Setting
Delegates reasoning and tools to Sol
Voice
marin
Turn taking
Service turn detection
Audio
24 kHz
Cost
Includes the delegated backend model.

Amazon Nova 2 Sonic

  • Misses caller names and phone numbers: blank or wrong.
  • Long forms rarely pass; more scenarios fail every run than for any other model.
  • Repeats the same tool call more than any other model.
  • Best voice score of the field, and the longest replies.
Model
amazon.nova-2-sonic-v1:0
Setting
Endpointing medium
Voice
matthew
Turn taking
Shared local turn detection
Audio
16 kHz
Cost
The vendor's rates for this model are not yet verified.

Frequently asked questions

Is GPT-Live 1 + GPT-6 Sol (low thinking) or Amazon Nova 2 Sonic more reliable?

GPT-Live 1 + GPT-6 Sol (low thinking) passed 62 of 82 scenarios on all three runs, and Amazon Nova 2 Sonic passed 41. GPT-Live 1 + GPT-6 Sol (low thinking) is ahead, by 25.6 points.

Which saves caller data more accurately, GPT-Live 1 + GPT-6 Sol (low thinking) or Amazon Nova 2 Sonic?

GPT-Live 1 + GPT-6 Sol (low thinking) saved everything exactly right on 89.0% of the calls that save data, and Amazon Nova 2 Sonic on 59.9%. On Appointments, 100.0% against 82.5%. On Medicare, 60.6% against 1.5%. A call counts only when nothing is missing, wrong, extra or saved twice.

Which responds faster, GPT-Live 1 + GPT-6 Sol (low thinking) or Amazon Nova 2 Sonic?

GPT-Live 1 + GPT-6 Sol (low thinking) starts replying 0.16s sooner at the median. GPT-Live 1 + GPT-6 Sol (low thinking) takes 2.02s at the median and 2.42s at p90; Amazon Nova 2 Sonic takes 2.18s and 2.76s. Both are measured on the call audio, from the caller finishing to the agent starting to reply.

Which costs less, GPT-Live 1 + GPT-6 Sol (low thinking) or Amazon Nova 2 Sonic?

There is no verified per-minute rate for Amazon Nova 2 Sonic, so cost cannot be compared yet.

How were GPT-Live 1 + GPT-6 Sol (low thinking) and Amazon Nova 2 Sonic compared?

Both ran as the whole agent in the same open-source Pipecat pipeline, with the same prompt, tools and phone line. A simulated caller worked through all 82 scenarios on live calls, three times each, 2,214 calls across the benchmark. A call passes only when the agent said the right things and saved the right data.