All comparisons
AmazonOpenAI

Amazon Nova 2 Sonic vs GPT Realtime 2.1

GPT Realtime 2.1 is ahead on reliability, success rate, agent responses, data accuracy and response time, and Amazon Nova 2 Sonic on stalled calls and interruption.

Compare

Success rate by suite

Appointments59 scenariosMedicare23 scenarios
  • GPT Realtime 2.1OpenAI
    92%
    81%
  • GPT-Live 1 + Sol lowOpenAI
    97%
    49%
  • Gemini 3.8 Live ThinkingGoogle
    82%
    61%
  • Grok Think Fast 2.0xAI
    96%
    25%
  • Gemini 3.1 Flash LiveGoogle
    93%
    32%
  • Phonic v1Phonic
    97%
    14%
  • Realtime 2.1 MiniOpenAI
    84%
    4%
  • Nova 2 SonicAmazon
    80%
    6%
  • Cascade, GPT-4.1Not ranked
    97%
    83%
Each model's share of calls that passed every check, on each suite.

Every metric, side by side

Reliability

Scenarios that passed on all three runs.

Nova 2 Sonic
50.0%7th
GPT Realtime 2.1
79.3%1st

GPT Realtime 2.1, 29.3 points higher

Success rate

Calls that passed every check.

Nova 2 Sonic
59.3%8th
GPT Realtime 2.1
89.0%1st

GPT Realtime 2.1, 29.7 points higher

Agent responses

Calls where the agent's replies reached the expected outcome.

Nova 2 Sonic
93.9%7th
GPT Realtime 2.1
99.2%2nd

GPT Realtime 2.1, 5.3 points higher

Data accuracy

Calls where everything the agent saved matched exactly.

Nova 2 Sonic
59.9%8th
GPT Realtime 2.1
92.0%1st

GPT Realtime 2.1, 32.1 points higher

Stalled calls

Calls where the agent went silent for 10 seconds or more.

Nova 2 Sonic
1.2%2nd
GPT Realtime 2.1
3.3%4th

Nova 2 Sonic, 2.0 points lower

Interruption

How well it handled being talked over, 0 to 5.

Nova 2 Sonic
4.991st
GPT Realtime 2.1
4.983rd

Nova 2 Sonic, 0.01 higher

Response time

Median time to start replying after the caller finishes.

Nova 2 Sonic
2.18s6th
GPT Realtime 2.1
1.94s3rd

GPT Realtime 2.1, 0.25s sooner

Cost / min

One minute of call at the vendor's published rates.

Nova 2 Sonic
No verified rate
GPT Realtime 2.1
$0.0855th

Not comparable

The small figure is the place among ranked models. The cascade is not ranked.

GPT Realtime 2.1 answers first

GPT Realtime 2.1 1.94s, Nova 2 Sonic 2.18s at the median.

Response time, median to p90

MedianTo p90
Medianp90
  • Phonic v1Phonic
    Phonic v1: median 1.59s, p90 1.86s
    1.59s1.86s
  • Grok Think Fast 2.0xAI
    Grok Think Fast 2.0: median 1.62s, p90 1.94s
    1.62s1.94s
  • GPT Realtime 2.1OpenAI
    GPT Realtime 2.1: median 1.94s, p90 2.46s
    1.94s2.46s
  • GPT-Live 1 + Sol lowOpenAI
    GPT-Live 1 + Sol low: median 2.02s, p90 2.42s
    2.02s2.42s
  • Realtime 2.1 MiniOpenAI
    Realtime 2.1 Mini: median 2.08s, p90 2.75s
    2.08s2.75s
  • Nova 2 SonicAmazon
    Nova 2 Sonic: median 2.18s, p90 2.76s
    2.18s2.76s
  • Gemini 3.8 Live ThinkingGoogle
    Gemini 3.8 Live Thinking: median 2.32s, p90 3.06s
    2.32s3.06s
  • Gemini 3.1 Flash LiveGoogle
    Gemini 3.1 Flash Live: median 2.88s, p90 4.00s
    2.88s4.00s
  • Cascade, GPT-4.1Not ranked
    Cascade, GPT-4.1: median 2.17s, p90 2.51s
    2.17s2.51s
2.0s3.0s4.0s
Seconds from the caller finishing to the agent starting to reply, per model.

What we saw on the calls

Amazon Nova 2 Sonic

  • Misses caller names and phone numbers: blank or wrong.
  • Long forms rarely pass; more scenarios fail every run than for any other model.
  • Repeats the same tool call more than any other model.
  • Best voice score of the field, and the longest replies.
Model
amazon.nova-2-sonic-v1:0
Setting
Endpointing medium
Voice
matthew
Turn taking
Shared local turn detection
Audio
16 kHz
Cost
The vendor's rates for this model are not yet verified.

GPT Realtime 2.1

  • Only a handful of scenarios ever fail, and rarely twice.
  • Best realtime model on the long form, level with the cascade.
  • Few data errors: a phone number or an earlier reference, wrong or left out.
  • Weakest on noisy lines.
Model
gpt-realtime-2.1
Setting
Reasoning high
Voice
marin
Turn taking
Service turn detection
Audio
24 kHz
Cost
Excludes caller transcription, which the service does not meter per call.

Frequently asked questions

Is Amazon Nova 2 Sonic or GPT Realtime 2.1 more reliable?

Amazon Nova 2 Sonic passed 41 of 82 scenarios on all three runs, and GPT Realtime 2.1 passed 65. GPT Realtime 2.1 is ahead, by 29.3 points.

Which saves caller data more accurately, Amazon Nova 2 Sonic or GPT Realtime 2.1?

Amazon Nova 2 Sonic saved everything exactly right on 59.9% of the calls that save data, and GPT Realtime 2.1 on 92.0%. On Appointments, 82.5% against 95.9%. On Medicare, 1.5% against 81.8%. A call counts only when nothing is missing, wrong, extra or saved twice.

Which responds faster, Amazon Nova 2 Sonic or GPT Realtime 2.1?

GPT Realtime 2.1 starts replying 0.25s sooner at the median. Amazon Nova 2 Sonic takes 2.18s at the median and 2.76s at p90; GPT Realtime 2.1 takes 1.94s and 2.46s. Both are measured on the call audio, from the caller finishing to the agent starting to reply.

Which costs less, Amazon Nova 2 Sonic or GPT Realtime 2.1?

There is no verified per-minute rate for Amazon Nova 2 Sonic, so cost cannot be compared yet.

How were Amazon Nova 2 Sonic and GPT Realtime 2.1 compared?

Both ran as the whole agent in the same open-source Pipecat pipeline, with the same prompt, tools and phone line. A simulated caller worked through all 82 scenarios on live calls, three times each, 2,214 calls across the benchmark. A call passes only when the agent said the right things and saved the right data.