All comparisons
AmazonCascade baseline

Amazon Nova 2 Sonic vs Flux → GPT-4.1 → ElevenLabs Flash

Flux → GPT-4.1 → ElevenLabs Flash is ahead on reliability, success rate, agent responses, data accuracy, stalled calls and response time, and Amazon Nova 2 Sonic on interruption.

Compare

Success rate by suite

Appointments59 scenariosMedicare23 scenarios
  • GPT Realtime 2.1OpenAI
    92%
    81%
  • GPT-Live 1 + Sol lowOpenAI
    97%
    49%
  • Gemini 3.8 Live ThinkingGoogle
    82%
    61%
  • Grok Think Fast 2.0xAI
    96%
    25%
  • Gemini 3.1 Flash LiveGoogle
    93%
    32%
  • Phonic v1Phonic
    97%
    14%
  • Realtime 2.1 MiniOpenAI
    84%
    4%
  • Nova 2 SonicAmazon
    80%
    6%
  • Cascade, GPT-4.1Not ranked
    97%
    83%
Each model's share of calls that passed every check, on each suite.

Every metric, side by side

Reliability

Scenarios that passed on all three runs.

Nova 2 Sonic
50.0%7th
Cascade, GPT-4.1
82.9%

Cascade, GPT-4.1, 32.9 points higher

Success rate

Calls that passed every check.

Nova 2 Sonic
59.3%8th
Cascade, GPT-4.1
93.1%

Cascade, GPT-4.1, 33.7 points higher

Agent responses

Calls where the agent's replies reached the expected outcome.

Nova 2 Sonic
93.9%7th
Cascade, GPT-4.1
99.2%

Cascade, GPT-4.1, 5.3 points higher

Data accuracy

Calls where everything the agent saved matched exactly.

Nova 2 Sonic
59.9%8th
Cascade, GPT-4.1
94.1%

Cascade, GPT-4.1, 34.2 points higher

Stalled calls

Calls where the agent went silent for 10 seconds or more.

Nova 2 Sonic
1.2%2nd
Cascade, GPT-4.1
0.8%

Cascade, GPT-4.1, 0.4 points lower

Interruption

How well it handled being talked over, 0 to 5.

Nova 2 Sonic
4.991st
Cascade, GPT-4.1
4.96

Nova 2 Sonic, 0.03 higher

Response time

Median time to start replying after the caller finishes.

Nova 2 Sonic
2.18s6th
Cascade, GPT-4.1
2.17s

Cascade, GPT-4.1, 0.01s sooner

Cost / min

One minute of call at the vendor's published rates.

Nova 2 Sonic
No verified rate
Cascade, GPT-4.1
No verified rate

Not comparable

The small figure is the place among ranked models. The cascade is not ranked.

Cascade, GPT-4.1 answers first

Cascade, GPT-4.1 2.17s, Nova 2 Sonic 2.18s at the median.

Response time, median to p90

MedianTo p90
Medianp90
  • Phonic v1Phonic
    Phonic v1: median 1.59s, p90 1.86s
    1.59s1.86s
  • Grok Think Fast 2.0xAI
    Grok Think Fast 2.0: median 1.62s, p90 1.94s
    1.62s1.94s
  • GPT Realtime 2.1OpenAI
    GPT Realtime 2.1: median 1.94s, p90 2.46s
    1.94s2.46s
  • GPT-Live 1 + Sol lowOpenAI
    GPT-Live 1 + Sol low: median 2.02s, p90 2.42s
    2.02s2.42s
  • Realtime 2.1 MiniOpenAI
    Realtime 2.1 Mini: median 2.08s, p90 2.75s
    2.08s2.75s
  • Nova 2 SonicAmazon
    Nova 2 Sonic: median 2.18s, p90 2.76s
    2.18s2.76s
  • Gemini 3.8 Live ThinkingGoogle
    Gemini 3.8 Live Thinking: median 2.32s, p90 3.06s
    2.32s3.06s
  • Gemini 3.1 Flash LiveGoogle
    Gemini 3.1 Flash Live: median 2.88s, p90 4.00s
    2.88s4.00s
  • Cascade, GPT-4.1Not ranked
    Cascade, GPT-4.1: median 2.17s, p90 2.51s
    2.17s2.51s
2.0s3.0s4.0s
Seconds from the caller finishing to the agent starting to reply, per model.

What we saw on the calls

Amazon Nova 2 Sonic

  • Misses caller names and phone numbers: blank or wrong.
  • Long forms rarely pass; more scenarios fail every run than for any other model.
  • Repeats the same tool call more than any other model.
  • Best voice score of the field, and the longest replies.
Model
amazon.nova-2-sonic-v1:0
Setting
Endpointing medium
Voice
matthew
Turn taking
Shared local turn detection
Audio
16 kHz
Cost
The vendor's rates for this model are not yet verified.

Flux → GPT-4.1 → ElevenLabs Flash

  • Most scenarios pass every run; none fails every run.
  • Rare long-form slips: a reference from an earlier step left out.
  • Not the quickest to reply, but almost never leaves the caller waiting.
Model
flux-general-en → gpt-4.1 → eleven_flash_v2_5
Setting
No reasoning setting
Voice
ElevenLabs 21m00Tcm4TlvDq8ikWAM
Turn taking
Speech-to-text turn detection
Audio
16 kHz
Cost
Not published: no verified combined rate for the three services.

Frequently asked questions

Is Amazon Nova 2 Sonic or Flux → GPT-4.1 → ElevenLabs Flash more reliable?

Amazon Nova 2 Sonic passed 41 of 82 scenarios on all three runs, and Flux → GPT-4.1 → ElevenLabs Flash passed 68. Flux → GPT-4.1 → ElevenLabs Flash is ahead, by 32.9 points.

Which saves caller data more accurately, Amazon Nova 2 Sonic or Flux → GPT-4.1 → ElevenLabs Flash?

Amazon Nova 2 Sonic saved everything exactly right on 59.9% of the calls that save data, and Flux → GPT-4.1 → ElevenLabs Flash on 94.1%. On Appointments, 82.5% against 98.8%. On Medicare, 1.5% against 81.8%. A call counts only when nothing is missing, wrong, extra or saved twice.

Which responds faster, Amazon Nova 2 Sonic or Flux → GPT-4.1 → ElevenLabs Flash?

Flux → GPT-4.1 → ElevenLabs Flash starts replying 0.01s sooner at the median. Amazon Nova 2 Sonic takes 2.18s at the median and 2.76s at p90; Flux → GPT-4.1 → ElevenLabs Flash takes 2.17s and 2.51s. Both are measured on the call audio, from the caller finishing to the agent starting to reply.

Which costs less, Amazon Nova 2 Sonic or Flux → GPT-4.1 → ElevenLabs Flash?

There is no verified per-minute rate for Amazon Nova 2 Sonic and Flux → GPT-4.1 → ElevenLabs Flash, so cost cannot be compared yet.

How were Amazon Nova 2 Sonic and Flux → GPT-4.1 → ElevenLabs Flash compared?

Both ran as the whole agent in the same open-source Pipecat pipeline, with the same prompt, tools and phone line. A simulated caller worked through all 82 scenarios on live calls, three times each, 2,214 calls across the benchmark. A call passes only when the agent said the right things and saved the right data. The cascade runs the same agent for reference and is never ranked.