All comparisons
GoogleAmazon

Gemini 3.8 Live Extended Thinking vs Amazon Nova 2 Sonic

Gemini 3.8 Live Extended Thinking is ahead on reliability, success rate, agent responses and data accuracy, and Amazon Nova 2 Sonic on stalled calls, interruption and response time.

Compare

Success rate by suite

Appointments59 scenariosMedicare23 scenarios
  • GPT Realtime 2.1OpenAI
    92%
    81%
  • GPT-Live 1 + Sol lowOpenAI
    97%
    49%
  • Gemini 3.8 Live ThinkingGoogle
    82%
    61%
  • Grok Think Fast 2.0xAI
    96%
    25%
  • Gemini 3.1 Flash LiveGoogle
    93%
    32%
  • Phonic v1Phonic
    97%
    14%
  • Realtime 2.1 MiniOpenAI
    84%
    4%
  • Nova 2 SonicAmazon
    80%
    6%
  • Cascade, GPT-4.1Not ranked
    97%
    83%
Each model's share of calls that passed every check, on each suite.

Every metric, side by side

Reliability

Scenarios that passed on all three runs.

Gemini 3.8 Live Thinking
56.1%6th
Nova 2 Sonic
50.0%7th

Gemini 3.8 Live Thinking, 6.1 points higher

Success rate

Calls that passed every check.

Gemini 3.8 Live Thinking
76.4%3rd
Nova 2 Sonic
59.3%8th

Gemini 3.8 Live Thinking, 17.1 points higher

Agent responses

Calls where the agent's replies reached the expected outcome.

Gemini 3.8 Live Thinking
94.7%6th
Nova 2 Sonic
93.9%7th

Gemini 3.8 Live Thinking, 0.8 points higher

Data accuracy

Calls where everything the agent saved matched exactly.

Gemini 3.8 Live Thinking
80.2%3rd
Nova 2 Sonic
59.9%8th

Gemini 3.8 Live Thinking, 20.3 points higher

Stalled calls

Calls where the agent went silent for 10 seconds or more.

Gemini 3.8 Live Thinking
3.7%6th
Nova 2 Sonic
1.2%2nd

Nova 2 Sonic, 2.4 points lower

Interruption

How well it handled being talked over, 0 to 5.

Gemini 3.8 Live Thinking
4.906th
Nova 2 Sonic
4.991st

Nova 2 Sonic, 0.09 higher

Response time

Median time to start replying after the caller finishes.

Gemini 3.8 Live Thinking
2.32s7th
Nova 2 Sonic
2.18s6th

Nova 2 Sonic, 0.13s sooner

Cost / min

One minute of call at the vendor's published rates.

Gemini 3.8 Live Thinking
No verified rate
Nova 2 Sonic
No verified rate

Not comparable

The small figure is the place among ranked models. The cascade is not ranked.

Nova 2 Sonic answers first

Nova 2 Sonic 2.18s, Gemini 3.8 Live Thinking 2.32s at the median.

Response time, median to p90

MedianTo p90
Medianp90
  • Phonic v1Phonic
    Phonic v1: median 1.59s, p90 1.86s
    1.59s1.86s
  • Grok Think Fast 2.0xAI
    Grok Think Fast 2.0: median 1.62s, p90 1.94s
    1.62s1.94s
  • GPT Realtime 2.1OpenAI
    GPT Realtime 2.1: median 1.94s, p90 2.46s
    1.94s2.46s
  • GPT-Live 1 + Sol lowOpenAI
    GPT-Live 1 + Sol low: median 2.02s, p90 2.42s
    2.02s2.42s
  • Realtime 2.1 MiniOpenAI
    Realtime 2.1 Mini: median 2.08s, p90 2.75s
    2.08s2.75s
  • Nova 2 SonicAmazon
    Nova 2 Sonic: median 2.18s, p90 2.76s
    2.18s2.76s
  • Gemini 3.8 Live ThinkingGoogle
    Gemini 3.8 Live Thinking: median 2.32s, p90 3.06s
    2.32s3.06s
  • Gemini 3.1 Flash LiveGoogle
    Gemini 3.1 Flash Live: median 2.88s, p90 4.00s
    2.88s4.00s
  • Cascade, GPT-4.1Not ranked
    Cascade, GPT-4.1: median 2.17s, p90 2.51s
    2.17s2.51s
2.0s3.0s4.0s
Seconds from the caller finishing to the agent starting to reply, per model.

What we saw on the calls

Gemini 3.8 Live Extended Thinking

  • Strong on the long form; when it slips, it skips a save it said it made.
  • Weak on the short form: phone numbers saved wrong, lookups never made.
  • Acts on half an answer instead of asking; weakest when the caller corrects themselves.
  • Least predictable: the same scenario passes one run and fails the next.
  • Replies fall behind or stop; callers hang up waiting.
Model
gemini-3.8-live-extended-thinking
Setting
Thinking level high
Voice
Charon
Turn taking
Shared local turn detection
Audio
16 kHz
Cost
Some calls used cached input the vendor meters at a rate we have not verified.

Amazon Nova 2 Sonic

  • Misses caller names and phone numbers: blank or wrong.
  • Long forms rarely pass; more scenarios fail every run than for any other model.
  • Repeats the same tool call more than any other model.
  • Best voice score of the field, and the longest replies.
Model
amazon.nova-2-sonic-v1:0
Setting
Endpointing medium
Voice
matthew
Turn taking
Shared local turn detection
Audio
16 kHz
Cost
The vendor's rates for this model are not yet verified.

Frequently asked questions

Is Gemini 3.8 Live Extended Thinking or Amazon Nova 2 Sonic more reliable?

Gemini 3.8 Live Extended Thinking passed 46 of 82 scenarios on all three runs, and Amazon Nova 2 Sonic passed 41. Gemini 3.8 Live Extended Thinking is ahead, by 6.1 points.

Which saves caller data more accurately, Gemini 3.8 Live Extended Thinking or Amazon Nova 2 Sonic?

Gemini 3.8 Live Extended Thinking saved everything exactly right on 80.2% of the calls that save data, and Amazon Nova 2 Sonic on 59.9%. On Appointments, 86.5% against 82.5%. On Medicare, 63.6% against 1.5%. A call counts only when nothing is missing, wrong, extra or saved twice.

Which responds faster, Gemini 3.8 Live Extended Thinking or Amazon Nova 2 Sonic?

Amazon Nova 2 Sonic starts replying 0.13s sooner at the median. Gemini 3.8 Live Extended Thinking takes 2.32s at the median and 3.06s at p90; Amazon Nova 2 Sonic takes 2.18s and 2.76s. Both are measured on the call audio, from the caller finishing to the agent starting to reply.

Which costs less, Gemini 3.8 Live Extended Thinking or Amazon Nova 2 Sonic?

There is no verified per-minute rate for Gemini 3.8 Live Extended Thinking and Amazon Nova 2 Sonic, so cost cannot be compared yet.

How were Gemini 3.8 Live Extended Thinking and Amazon Nova 2 Sonic compared?

Both ran as the whole agent in the same open-source Pipecat pipeline, with the same prompt, tools and phone line. A simulated caller worked through all 82 scenarios on live calls, three times each, 2,214 calls across the benchmark. A call passes only when the agent said the right things and saved the right data.