All comparisons
GoogleGoogle

Gemini 3.8 Live Extended Thinking vs Gemini 3.8 Live

Gemini 3.8 Live is ahead on agent responses, stalled calls, interruption and response time, and Gemini 3.8 Live Extended Thinking on reliability, success rate and data accuracy.

Compare

Success rate by suite

Appointments59 scenariosMedicare23 scenarios
  • GPT Realtime 2.1OpenAI
    92%
    81%
  • GPT-Live 1 + Sol lowOpenAI
    97%
    49%
  • Phonic v1Phonic
    97%
    26%
  • Gemini 3.8 Live ThinkingGoogle
    79%
    70%
  • Grok Think Fast 2.0xAI
    96%
    25%
  • Gemini 3.1 Flash LiveGoogle
    93%
    32%
  • Gemini 3.8 LiveGoogle
    76%
    52%
  • Realtime 2.1 MiniOpenAI
    84%
    4%
  • Nova 2 SonicAmazon
    80%
    6%
  • Cascade, GPT-4.1Not ranked
    97%
    83%
Each model's share of calls that passed every check, on each suite.

Every metric, side by side

Reliability

Scenarios that passed on all three runs.

Gemini 3.8 Live Thinking
52.4%6th
Gemini 3.8 Live
41.5%9th

Gemini 3.8 Live Thinking, 11.0 points higher

Success rate

Calls that passed every check.

Gemini 3.8 Live Thinking
76.0%4th
Gemini 3.8 Live
69.1%7th

Gemini 3.8 Live Thinking, 6.9 points higher

Agent responses

Calls where the agent's replies reached the expected outcome.

Gemini 3.8 Live Thinking
88.6%9th
Gemini 3.8 Live
94.3%6th

Gemini 3.8 Live, 5.7 points higher

Data accuracy

Calls where everything the agent saved matched exactly.

Gemini 3.8 Live Thinking
84.4%3rd
Gemini 3.8 Live
70.5%7th

Gemini 3.8 Live Thinking, 13.9 points higher

Stalled calls

Calls where the agent went silent for 10 seconds or more.

Gemini 3.8 Live Thinking
5.3%8th
Gemini 3.8 Live
2.0%3rd

Gemini 3.8 Live, 3.3 points lower

Interruption

How well it handled being talked over, 0 to 5.

Gemini 3.8 Live Thinking
4.956th
Gemini 3.8 Live
4.983rd

Gemini 3.8 Live, 0.03 higher

Response time

Median time to start replying after the caller finishes.

Gemini 3.8 Live Thinking
2.07s6th
Gemini 3.8 Live
1.91s3rd

Gemini 3.8 Live, 0.16s sooner

Cost / min

One minute of call at the vendor's published rates.

Gemini 3.8 Live Thinking
No verified rate
Gemini 3.8 Live
No verified rate

Not comparable

The small figure is the place among ranked models. The cascade is not ranked.

Gemini 3.8 Live answers first

Gemini 3.8 Live 1.91s, Gemini 3.8 Live Thinking 2.07s at the median.

Response time, median to p90

MedianTo p90
Medianp90
  • Phonic v1Phonic
    Phonic v1: median 1.60s, p90 1.88s
    1.60s1.88s
  • Grok Think Fast 2.0xAI
    Grok Think Fast 2.0: median 1.62s, p90 1.94s
    1.62s1.94s
  • Gemini 3.8 LiveGoogle
    Gemini 3.8 Live: median 1.91s, p90 2.23s
    1.91s2.23s
  • GPT Realtime 2.1OpenAI
    GPT Realtime 2.1: median 1.94s, p90 2.46s
    1.94s2.46s
  • GPT-Live 1 + Sol lowOpenAI
    GPT-Live 1 + Sol low: median 2.02s, p90 2.42s
    2.02s2.42s
  • Gemini 3.8 Live ThinkingGoogle
    Gemini 3.8 Live Thinking: median 2.07s, p90 2.74s
    2.07s2.74s
  • Realtime 2.1 MiniOpenAI
    Realtime 2.1 Mini: median 2.08s, p90 2.75s
    2.08s2.75s
  • Nova 2 SonicAmazon
    Nova 2 Sonic: median 2.18s, p90 2.76s
    2.18s2.76s
  • Gemini 3.1 Flash LiveGoogle
    Gemini 3.1 Flash Live: median 2.88s, p90 4.00s
    2.88s4.00s
  • Cascade, GPT-4.1Not ranked
    Cascade, GPT-4.1: median 2.17s, p90 2.51s
    2.17s2.51s
2.0s3.0s4.0s
Seconds from the caller finishing to the agent starting to reply, per model.

What we saw on the calls

Gemini 3.8 Live Extended Thinking

  • Second only to GPT Realtime 2.1 on the long form.
  • Weakest of the field on plain bookings: the call misses the outcome, often with every tool call right.
  • Unpredictable: the same scenario passes one run and fails the next.
  • Google's turn detection stalls it under background conversation: the caller finishes and no reply comes (AS14, AS50, AS52).
Model
gemini-3.8-live-extended-thinking
Setting
Thinking level high
Voice
Charon
Turn taking
Service turn detection
Audio
16 kHz
Cost
The export carries audio tokens but not text or cached input, so a per-call cost is not derived.

Gemini 3.8 Live

  • Without extended thinking it replies sooner, and short-form calls reach the outcome more often.
  • On the long form the conversation still reaches the outcome, but the saved record is wrong far more often.
  • Weakest of the field when the caller corrects themselves.
  • Google's turn detection stalls it under background conversation: the caller finishes and no reply comes (AS50, AS52, AS14).
  • Least predictable of the field.
Model
gemini-3.8-live
Setting
No thinking level
Voice
Charon
Turn taking
Service turn detection
Audio
16 kHz
Cost
The export carries audio tokens but not text or cached input, so a per-call cost is not derived.

Frequently asked questions

Is Gemini 3.8 Live Extended Thinking or Gemini 3.8 Live more reliable?

Gemini 3.8 Live Extended Thinking passed 43 of 82 scenarios on all three runs, and Gemini 3.8 Live passed 34. Gemini 3.8 Live Extended Thinking is ahead, by 11.0 points.

Which saves caller data more accurately, Gemini 3.8 Live Extended Thinking or Gemini 3.8 Live?

Gemini 3.8 Live Extended Thinking saved everything exactly right on 84.4% of the calls that save data, and Gemini 3.8 Live on 70.5%. On Appointments, 88.9% against 78.4%. On Medicare, 72.7% against 50.0%. A call counts only when nothing is missing, wrong, extra or saved twice.

Which responds faster, Gemini 3.8 Live Extended Thinking or Gemini 3.8 Live?

Gemini 3.8 Live starts replying 0.16s sooner at the median. Gemini 3.8 Live Extended Thinking takes 2.07s at the median and 2.74s at p90; Gemini 3.8 Live takes 1.91s and 2.23s. Both are measured on the call audio, from the caller finishing to the agent starting to reply.

Which costs less, Gemini 3.8 Live Extended Thinking or Gemini 3.8 Live?

There is no verified per-minute rate for Gemini 3.8 Live Extended Thinking and Gemini 3.8 Live, so cost cannot be compared yet.

How were Gemini 3.8 Live Extended Thinking and Gemini 3.8 Live compared?

Both ran as the whole agent in the same open-source Pipecat pipeline, with the same prompt, tools and connection. A simulated caller worked through all 82 scenarios on live calls, three times each, 2,460 calls across the benchmark. A call passes only when the agent said the right things and saved the right data.