All comparisons
GoogleOpenAI

Gemini 3.1 Flash Live Preview vs GPT Realtime 2.1 Mini

Gemini 3.1 Flash Live Preview is ahead on reliability, success rate, agent responses and data accuracy, and GPT Realtime 2.1 Mini on stalled calls, interruption, response time and cost.

Compare

Success rate by suite

Appointments59 scenariosMedicare23 scenarios
  • GPT Realtime 2.1OpenAI
    92%
    81%
  • GPT-Live 1 + Sol lowOpenAI
    97%
    49%
  • Gemini 3.8 Live ThinkingGoogle
    82%
    61%
  • Grok Think Fast 2.0xAI
    96%
    25%
  • Gemini 3.1 Flash LiveGoogle
    93%
    32%
  • Phonic v1Phonic
    97%
    14%
  • Realtime 2.1 MiniOpenAI
    84%
    4%
  • Nova 2 SonicAmazon
    80%
    6%
  • Cascade, GPT-4.1Not ranked
    97%
    83%
Each model's share of calls that passed every check, on each suite.

Every metric, side by side

Reliability

Scenarios that passed on all three runs.

Gemini 3.1 Flash Live
62.2%5th
Realtime 2.1 Mini
50.0%7th

Gemini 3.1 Flash Live, 12.2 points higher

Success rate

Calls that passed every check.

Gemini 3.1 Flash Live
76.0%4th
Realtime 2.1 Mini
61.4%7th

Gemini 3.1 Flash Live, 14.6 points higher

Agent responses

Calls where the agent's replies reached the expected outcome.

Gemini 3.1 Flash Live
95.9%5th
Realtime 2.1 Mini
93.5%8th

Gemini 3.1 Flash Live, 2.4 points higher

Data accuracy

Calls where everything the agent saved matched exactly.

Gemini 3.1 Flash Live
79.2%4th
Realtime 2.1 Mini
64.5%7th

Gemini 3.1 Flash Live, 14.7 points higher

Stalled calls

Calls where the agent went silent for 10 seconds or more.

Gemini 3.1 Flash Live
3.3%5th
Realtime 2.1 Mini
2.5%3rd

Realtime 2.1 Mini, 0.8 points lower

Interruption

How well it handled being talked over, 0 to 5.

Gemini 3.1 Flash Live
4.984th
Realtime 2.1 Mini
4.992nd

Realtime 2.1 Mini, 0.01 higher

Response time

Median time to start replying after the caller finishes.

Gemini 3.1 Flash Live
2.88s8th
Realtime 2.1 Mini
2.08s5th

Realtime 2.1 Mini, 0.80s sooner

Cost / min

One minute of call at the vendor's published rates.

Gemini 3.1 Flash Live
$0.0703rd
Realtime 2.1 Mini
$0.0201st

Realtime 2.1 Mini, $0.050 cheaper

The small figure is the place among ranked models. The cascade is not ranked.

Realtime 2.1 Mini answers first

Realtime 2.1 Mini 2.08s, Gemini 3.1 Flash Live 2.88s at the median.

Response time, median to p90

MedianTo p90
Medianp90
  • Phonic v1Phonic
    Phonic v1: median 1.59s, p90 1.86s
    1.59s1.86s
  • Grok Think Fast 2.0xAI
    Grok Think Fast 2.0: median 1.62s, p90 1.94s
    1.62s1.94s
  • GPT Realtime 2.1OpenAI
    GPT Realtime 2.1: median 1.94s, p90 2.46s
    1.94s2.46s
  • GPT-Live 1 + Sol lowOpenAI
    GPT-Live 1 + Sol low: median 2.02s, p90 2.42s
    2.02s2.42s
  • Realtime 2.1 MiniOpenAI
    Realtime 2.1 Mini: median 2.08s, p90 2.75s
    2.08s2.75s
  • Nova 2 SonicAmazon
    Nova 2 Sonic: median 2.18s, p90 2.76s
    2.18s2.76s
  • Gemini 3.8 Live ThinkingGoogle
    Gemini 3.8 Live Thinking: median 2.32s, p90 3.06s
    2.32s3.06s
  • Gemini 3.1 Flash LiveGoogle
    Gemini 3.1 Flash Live: median 2.88s, p90 4.00s
    2.88s4.00s
  • Cascade, GPT-4.1Not ranked
    Cascade, GPT-4.1: median 2.17s, p90 2.51s
    2.17s2.51s
2.0s3.0s4.0s
Seconds from the caller finishing to the agent starting to reply, per model.

What we saw on the calls

Gemini 3.1 Flash Live Preview

  • Steady on the short form, though it repeats lookups it already made.
  • On the long form it hands the caller over with the wrong details.
  • Slowest to reply; after several tool calls its replies fall behind or go missing.
  • Sometimes says goodbye without hanging up.
Model
models/gemini-3.1-flash-live-preview
Setting
Thinking level high
Voice
Charon
Turn taking
Shared local turn detection
Audio
16 kHz
Cost
Averaged over the calls with a usage record.

GPT Realtime 2.1 Mini

  • Gets most short-form calls right, at a fraction of the cost.
  • Sometimes mishears a phone number.
  • Drops details given earlier: not one long-form record was exact.
  • Rarely recovers when a tool returns an error.
Model
gpt-realtime-2.1-mini
Setting
Reasoning high
Voice
marin
Turn taking
Service turn detection
Audio
24 kHz
Cost
Averaged over the calls with a usage record.

Frequently asked questions

Is Gemini 3.1 Flash Live Preview or GPT Realtime 2.1 Mini more reliable?

Gemini 3.1 Flash Live Preview passed 51 of 82 scenarios on all three runs, and GPT Realtime 2.1 Mini passed 41. Gemini 3.1 Flash Live Preview is ahead, by 12.2 points.

Which saves caller data more accurately, Gemini 3.1 Flash Live Preview or GPT Realtime 2.1 Mini?

Gemini 3.1 Flash Live Preview saved everything exactly right on 79.2% of the calls that save data, and GPT Realtime 2.1 Mini on 64.5%. On Appointments, 96.5% against 88.8%. On Medicare, 34.8% against 0.0%. A call counts only when nothing is missing, wrong, extra or saved twice.

Which responds faster, Gemini 3.1 Flash Live Preview or GPT Realtime 2.1 Mini?

GPT Realtime 2.1 Mini starts replying 0.80s sooner at the median. Gemini 3.1 Flash Live Preview takes 2.88s at the median and 4.00s at p90; GPT Realtime 2.1 Mini takes 2.08s and 2.75s. Both are measured on the call audio, from the caller finishing to the agent starting to reply.

Which costs less, Gemini 3.1 Flash Live Preview or GPT Realtime 2.1 Mini?

Gemini 3.1 Flash Live Preview costs $0.070 a minute of call and GPT Realtime 2.1 Mini $0.020, at each vendor's published rates. GPT Realtime 2.1 Mini is $0.050 cheaper a minute.

How were Gemini 3.1 Flash Live Preview and GPT Realtime 2.1 Mini compared?

Both ran as the whole agent in the same open-source Pipecat pipeline, with the same prompt, tools and phone line. A simulated caller worked through all 82 scenarios on live calls, three times each, 2,214 calls across the benchmark. A call passes only when the agent said the right things and saved the right data.