All comparisons
OpenAICascade baseline

GPT Realtime 2.1 Mini vs Flux → GPT-4.1 → ElevenLabs Flash

Flux → GPT-4.1 → ElevenLabs Flash is ahead on reliability, success rate, agent responses, data accuracy and stalled calls, and GPT Realtime 2.1 Mini on interruption and response time.

Compare

Success rate by suite

Appointments59 scenariosMedicare23 scenarios
  • GPT Realtime 2.1OpenAI
    92%
    81%
  • GPT-Live 1 + Sol lowOpenAI
    97%
    49%
  • Gemini 3.8 Live ThinkingGoogle
    82%
    61%
  • Grok Think Fast 2.0xAI
    96%
    25%
  • Gemini 3.1 Flash LiveGoogle
    93%
    32%
  • Phonic v1Phonic
    97%
    14%
  • Realtime 2.1 MiniOpenAI
    84%
    4%
  • Nova 2 SonicAmazon
    80%
    6%
  • Cascade, GPT-4.1Not ranked
    97%
    83%
Each model's share of calls that passed every check, on each suite.

Every metric, side by side

Reliability

Scenarios that passed on all three runs.

Realtime 2.1 Mini
50.0%7th
Cascade, GPT-4.1
82.9%

Cascade, GPT-4.1, 32.9 points higher

Success rate

Calls that passed every check.

Realtime 2.1 Mini
61.4%7th
Cascade, GPT-4.1
93.1%

Cascade, GPT-4.1, 31.7 points higher

Agent responses

Calls where the agent's replies reached the expected outcome.

Realtime 2.1 Mini
93.5%8th
Cascade, GPT-4.1
99.2%

Cascade, GPT-4.1, 5.7 points higher

Data accuracy

Calls where everything the agent saved matched exactly.

Realtime 2.1 Mini
64.5%7th
Cascade, GPT-4.1
94.1%

Cascade, GPT-4.1, 29.6 points higher

Stalled calls

Calls where the agent went silent for 10 seconds or more.

Realtime 2.1 Mini
2.5%3rd
Cascade, GPT-4.1
0.8%

Cascade, GPT-4.1, 1.6 points lower

Interruption

How well it handled being talked over, 0 to 5.

Realtime 2.1 Mini
4.992nd
Cascade, GPT-4.1
4.96

Realtime 2.1 Mini, 0.03 higher

Response time

Median time to start replying after the caller finishes.

Realtime 2.1 Mini
2.08s5th
Cascade, GPT-4.1
2.17s

Realtime 2.1 Mini, 0.09s sooner

Cost / min

One minute of call at the vendor's published rates.

Realtime 2.1 Mini
$0.0201st
Cascade, GPT-4.1
No verified rate

Not comparable

The small figure is the place among ranked models. The cascade is not ranked.

Realtime 2.1 Mini answers first

Realtime 2.1 Mini 2.08s, Cascade, GPT-4.1 2.17s at the median.

Response time, median to p90

MedianTo p90
Medianp90
  • Phonic v1Phonic
    Phonic v1: median 1.59s, p90 1.86s
    1.59s1.86s
  • Grok Think Fast 2.0xAI
    Grok Think Fast 2.0: median 1.62s, p90 1.94s
    1.62s1.94s
  • GPT Realtime 2.1OpenAI
    GPT Realtime 2.1: median 1.94s, p90 2.46s
    1.94s2.46s
  • GPT-Live 1 + Sol lowOpenAI
    GPT-Live 1 + Sol low: median 2.02s, p90 2.42s
    2.02s2.42s
  • Realtime 2.1 MiniOpenAI
    Realtime 2.1 Mini: median 2.08s, p90 2.75s
    2.08s2.75s
  • Nova 2 SonicAmazon
    Nova 2 Sonic: median 2.18s, p90 2.76s
    2.18s2.76s
  • Gemini 3.8 Live ThinkingGoogle
    Gemini 3.8 Live Thinking: median 2.32s, p90 3.06s
    2.32s3.06s
  • Gemini 3.1 Flash LiveGoogle
    Gemini 3.1 Flash Live: median 2.88s, p90 4.00s
    2.88s4.00s
  • Cascade, GPT-4.1Not ranked
    Cascade, GPT-4.1: median 2.17s, p90 2.51s
    2.17s2.51s
2.0s3.0s4.0s
Seconds from the caller finishing to the agent starting to reply, per model.

What we saw on the calls

GPT Realtime 2.1 Mini

  • Gets most short-form calls right, at a fraction of the cost.
  • Sometimes mishears a phone number.
  • Drops details given earlier: not one long-form record was exact.
  • Rarely recovers when a tool returns an error.
Model
gpt-realtime-2.1-mini
Setting
Reasoning high
Voice
marin
Turn taking
Service turn detection
Audio
24 kHz
Cost
Averaged over the calls with a usage record.

Flux → GPT-4.1 → ElevenLabs Flash

  • Most scenarios pass every run; none fails every run.
  • Rare long-form slips: a reference from an earlier step left out.
  • Not the quickest to reply, but almost never leaves the caller waiting.
Model
flux-general-en → gpt-4.1 → eleven_flash_v2_5
Setting
No reasoning setting
Voice
ElevenLabs 21m00Tcm4TlvDq8ikWAM
Turn taking
Speech-to-text turn detection
Audio
16 kHz
Cost
Not published: no verified combined rate for the three services.

Frequently asked questions

Is GPT Realtime 2.1 Mini or Flux → GPT-4.1 → ElevenLabs Flash more reliable?

GPT Realtime 2.1 Mini passed 41 of 82 scenarios on all three runs, and Flux → GPT-4.1 → ElevenLabs Flash passed 68. Flux → GPT-4.1 → ElevenLabs Flash is ahead, by 32.9 points.

Which saves caller data more accurately, GPT Realtime 2.1 Mini or Flux → GPT-4.1 → ElevenLabs Flash?

GPT Realtime 2.1 Mini saved everything exactly right on 64.5% of the calls that save data, and Flux → GPT-4.1 → ElevenLabs Flash on 94.1%. On Appointments, 88.8% against 98.8%. On Medicare, 0.0% against 81.8%. A call counts only when nothing is missing, wrong, extra or saved twice.

Which responds faster, GPT Realtime 2.1 Mini or Flux → GPT-4.1 → ElevenLabs Flash?

GPT Realtime 2.1 Mini starts replying 0.09s sooner at the median. GPT Realtime 2.1 Mini takes 2.08s at the median and 2.75s at p90; Flux → GPT-4.1 → ElevenLabs Flash takes 2.17s and 2.51s. Both are measured on the call audio, from the caller finishing to the agent starting to reply.

Which costs less, GPT Realtime 2.1 Mini or Flux → GPT-4.1 → ElevenLabs Flash?

There is no verified per-minute rate for Flux → GPT-4.1 → ElevenLabs Flash, so cost cannot be compared yet.

How were GPT Realtime 2.1 Mini and Flux → GPT-4.1 → ElevenLabs Flash compared?

Both ran as the whole agent in the same open-source Pipecat pipeline, with the same prompt, tools and phone line. A simulated caller worked through all 82 scenarios on live calls, three times each, 2,214 calls across the benchmark. A call passes only when the agent said the right things and saved the right data. The cascade runs the same agent for reference and is never ranked.