All comparisons
OpenAICascade baseline

GPT-Live 1 + GPT-6 Sol (low thinking) vs Flux → GPT-4.1 → ElevenLabs Flash

Flux → GPT-4.1 → ElevenLabs Flash is ahead on reliability, success rate, data accuracy, stalled calls and interruption, and GPT-Live 1 + GPT-6 Sol (low thinking) on agent responses and response time.

Compare

Success rate by suite

Appointments59 scenariosMedicare23 scenarios
  • GPT Realtime 2.1OpenAI
    92%
    81%
  • GPT-Live 1 + Sol lowOpenAI
    97%
    49%
  • Gemini 3.8 Live ThinkingGoogle
    82%
    61%
  • Grok Think Fast 2.0xAI
    96%
    25%
  • Gemini 3.1 Flash LiveGoogle
    93%
    32%
  • Phonic v1Phonic
    97%
    14%
  • Realtime 2.1 MiniOpenAI
    84%
    4%
  • Nova 2 SonicAmazon
    80%
    6%
  • Cascade, GPT-4.1Not ranked
    97%
    83%
Each model's share of calls that passed every check, on each suite.

Every metric, side by side

Reliability

Scenarios that passed on all three runs.

GPT-Live 1 + Sol low
75.6%2nd
Cascade, GPT-4.1
82.9%

Cascade, GPT-4.1, 7.3 points higher

Success rate

Calls that passed every check.

GPT-Live 1 + Sol low
83.7%2nd
Cascade, GPT-4.1
93.1%

Cascade, GPT-4.1, 9.3 points higher

Agent responses

Calls where the agent's replies reached the expected outcome.

GPT-Live 1 + Sol low
100.0%1st
Cascade, GPT-4.1
99.2%

GPT-Live 1 + Sol low, 0.8 points higher

Data accuracy

Calls where everything the agent saved matched exactly.

GPT-Live 1 + Sol low
89.0%2nd
Cascade, GPT-4.1
94.1%

Cascade, GPT-4.1, 5.1 points higher

Stalled calls

Calls where the agent went silent for 10 seconds or more.

GPT-Live 1 + Sol low
12.2%8th
Cascade, GPT-4.1
0.8%

Cascade, GPT-4.1, 11.4 points lower

Interruption

How well it handled being talked over, 0 to 5.

GPT-Live 1 + Sol low
4.897th
Cascade, GPT-4.1
4.96

Cascade, GPT-4.1, 0.07 higher

Response time

Median time to start replying after the caller finishes.

GPT-Live 1 + Sol low
2.02s4th
Cascade, GPT-4.1
2.17s

GPT-Live 1 + Sol low, 0.15s sooner

Cost / min

One minute of call at the vendor's published rates.

GPT-Live 1 + Sol low
$0.0552nd
Cascade, GPT-4.1
No verified rate

Not comparable

The small figure is the place among ranked models. The cascade is not ranked.

GPT-Live 1 + Sol low answers first

GPT-Live 1 + Sol low 2.02s, Cascade, GPT-4.1 2.17s at the median.

Response time, median to p90

MedianTo p90
Medianp90
  • Phonic v1Phonic
    Phonic v1: median 1.59s, p90 1.86s
    1.59s1.86s
  • Grok Think Fast 2.0xAI
    Grok Think Fast 2.0: median 1.62s, p90 1.94s
    1.62s1.94s
  • GPT Realtime 2.1OpenAI
    GPT Realtime 2.1: median 1.94s, p90 2.46s
    1.94s2.46s
  • GPT-Live 1 + Sol lowOpenAI
    GPT-Live 1 + Sol low: median 2.02s, p90 2.42s
    2.02s2.42s
  • Realtime 2.1 MiniOpenAI
    Realtime 2.1 Mini: median 2.08s, p90 2.75s
    2.08s2.75s
  • Nova 2 SonicAmazon
    Nova 2 Sonic: median 2.18s, p90 2.76s
    2.18s2.76s
  • Gemini 3.8 Live ThinkingGoogle
    Gemini 3.8 Live Thinking: median 2.32s, p90 3.06s
    2.32s3.06s
  • Gemini 3.1 Flash LiveGoogle
    Gemini 3.1 Flash Live: median 2.88s, p90 4.00s
    2.88s4.00s
  • Cascade, GPT-4.1Not ranked
    Cascade, GPT-4.1: median 2.17s, p90 2.51s
    2.17s2.51s
2.0s3.0s4.0s
Seconds from the caller finishing to the agent starting to reply, per model.

What we saw on the calls

GPT-Live 1 + GPT-6 Sol (low thinking)

  • Near perfect on the short form, run after run.
  • Every failure is in the saved record; the conversation itself always reaches the outcome.
  • On the long form it routes the caller wrong or skips a step.
  • Says goodbye without hanging up, and stalls more than any other model.
  • Lowest voice score of the field.
Model
gpt-live-1
Setting
Delegates reasoning and tools to Sol
Voice
marin
Turn taking
Service turn detection
Audio
24 kHz
Cost
Includes the delegated backend model.

Flux → GPT-4.1 → ElevenLabs Flash

  • Most scenarios pass every run; none fails every run.
  • Rare long-form slips: a reference from an earlier step left out.
  • Not the quickest to reply, but almost never leaves the caller waiting.
Model
flux-general-en → gpt-4.1 → eleven_flash_v2_5
Setting
No reasoning setting
Voice
ElevenLabs 21m00Tcm4TlvDq8ikWAM
Turn taking
Speech-to-text turn detection
Audio
16 kHz
Cost
Not published: no verified combined rate for the three services.

Frequently asked questions

Is GPT-Live 1 + GPT-6 Sol (low thinking) or Flux → GPT-4.1 → ElevenLabs Flash more reliable?

GPT-Live 1 + GPT-6 Sol (low thinking) passed 62 of 82 scenarios on all three runs, and Flux → GPT-4.1 → ElevenLabs Flash passed 68. Flux → GPT-4.1 → ElevenLabs Flash is ahead, by 7.3 points.

Which saves caller data more accurately, GPT-Live 1 + GPT-6 Sol (low thinking) or Flux → GPT-4.1 → ElevenLabs Flash?

GPT-Live 1 + GPT-6 Sol (low thinking) saved everything exactly right on 89.0% of the calls that save data, and Flux → GPT-4.1 → ElevenLabs Flash on 94.1%. On Appointments, 100.0% against 98.8%. On Medicare, 60.6% against 81.8%. A call counts only when nothing is missing, wrong, extra or saved twice.

Which responds faster, GPT-Live 1 + GPT-6 Sol (low thinking) or Flux → GPT-4.1 → ElevenLabs Flash?

GPT-Live 1 + GPT-6 Sol (low thinking) starts replying 0.15s sooner at the median. GPT-Live 1 + GPT-6 Sol (low thinking) takes 2.02s at the median and 2.42s at p90; Flux → GPT-4.1 → ElevenLabs Flash takes 2.17s and 2.51s. Both are measured on the call audio, from the caller finishing to the agent starting to reply.

Which costs less, GPT-Live 1 + GPT-6 Sol (low thinking) or Flux → GPT-4.1 → ElevenLabs Flash?

There is no verified per-minute rate for Flux → GPT-4.1 → ElevenLabs Flash, so cost cannot be compared yet.

How were GPT-Live 1 + GPT-6 Sol (low thinking) and Flux → GPT-4.1 → ElevenLabs Flash compared?

Both ran as the whole agent in the same open-source Pipecat pipeline, with the same prompt, tools and phone line. A simulated caller worked through all 82 scenarios on live calls, three times each, 2,214 calls across the benchmark. A call passes only when the agent said the right things and saved the right data. The cascade runs the same agent for reference and is never ranked.