All comparisons
OpenAIxAI

GPT-Live 1 + GPT-6 Sol (low thinking) vs Grok Voice Think Fast 2.0

GPT-Live 1 + GPT-6 Sol (low thinking) is ahead on reliability, success rate, agent responses, data accuracy and cost, and Grok Voice Think Fast 2.0 on stalled calls, interruption and response time.

Compare

Success rate by suite

Appointments59 scenariosMedicare23 scenarios
  • GPT Realtime 2.1OpenAI
    92%
    81%
  • GPT-Live 1 + Sol lowOpenAI
    97%
    49%
  • Gemini 3.8 Live ThinkingGoogle
    82%
    61%
  • Grok Think Fast 2.0xAI
    96%
    25%
  • Gemini 3.1 Flash LiveGoogle
    93%
    32%
  • Phonic v1Phonic
    97%
    14%
  • Realtime 2.1 MiniOpenAI
    84%
    4%
  • Nova 2 SonicAmazon
    80%
    6%
  • Cascade, GPT-4.1Not ranked
    97%
    83%
Each model's share of calls that passed every check, on each suite.

Every metric, side by side

Reliability

Scenarios that passed on all three runs.

GPT-Live 1 + Sol low
75.6%2nd
Grok Think Fast 2.0
67.1%3rd

GPT-Live 1 + Sol low, 8.5 points higher

Success rate

Calls that passed every check.

GPT-Live 1 + Sol low
83.7%2nd
Grok Think Fast 2.0
76.0%4th

GPT-Live 1 + Sol low, 7.7 points higher

Agent responses

Calls where the agent's replies reached the expected outcome.

GPT-Live 1 + Sol low
100.0%1st
Grok Think Fast 2.0
96.7%4th

GPT-Live 1 + Sol low, 3.3 points higher

Data accuracy

Calls where everything the agent saved matched exactly.

GPT-Live 1 + Sol low
89.0%2nd
Grok Think Fast 2.0
78.9%5th

GPT-Live 1 + Sol low, 10.1 points higher

Stalled calls

Calls where the agent went silent for 10 seconds or more.

GPT-Live 1 + Sol low
12.2%8th
Grok Think Fast 2.0
4.5%7th

Grok Think Fast 2.0, 7.7 points lower

Interruption

How well it handled being talked over, 0 to 5.

GPT-Live 1 + Sol low
4.897th
Grok Think Fast 2.0
4.925th

Grok Think Fast 2.0, 0.03 higher

Response time

Median time to start replying after the caller finishes.

GPT-Live 1 + Sol low
2.02s4th
Grok Think Fast 2.0
1.62s2nd

Grok Think Fast 2.0, 0.40s sooner

Cost / min

One minute of call at the vendor's published rates.

GPT-Live 1 + Sol low
$0.0552nd
Grok Think Fast 2.0
$0.0804th

GPT-Live 1 + Sol low, $0.025 cheaper

The small figure is the place among ranked models. The cascade is not ranked.

Grok Think Fast 2.0 answers first

Grok Think Fast 2.0 1.62s, GPT-Live 1 + Sol low 2.02s at the median.

Response time, median to p90

MedianTo p90
Medianp90
  • Phonic v1Phonic
    Phonic v1: median 1.59s, p90 1.86s
    1.59s1.86s
  • Grok Think Fast 2.0xAI
    Grok Think Fast 2.0: median 1.62s, p90 1.94s
    1.62s1.94s
  • GPT Realtime 2.1OpenAI
    GPT Realtime 2.1: median 1.94s, p90 2.46s
    1.94s2.46s
  • GPT-Live 1 + Sol lowOpenAI
    GPT-Live 1 + Sol low: median 2.02s, p90 2.42s
    2.02s2.42s
  • Realtime 2.1 MiniOpenAI
    Realtime 2.1 Mini: median 2.08s, p90 2.75s
    2.08s2.75s
  • Nova 2 SonicAmazon
    Nova 2 Sonic: median 2.18s, p90 2.76s
    2.18s2.76s
  • Gemini 3.8 Live ThinkingGoogle
    Gemini 3.8 Live Thinking: median 2.32s, p90 3.06s
    2.32s3.06s
  • Gemini 3.1 Flash LiveGoogle
    Gemini 3.1 Flash Live: median 2.88s, p90 4.00s
    2.88s4.00s
  • Cascade, GPT-4.1Not ranked
    Cascade, GPT-4.1: median 2.17s, p90 2.51s
    2.17s2.51s
2.0s3.0s4.0s
Seconds from the caller finishing to the agent starting to reply, per model.

What we saw on the calls

GPT-Live 1 + GPT-6 Sol (low thinking)

  • Near perfect on the short form, run after run.
  • Every failure is in the saved record; the conversation itself always reaches the outcome.
  • On the long form it routes the caller wrong or skips a step.
  • Says goodbye without hanging up, and stalls more than any other model.
  • Lowest voice score of the field.
Model
gpt-live-1
Setting
Delegates reasoning and tools to Sol
Voice
marin
Turn taking
Service turn detection
Audio
24 kHz
Cost
Includes the delegated backend model.

Grok Voice Think Fast 2.0

  • Quick, with the shortest replies of the field; steady on the short form.
  • Struggles on the long form: fields left out or wrong, usually an earlier reference.
  • Every stall was on a long form.
Model
grok-voice-think-fast-2.0
Setting
Reasoning high
Voice
eve
Turn taking
Service turn detection
Audio
24 kHz
Cost
Excludes a text-input charge the vendor lists without a unit.

Frequently asked questions

Is GPT-Live 1 + GPT-6 Sol (low thinking) or Grok Voice Think Fast 2.0 more reliable?

GPT-Live 1 + GPT-6 Sol (low thinking) passed 62 of 82 scenarios on all three runs, and Grok Voice Think Fast 2.0 passed 55. GPT-Live 1 + GPT-6 Sol (low thinking) is ahead, by 8.5 points.

Which saves caller data more accurately, GPT-Live 1 + GPT-6 Sol (low thinking) or Grok Voice Think Fast 2.0?

GPT-Live 1 + GPT-6 Sol (low thinking) saved everything exactly right on 89.0% of the calls that save data, and Grok Voice Think Fast 2.0 on 78.9%. On Appointments, 100.0% against 97.1%. On Medicare, 60.6% against 31.8%. A call counts only when nothing is missing, wrong, extra or saved twice.

Which responds faster, GPT-Live 1 + GPT-6 Sol (low thinking) or Grok Voice Think Fast 2.0?

Grok Voice Think Fast 2.0 starts replying 0.40s sooner at the median. GPT-Live 1 + GPT-6 Sol (low thinking) takes 2.02s at the median and 2.42s at p90; Grok Voice Think Fast 2.0 takes 1.62s and 1.94s. Both are measured on the call audio, from the caller finishing to the agent starting to reply.

Which costs less, GPT-Live 1 + GPT-6 Sol (low thinking) or Grok Voice Think Fast 2.0?

GPT-Live 1 + GPT-6 Sol (low thinking) costs $0.055 a minute of call and Grok Voice Think Fast 2.0 $0.080, at each vendor's published rates. GPT-Live 1 + GPT-6 Sol (low thinking) is $0.025 cheaper a minute.

How were GPT-Live 1 + GPT-6 Sol (low thinking) and Grok Voice Think Fast 2.0 compared?

Both ran as the whole agent in the same open-source Pipecat pipeline, with the same prompt, tools and phone line. A simulated caller worked through all 82 scenarios on live calls, three times each, 2,214 calls across the benchmark. A call passes only when the agent said the right things and saved the right data.