All comparisons
OpenAIPhonic

GPT Realtime 2.1 vs Phonic v1

GPT Realtime 2.1 is ahead on reliability, success rate, agent responses, data accuracy, interruption and cost, and Phonic v1 on stalled calls and response time.

Compare

Success rate by suite

Appointments59 scenariosMedicare23 scenarios
  • GPT Realtime 2.1OpenAI
    92%
    81%
  • GPT-Live 1 + Sol lowOpenAI
    97%
    49%
  • Gemini 3.8 Live ThinkingGoogle
    82%
    61%
  • Grok Think Fast 2.0xAI
    96%
    25%
  • Gemini 3.1 Flash LiveGoogle
    93%
    32%
  • Phonic v1Phonic
    97%
    14%
  • Realtime 2.1 MiniOpenAI
    84%
    4%
  • Nova 2 SonicAmazon
    80%
    6%
  • Cascade, GPT-4.1Not ranked
    97%
    83%
Each model's share of calls that passed every check, on each suite.

Every metric, side by side

Reliability

Scenarios that passed on all three runs.

GPT Realtime 2.1
79.3%1st
Phonic v1
67.1%3rd

GPT Realtime 2.1, 12.2 points higher

Success rate

Calls that passed every check.

GPT Realtime 2.1
89.0%1st
Phonic v1
74.0%6th

GPT Realtime 2.1, 15.0 points higher

Agent responses

Calls where the agent's replies reached the expected outcome.

GPT Realtime 2.1
99.2%2nd
Phonic v1
98.0%3rd

GPT Realtime 2.1, 1.2 points higher

Data accuracy

Calls where everything the agent saved matched exactly.

GPT Realtime 2.1
92.0%1st
Phonic v1
74.3%6th

GPT Realtime 2.1, 17.7 points higher

Stalled calls

Calls where the agent went silent for 10 seconds or more.

GPT Realtime 2.1
3.3%4th
Phonic v1
0.0%1st

Phonic v1, 3.3 points lower

Interruption

How well it handled being talked over, 0 to 5.

GPT Realtime 2.1
4.983rd
Phonic v1
4.888th

GPT Realtime 2.1, 0.10 higher

Response time

Median time to start replying after the caller finishes.

GPT Realtime 2.1
1.94s3rd
Phonic v1
1.59s1st

Phonic v1, 0.35s sooner

Cost / min

One minute of call at the vendor's published rates.

GPT Realtime 2.1
$0.0855th
Phonic v1
$0.1406th

GPT Realtime 2.1, $0.055 cheaper

The small figure is the place among ranked models. The cascade is not ranked.

Phonic v1 answers first

Phonic v1 1.59s, GPT Realtime 2.1 1.94s at the median.

Response time, median to p90

MedianTo p90
Medianp90
  • Phonic v1Phonic
    Phonic v1: median 1.59s, p90 1.86s
    1.59s1.86s
  • Grok Think Fast 2.0xAI
    Grok Think Fast 2.0: median 1.62s, p90 1.94s
    1.62s1.94s
  • GPT Realtime 2.1OpenAI
    GPT Realtime 2.1: median 1.94s, p90 2.46s
    1.94s2.46s
  • GPT-Live 1 + Sol lowOpenAI
    GPT-Live 1 + Sol low: median 2.02s, p90 2.42s
    2.02s2.42s
  • Realtime 2.1 MiniOpenAI
    Realtime 2.1 Mini: median 2.08s, p90 2.75s
    2.08s2.75s
  • Nova 2 SonicAmazon
    Nova 2 Sonic: median 2.18s, p90 2.76s
    2.18s2.76s
  • Gemini 3.8 Live ThinkingGoogle
    Gemini 3.8 Live Thinking: median 2.32s, p90 3.06s
    2.32s3.06s
  • Gemini 3.1 Flash LiveGoogle
    Gemini 3.1 Flash Live: median 2.88s, p90 4.00s
    2.88s4.00s
  • Cascade, GPT-4.1Not ranked
    Cascade, GPT-4.1: median 2.17s, p90 2.51s
    2.17s2.51s
2.0s3.0s4.0s
Seconds from the caller finishing to the agent starting to reply, per model.

What we saw on the calls

GPT Realtime 2.1

  • Only a handful of scenarios ever fail, and rarely twice.
  • Best realtime model on the long form, level with the cascade.
  • Few data errors: a phone number or an earlier reference, wrong or left out.
  • Weakest on noisy lines.
Model
gpt-realtime-2.1
Setting
Reasoning high
Voice
marin
Turn taking
Service turn detection
Audio
24 kHz
Cost
Excludes caller transcription, which the service does not meter per call.

Phonic v1

  • Level with the best on the short form, run after run.
  • Quickest to reply: no late reply, no stalled call.
  • Fails the long form's exact check: it fills in optional fields it should leave out, though the record would usually match.
  • Short-form slips are an extra tool call; it repeats calls more than most.
Model
phonic_v1
Setting
Intelligence high
Voice
sabrina
Turn taking
Service turn detection
Audio
16 kHz
Cost
The vendor's Starter plan rate per conversation minute, applied to each call's length; the Pro plan is $0.12.

Frequently asked questions

Is GPT Realtime 2.1 or Phonic v1 more reliable?

GPT Realtime 2.1 passed 65 of 82 scenarios on all three runs, and Phonic v1 passed 55. GPT Realtime 2.1 is ahead, by 12.2 points.

Which saves caller data more accurately, GPT Realtime 2.1 or Phonic v1?

GPT Realtime 2.1 saved everything exactly right on 92.0% of the calls that save data, and Phonic v1 on 74.3%. On Appointments, 95.9% against 98.2%. On Medicare, 81.8% against 12.1%. A call counts only when nothing is missing, wrong, extra or saved twice.

Which responds faster, GPT Realtime 2.1 or Phonic v1?

Phonic v1 starts replying 0.35s sooner at the median. GPT Realtime 2.1 takes 1.94s at the median and 2.46s at p90; Phonic v1 takes 1.59s and 1.86s. Both are measured on the call audio, from the caller finishing to the agent starting to reply.

Which costs less, GPT Realtime 2.1 or Phonic v1?

GPT Realtime 2.1 costs $0.085 a minute of call and Phonic v1 $0.140, at each vendor's published rates. GPT Realtime 2.1 is $0.055 cheaper a minute.

How were GPT Realtime 2.1 and Phonic v1 compared?

Both ran as the whole agent in the same open-source Pipecat pipeline, with the same prompt, tools and phone line. A simulated caller worked through all 82 scenarios on live calls, three times each, 2,214 calls across the benchmark. A call passes only when the agent said the right things and saved the right data.