Speech-to-speech

Realtime voice models running a complete phone agent on live calls. Same prompt, same tools, same line: only the model changes.

GPT Realtime 2.1 is the most reliable, passing 65 of 82 scenarios on all three runs. Phonic v1 answers fastest, at 1.59 seconds.

8
models
82
scenarios
3
runs per scenario
2,214
calls

Faster, or more reliable

Pareto line
Reliability, pass³ (%)
◤ better
Median response time (s)
Reliability against median response time. The dotted line joins the models nothing beats on both: two of them.
Kinds of call

The short form is easy. The long form is not.

5 of 8 models pass 85% of the short form, a clinic booking with a few fields a call. 2 pass half of the long form, a Medicare intake with up to 16 fields.

Short formLong form
  • GPT Realtime 2.186% → 61%
  • GPT-Live 1 + Sol low92% → 35%
  • Gemini 3.8 Live Thinking66% → 30%
  • Gemini 3.1 Flash Live83% → 9%
  • Grok Think Fast 2.092% → 4%
  • Nova 2 Sonic68% → 4%
  • Realtime 2.1 Mini68% → 4%
  • Phonic v193% → 0%
  • Cascade (reference)93% → 57%
0%50%100%
Scenarios that passed on all 3 runs.
Leaderboard

Every model, side by side

Ranked on reliability. A call passes only when the agent said the right things and saved the right data.

  • What we saw
    • Only a handful of scenarios ever fail, and rarely twice.
    • Best realtime model on the long form, level with the cascade.
    • Few data errors: a phone number or an earlier reference, wrong or left out.
    • Weakest on noisy lines.
    More figures
    Scenarios passed every run
    65 of 82
    Agent responses
    99.2%
    Interruption
    4.98 of 5
    Response time, p90
    2.46s
    Appointments success
    92.1%
    Medicare success
    81.2%
    Setup
    Model
    gpt-realtime-2.1
    Setting
    Reasoning high
    Voice
    marin
    Turn taking
    Service turn detection
  • What we saw
    • Near perfect on the short form, run after run.
    • Every failure is in the saved record; the conversation itself always reaches the outcome.
    • On the long form it routes the caller wrong or skips a step.
    • Says goodbye without hanging up, and stalls more than any other model.
    • Lowest voice score of the field.
    More figures
    Scenarios passed every run
    62 of 82
    Agent responses
    100.0%
    Interruption
    4.89 of 5
    Response time, p90
    2.42s
    Appointments success
    97.2%
    Medicare success
    49.3%
    Setup
    Model
    gpt-live-1
    Setting
    Delegates reasoning and tools to Sol
    Voice
    marin
    Turn taking
    Service turn detection
  • What we saw
    • Quick, with the shortest replies of the field; steady on the short form.
    • Struggles on the long form: fields left out or wrong, usually an earlier reference.
    • Every stall was on a long form.
    More figures
    Scenarios passed every run
    55 of 82
    Agent responses
    96.7%
    Interruption
    4.92 of 5
    Response time, p90
    1.94s
    Appointments success
    96.0%
    Medicare success
    24.6%
    Setup
    Model
    grok-voice-think-fast-2.0
    Setting
    Reasoning high
    Voice
    eve
    Turn taking
    Service turn detection
  • What we saw
    • Level with the best on the short form, run after run.
    • Quickest to reply: no late reply, no stalled call.
    • Fails the long form's exact check: it fills in optional fields it should leave out, though the record would usually match.
    • Short-form slips are an extra tool call; it repeats calls more than most.
    More figures
    Scenarios passed every run
    55 of 82
    Agent responses
    98.0%
    Interruption
    4.88 of 5
    Response time, p90
    1.86s
    Appointments success
    97.2%
    Medicare success
    14.5%
    Setup
    Model
    phonic_v1
    Setting
    Intelligence high
    Voice
    sabrina
    Turn taking
    Service turn detection
  • What we saw
    • Steady on the short form, though it repeats lookups it already made.
    • On the long form it hands the caller over with the wrong details.
    • Slowest to reply; after several tool calls its replies fall behind or go missing.
    • Sometimes says goodbye without hanging up.
    More figures
    Scenarios passed every run
    51 of 82
    Agent responses
    95.9%
    Interruption
    4.98 of 5
    Response time, p90
    4.00s
    Appointments success
    93.2%
    Medicare success
    31.9%
    Setup
    Model
    models/gemini-3.1-flash-live-preview
    Setting
    Thinking level high
    Voice
    Charon
    Turn taking
    Shared local turn detection
  • What we saw
    • Strong on the long form; when it slips, it skips a save it said it made.
    • Weak on the short form: phone numbers saved wrong, lookups never made.
    • Acts on half an answer instead of asking; weakest when the caller corrects themselves.
    • Least predictable: the same scenario passes one run and fails the next.
    • Replies fall behind or stop; callers hang up waiting.
    More figures
    Scenarios passed every run
    46 of 82
    Agent responses
    94.7%
    Interruption
    4.90 of 5
    Response time, p90
    3.06s
    Appointments success
    82.5%
    Medicare success
    60.9%
    Setup
    Model
    gemini-3.8-live-extended-thinking
    Setting
    Thinking level high
    Voice
    Charon
    Turn taking
    Shared local turn detection
  • What we saw
    • Misses caller names and phone numbers: blank or wrong.
    • Long forms rarely pass; more scenarios fail every run than for any other model.
    • Repeats the same tool call more than any other model.
    • Best voice score of the field, and the longest replies.
    More figures
    Scenarios passed every run
    41 of 82
    Agent responses
    93.9%
    Interruption
    4.99 of 5
    Response time, p90
    2.76s
    Appointments success
    80.2%
    Medicare success
    5.8%
    Setup
    Model
    amazon.nova-2-sonic-v1:0
    Setting
    Endpointing medium
    Voice
    matthew
    Turn taking
    Shared local turn detection
  • What we saw
    • Gets most short-form calls right, at a fraction of the cost.
    • Sometimes mishears a phone number.
    • Drops details given earlier: not one long-form record was exact.
    • Rarely recovers when a tool returns an error.
    More figures
    Scenarios passed every run
    41 of 82
    Agent responses
    93.5%
    Interruption
    4.99 of 5
    Response time, p90
    2.75s
    Appointments success
    83.6%
    Medicare success
    4.3%
    Setup
    Model
    gpt-realtime-2.1-mini
    Setting
    Reasoning high
    Voice
    marin
    Turn taking
    Service turn detection
  • Reference · not ranked
  • What we saw
    • Most scenarios pass every run; none fails every run.
    • Rare long-form slips: a reference from an earlier step left out.
    • Not the quickest to reply, but almost never leaves the caller waiting.
    More figures
    Scenarios passed every run
    68 of 82
    Agent responses
    99.2%
    Interruption
    4.96 of 5
    Response time, p90
    2.51s
    Appointments success
    97.2%
    Medicare success
    82.6%
    Setup
    Model
    flux-general-en → gpt-4.1 → eleven_flash_v2_5
    Setting
    No reasoning setting
    Voice
    ElevenLabs 21m00Tcm4TlvDq8ikWAM
    Turn taking
    Speech-to-text turn detection
Reliability
Scenarios that passed on all 3 runs (pass³).
Success rate
Calls that passed every check: what the agent said, what it saved and how the call went.
Data accuracy
Calls where everything the agent saved matched exactly. Calls that save nothing are left out.
Stalled calls
Calls where the agent went silent for 10 seconds or more after the caller spoke, at least once. Checked on the recording.
Response time
Median time from the caller finishing to the agent starting to reply, on the call audio.
Cost / min
One minute of call at the vendor's published rates. Blank where no rate is verified.
Rankings

One metric at a time

Scenarios that passed on all 3 runs, pass³. Higher is better.

  • 1GPT Realtime 2.179.3%
  • 2GPT-Live 1 + Sol low75.6%
  • 3Grok Think Fast 2.067.1%
  • 4Phonic v167.1%
  • 5Gemini 3.1 Flash Live62.2%
  • 6Gemini 3.8 Live Thinking56.1%
  • 7Nova 2 Sonic50.0%
  • 8Realtime 2.1 Mini50.0%
  • —Cascade, GPT-4.182.9%
Differences

How each model fails

Two models can share a score and fail in different ways.

Kind of scenarioGPT Realtime 2.1GPT-Live 1 + Sol lowGrok Think Fast 2.0Phonic v1Gemini 3.1 Flash LiveGemini 3.8 Live ThinkingNova 2 SonicRealtime 2.1 MiniCascade
Plain booking & lookup15
98%
100%
100%
100%
100%
89%
89%
89%
98%
Several requests in one call7
100%
95%
95%
100%
100%
95%
81%
90%
100%
Caller corrects or changes mind8
83%
88%
71%
83%
83%
50%
63%
79%
88%
Background noise7
76%
100%
100%
100%
86%
86%
81%
81%
95%
Interruptions & barge-in5
100%
100%
80%
93%
87%
93%
100%
87%
100%
Silence, cough & packet loss8
88%
96%
100%
92%
83%
75%
63%
88%
100%
Accent & slow speech2
100%
83%
100%
83%
83%
50%
50%
83%
100%
Tool or system errors4
100%
92%
92%
100%
92%
83%
92%
25%
100%
Safety & scope6
94%
83%
89%
78%
83%
78%
67%
78%
100%
Medicare compliance & routing20
78%
50%
25%
13%
32%
65%
2%
0%
80%

Success rate per kind of scenario. Deeper blue is better; the small number is how many scenarios it holds.

Models

What we saw on each model's calls

  • GPT Realtime 2.1

    OpenAI · Reasoning high · service turn detection

    Reliability
    79%
    Response
    1.94s
    Per minute
    $0.085
    • Only a handful of scenarios ever fail, and rarely twice.
    • Best realtime model on the long form, level with the cascade.
    • Few data errors: a phone number or an earlier reference, wrong or left out.
    • Weakest on noisy lines.

    Most often wrong: route_idphonedestination

    Calls:All Short form Long form

  • GPT-Live 1 + GPT-6 Sol (low thinking)

    OpenAI · Delegates reasoning and tools to Sol · service turn detection

    Reliability
    76%
    Response
    2.02s
    Per minute
    $0.055
    • Near perfect on the short form, run after run.
    • Every failure is in the saved record; the conversation itself always reaches the outcome.
    • On the long form it routes the caller wrong or skips a step.
    • Says goodbye without hanging up, and stalls more than any other model.
    • Lowest voice score of the field.

    Most often wrong: destinationroute_statusstate

    Calls:All Short form Long form

  • Grok Voice Think Fast 2.0

    xAI · Reasoning high · service turn detection

    Reliability
    67%
    Response
    1.62s
    Per minute
    $0.080
    • Quick, with the shortest replies of the field; steady on the short form.
    • Struggles on the long form: fields left out or wrong, usually an earlier reference.
    • Every stall was on a long form.

    Most often wrong: lead_idroute_idconsent_id

    Calls:All Short form Long form

  • Phonic v1

    Phonic · Intelligence high · service turn detection

    Reliability
    67%
    Response
    1.59s
    Per minute
    $0.140
    • Level with the best on the short form, run after run.
    • Quickest to reply: no late reply, no stalled call.
    • Fails the long form's exact check: it fills in optional fields it should leave out, though the record would usually match.
    • Short-form slips are an extra tool call; it repeats calls more than most.

    Most often wrong: beneficiary_namelead_idstate

    Calls:All Short form Long form

  • Gemini 3.1 Flash Live Preview

    Google · Thinking level high · shared local turn detection

    Reliability
    62%
    Response
    2.88s
    Per minute
    $0.070
    • Steady on the short form, though it repeats lookups it already made.
    • On the long form it hands the caller over with the wrong details.
    • Slowest to reply; after several tool calls its replies fall behind or go missing.
    • Sometimes says goodbye without hanging up.

    Most often wrong: lead_iddestinationroute_id

    Calls:All Short form Long form

  • Gemini 3.8 Live Extended Thinking

    Google · Thinking level high · shared local turn detection

    Reliability
    56%
    Response
    2.32s
    Per minute
    —
    • Strong on the long form; when it slips, it skips a save it said it made.
    • Weak on the short form: phone numbers saved wrong, lookups never made.
    • Acts on half an answer instead of asking; weakest when the caller corrects themselves.
    • Least predictable: the same scenario passes one run and fails the next.
    • Replies fall behind or stop; callers hang up waiting.

    Most often wrong: phonecaller_namestate

    Calls:All Short form Long form

  • Amazon Nova 2 Sonic

    Amazon · Endpointing medium · shared local turn detection

    Reliability
    50%
    Response
    2.18s
    Per minute
    —
    • Misses caller names and phone numbers: blank or wrong.
    • Long forms rarely pass; more scenarios fail every run than for any other model.
    • Repeats the same tool call more than any other model.
    • Best voice score of the field, and the longest replies.

    Most often wrong: caller_nameconsent_idcallback_phone

    Calls:All Short form Long form

  • GPT Realtime 2.1 Mini

    OpenAI · Reasoning high · service turn detection

    Reliability
    50%
    Response
    2.08s
    Per minute
    $0.020
    • Gets most short-form calls right, at a fraction of the cost.
    • Sometimes mishears a phone number.
    • Drops details given earlier: not one long-form record was exact.
    • Rarely recovers when a tool returns an error.

    Most often wrong: route_idzip_codeconsent_id

    Calls:All Short form Long form

  • Flux → GPT-4.1 → ElevenLabs Flash

    Cascade baseline, not ranked

    Reliability
    83%
    Response
    2.17s
    Per minute
    —
    • Most scenarios pass every run; none fails every run.
    • Rare long-form slips: a reference from an earlier step left out.
    • Not the quickest to reply, but almost never leaves the caller waiting.

    Most often wrong: route_idintake_intentservice_issue_type

    Calls:All Short form Long form

Compare two models

Compare
GPT Realtime 2.1OpenAI
Flux → GPT-4.1 → ElevenLabs FlashBaseline, not ranked
Reliability
79.3%1st
82.9%
3.7 points higher
Data accuracy
92.0%1st
94.1%
2.1 points higher
Response time
1.94s3rd
0.23s sooner
2.17s
Cost / min
$0.0855th
No verified rate
All suites. The small figure is the place among ranked models.All 36 comparisons
Methodology

How the calls were run

  1. 01

    Only the model changes

    One open-source Pipecat agent: the realtime model hears the caller, calls the tools and speaks. Prompt, tools and phone line are the same for every model. A cascade (Flux → GPT-4.1 → ElevenLabs Flash) runs the same agent for reference, unranked.

  2. 02

    Live calls, three runs each

    A simulated caller works through all 82 scenarios on live calls, three times over. Every call counts, including ones that stalled or ended early.

  3. 03

    Words and data both count

    A call passes only when the agent reached the expected outcome and every field it saved matched, checked on the Cekura platform.

  4. 04

    Every call is public

    Each row links to the Cekura report of its calls, with transcripts, tool calls and scores, so any figure here can be checked.

  • Appointments

    the short form · 59 scenarios · 6 fields · 4 tools

    A patient calling a clinic. Find the patient, check openings, then book, move or cancel a visit.

    • lookup_patient
    • check_availability
    • book_appointment
    • cancel_appointment
  • Medicare

    the long form · 23 scenarios · 42 fields · 4 tools

    Someone calling a Medicare insurance line. Record consent first, qualify the caller, route them to the right team and write a handoff note.

    • record_medicare_permissions
    • save_medicare_qualification
    • route_medicare_call
    • create_handoff_summary