Speech-to-speech
Realtime voice models running a complete phone agent on live calls. Same prompt, same tools, same line: only the model changes.
GPT Realtime 2.1 is the most reliable, passing 65 of 82 scenarios on all three runs. Phonic v1 answers fastest, at 1.59 seconds.
- 8
- models
- 82
- scenarios
- 3
- runs per scenario
- 2,214
- calls
Faster, or more reliable
The short form is easy. The long form is not.
5 of 8 models pass 85% of the short form, a clinic booking with a few fields a call. 2 pass half of the long form, a Medicare intake with up to 16 fields.
- GPT Realtime 2.186% → 61%
- GPT-Live 1 + Sol low92% → 35%
- Gemini 3.8 Live Thinking66% → 30%
- Gemini 3.1 Flash Live83% → 9%
- Grok Think Fast 2.092% → 4%
- Nova 2 Sonic68% → 4%
- Realtime 2.1 Mini68% → 4%
- Phonic v193% → 0%
- Cascade (reference)93% → 57%
Every model, side by side
Ranked on reliability. A call passes only when the agent said the right things and saved the right data.
- What we saw
- Only a handful of scenarios ever fail, and rarely twice.
- Best realtime model on the long form, level with the cascade.
- Few data errors: a phone number or an earlier reference, wrong or left out.
- Weakest on noisy lines.
More figures- Scenarios passed every run
- 65 of 82
- Agent responses
- 99.2%
- Interruption
- 4.98 of 5
- Response time, p90
- 2.46s
- Appointments success
- 92.1%
- Medicare success
- 81.2%
Setup- Model
- gpt-realtime-2.1
- Setting
- Reasoning high
- Voice
- marin
- Turn taking
- Service turn detection
- What we saw
- Near perfect on the short form, run after run.
- Every failure is in the saved record; the conversation itself always reaches the outcome.
- On the long form it routes the caller wrong or skips a step.
- Says goodbye without hanging up, and stalls more than any other model.
- Lowest voice score of the field.
More figures- Scenarios passed every run
- 62 of 82
- Agent responses
- 100.0%
- Interruption
- 4.89 of 5
- Response time, p90
- 2.42s
- Appointments success
- 97.2%
- Medicare success
- 49.3%
Setup- Model
- gpt-live-1
- Setting
- Delegates reasoning and tools to Sol
- Voice
- marin
- Turn taking
- Service turn detection
- What we saw
- Quick, with the shortest replies of the field; steady on the short form.
- Struggles on the long form: fields left out or wrong, usually an earlier reference.
- Every stall was on a long form.
More figures- Scenarios passed every run
- 55 of 82
- Agent responses
- 96.7%
- Interruption
- 4.92 of 5
- Response time, p90
- 1.94s
- Appointments success
- 96.0%
- Medicare success
- 24.6%
Setup- Model
- grok-voice-think-fast-2.0
- Setting
- Reasoning high
- Voice
- eve
- Turn taking
- Service turn detection
- What we saw
- Level with the best on the short form, run after run.
- Quickest to reply: no late reply, no stalled call.
- Fails the long form's exact check: it fills in optional fields it should leave out, though the record would usually match.
- Short-form slips are an extra tool call; it repeats calls more than most.
More figures- Scenarios passed every run
- 55 of 82
- Agent responses
- 98.0%
- Interruption
- 4.88 of 5
- Response time, p90
- 1.86s
- Appointments success
- 97.2%
- Medicare success
- 14.5%
Setup- Model
- phonic_v1
- Setting
- Intelligence high
- Voice
- sabrina
- Turn taking
- Service turn detection
- What we saw
- Steady on the short form, though it repeats lookups it already made.
- On the long form it hands the caller over with the wrong details.
- Slowest to reply; after several tool calls its replies fall behind or go missing.
- Sometimes says goodbye without hanging up.
More figures- Scenarios passed every run
- 51 of 82
- Agent responses
- 95.9%
- Interruption
- 4.98 of 5
- Response time, p90
- 4.00s
- Appointments success
- 93.2%
- Medicare success
- 31.9%
Setup- Model
- models/gemini-3.1-flash-live-preview
- Setting
- Thinking level high
- Voice
- Charon
- Turn taking
- Shared local turn detection
- What we saw
- Strong on the long form; when it slips, it skips a save it said it made.
- Weak on the short form: phone numbers saved wrong, lookups never made.
- Acts on half an answer instead of asking; weakest when the caller corrects themselves.
- Least predictable: the same scenario passes one run and fails the next.
- Replies fall behind or stop; callers hang up waiting.
More figures- Scenarios passed every run
- 46 of 82
- Agent responses
- 94.7%
- Interruption
- 4.90 of 5
- Response time, p90
- 3.06s
- Appointments success
- 82.5%
- Medicare success
- 60.9%
Setup- Model
- gemini-3.8-live-extended-thinking
- Setting
- Thinking level high
- Voice
- Charon
- Turn taking
- Shared local turn detection
- What we saw
- Misses caller names and phone numbers: blank or wrong.
- Long forms rarely pass; more scenarios fail every run than for any other model.
- Repeats the same tool call more than any other model.
- Best voice score of the field, and the longest replies.
More figures- Scenarios passed every run
- 41 of 82
- Agent responses
- 93.9%
- Interruption
- 4.99 of 5
- Response time, p90
- 2.76s
- Appointments success
- 80.2%
- Medicare success
- 5.8%
Setup- Model
- amazon.nova-2-sonic-v1:0
- Setting
- Endpointing medium
- Voice
- matthew
- Turn taking
- Shared local turn detection
- What we saw
- Gets most short-form calls right, at a fraction of the cost.
- Sometimes mishears a phone number.
- Drops details given earlier: not one long-form record was exact.
- Rarely recovers when a tool returns an error.
More figures- Scenarios passed every run
- 41 of 82
- Agent responses
- 93.5%
- Interruption
- 4.99 of 5
- Response time, p90
- 2.75s
- Appointments success
- 83.6%
- Medicare success
- 4.3%
Setup- Model
- gpt-realtime-2.1-mini
- Setting
- Reasoning high
- Voice
- marin
- Turn taking
- Service turn detection
- Reference · not ranked
- What we saw
- Most scenarios pass every run; none fails every run.
- Rare long-form slips: a reference from an earlier step left out.
- Not the quickest to reply, but almost never leaves the caller waiting.
More figures- Scenarios passed every run
- 68 of 82
- Agent responses
- 99.2%
- Interruption
- 4.96 of 5
- Response time, p90
- 2.51s
- Appointments success
- 97.2%
- Medicare success
- 82.6%
Setup- Model
- flux-general-en → gpt-4.1 → eleven_flash_v2_5
- Setting
- No reasoning setting
- Voice
- ElevenLabs 21m00Tcm4TlvDq8ikWAM
- Turn taking
- Speech-to-text turn detection
- Reliability
- Scenarios that passed on all 3 runs (pass³).
- Success rate
- Calls that passed every check: what the agent said, what it saved and how the call went.
- Data accuracy
- Calls where everything the agent saved matched exactly. Calls that save nothing are left out.
- Stalled calls
- Calls where the agent went silent for 10 seconds or more after the caller spoke, at least once. Checked on the recording.
- Response time
- Median time from the caller finishing to the agent starting to reply, on the call audio.
- Cost / min
- One minute of call at the vendor's published rates. Blank where no rate is verified.
One metric at a time
Scenarios that passed on all 3 runs, pass³. Higher is better.
- 1GPT Realtime 2.1Reasoning high79.3%
- 2GPT-Live 1 + Sol lowDelegates reasoning and tools to Sol75.6%
- 3Grok Think Fast 2.0Reasoning high67.1%
- 4Phonic v1Intelligence high67.1%
- 5Gemini 3.1 Flash LiveThinking level high62.2%
- 6Gemini 3.8 Live ThinkingThinking level high56.1%
- 7Nova 2 SonicEndpointing medium50.0%
- 8Realtime 2.1 MiniReasoning high50.0%
- —Cascade, GPT-4.1Cascade baseline, not ranked82.9%
How each model fails
Two models can share a score and fail in different ways.
| Kind of scenario | GPT Realtime 2.1 | GPT-Live 1 + Sol low | Grok Think Fast 2.0 | Phonic v1 | Gemini 3.1 Flash Live | Gemini 3.8 Live Thinking | Nova 2 Sonic | Realtime 2.1 Mini | Cascade |
|---|---|---|---|---|---|---|---|---|---|
| Plain booking & lookup15 | 98% | 100% | 100% | 100% | 100% | 89% | 89% | 89% | 98% |
| Several requests in one call7 | 100% | 95% | 95% | 100% | 100% | 95% | 81% | 90% | 100% |
| Caller corrects or changes mind8 | 83% | 88% | 71% | 83% | 83% | 50% | 63% | 79% | 88% |
| Background noise7 | 76% | 100% | 100% | 100% | 86% | 86% | 81% | 81% | 95% |
| Interruptions & barge-in5 | 100% | 100% | 80% | 93% | 87% | 93% | 100% | 87% | 100% |
| Silence, cough & packet loss8 | 88% | 96% | 100% | 92% | 83% | 75% | 63% | 88% | 100% |
| Accent & slow speech2 | 100% | 83% | 100% | 83% | 83% | 50% | 50% | 83% | 100% |
| Tool or system errors4 | 100% | 92% | 92% | 100% | 92% | 83% | 92% | 25% | 100% |
| Safety & scope6 | 94% | 83% | 89% | 78% | 83% | 78% | 67% | 78% | 100% |
| Medicare compliance & routing20 | 78% | 50% | 25% | 13% | 32% | 65% | 2% | 0% | 80% |
Success rate per kind of scenario. Deeper blue is better; the small number is how many scenarios it holds.
What we saw on each model's calls
GPT Realtime 2.1
OpenAI · Reasoning high · service turn detection
- Reliability
- 79%
- Response
- 1.94s
- Per minute
- $0.085
- Only a handful of scenarios ever fail, and rarely twice.
- Best realtime model on the long form, level with the cascade.
- Few data errors: a phone number or an earlier reference, wrong or left out.
- Weakest on noisy lines.
Most often wrong:
route_idphonedestinationCalls:All Short form Long form
GPT-Live 1 + GPT-6 Sol (low thinking)
OpenAI · Delegates reasoning and tools to Sol · service turn detection
- Reliability
- 76%
- Response
- 2.02s
- Per minute
- $0.055
- Near perfect on the short form, run after run.
- Every failure is in the saved record; the conversation itself always reaches the outcome.
- On the long form it routes the caller wrong or skips a step.
- Says goodbye without hanging up, and stalls more than any other model.
- Lowest voice score of the field.
Most often wrong:
destinationroute_statusstateCalls:All Short form Long form
Grok Voice Think Fast 2.0
xAI · Reasoning high · service turn detection
- Reliability
- 67%
- Response
- 1.62s
- Per minute
- $0.080
- Quick, with the shortest replies of the field; steady on the short form.
- Struggles on the long form: fields left out or wrong, usually an earlier reference.
- Every stall was on a long form.
Most often wrong:
lead_idroute_idconsent_idCalls:All Short form Long form
Phonic v1
Phonic · Intelligence high · service turn detection
- Reliability
- 67%
- Response
- 1.59s
- Per minute
- $0.140
- Level with the best on the short form, run after run.
- Quickest to reply: no late reply, no stalled call.
- Fails the long form's exact check: it fills in optional fields it should leave out, though the record would usually match.
- Short-form slips are an extra tool call; it repeats calls more than most.
Most often wrong:
beneficiary_namelead_idstateCalls:All Short form Long form
Gemini 3.1 Flash Live Preview
Google · Thinking level high · shared local turn detection
- Reliability
- 62%
- Response
- 2.88s
- Per minute
- $0.070
- Steady on the short form, though it repeats lookups it already made.
- On the long form it hands the caller over with the wrong details.
- Slowest to reply; after several tool calls its replies fall behind or go missing.
- Sometimes says goodbye without hanging up.
Most often wrong:
lead_iddestinationroute_idCalls:All Short form Long form
Gemini 3.8 Live Extended Thinking
Google · Thinking level high · shared local turn detection
- Reliability
- 56%
- Response
- 2.32s
- Per minute
- —
- Strong on the long form; when it slips, it skips a save it said it made.
- Weak on the short form: phone numbers saved wrong, lookups never made.
- Acts on half an answer instead of asking; weakest when the caller corrects themselves.
- Least predictable: the same scenario passes one run and fails the next.
- Replies fall behind or stop; callers hang up waiting.
Most often wrong:
phonecaller_namestateCalls:All Short form Long form
Amazon Nova 2 Sonic
Amazon · Endpointing medium · shared local turn detection
- Reliability
- 50%
- Response
- 2.18s
- Per minute
- —
- Misses caller names and phone numbers: blank or wrong.
- Long forms rarely pass; more scenarios fail every run than for any other model.
- Repeats the same tool call more than any other model.
- Best voice score of the field, and the longest replies.
Most often wrong:
caller_nameconsent_idcallback_phoneCalls:All Short form Long form
GPT Realtime 2.1 Mini
OpenAI · Reasoning high · service turn detection
- Reliability
- 50%
- Response
- 2.08s
- Per minute
- $0.020
- Gets most short-form calls right, at a fraction of the cost.
- Sometimes mishears a phone number.
- Drops details given earlier: not one long-form record was exact.
- Rarely recovers when a tool returns an error.
Most often wrong:
route_idzip_codeconsent_idCalls:All Short form Long form
Flux → GPT-4.1 → ElevenLabs Flash
Cascade baseline, not ranked
- Reliability
- 83%
- Response
- 2.17s
- Per minute
- —
- Most scenarios pass every run; none fails every run.
- Rare long-form slips: a reference from an earlier step left out.
- Not the quickest to reply, but almost never leaves the caller waiting.
Most often wrong:
route_idintake_intentservice_issue_typeCalls:All Short form Long form
Compare two models
- Reliability
- 79.3%1st
- 82.9%3.7 points higher
- Data accuracy
- 92.0%1st
- 94.1%2.1 points higher
- Response time
- 1.94s3rd0.23s sooner
- 2.17s
- Cost / min
- $0.0855th
- No verified rate
How the calls were run
- 01
Only the model changes
One open-source Pipecat agent: the realtime model hears the caller, calls the tools and speaks. Prompt, tools and phone line are the same for every model. A cascade (Flux → GPT-4.1 → ElevenLabs Flash) runs the same agent for reference, unranked.
- 02
Live calls, three runs each
A simulated caller works through all 82 scenarios on live calls, three times over. Every call counts, including ones that stalled or ended early.
- 03
Words and data both count
A call passes only when the agent reached the expected outcome and every field it saved matched, checked on the Cekura platform.
- 04
Every call is public
Each row links to the Cekura report of its calls, with transcripts, tool calls and scores, so any figure here can be checked.
Appointments
the short form · 59 scenarios · 6 fields · 4 toolsA patient calling a clinic. Find the patient, check openings, then book, move or cancel a visit.
- lookup_patient
- check_availability
- book_appointment
- cancel_appointment
Medicare
the long form · 23 scenarios · 42 fields · 4 toolsSomeone calling a Medicare insurance line. Record consent first, qualify the caller, route them to the right team and write a handoff note.
- record_medicare_permissions
- save_medicare_qualification
- route_medicare_call
- create_handoff_summary