All comparisons
DeepslateOpenAI

Deepslate Opal vs GPT-Live 1 + GPT-6 Sol (low thinking)

GPT-Live 1 + GPT-6 Sol (low thinking) is ahead on reliability, success rate, agent responses, data accuracy, response time and cost, and Deepslate Opal on stalled calls and interruption.

Last updated Methodology by Dileep Chagam

Compare

Success rate by suite

Higher is better
Appointments59 scenariosMedicare23 scenarios
  • GPT-Live 1 + Sol lowOpenAI
    97%
    49%
  • Deepslate OpalDeepslate
    84%
    35%
Each model's share of calls that passed every check, on each suite.
Verdict

Which to pick

Pick Deepslate Opal for fewer stalled calls (6.9 points lower) and handling interruptions (0.07 higher). Pick GPT-Live 1 + GPT-6 Sol (low thinking) for reliability across repeated calls (18.3 points higher), calls that pass every check (13.8 points higher), getting the agent's replies right (4.1 points higher), saving the caller's data correctly (16.5 points higher), fast replies (0.25s sooner) and lower cost ($0.057 cheaper).

Every metric, side by side

Reliability

Scenarios that passed on all three runs.

Deepslate Opal
57.3%7th
GPT-Live 1 + Sol low
75.6%2nd

GPT-Live 1 + Sol low, 18.3 points higher

Success rate

Calls that passed every check.

Deepslate Opal
69.9%8th
GPT-Live 1 + Sol low
83.7%3rd

GPT-Live 1 + Sol low, 13.8 points higher

Agent responses

Calls where the agent's replies reached the expected outcome.

Deepslate Opal
95.9%6th
GPT-Live 1 + Sol low
100.0%1st

GPT-Live 1 + Sol low, 4.1 points higher

Data accuracy

Calls where everything the agent saved matched exactly.

Deepslate Opal
72.6%8th
GPT-Live 1 + Sol low
89.0%2nd

GPT-Live 1 + Sol low, 16.5 points higher

Stalled calls

Calls where the agent went silent for 10 seconds or more.

Deepslate Opal
5.3%9th
GPT-Live 1 + Sol low
12.2%11th

Deepslate Opal, 6.9 points lower

Interruption

How well it handled being talked over, 0 to 5.

Deepslate Opal
4.967th
GPT-Live 1 + Sol low
4.8910th

Deepslate Opal, 0.07 higher

Response time

Median time to start replying after the caller finishes.

Deepslate Opal
2.27s10th
GPT-Live 1 + Sol low
2.02s6th

GPT-Live 1 + Sol low, 0.25s sooner

Cost / min

One minute of call at the vendor's published rates.

Deepslate Opal
$0.1127th
GPT-Live 1 + Sol low
$0.0552nd

GPT-Live 1 + Sol low, $0.057 cheaper

The small figure is the place among ranked models. The cascade is not ranked.

GPT-Live 1 + Sol low answers first

GPT-Live 1 + Sol low 2.02s, Deepslate Opal 2.27s at the median.

Response time, median to p90

Lower is betterMedianTo p90
Medianp90
  • GPT-Live 1 + Sol lowOpenAI
    GPT-Live 1 + Sol low: median 2.02s, p90 2.42s
    2.02s2.42s
  • Deepslate OpalDeepslate
    Deepslate Opal: median 2.27s, p90 2.83s
    2.27s2.83s
2.0s3.0s
Seconds from the caller finishing to the agent starting to reply, per model.

What we saw on the calls

Deepslate Opal

  • Every interrupted call and every call with several requests passed.
All 3 notes on Deepslate Opal
Model
opal
Setting
No reasoning setting
Voice
ElevenLabs 21m00Tcm4TlvDq8ikWAM
Turn taking
Service turn detection
Audio
16 kHz
Cost
The vendor's regular rate of €0.10 a minute, at the ECB rate of 5 Oct 2026, applied to each call's length; its launch offer is €0.02. Excludes the ElevenLabs voice, billed on the customer's own ElevenLabs key.

GPT-Live 1 + GPT-6 Sol (low thinking)

  • Near perfect on bookings, run after run.
All 4 notes on GPT-Live 1 + Sol low
Model
gpt-live-1
Setting
Delegates reasoning and tools to Sol
Voice
marin
Turn taking
Service turn detection
Audio
24 kHz
Cost
Includes the delegated backend model.

Frequently asked questions

Is Deepslate Opal or GPT-Live 1 + GPT-6 Sol (low thinking) more reliable?

GPT-Live 1 + GPT-6 Sol (low thinking) is more reliable, by 18.3 points. Deepslate Opal passed 47 of 82 scenarios on all three runs, and GPT-Live 1 + GPT-6 Sol (low thinking) passed 62. A scenario only counts when every run of it passed.

Which should I pick, Deepslate Opal or GPT-Live 1 + GPT-6 Sol (low thinking)?

For a phone agent, GPT-Live 1 + GPT-6 Sol (low thinking), since it gets more scenarios right on every run. GPT-Live 1 + GPT-6 Sol (low thinking) is ahead on reliability, success rate, agent responses, data accuracy, response time and cost, and Deepslate Opal on stalled calls and interruption.

Which saves caller data more accurately, Deepslate Opal or GPT-Live 1 + GPT-6 Sol (low thinking)?

GPT-Live 1 + GPT-6 Sol (low thinking), by 16.5 points. Deepslate Opal saved everything exactly right on 72.6% of the calls that save data, and GPT-Live 1 + GPT-6 Sol (low thinking) on 89.0%. On Appointments, 85.4% against 100.0%. On Medicare, 39.4% against 60.6%. A call counts only when nothing is missing, wrong, extra or saved twice.

Is Deepslate Opal faster than GPT-Live 1 + GPT-6 Sol (low thinking)?

No. GPT-Live 1 + GPT-6 Sol (low thinking) starts replying 0.25s sooner at the median. Deepslate Opal takes 2.27s at the median and 2.83s at p90; GPT-Live 1 + GPT-6 Sol (low thinking) takes 2.02s and 2.42s. Both are measured on the call audio, from the caller finishing to the agent starting to reply.

Which handles interruptions better, Deepslate Opal or GPT-Live 1 + GPT-6 Sol (low thinking)?

Deepslate Opal, by 0.07 on a 0 to 5 scale. Deepslate Opal scores 4.96 and GPT-Live 1 + GPT-6 Sol (low thinking) 4.89 for how well the agent handled being talked over.

Which goes silent on callers less often, Deepslate Opal or GPT-Live 1 + GPT-6 Sol (low thinking)?

Deepslate Opal, by 6.9 points. The agent went silent for 10 seconds or more on 5.3% of Deepslate Opal calls and 12.2% of GPT-Live 1 + GPT-6 Sol (low thinking) calls.

Which is cheaper, Deepslate Opal or GPT-Live 1 + GPT-6 Sol (low thinking)?

GPT-Live 1 + GPT-6 Sol (low thinking) is $0.057 cheaper a minute. Deepslate Opal costs $0.112 a minute of call and GPT-Live 1 + GPT-6 Sol (low thinking) $0.055, at each vendor's published rates.

Where does the Deepslate Opal vs GPT-Live 1 + GPT-6 Sol (low thinking) data come from?

Cekura ran both as the whole agent in the same open-source Pipecat pipeline, with the same prompt, tools and connection. A simulated caller worked through all 82 scenarios on live calls, three times each, 2,952 calls across the benchmark. A call passes only when the agent said the right things and saved the right data.