All comparisons
DeepslateOpenAI

Deepslate Opal vs GPT Realtime 2.1 Mini

Deepslate Opal is ahead on reliability, success rate, agent responses and data accuracy, and GPT Realtime 2.1 Mini on stalled calls, interruption, response time and cost.

Last updated Methodology by Dileep Chagam

Compare

Success rate by suite

Higher is better
Appointments59 scenariosMedicare23 scenarios
  • Deepslate OpalDeepslate
    84%
    35%
  • Realtime 2.1 MiniOpenAI
    84%
    4%
Each model's share of calls that passed every check, on each suite.
Verdict

Which to pick

Pick Deepslate Opal for reliability across repeated calls (7.3 points higher), calls that pass every check (8.5 points higher), getting the agent's replies right (2.4 points higher) and saving the caller's data correctly (8.0 points higher). Pick GPT Realtime 2.1 Mini for fewer stalled calls (2.8 points lower), handling interruptions (0.03 higher), fast replies (0.19s sooner) and lower cost ($0.092 cheaper).

Every metric, side by side

Reliability

Scenarios that passed on all three runs.

Deepslate Opal
57.3%7th
Realtime 2.1 Mini
50.0%9th

Deepslate Opal, 7.3 points higher

Success rate

Calls that passed every check.

Deepslate Opal
69.9%8th
Realtime 2.1 Mini
61.4%10th

Deepslate Opal, 8.5 points higher

Agent responses

Calls where the agent's replies reached the expected outcome.

Deepslate Opal
95.9%6th
Realtime 2.1 Mini
93.5%10th

Deepslate Opal, 2.4 points higher

Data accuracy

Calls where everything the agent saved matched exactly.

Deepslate Opal
72.6%8th
Realtime 2.1 Mini
64.5%10th

Deepslate Opal, 8.0 points higher

Stalled calls

Calls where the agent went silent for 10 seconds or more.

Deepslate Opal
5.3%9th
Realtime 2.1 Mini
2.5%4th

Realtime 2.1 Mini, 2.8 points lower

Interruption

How well it handled being talked over, 0 to 5.

Deepslate Opal
4.967th
Realtime 2.1 Mini
4.992nd

Realtime 2.1 Mini, 0.03 higher

Response time

Median time to start replying after the caller finishes.

Deepslate Opal
2.27s10th
Realtime 2.1 Mini
2.08s8th

Realtime 2.1 Mini, 0.19s sooner

Cost / min

One minute of call at the vendor's published rates.

Deepslate Opal
$0.1127th
Realtime 2.1 Mini
$0.0201st

Realtime 2.1 Mini, $0.092 cheaper

The small figure is the place among ranked models. The cascade is not ranked.

Realtime 2.1 Mini answers first

Realtime 2.1 Mini 2.08s, Deepslate Opal 2.27s at the median.

Response time, median to p90

Lower is betterMedianTo p90
Medianp90
  • Realtime 2.1 MiniOpenAI
    Realtime 2.1 Mini: median 2.08s, p90 2.75s
    2.08s2.75s
  • Deepslate OpalDeepslate
    Deepslate Opal: median 2.27s, p90 2.83s
    2.27s2.83s
2.0s3.0s
Seconds from the caller finishing to the agent starting to reply, per model.

What we saw on the calls

Deepslate Opal

  • Every interrupted call and every call with several requests passed.
All 3 notes on Deepslate Opal
Model
opal
Setting
No reasoning setting
Voice
ElevenLabs 21m00Tcm4TlvDq8ikWAM
Turn taking
Service turn detection
Audio
16 kHz
Cost
The vendor's regular rate of €0.10 a minute, at the ECB rate of 5 Oct 2026, applied to each call's length; its launch offer is €0.02. Excludes the ElevenLabs voice, billed on the customer's own ElevenLabs key.

GPT Realtime 2.1 Mini

  • Gets most bookings right, at a fraction of the cost.
All 3 notes on Realtime 2.1 Mini
Model
gpt-realtime-2.1-mini
Setting
Reasoning high
Voice
marin
Turn taking
Service turn detection
Audio
24 kHz
Cost
Averaged over the calls with a usage record.

Frequently asked questions

Is Deepslate Opal or GPT Realtime 2.1 Mini more reliable?

Deepslate Opal is more reliable, by 7.3 points. Deepslate Opal passed 47 of 82 scenarios on all three runs, and GPT Realtime 2.1 Mini passed 41. A scenario only counts when every run of it passed.

Which should I pick, Deepslate Opal or GPT Realtime 2.1 Mini?

For a phone agent, Deepslate Opal, since it gets more scenarios right on every run. Deepslate Opal is ahead on reliability, success rate, agent responses and data accuracy, and GPT Realtime 2.1 Mini on stalled calls, interruption, response time and cost.

Which saves caller data more accurately, Deepslate Opal or GPT Realtime 2.1 Mini?

Deepslate Opal, by 8.0 points. Deepslate Opal saved everything exactly right on 72.6% of the calls that save data, and GPT Realtime 2.1 Mini on 64.5%. On Appointments, 85.4% against 88.8%. On Medicare, 39.4% against 0.0%. A call counts only when nothing is missing, wrong, extra or saved twice.

Is Deepslate Opal faster than GPT Realtime 2.1 Mini?

No. GPT Realtime 2.1 Mini starts replying 0.19s sooner at the median. Deepslate Opal takes 2.27s at the median and 2.83s at p90; GPT Realtime 2.1 Mini takes 2.08s and 2.75s. Both are measured on the call audio, from the caller finishing to the agent starting to reply.

Which handles interruptions better, Deepslate Opal or GPT Realtime 2.1 Mini?

GPT Realtime 2.1 Mini, by 0.03 on a 0 to 5 scale. Deepslate Opal scores 4.96 and GPT Realtime 2.1 Mini 4.99 for how well the agent handled being talked over.

Which goes silent on callers less often, Deepslate Opal or GPT Realtime 2.1 Mini?

GPT Realtime 2.1 Mini, by 2.8 points. The agent went silent for 10 seconds or more on 5.3% of Deepslate Opal calls and 2.5% of GPT Realtime 2.1 Mini calls.

Which is cheaper, Deepslate Opal or GPT Realtime 2.1 Mini?

GPT Realtime 2.1 Mini is $0.092 cheaper a minute. Deepslate Opal costs $0.112 a minute of call and GPT Realtime 2.1 Mini $0.020, at each vendor's published rates.

Where does the Deepslate Opal vs GPT Realtime 2.1 Mini data come from?

Cekura ran both as the whole agent in the same open-source Pipecat pipeline, with the same prompt, tools and connection. A simulated caller worked through all 82 scenarios on live calls, three times each, 2,952 calls across the benchmark. A call passes only when the agent said the right things and saved the right data.