All comparisons
DeepslateOpenAI

Deepslate Opal vs GPT Realtime 2.1

GPT Realtime 2.1 is ahead of Deepslate Opal on all eight metrics.

Last updated Methodology by Dileep Chagam

Compare

Success rate by suite

Higher is better
Appointments59 scenariosMedicare23 scenarios
  • GPT Realtime 2.1OpenAI
    92%
    81%
  • Deepslate OpalDeepslate
    84%
    35%
Each model's share of calls that passed every check, on each suite.
Verdict

Which to pick

GPT Realtime 2.1 is the better pick over Deepslate Opal on every measure that separates them: reliability across repeated calls (22.0 points higher), calls that pass every check (19.1 points higher), getting the agent's replies right (3.3 points higher), saving the caller's data correctly (19.4 points higher), fewer stalled calls (2.0 points lower), handling interruptions (0.02 higher), fast replies (0.34s sooner) and lower cost ($0.027 cheaper).

Every metric, side by side

Reliability

Scenarios that passed on all three runs.

Deepslate Opal
57.3%7th
GPT Realtime 2.1
79.3%1st

GPT Realtime 2.1, 22.0 points higher

Success rate

Calls that passed every check.

Deepslate Opal
69.9%8th
GPT Realtime 2.1
89.0%1st

GPT Realtime 2.1, 19.1 points higher

Agent responses

Calls where the agent's replies reached the expected outcome.

Deepslate Opal
95.9%6th
GPT Realtime 2.1
99.2%2nd

GPT Realtime 2.1, 3.3 points higher

Data accuracy

Calls where everything the agent saved matched exactly.

Deepslate Opal
72.6%8th
GPT Realtime 2.1
92.0%1st

GPT Realtime 2.1, 19.4 points higher

Stalled calls

Calls where the agent went silent for 10 seconds or more.

Deepslate Opal
5.3%9th
GPT Realtime 2.1
3.3%5th

GPT Realtime 2.1, 2.0 points lower

Interruption

How well it handled being talked over, 0 to 5.

Deepslate Opal
4.967th
GPT Realtime 2.1
4.984th

GPT Realtime 2.1, 0.02 higher

Response time

Median time to start replying after the caller finishes.

Deepslate Opal
2.27s10th
GPT Realtime 2.1
1.94s5th

GPT Realtime 2.1, 0.34s sooner

Cost / min

One minute of call at the vendor's published rates.

Deepslate Opal
$0.1127th
GPT Realtime 2.1
$0.0856th

GPT Realtime 2.1, $0.027 cheaper

The small figure is the place among ranked models. The cascade is not ranked.

GPT Realtime 2.1 answers first

GPT Realtime 2.1 1.94s, Deepslate Opal 2.27s at the median.

Response time, median to p90

Lower is betterMedianTo p90
Medianp90
  • GPT Realtime 2.1OpenAI
    GPT Realtime 2.1: median 1.94s, p90 2.46s
    1.94s2.46s
  • Deepslate OpalDeepslate
    Deepslate Opal: median 2.27s, p90 2.83s
    2.27s2.83s
2.0s3.0s
Seconds from the caller finishing to the agent starting to reply, per model.

What we saw on the calls

Deepslate Opal

  • Every interrupted call and every call with several requests passed.
All 3 notes on Deepslate Opal
Model
opal
Setting
No reasoning setting
Voice
ElevenLabs 21m00Tcm4TlvDq8ikWAM
Turn taking
Service turn detection
Audio
16 kHz
Cost
The vendor's regular rate of €0.10 a minute, at the ECB rate of 5 Oct 2026, applied to each call's length; its launch offer is €0.02. Excludes the ElevenLabs voice, billed on the customer's own ElevenLabs key.

GPT Realtime 2.1

  • Almost never fails, and rarely twice in a row.
All 4 notes on GPT Realtime 2.1
Model
gpt-realtime-2.1
Setting
Reasoning high
Voice
marin
Turn taking
Service turn detection
Audio
24 kHz
Cost
Excludes caller transcription, which the service does not meter per call.

Frequently asked questions

Is Deepslate Opal or GPT Realtime 2.1 more reliable?

GPT Realtime 2.1 is more reliable, by 22.0 points. Deepslate Opal passed 47 of 82 scenarios on all three runs, and GPT Realtime 2.1 passed 65. A scenario only counts when every run of it passed.

Which should I pick, Deepslate Opal or GPT Realtime 2.1?

For a phone agent, GPT Realtime 2.1, since it gets more scenarios right on every run. GPT Realtime 2.1 is ahead of Deepslate Opal on all eight metrics.

Which saves caller data more accurately, Deepslate Opal or GPT Realtime 2.1?

GPT Realtime 2.1, by 19.4 points. Deepslate Opal saved everything exactly right on 72.6% of the calls that save data, and GPT Realtime 2.1 on 92.0%. On Appointments, 85.4% against 95.9%. On Medicare, 39.4% against 81.8%. A call counts only when nothing is missing, wrong, extra or saved twice.

Is Deepslate Opal faster than GPT Realtime 2.1?

No. GPT Realtime 2.1 starts replying 0.34s sooner at the median. Deepslate Opal takes 2.27s at the median and 2.83s at p90; GPT Realtime 2.1 takes 1.94s and 2.46s. Both are measured on the call audio, from the caller finishing to the agent starting to reply.

Which handles interruptions better, Deepslate Opal or GPT Realtime 2.1?

GPT Realtime 2.1, by 0.02 on a 0 to 5 scale. Deepslate Opal scores 4.96 and GPT Realtime 2.1 4.98 for how well the agent handled being talked over.

Which goes silent on callers less often, Deepslate Opal or GPT Realtime 2.1?

GPT Realtime 2.1, by 2.0 points. The agent went silent for 10 seconds or more on 5.3% of Deepslate Opal calls and 3.3% of GPT Realtime 2.1 calls.

Which is cheaper, Deepslate Opal or GPT Realtime 2.1?

GPT Realtime 2.1 is $0.027 cheaper a minute. Deepslate Opal costs $0.112 a minute of call and GPT Realtime 2.1 $0.085, at each vendor's published rates.

Where does the Deepslate Opal vs GPT Realtime 2.1 data come from?

Cekura ran both as the whole agent in the same open-source Pipecat pipeline, with the same prompt, tools and connection. A simulated caller worked through all 82 scenarios on live calls, three times each, 2,952 calls across the benchmark. A call passes only when the agent said the right things and saved the right data.