Compare speech-to-speech models

Any two of the 8 models, or any one against the cascade, on the same phone agent and the same calls.

Compare
GPT Realtime 2.1OpenAI
Flux → GPT-4.1 → ElevenLabs FlashBaseline, not ranked
Reliability
79.3%1st
82.9%
3.7 points higher
Data accuracy
92.0%1st
94.1%
2.1 points higher
Response time
1.94s3rd
0.23s sooner
2.17s
Cost / min
$0.0855th
No verified rate
All suites. The small figure is the place among ranked models.All 36 comparisons

Every comparison

GPT Realtime 2.1

GPT-Live 1 + Sol low

Grok Think Fast 2.0

Phonic v1

Gemini 3.1 Flash Live

Gemini 3.8 Live Thinking

Nova 2 Sonic

Realtime 2.1 Mini

Cascade, GPT-4.1Not ranked