OpenAI is ahead on reliability, success rate, agent responses, data accuracy, interruption and cost, and Google on stalled calls and response time.
Each row compares the best Google model with the best OpenAI model on that metric, from the same tests.
Last updated Methodology by Dileep Chagam
| Metric | OpenAI | |
|---|---|---|
| Reliability | 62.2%Gemini 3.1 Flash Live Preview | 79.3%GPT Realtime 2.1 |
| Success rate | 76.0%Gemini 3.8 Live Extended Thinking | 89.0%GPT Realtime 2.1 |
| Agent responses | 95.9%Gemini 3.1 Flash Live Preview | 100.0%GPT-Live 1 + GPT-6 Sol (low thinking) |
| Data accuracy | 84.4%Gemini 3.8 Live Extended Thinking | 92.0%GPT Realtime 2.1 |
| Stalled calls | 2.0%Gemini 3.8 Live | 2.5%GPT Realtime 2.1 Mini |
| Interruption | 4.98Gemini 3.8 Live | 4.99GPT Realtime 2.1 Mini |
| Response time | 1.91sGemini 3.8 Live | 1.94sGPT Realtime 2.1 |
| Cost / min | $0.070Gemini 3.1 Flash Live Preview | $0.020GPT Realtime 2.1 Mini |
Google has 3 models on Speech-to-speech: Gemini 3.8 Live Extended Thinking, Gemini 3.8 Live and Gemini 3.1 Flash Live Preview. Its best on reliability is Gemini 3.1 Flash Live Preview (62.2%). Its fastest is Gemini 3.8 Live (1.91s).
OpenAI has 3 models on Speech-to-speech: GPT Realtime 2.1, GPT-Live 1 + GPT-6 Sol (low thinking) and GPT Realtime 2.1 Mini. Its best on reliability is GPT Realtime 2.1 (79.3%). Its fastest is GPT Realtime 2.1 (1.94s).
Pick Google for stalled calls (0.4 points lower) and response time (0.03s sooner). Pick OpenAI for reliability (17.1 points higher), success rate (13.0 points higher), agent responses (4.1 points higher), data accuracy (7.6 points higher), interruption (0.01 higher) and cost ($0.050 cheaper).
OpenAI is ahead, 17.1 points higher. The best Google result is Gemini 3.1 Flash Live Preview at 62.2%, and the best OpenAI result is GPT Realtime 2.1 at 79.3%, on reliability.
Google is faster, 0.03s sooner. The best Google result is Gemini 3.8 Live at 1.91s, and the best OpenAI result is GPT Realtime 2.1 at 1.94s, on response time.
Google: Gemini 3.8 Live Extended Thinking, Gemini 3.8 Live, Gemini 3.1 Flash Live Preview. OpenAI: GPT Realtime 2.1, GPT-Live 1 + GPT-6 Sol (low thinking), GPT Realtime 2.1 Mini. Every model ran the same tests on Speech-to-speech.