Google is ahead on final-text delay, and OpenAI on time to first text. The Pipecat Dataset and The Ocular Dataset are too close to call.
Each row compares the best Google model with the best OpenAI model on that metric, from the same tests.
Last updated Methodology by Dileep Chagam
| Metric | OpenAI | |
|---|---|---|
| Pipecat Dataset WER | 2.00%Google Chirp 3 | 2.16%GPT Realtime Whisper |
| Ocular Dataset WER | 3.80%Google Chirp 2 | 3.53%GPT Realtime Whisper |
| Time to first text | 4.48sGoogle Chirp 3 | 1.38sGPT Realtime Whisper |
| Final-text delay | 475msGoogle Chirp 3 | 543msGPT Realtime Whisper |
Google has 2 models on Converse-STT: Google Chirp 2 and Google Chirp 3. Its best on accuracy is Google Chirp 3 (2.00%). Its fastest is Google Chirp 3 (475ms).
OpenAI has 3 models on Converse-STT: GPT-4o Mini Transcribe, GPT-4o Transcribe and GPT Realtime Whisper. Its best on accuracy is GPT Realtime Whisper (2.16%). Its fastest is GPT Realtime Whisper (543ms).
Pick Google for final-text delay (68ms sooner). Pick OpenAI for time to first text (3.10s sooner). For the Pipecat Dataset and the Ocular Dataset, the gap is inside the margin of error, so neither is ahead.
It is too close to call: the gap is inside the margin of error. The best Google result is Google Chirp 3 at 2.00%, and the best OpenAI result is GPT Realtime Whisper at 2.16%, on the Pipecat Dataset.
Google is faster, 68ms sooner. The best Google result is Google Chirp 3 at 475ms, and the best OpenAI result is GPT Realtime Whisper at 543ms, on final-text delay.
Google: Google Chirp 2, Google Chirp 3. OpenAI: GPT-4o Mini Transcribe, GPT-4o Transcribe, GPT Realtime Whisper. Every model ran the same tests on Converse-STT.