Deepgram is ahead on time to first text and final-text delay, and OpenAI on the Pipecat Dataset. The Ocular Dataset is too close to call.
Each row compares the best Deepgram model with the best OpenAI model on that metric, from the same tests.
Last updated Methodology by Dileep Chagam
| Metric | Deepgram | OpenAI |
|---|---|---|
| Pipecat Dataset WER | 3.46%Deepgram Nova-3 | 2.16%GPT Realtime Whisper |
| Ocular Dataset WER | 4.18%Deepgram Flux English | 3.53%GPT Realtime Whisper |
| Time to first text | 791msDeepgram Flux English | 1.38sGPT Realtime Whisper |
| Final-text delay | 103msDeepgram Nova-3 | 543msGPT Realtime Whisper |
Deepgram has 3 models on Converse-STT: Deepgram Flux English, Deepgram Flux Multilingual and Deepgram Nova-3. Its best on accuracy is Deepgram Nova-3 (3.46%). Its fastest is Deepgram Nova-3 (103ms).
OpenAI has 3 models on Converse-STT: GPT-4o Mini Transcribe, GPT-4o Transcribe and GPT Realtime Whisper. Its best on accuracy is GPT Realtime Whisper (2.16%). Its fastest is GPT Realtime Whisper (543ms).
Pick Deepgram for time to first text (588ms sooner) and final-text delay (440ms sooner). Pick OpenAI for the Pipecat Dataset (1.31 points lower). For the Ocular Dataset, the gap is inside the margin of error, so neither is ahead.
OpenAI is ahead, 1.31 points lower. The best Deepgram result is Deepgram Nova-3 at 3.46%, and the best OpenAI result is GPT Realtime Whisper at 2.16%, on the Pipecat Dataset.
Deepgram is faster, 440ms sooner. The best Deepgram result is Deepgram Nova-3 at 103ms, and the best OpenAI result is GPT Realtime Whisper at 543ms, on final-text delay.
Deepgram: Deepgram Flux English, Deepgram Flux Multilingual, Deepgram Nova-3. OpenAI: GPT-4o Mini Transcribe, GPT-4o Transcribe, GPT Realtime Whisper. Every model ran the same tests on Converse-STT.