Deepgram is ahead on time to first text and final-text delay, and Google on the Pipecat Dataset. The Ocular Dataset is too close to call.
Each row compares the best Deepgram model with the best Google model on that metric, from the same tests.
Last updated Methodology by Dileep Chagam
| Metric | Deepgram | |
|---|---|---|
| Pipecat Dataset WER | 3.46%Deepgram Nova-3 | 2.00%Google Chirp 3 |
| Ocular Dataset WER | 4.18%Deepgram Flux English | 3.80%Google Chirp 2 |
| Time to first text | 791msDeepgram Flux English | 4.48sGoogle Chirp 3 |
| Final-text delay | 103msDeepgram Nova-3 | 475msGoogle Chirp 3 |
Deepgram has 3 models on Converse-STT: Deepgram Flux English, Deepgram Flux Multilingual and Deepgram Nova-3. Its best on accuracy is Deepgram Nova-3 (3.46%). Its fastest is Deepgram Nova-3 (103ms).
Google has 2 models on Converse-STT: Google Chirp 2 and Google Chirp 3. Its best on accuracy is Google Chirp 3 (2.00%). Its fastest is Google Chirp 3 (475ms).
Pick Deepgram for time to first text (3.69s sooner) and final-text delay (372ms sooner). Pick Google for the Pipecat Dataset (1.46 points lower). For the Ocular Dataset, the gap is inside the margin of error, so neither is ahead.
Google is ahead, 1.46 points lower. The best Deepgram result is Deepgram Nova-3 at 3.46%, and the best Google result is Google Chirp 3 at 2.00%, on the Pipecat Dataset.
Deepgram is faster, 372ms sooner. The best Deepgram result is Deepgram Nova-3 at 103ms, and the best Google result is Google Chirp 3 at 475ms, on final-text delay.
Deepgram: Deepgram Flux English, Deepgram Flux Multilingual, Deepgram Nova-3. Google: Google Chirp 2, Google Chirp 3. Every model ran the same tests on Converse-STT.