All comparisons

Deepgram vs Google speech-to-text

Deepgram is ahead on time to first text and final-text delay, and Google on the Pipecat Dataset. The Ocular Dataset is too close to call.

Each row compares the best Deepgram model with the best Google model on that metric, from the same tests.

Last updated Methodology by Dileep Chagam

Best of each

Deepgram and Google, metric by metric

MetricDeepgramGoogle
Pipecat Dataset WER3.46%Deepgram Nova-32.00%Google Chirp 3
Ocular Dataset WER4.18%Deepgram Flux English3.80%Google Chirp 2
Time to first text791msDeepgram Flux English4.48sGoogle Chirp 3
Final-text delay103msDeepgram Nova-3475msGoogle Chirp 3
Deepgram line-up

Deepgram has 3 models on Converse-STT: Deepgram Flux English, Deepgram Flux Multilingual and Deepgram Nova-3. Its best on accuracy is Deepgram Nova-3 (3.46%). Its fastest is Deepgram Nova-3 (103ms).

Google line-up

Google has 2 models on Converse-STT: Google Chirp 2 and Google Chirp 3. Its best on accuracy is Google Chirp 3 (2.00%). Its fastest is Google Chirp 3 (475ms).

Verdict

Which to pick

Pick Deepgram for time to first text (3.69s sooner) and final-text delay (372ms sooner). Pick Google for the Pipecat Dataset (1.46 points lower). For the Ocular Dataset, the gap is inside the margin of error, so neither is ahead.

Frequently asked questions

Is Deepgram or Google better for speech-to-text?

Google is ahead, 1.46 points lower. The best Deepgram result is Deepgram Nova-3 at 3.46%, and the best Google result is Google Chirp 3 at 2.00%, on the Pipecat Dataset.

Which is faster, Deepgram or Google?

Deepgram is faster, 372ms sooner. The best Deepgram result is Deepgram Nova-3 at 103ms, and the best Google result is Google Chirp 3 at 475ms, on final-text delay.

Which Deepgram and Google models were tested?

Deepgram: Deepgram Flux English, Deepgram Flux Multilingual, Deepgram Nova-3. Google: Google Chirp 2, Google Chirp 3. Every model ran the same tests on Converse-STT.