All comparisons

AssemblyAI vs Deepgram speech-to-text

AssemblyAI is ahead of Deepgram on the Pipecat Dataset, time to first text and final-text delay. The Ocular Dataset is too close to call.

Each row compares the best AssemblyAI model with the best Deepgram model on that metric, from the same tests.

Last updated Methodology by Dileep Chagam

Best of each

AssemblyAI and Deepgram, metric by metric

MetricAssemblyAIDeepgram
Pipecat Dataset WER1.77%AssemblyAI Universal 3.6 Pro3.46%Deepgram Nova-3
Ocular Dataset WER4.23%AssemblyAI Universal 3.5 Pro4.18%Deepgram Flux English
Time to first text485msAssemblyAI Universal 3.6 Pro791msDeepgram Flux English
Final-text delay91msAssemblyAI Universal 3.6 Pro103msDeepgram Nova-3
WER · Word error rate
Share of words the model got wrong: replaced, missing or added. Lower is better.
Pipecat Dataset
1,000 short public clips from Pipecat's STT benchmark dataset.
Ocular Dataset
8 recordings from 4 real phone conversations, licensed from Ocular.
TTFT · Time to first text
How long after the audio starts streaming until the first words appear.
TTFS · Final-text delay
How long after the caller stops until the full transcript arrives. A voice agent waits for this before it replies.
AssemblyAI line-up

AssemblyAI has 2 models on Converse-STT: AssemblyAI Universal 3.5 Pro and AssemblyAI Universal 3.6 Pro. Its best on accuracy is AssemblyAI Universal 3.6 Pro (1.77%). Its fastest is AssemblyAI Universal 3.6 Pro (91ms).

Deepgram line-up

Deepgram has 3 models on Converse-STT: Deepgram Flux English, Deepgram Flux Multilingual and Deepgram Nova-3. Its best on accuracy is Deepgram Nova-3 (3.46%). Its fastest is Deepgram Nova-3 (103ms).

Verdict

Which to pick

AssemblyAI is the better pick over Deepgram on every measure that separates them: the Pipecat Dataset (1.69 points lower), time to first text (306ms sooner) and final-text delay (12ms sooner). For the Ocular Dataset, the gap is inside the margin of error, so neither is ahead.

Frequently asked questions

Is AssemblyAI or Deepgram better for speech-to-text?

AssemblyAI is ahead, 1.69 points lower. The best AssemblyAI result is AssemblyAI Universal 3.6 Pro at 1.77%, and the best Deepgram result is Deepgram Nova-3 at 3.46%, on the Pipecat Dataset.

Which is faster, AssemblyAI or Deepgram?

AssemblyAI is faster, 12ms sooner. The best AssemblyAI result is AssemblyAI Universal 3.6 Pro at 91ms, and the best Deepgram result is Deepgram Nova-3 at 103ms, on final-text delay.

Which AssemblyAI and Deepgram models were tested?

AssemblyAI: AssemblyAI Universal 3.5 Pro, AssemblyAI Universal 3.6 Pro. Deepgram: Deepgram Flux English, Deepgram Flux Multilingual, Deepgram Nova-3. Every model ran the same tests on Converse-STT.