Compare speech-to-text models

Any two of the 15 models, head to head on accuracy and latency, measured under the same conditions.

Compare
AssemblyAI Universal 3.5 ProAssemblyAI
Inworld STT-1Inworld
Pipecat Dataset WER
1.93%1st
0.79 points lower
2.72%8th
Ocular Dataset WER
4.23%9th
6.38 points lower
10.60%14th
Time to first text
489ms1st
574ms sooner
1.06s5th
Final-text delay
180ms7th
41ms1st
139ms sooner
Lower is better on every metric. The small figure is the place in the whole field.All 105 comparisons

Every comparison

AssemblyAI Universal 3.5 Pro

Google Chirp 3

Reson8

GPT Realtime Whisper

Cartesia Ink 2

Speechmatics Linden

Google Chirp 2

Inworld STT-1

GPT-4o Mini Transcribe

Deepgram Nova-3

Smallest Pulse

GPT-4o Transcribe

Deepgram Flux English

Deepgram Flux Multilingual

Gradium