Compare speech-to-text models
Any two of the 15 models, head to head on accuracy and latency, measured under the same conditions.
AssemblyAI Universal 3.5 ProAssemblyAI
Inworld STT-1Inworld
- Pipecat Dataset WER
- 1.93%1st0.79 points lower
- 2.72%8th
- Ocular Dataset WER
- 4.23%9th6.38 points lower
- 10.60%14th
- Time to first text
- 489ms1st574ms sooner
- 1.06s5th
- Final-text delay
- 180ms7th
- 41ms1st139ms sooner
Lower is better on every metric. The small figure is the place in the whole field.All 105 comparisons