AssemblyAI is ahead of Google on time to first text and final-text delay. The Pipecat Dataset and The Ocular Dataset are too close to call.
Each row compares the best AssemblyAI model with the best Google model on that metric, from the same tests.
Last updated Methodology by Dileep Chagam
| Metric | AssemblyAI | |
|---|---|---|
| Pipecat Dataset WER | 1.77%AssemblyAI Universal 3.6 Pro | 2.00%Google Chirp 3 |
| Ocular Dataset WER | 4.23%AssemblyAI Universal 3.5 Pro | 3.80%Google Chirp 2 |
| Time to first text | 485msAssemblyAI Universal 3.6 Pro | 4.48sGoogle Chirp 3 |
| Final-text delay | 91msAssemblyAI Universal 3.6 Pro | 475msGoogle Chirp 3 |
AssemblyAI has 2 models on Converse-STT: AssemblyAI Universal 3.5 Pro and AssemblyAI Universal 3.6 Pro. Its best on accuracy is AssemblyAI Universal 3.6 Pro (1.77%). Its fastest is AssemblyAI Universal 3.6 Pro (91ms).
Google has 2 models on Converse-STT: Google Chirp 2 and Google Chirp 3. Its best on accuracy is Google Chirp 3 (2.00%). Its fastest is Google Chirp 3 (475ms).
AssemblyAI is the better pick over Google on every measure that separates them: time to first text (4.00s sooner) and final-text delay (384ms sooner). For the Pipecat Dataset and the Ocular Dataset, the gap is inside the margin of error, so neither is ahead.
It is too close to call: the gap is inside the margin of error. The best AssemblyAI result is AssemblyAI Universal 3.6 Pro at 1.77%, and the best Google result is Google Chirp 3 at 2.00%, on the Pipecat Dataset.
AssemblyAI is faster, 384ms sooner. The best AssemblyAI result is AssemblyAI Universal 3.6 Pro at 91ms, and the best Google result is Google Chirp 3 at 475ms, on final-text delay.
AssemblyAI: AssemblyAI Universal 3.5 Pro, AssemblyAI Universal 3.6 Pro. Google: Google Chirp 2, Google Chirp 3. Every model ran the same tests on Converse-STT.