AssemblyAI is ahead of OpenAI on time to first text and final-text delay. The Pipecat Dataset and The Ocular Dataset are too close to call.
Each row compares the best AssemblyAI model with the best OpenAI model on that metric, from the same tests.
Last updated Methodology by Dileep Chagam
| Metric | AssemblyAI | OpenAI |
|---|---|---|
| Pipecat Dataset WER | 1.77%AssemblyAI Universal 3.6 Pro | 2.16%GPT Realtime Whisper |
| Ocular Dataset WER | 4.23%AssemblyAI Universal 3.5 Pro | 3.53%GPT Realtime Whisper |
| Time to first text | 485msAssemblyAI Universal 3.6 Pro | 1.38sGPT Realtime Whisper |
| Final-text delay | 91msAssemblyAI Universal 3.6 Pro | 543msGPT Realtime Whisper |
AssemblyAI has 2 models on Converse-STT: AssemblyAI Universal 3.5 Pro and AssemblyAI Universal 3.6 Pro. Its best on accuracy is AssemblyAI Universal 3.6 Pro (1.77%). Its fastest is AssemblyAI Universal 3.6 Pro (91ms).
OpenAI has 3 models on Converse-STT: GPT-4o Mini Transcribe, GPT-4o Transcribe and GPT Realtime Whisper. Its best on accuracy is GPT Realtime Whisper (2.16%). Its fastest is GPT Realtime Whisper (543ms).
AssemblyAI is the better pick over OpenAI on every measure that separates them: time to first text (894ms sooner) and final-text delay (452ms sooner). For the Pipecat Dataset and the Ocular Dataset, the gap is inside the margin of error, so neither is ahead.
It is too close to call: the gap is inside the margin of error. The best AssemblyAI result is AssemblyAI Universal 3.6 Pro at 1.77%, and the best OpenAI result is GPT Realtime Whisper at 2.16%, on the Pipecat Dataset.
AssemblyAI is faster, 452ms sooner. The best AssemblyAI result is AssemblyAI Universal 3.6 Pro at 91ms, and the best OpenAI result is GPT Realtime Whisper at 543ms, on final-text delay.
AssemblyAI: AssemblyAI Universal 3.5 Pro, AssemblyAI Universal 3.6 Pro. OpenAI: GPT-4o Mini Transcribe, GPT-4o Transcribe, GPT Realtime Whisper. Every model ran the same tests on Converse-STT.