Converse-STT
In partnership with Ocular
Compare English speech models on transcription accuracy, final-text speed, and reliability.
- 15
- models
- 1,000
- Pipecat clips
- 8
- Ocular recordings
- 206
- turn clips
Quality and speed frontier
Every model's error rate on the public clips, against how long its transcript takes to settle.
Speech recognition results
Four metrics, each measured the same way for every model. Lower is better on all of them.
| AssemblyAI Universal 3.5 Pro | 1.93% | 4.23% | P50 489msP90 966msP95 1.08s | P50 180msP90 532msP95 692ms |
| Google Chirp 3 | 2.00% | 4.19% | P50 4.48sP90 6.21sP95 7.97s | P50 475ms*P90 941ms*P95 1.41s* |
| Reson8 | 2.10% | 2.93% | P50 2.40sP90 7.97sP95 10.20s | P50 235msP90 357msP95 398ms |
| GPT Realtime Whisper | 2.16% | 3.53% | P50 1.38sP90 1.84sP95 2.86s | P50 543msP90 647msP95 669ms |
| Cartesia Ink 2 | 2.39% | 3.09% | P50 1.51sP90 2.04sP95 3.23s | P50 123msP90 139msP95 145ms |
| Speechmatics Linden | 2.46% | 3.93% | P50 1.19sP90 1.56sP95 2.52s | P50 147msP90 192msP95 213ms |
| Google Chirp 2 | 2.59% | 3.80% | P50 4.50sP90 6.40sP95 8.15s | P50 568ms*P90 1.04s*P95 1.22s* |
| Inworld STT-1 | 2.72% | 10.60% | P50 1.06sP90 1.07sP95 1.07s | P50 41msP90 60msP95 65ms |
| GPT-4o Mini Transcribe | 3.35% | 4.45% | At speech end ‡ | P50 630msP90 1.34sP95 1.53s |
| Deepgram Nova-3 | 3.46% | 4.39% | P50 1.06sP90 1.10sP95 2.07s | P50 103msP90 143msP95 155ms |
| Smallest Pulse | 3.66% | 3.46% | P50 1.06sP90 2.25sP95 3.03s | P50 201msP90 356msP95 773ms |
| GPT-4o Transcribe | 3.90% | 12.45% | At speech end ‡ | P50 652msP90 1.73sP95 1.97s |
| Deepgram Flux English | 3.90% | 4.18% | P50 791msP90 1.23sP95 1.51s | P50 108msP90 132msP95 141ms |
| Deepgram Flux Multilingual | 4.96% | 4.53% | P50 793msP90 2.00sP95 3.41s | P50 114msP90 135msP95 141ms |
| Gradium | 6.48% | 4.45% | P50 1.71sP90 4.05sP95 7.32s | P50 212msP90 266msP95 280ms |
Pipecat Dataset WER covers the 1,000 public benchmark clips. Ocular Dataset WER covers the eight licensed recordings. TTFT is time to first text. TTFS is the final-text delay after speech ends. Lower is better on all four.
There is no single fastest and most accurate model
AssemblyAI Universal 3.5 Pro is the most accurate at 1.93%, and Inworld STT-1 is the fastest to a final transcript at 41ms. Neither is both.
Pipecat vs Ocular accuracy
Every model's error rate on the public clips, against the same model on real conversations.
Transcription response time
How soon each model returns its first words, against how soon its transcript is final.
Model ranking
Word error rate on 1,000 clips from Pipecat's STT benchmark dataset. Lower is better.
| # | Model | Pipecat Dataset | clips | |
|---|---|---|---|---|
| 1 | AssemblyAI Universal 3.5 Pro | 1.93% | 1,000 | |
| 2 | Google Chirp 3 | 2.00% | 1,000 | |
| 3 | Reson8 | 2.10% | 1,000 | |
| 4 | GPT Realtime Whisper | 2.16% | 1,000 | |
| 5 | Cartesia Ink 2 | 2.39% | 1,000 | |
| 6 | Speechmatics Linden | 2.46% | 998 | |
| 7 | Google Chirp 2 | 2.59% | 1,000 | |
| 8 | Inworld STT-1 | 2.72% | 1,000 | |
| 9 | GPT-4o Mini Transcribe | 3.35% | 1,000 | |
| 10 | Deepgram Nova-3 | 3.46% | 1,000 | |
| 11 | Smallest Pulse | 3.66% | 1,000 | |
| 12 | GPT-4o Transcribe | 3.90% | 1,000 | |
| 13 | Deepgram Flux English | 3.90% | 999 | |
| 14 | Deepgram Flux Multilingual | 4.96% | 999 | |
| 15 | Gradium | 6.48% | 1,000 |
Counts are the clips each model returned usable text for, so they differ between models and between datasets.
Final-text timing uses each provider's supported finalization contract. Compare models with similar contracts before treating small latency gaps as meaningful.
How to read the results
- 01
Accuracy
WER counts incorrect, missing, and extra words against the reference transcript. Lower is better.
- 02
Coverage
Pipecat and Ocular word error rates stay separate, because a model can perform differently across the two datasets.
- 03
Speed
TTFT and TTFS are shown at P50, P90, and P95, measured on the same 206 Ocular Dataset turns for every model.
Pipecat Dataset
1,000 public voice-agent clips from Pipecat's STT benchmark dataset. Its reference transcripts are generated with Gemini and described by Pipecat as human-reviewed.
View Pipecat datasetOcular Dataset
Eight licensed speaker recordings from four conversations, provided by Ocular for this benchmark.
About Ocular