Converse-STT

In partnership with Ocular

Compare English speech models on transcription accuracy, final-text speed, and reliability.

15
models
1,000
Pipecat clips
8
Ocular recordings
206
turn clips

Quality and speed frontier

Every model's error rate on the public clips, against how long its transcript takes to settle.

Most attractive quadrantPareto line
Final-text delay P50 (ms)
◣ better
Pipecat Dataset WER (%)
The line is the frontier: no model beats these on one measure without giving up the other.
Latest results

Speech recognition results

Four metrics, each measured the same way for every model. Lower is better on all of them.

AssemblyAI Universal 3.5 Pro1.93%4.23%P50 489msP90 966msP95 1.08sP50 180msP90 532msP95 692ms
Google Chirp 32.00%4.19%P50 4.48sP90 6.21sP95 7.97sP50 475ms*P90 941ms*P95 1.41s*
Reson82.10%2.93%P50 2.40sP90 7.97sP95 10.20sP50 235msP90 357msP95 398ms
GPT Realtime Whisper2.16%3.53%P50 1.38sP90 1.84sP95 2.86sP50 543msP90 647msP95 669ms
Cartesia Ink 22.39%3.09%P50 1.51sP90 2.04sP95 3.23sP50 123msP90 139msP95 145ms
Speechmatics Linden2.46%3.93%P50 1.19sP90 1.56sP95 2.52sP50 147msP90 192msP95 213ms
Google Chirp 22.59%3.80%P50 4.50sP90 6.40sP95 8.15sP50 568ms*P90 1.04s*P95 1.22s*
Inworld STT-12.72%10.60%P50 1.06sP90 1.07sP95 1.07sP50 41msP90 60msP95 65ms
GPT-4o Mini Transcribe3.35%4.45%At speech end ‡P50 630msP90 1.34sP95 1.53s
Deepgram Nova-33.46%4.39%P50 1.06sP90 1.10sP95 2.07sP50 103msP90 143msP95 155ms
Smallest Pulse3.66%3.46%P50 1.06sP90 2.25sP95 3.03sP50 201msP90 356msP95 773ms
GPT-4o Transcribe3.90%12.45%At speech end ‡P50 652msP90 1.73sP95 1.97s
Deepgram Flux English3.90%4.18%P50 791msP90 1.23sP95 1.51sP50 108msP90 132msP95 141ms
Deepgram Flux Multilingual4.96%4.53%P50 793msP90 2.00sP95 3.41sP50 114msP90 135msP95 141ms
Gradium6.48%4.45%P50 1.71sP90 4.05sP95 7.32sP50 212msP90 266msP95 280ms

Pipecat Dataset WER covers the 1,000 public benchmark clips. Ocular Dataset WER covers the eight licensed recordings. TTFT is time to first text. TTFS is the final-text delay after speech ends. Lower is better on all four.

* Google Chirp: TTFS shows observed final-text delay after stream close because controlled finalization is unavailable.‡ GPT-4o Transcribe and GPT-4o Mini Transcribe: no text arrives until the speaker stops, so first-text time is not comparable.
Speed and accuracy

There is no single fastest and most accurate model

AssemblyAI Universal 3.5 Pro is the most accurate at 1.93%, and Inworld STT-1 is the fastest to a final transcript at 41ms. Neither is both.

Pipecat vs Ocular accuracy

Every model's error rate on the public clips, against the same model on real conversations.

Most attractive quadrant
Ocular Dataset WER (%)
◣ better
Pipecat Dataset WER (%)

Transcription response time

How soon each model returns its first words, against how soon its transcript is final.

Most attractive quadrant
Final-text delay P50 (ms)
◣ better
Time to first text P50 (ms)

Model ranking

Word error rate on 1,000 clips from Pipecat's STT benchmark dataset. Lower is better.

#ModelPipecat Datasetclips
1AssemblyAI Universal 3.5 Pro1.93%1,000
2Google Chirp 32.00%1,000
3Reson82.10%1,000
4GPT Realtime Whisper2.16%1,000
5Cartesia Ink 22.39%1,000
6Speechmatics Linden2.46%998
7Google Chirp 22.59%1,000
8Inworld STT-12.72%1,000
9GPT-4o Mini Transcribe3.35%1,000
10Deepgram Nova-33.46%1,000
11Smallest Pulse3.66%1,000
12GPT-4o Transcribe3.90%1,000
13Deepgram Flux English3.90%999
14Deepgram Flux Multilingual4.96%999
15Gradium6.48%1,000

Counts are the clips each model returned usable text for, so they differ between models and between datasets.

Final-text timing uses each provider's supported finalization contract. Compare models with similar contracts before treating small latency gaps as meaningful.

Methodology

How to read the results

  1. 01

    Accuracy

    WER counts incorrect, missing, and extra words against the reference transcript. Lower is better.

  2. 02

    Coverage

    Pipecat and Ocular word error rates stay separate, because a model can perform differently across the two datasets.

  3. 03

    Speed

    TTFT and TTFS are shown at P50, P90, and P95, measured on the same 206 Ocular Dataset turns for every model.

Dataset sources

Pipecat Dataset

1,000 public voice-agent clips from Pipecat's STT benchmark dataset. Its reference transcripts are generated with Gemini and described by Pipecat as human-reviewed.

View Pipecat dataset

Ocular Dataset

Eight licensed speaker recordings from four conversations, provided by Ocular for this benchmark.

About Ocular