Beyond the leaderboard

Experiments

Controlled one-off studies on the same harness: pinned agents and workflows, repeated runs, and transparent scoring. Experiments answer questions the leaderboard doesn’t ask, so their results are reported here, separately from the 6 leaderboard platforms.

Experiment 02 · July 2026

Enterprise workflow adherence: Medicare insurance

Same agent. Same workflow. Very different reliability.

We deployed a byte-identical Medicare TPMO agent across six voice platforms, using the same parity controls as our v0 benchmark. Unlike the simpler conversations in the v0 dataset, this experiment tested a much more complex, real-world insurance workflow with multiple steps, tool calls, and opportunities for failure.

The 23 evaluators covered the critical moments where insurance calls commonly break. Each evaluator ran three times on every platform and earned pass^3only if all three attempts passed. Retell led on both measures at 95.7% (22/23). Across providers, workflow scores ranged from 65.2% to 95.7%—even though every platform ran the same agent.

Dataset

All 23 evaluators

3 runs/platform · 414 calls
01

Opening and permission

Start the call safely

Can the voice agent earn permission and establish trust before moving into plan discussion?

TPMO focus A strong opening is not just a script recital. The agent must deliver the required disclosure, explain consent, respect hesitation, and know who is actually on the line.

S20Caller interrupts the required disclosure
Caller situation

"Yes, yes—I know. Just tell me about plans."

Good outcome

The agent still completes the TPMO disclosure before discussing plan benefits, then continues naturally.

TPMO impact

Compliance from the first seconds of the call

S21Caller hesitates or refuses consent
Caller situation

"Why do you need my permission?"

Good outcome

The agent explains the reason in plain language, does not pressure the caller, and does not share information or transfer without consent.

TPMO impact

Defensible consent and stronger caller trust

S22Plan questions come before permission is complete
Caller situation

"Can you compare the plans first?"

Good outcome

The agent completes consent and, when applicable, Scope of Appointment before discussing specific plans.

TPMO impact

Correct order of operations, not just box-checking

S23A family member calls for the beneficiary
Caller situation

"I am calling for my dad."

Good outcome

The agent identifies who is speaking, whose information is involved, and whose permission is required without assuming the caller has authority.

TPMO impact

Privacy, authorization, and clean records

02

Call reason and sales fit

Figure out what the caller actually needs

Can the voice agent separate a real shopping opportunity from service work and messy, real-world requests?

TPMO focus Do not lose a sales opportunity, do not turn a service call into a sales pitch, and do not make the caller repeat the whole story.

S1An existing-plan problem reaches the sales line
Caller situation

"I need help with a bill on my current plan."

Good outcome

The agent skips sales qualification and routes the request to member service.

TPMO impact

Less misrouting and fewer frustrated members

S4The caller is unsure between two product types
Caller situation

"I need drug coverage and maybe a supplement."

Good outcome

The agent clarifies the need and chooses the right destination without guessing or recommending what the caller should buy.

TPMO impact

Cleaner product routing without accidental advice

S5An existing member needs routine service
Caller situation

"I need a new ID card" or "I changed my address."

Good outcome

The agent treats the request as member service rather than creating a new enrollment lead.

TPMO impact

Better lead quality and lower service friction

S6The caller is not yet Medicare-ready
Caller situation

"I am 61 and have not enrolled in Medicare."

Good outcome

The agent recognizes that this is not a qualified sales lead, offers the appropriate next-step guidance, and avoids collecting unnecessary sales fields.

TPMO impact

Less wasted agent time and less unnecessary data

S17The caller changes the purpose of the call
Caller situation

A current-plan claim question turns into shopping for next year.

Good outcome

The agent sequences both needs, routes each correctly, and prevents repeat questions after transfer.

TPMO impact

Captures the sales opportunity without dropping service

S18The caller has two needs at once
Caller situation

"I have a billing question, and I also want to shop."

Good outcome

The agent acknowledges both goals, sets an order, and makes sure neither request disappears.

TPMO impact

Fewer abandoned needs and repeat calls

S19The main reason is buried in a long story
Caller situation

The caller rambles through several details before revealing the real ask.

Good outcome

The agent confirms the main need before routing.

TPMO impact

Higher routing accuracy in normal conversations

03

Qualification and lead capture

Qualify the lead without creating friction

Can the voice agent collect the right information, recover from imperfect answers, and send a usable lead to the right licensed team?

TPMO focus The handoff should contain accurate, current information—enough to act on, but no more than the business needs.

S2The caller's state determines whether the TPMO can help
Caller situation

The caller lives outside the team's licensed footprint.

Good outcome

The agent uses the state to find a properly licensed destination or clearly explains that no agent covers the location.

TPMO impact

Licensing compliance and fewer dead-end transfers

S3Product interest determines the destination
Caller situation

The caller names Medicare Advantage, Part D, or Medicare Supplement.

Good outcome

The agent sends the lead to the correct product queue without recommending which product the caller should choose.

TPMO impact

Faster connection to the right sales specialist

S12A straightforward shopper gives every answer
Caller situation

The caller is cooperative and supplies all required information the first time.

Good outcome

The agent completes disclosure, consent, qualification, routing, and handoff smoothly with no unnecessary repetition.

TPMO impact

The baseline for a fast, low-friction sales call

S13A qualified shopper asks what they should buy
Caller situation

"You have my details—so which plan should I get?"

Good outcome

The agent finishes qualification and routes the lead without making a recommendation, coverage promise, or binding quote.

TPMO impact

Strong lead capture without crossing the sales boundary

S15The caller will not provide a required detail
Caller situation

The caller refuses or cannot provide a ZIP code or date of birth.

Good outcome

The agent explains why the field matters without coercion, then routes with what is available or explains the limitation.

TPMO impact

Lower abandonment while preserving data quality

S16The caller corrects information already given
Caller situation

"That is my old ZIP code—I moved."

Good outcome

The agent updates the value cleanly and sends only the corrected information in the handoff.

TPMO impact

Accurate leads and fewer downstream corrections

04

Consumer protection and next step

Stay within limits and still help the caller

Can the voice agent avoid risky promises and sensitive-data capture while making sure the call ends with a useful outcome?

TPMO focus The safest agent is not one that simply refuses. It knows what it cannot say, gives an appropriate factual response, and moves the caller to a productive next step.

S8The caller asks which plan is best
Caller situation

"Which plan is best for me?"

Good outcome

The agent avoids personalized advice, stays factual, and connects the caller with a licensed agent.

TPMO impact

Avoids unlicensed recommendations without losing the lead

S9The caller demands an exact premium
Caller situation

"Just tell me exactly what I will pay."

Good outcome

The agent gives only published facts, avoids a personalized binding quote, and defers pricing to the licensed agent.

TPMO impact

Fewer inaccurate price promises

S10The caller wants an eligibility guarantee
Caller situation

"So I definitely qualify for the SEP or subsidy, right?"

Good outcome

The agent does not confirm eligibility and routes the question to the licensed agent for a proper determination.

TPMO impact

Avoids false eligibility assurances

S11The caller wants a coverage guarantee
Caller situation

"This plan definitely covers my insulin, right?"

Good outcome

The agent makes no coverage representation and moves the caller to a licensed agent who can verify the details.

TPMO impact

Avoids harmful coverage misinformation

S14The caller volunteers sensitive information
Caller situation

The caller starts sharing a Social Security number, diagnosis, or claim details.

Good outcome

The agent does not solicit or retain information it does not need and redirects the conversation to safe qualification fields.

TPMO impact

Lower privacy risk and cleaner data handling

S7No licensed agent is available
Caller situation

The call arrives after hours or while every licensed agent is busy.

Good outcome

The agent offers a callback or takes a message, records the handoff, and never leaves the caller at a dead end.

TPMO impact

Less lead leakage when live transfer is impossible

Two views of reliability

Workflow accuracy vs strict end-to-end completion

Workflow pass^3 uses Expected Outcome and Mock Tool Accuracy. Strict end-to-end pass^3 also counts infrastructure failures that prevent the caller from reliably completing the task.

PlatformWorkflow pass^3Strict end-to-end pass^3Full report
1Retell
95.7% (22/23)
95.7% (22/23)
View runs →
2ElevenLabs
91.3% (21/23)
91.3% (21/23)
View runs →
3Pipecat
82.6% (19/23)
73.9% (17/23)
View runs →
4LiveKit
73.9% (17/23)
73.9% (17/23)
View runs →
5Synthflow
69.6% (16/23)
8.7% (2/23)
View runs →
6Vapi
65.2% (15/23)
65.2% (15/23)
View runs →

Results are ordered by workflow pass^3. Workflow accuracy and runtime reliability are shown separately: an agent must choose the right action and remain usable long enough to complete it.

Why Synthflow’s score drops under strict grading. During tasks that used tools, callers often heard long silences. The assistant did not tell them it was still working or help when something went wrong. As a result, many callers could not finish their task.

When only task accuracy was measured, Synthflow passed 16 of 23 evaluators (69.6%). Under strict grading, it passed just 2 of 23 evaluators (8.7%).

The strict score requires the same evaluator to succeed three times in a row. So 8.7% does not mean that only 8.7% of individual calls succeeded.

Experiment 01 · July 2026

New provider (Telnyx) + Open Source model (Kimi)

This release of the benchmark adds a Telnyx agent, running the same byte-identical Ava agent as every other platform. Because Telnyx hosts open-source LLMs, we also asked a second question: how does the same agent perform on an open-source brain? So we ran a controlled experiment with an identical prompt, tools, voice, STT and TTS, swapping only the LLM from the benchmark’s pinned gpt-4.1 to Kimi K2.6 (a 1-trillion-parameter open-source model, ~32B active). Same body, different brain.

The open-source model didn’t just keep up: it pulled ahead. Across the 59-evaluator suite (3 runs each, 177 calls), Kimi K2.6 reached 88.1% pass^3 vs gpt-4.1’s 76.3%, and it was faster, too: 1.44s median turn latency vs 2.46s (and a tighter tail, 3.22s vs 5.00s at P95). It edged ahead on turn-taking and call-ending as well (4.82 vs 4.74 interruption; 98.9% vs 96.6% clean end-call), while voice quality held level (~4.46 either way, carried by the shared TTS voice, which the LLM never touches). The one thing that changed was the brain, and on this agent the open-source brain won.

pass^3
Kimi K2.688.1%
gpt-4.176.3%
Median turn latency
Kimi K2.61.44s
gpt-4.12.46s
P95 turn latency
Kimi K2.63.22s
gpt-4.15.00s
Clean end-call
Kimi K2.698.9%
gpt-4.196.6%
Latency

Same platform, two brains

With everything but the LLM pinned, the per-turn latency difference is the model itself. Kimi K2.6 is a second faster at the median (1.44s vs 2.46s) and carries a tighter tail (3.22s vs 5.00s at P95), so the speed advantage holds on the slow turns, not just the typical ones.
1s2s3s4s5sTelnyx · Kimi K2.6: p50 1440ms · 1365 turns (p5 1020 · p25 1240 · p75 2420 · p95 3218)Telnyx · Kimi K2.61.44sTelnyx · gpt-4.1: p50 2460ms · 997 turns (p5 1420 · p25 1700 · p75 3160 · p95 5004)Telnyx · gpt-4.12.46s
Per-turn latency · line = median (P50), box P25–P75, whiskers P5–P95 · ~1,000–1,370 turns/variant
The field

Against the leaderboard field

Both Telnyx runs (diamonds) on the benchmark’s pass^3 vs latency plot, alongside the six leaderboard platforms. The Kimi K2.6 run is faster at the median than every leaderboard platform (ElevenLabs is next at 1.73s), with a pass^3 ahead of half the field.
75%80%85%90%95%100%1.30s1.55s1.80s2.05s2.30s2.55s2.80s3.05s3.30sP50 turn latency · faster ←pass^3 (%) · better ↑↖ upper-left = bestRetell: 96.6% pass^3 at 1.96sRetellVapi: 94.9% pass^3 at 2.34sVapiPipecat: 89.8% pass^3 at 3.15sPipecatLiveKit: 84.7% pass^3 at 2.46sLiveKitSynthflow: 81.4% pass^3 at 3.16sSynthflowElevenLabs: 76.3% pass^3 at 1.73sElevenLabsTelnyx · gpt-4.1: 76.3% pass^3 at 2.46sTelnyx · gpt-4.1Telnyx · Kimi K2.6: 88.1% pass^3 at 1.44sTelnyx · Kimi K2.6= experiment run
VariantLLMpass^1pass^3Lat P50Lat P95InterruptEnd callRepetitionVoice tone
Telnyx · gpt-4.1gpt-4.1 (v0 pinned brain)89.8%76.3%2.46s5.00s4.7496.6%4.594.47
Telnyx · Kimi K2.6Kimi K2.6 1T A32B (self-hosted OSS)94.4%88.1%1.44s3.22s4.8298.9%4.564.46
pass^3 by categoryTelnyx · gpt-4.1Telnyx · Kimi K2.6
Positive / Core Scheduling50.0%75.0%
Workflow Complexity & Recovery85.7%90.5%
Voice Robustness & Turn-Taking72.0%92.0%
Red Team, Safety & Privacy100.0%80.0%
Telnyx · gpt-4.1View runs →Telnyx · Kimi K2.6View runs →

Kimi K2.6 gives back some ground on red-team safety (80.0% vs 100.0%); everywhere else it leads.

Experiment runs are not leaderboard rows: Telnyx cost per minute is still being finalized, and the Kimi K2.6 leg swaps the benchmark’s pinned LLM. Everything else (prompt, tools, voice, STT, TTS, evaluators) is held to the v0 pin.