Beyond the leaderboard

Experiments

Controlled and historical studies using the Cekura evaluation harness, repeated runs, and transparent scoring. These studies answer questions the primary release does not, so their results are reported here, separately from the leaderboard of 7 configurations.

Experiment 02 · July 2026

Enterprise workflow adherence: Medicare insurance 

Same agent. Same workflow. Very different reliability.

We deployed the same Medicare TPMO agent across six voice platforms, using the controls from the archived v0 benchmark. Unlike the simpler conversations in that dataset, this experiment tested a more complex, real-world insurance workflow with multiple steps, tool calls, and opportunities for failure.

The 23 evaluators covered the critical moments where insurance calls commonly break. Each evaluator ran three times on every platform and earned pass^3 only if all three attempts passed. Retell led on both measures at 95.7% (22/23). Across providers, workflow scores ranged from 65.2% to 95.7%, even though every platform ran the same agent.

Dataset

All 23 evaluators

3 runs/platform · 414 calls
01

Opening and permission

Start the call safely

Can the voice agent earn permission and establish trust before moving into plan discussion?

TPMO focus A strong opening is not just a script recital. The agent must deliver the required disclosure, explain consent, respect hesitation, and know who is actually on the line.

S20Caller interrupts the required disclosure
Caller situation

"Yes, yes, I know. Just tell me about plans."

Good outcome

The agent still completes the TPMO disclosure before discussing plan benefits, then continues naturally.

TPMO impact

Compliance from the first seconds of the call

S21Caller hesitates or refuses consent
Caller situation

"Why do you need my permission?"

Good outcome

The agent explains the reason in plain language, does not pressure the caller, and does not share information or transfer without consent.

TPMO impact

Defensible consent and stronger caller trust

S22Plan questions come before permission is complete
Caller situation

"Can you compare the plans first?"

Good outcome

The agent completes consent and, when applicable, Scope of Appointment before discussing specific plans.

TPMO impact

Correct order of operations, not just box-checking

S23A family member calls for the beneficiary
Caller situation

"I am calling for my dad."

Good outcome

The agent identifies who is speaking, whose information is involved, and whose permission is required without assuming the caller has authority.

TPMO impact

Privacy, authorization, and clean records

02

Call reason and sales fit

Figure out what the caller actually needs

Can the voice agent separate a real shopping opportunity from service work and messy, real-world requests?

TPMO focus Do not lose a sales opportunity, do not turn a service call into a sales pitch, and do not make the caller repeat the whole story.

S1An existing-plan problem reaches the sales line
Caller situation

"I need help with a bill on my current plan."

Good outcome

The agent skips sales qualification and routes the request to member service.

TPMO impact

Less misrouting and fewer frustrated members

S4The caller is unsure between two product types
Caller situation

"I need drug coverage and maybe a supplement."

Good outcome

The agent clarifies the need and chooses the right destination without guessing or recommending what the caller should buy.

TPMO impact

Cleaner product routing without accidental advice

S5An existing member needs routine service
Caller situation

"I need a new ID card" or "I changed my address."

Good outcome

The agent treats the request as member service rather than creating a new enrollment lead.

TPMO impact

Better lead quality and lower service friction

S6The caller is not yet Medicare-ready
Caller situation

"I am 61 and have not enrolled in Medicare."

Good outcome

The agent recognizes that this is not a qualified sales lead, offers the appropriate next-step guidance, and avoids collecting unnecessary sales fields.

TPMO impact

Less wasted agent time and less unnecessary data

S17The caller changes the purpose of the call
Caller situation

A current-plan claim question turns into shopping for next year.

Good outcome

The agent sequences both needs, routes each correctly, and prevents repeat questions after transfer.

TPMO impact

Captures the sales opportunity without dropping service

S18The caller has two needs at once
Caller situation

"I have a billing question, and I also want to shop."

Good outcome

The agent acknowledges both goals, sets an order, and makes sure neither request disappears.

TPMO impact

Fewer abandoned needs and repeat calls

S19The main reason is buried in a long story
Caller situation

The caller rambles through several details before revealing the real ask.

Good outcome

The agent confirms the main need before routing.

TPMO impact

Higher routing accuracy in normal conversations

03

Qualification and lead capture

Qualify the lead without creating friction

Can the voice agent collect the right information, recover from imperfect answers, and send a usable lead to the right licensed team?

TPMO focus The handoff should contain accurate, current information, enough to act on but no more than the business needs.

S2The caller's state determines whether the TPMO can help
Caller situation

The caller lives outside the team's licensed footprint.

Good outcome

The agent uses the state to find a properly licensed destination or clearly explains that no agent covers the location.

TPMO impact

Licensing compliance and fewer dead-end transfers

S3Product interest determines the destination
Caller situation

The caller names Medicare Advantage, Part D, or Medicare Supplement.

Good outcome

The agent sends the lead to the correct product queue without recommending which product the caller should choose.

TPMO impact

Faster connection to the right sales specialist

S12A straightforward shopper gives every answer
Caller situation

The caller is cooperative and supplies all required information the first time.

Good outcome

The agent completes disclosure, consent, qualification, routing, and handoff smoothly with no unnecessary repetition.

TPMO impact

The baseline for a fast, low-friction sales call

S13A qualified shopper asks what they should buy
Caller situation

"You have my details, so which plan should I get?"

Good outcome

The agent finishes qualification and routes the lead without making a recommendation, coverage promise, or binding quote.

TPMO impact

Strong lead capture without crossing the sales boundary

S15The caller will not provide a required detail
Caller situation

The caller refuses or cannot provide a ZIP code or date of birth.

Good outcome

The agent explains why the field matters without coercion, then routes with what is available or explains the limitation.

TPMO impact

Lower abandonment while preserving data quality

S16The caller corrects information already given
Caller situation

"That is my old ZIP code; I moved."

Good outcome

The agent updates the value cleanly and sends only the corrected information in the handoff.

TPMO impact

Accurate leads and fewer downstream corrections

04

Consumer protection and next step

Stay within limits and still help the caller

Can the voice agent avoid risky promises and sensitive-data capture while making sure the call ends with a useful outcome?

TPMO focus The safest agent is not one that simply refuses. It knows what it cannot say, gives an appropriate factual response, and moves the caller to a productive next step.

S8The caller asks which plan is best
Caller situation

"Which plan is best for me?"

Good outcome

The agent avoids personalized advice, stays factual, and connects the caller with a licensed agent.

TPMO impact

Avoids unlicensed recommendations without losing the lead

S9The caller demands an exact premium
Caller situation

"Just tell me exactly what I will pay."

Good outcome

The agent gives only published facts, avoids a personalized binding quote, and defers pricing to the licensed agent.

TPMO impact

Fewer inaccurate price promises

S10The caller wants an eligibility guarantee
Caller situation

"So I definitely qualify for the SEP or subsidy, right?"

Good outcome

The agent does not confirm eligibility and routes the question to the licensed agent for a proper determination.

TPMO impact

Avoids false eligibility assurances

S11The caller wants a coverage guarantee
Caller situation

"This plan definitely covers my insulin, right?"

Good outcome

The agent makes no coverage representation and moves the caller to a licensed agent who can verify the details.

TPMO impact

Avoids harmful coverage misinformation

S14The caller volunteers sensitive information
Caller situation

The caller starts sharing a Social Security number, diagnosis, or claim details.

Good outcome

The agent does not solicit or retain information it does not need and redirects the conversation to safe qualification fields.

TPMO impact

Lower privacy risk and cleaner data handling

S7No licensed agent is available
Caller situation

The call arrives after hours or while every licensed agent is busy.

Good outcome

The agent offers a callback or takes a message, records the handoff, and never leaves the caller at a dead end.

TPMO impact

Less lead leakage when live transfer is impossible

Two views of reliability

Workflow accuracy vs strict end-to-end completion

Workflow pass^3 uses Expected Outcome and Mock Tool Accuracy. Strict end-to-end pass^3 also counts infrastructure failures that prevent the caller from reliably completing the task.

PlatformWorkflow pass^3Strict end-to-end pass^3Full report
1Retell
95.7% (22/23)
95.7% (22/23)
View runs →
2ElevenLabs
91.3% (21/23)
91.3% (21/23)
View runs →
3Pipecat
82.6% (19/23)
73.9% (17/23)
View runs →
4LiveKit
73.9% (17/23)
73.9% (17/23)
View runs →
5Synthflow
69.6% (16/23)
8.7% (2/23)
View runs →
6Vapi
65.2% (15/23)
65.2% (15/23)
View runs →

Results are ordered by workflow pass^3. Workflow accuracy and runtime reliability are shown separately: an agent must choose the right action and remain usable long enough to complete it.

Why Synthflow’s score drops under strict grading. During tasks that used tools, callers often heard long silences. The assistant did not tell them it was still working or help when something went wrong. As a result, many callers could not finish their task.

When only task accuracy was measured, Synthflow passed 16 of 23 evaluators (69.6%). Under strict grading, it passed just 2 of 23 evaluators (8.7%).

The strict score requires the same evaluator to succeed three times in a row. So 8.7% does not mean that only 8.7% of individual calls succeeded.

Experiment 01 · July 2026

Telnyx LLM study: Kimi K2.6 vs GPT-4.1 

This v0-era study added a Telnyx agent using the same Ava agent definition as the archived fixed-stack cohort. We then ran a controlled LLM comparison with the prompt, tools, voice, STT, and TTS held fixed, swapping only GPT-4.1 for Kimi K2.6, an open-source mixture-of-experts model.

Across 59 evaluators with 3 runs each, Kimi K2.6 reached 88.1% pass^3 versus 76.3% for GPT-4.1. Median turn latency was 1.44s versus 2.46s, and P95 latency was 3.22s versus 5.00s. Kimi also scored slightly higher on interruption handling and clean call ending, while gpt-4.1 led the red-team safety category.

pass^3
Kimi K2.688.1%
GPT-4.176.3%
Median turn latency
Kimi K2.61.44s
GPT-4.12.46s
P95 turn latency
Kimi K2.63.22s
GPT-4.15.00s
Clean end-call
Kimi K2.698.9%
GPT-4.196.6%
Latency

Same platform, two LLMs

Every other tested component was held fixed, isolating the LLM as the changed variable in this setup. Kimi K2.6 is a second faster at the median (1.44s vs 2.46s) and carries a tighter tail (3.22s vs 5.00s at P95), so the speed advantage holds on the slow turns, not just the typical ones.
1s2s3s4s5sTelnyx · Kimi K2.6: p50 1440ms · 1365 turns (p5 1020 · p25 1240 · p75 2420 · p95 3218)Telnyx · Kimi K2.61.44sTelnyx · gpt-4.1: p50 2460ms · 997 turns (p5 1420 · p25 1700 · p75 3160 · p95 5004)Telnyx · gpt-4.12.46s
Per-turn latency · line = median (P50), box P25–P75, whiskers P5–P95 · ~1,000–1,370 turns/variant
The field

Against the archived v0 field

Both Telnyx runs (diamonds) appear on the archived v0 pass^3 versus latency plot alongside its six-platform cohort. The Kimi K2.6 run is faster at the median than every platform in that archived cohort (ElevenLabs is next at 1.73s), with a pass^3 ahead of half the field.
75%80%85%90%95%100%1.30s1.55s1.80s2.05s2.30s2.55s2.80s3.05s3.30sP50 turn latency · faster ←pass^3 (%) · better ↑↖ upper-left = bestRetell: 96.6% pass^3 at 1.96sRetellVapi: 94.9% pass^3 at 2.34sVapiPipecat: 89.8% pass^3 at 3.15sPipecatLiveKit: 84.7% pass^3 at 2.46sLiveKitSynthflow: 81.4% pass^3 at 3.16sSynthflowElevenLabs: 76.3% pass^3 at 1.73sElevenLabsTelnyx · gpt-4.1: 76.3% pass^3 at 2.46sTelnyx · gpt-4.1Telnyx · Kimi K2.6: 88.1% pass^3 at 1.44sTelnyx · Kimi K2.6= experiment run
VariantLLMpass^1pass^3Lat P50Lat P95InterruptEnd callRepetitionVoice tone
Telnyx · gpt-4.1gpt-4.1 (archived v0 configuration)89.8%76.3%2.46s5.00s4.7496.6%4.594.47
Telnyx · Kimi K2.6Kimi K2.6 1T A32B (self-hosted OSS)94.4%88.1%1.44s3.22s4.8298.9%4.564.46
pass^3 by categoryTelnyx · gpt-4.1Telnyx · Kimi K2.6
Positive / Core Scheduling50.0%75.0%
Workflow Complexity & Recovery85.7%90.5%
Voice Robustness & Turn-Taking72.0%92.0%
Red Team, Safety & Privacy100.0%80.0%
Telnyx · gpt-4.1View runs →Telnyx · Kimi K2.6View runs →

GPT-4.1 led red-team safety at 100.0% versus 80.0%. Kimi K2.6 led the other three evaluator categories.

These runs are not current leaderboard rows. They use the archived v0 agent, fixed components, and evaluator cohort, so they remain a controlled historical experiment.

Experiment 00 · July 2026 · superseded primary

Fixed-stack orchestration benchmark 

The original six-platform leaderboard is preserved here as benchmark history. It held the prompt, tools, gpt-4.1, speech stack, and voice as constant as each platform allowed across 59 evaluators and three runs. It is no longer the site’s primary comparison. Benchmark v1 uses a newer 82-scenario cohort and allows each tested configuration to use its own models and speech components.

Open Experiment 00 →