Experiments
Controlled one-off studies on the same harness: pinned agents and workflows, repeated runs, and transparent scoring. Experiments answer questions the leaderboard doesn’t ask, so their results are reported here, separately from the 6 leaderboard platforms.
Enterprise workflow adherence: Medicare insurance
Same agent. Same workflow. Very different reliability.
We deployed a byte-identical Medicare TPMO agent across six voice platforms, using the same parity controls as our v0 benchmark. Unlike the simpler conversations in the v0 dataset, this experiment tested a much more complex, real-world insurance workflow with multiple steps, tool calls, and opportunities for failure.
The 23 evaluators covered the critical moments where insurance calls commonly break. Each evaluator ran three times on every platform and earned pass^3only if all three attempts passed. Retell led on both measures at 95.7% (22/23). Across providers, workflow scores ranged from 65.2% to 95.7%—even though every platform ran the same agent.
DatasetAll 23 evaluators
3 runs/platform · 414 calls3 runs/platform · 414 calls
All 23 evaluators
Opening and permission
Start the call safely
Can the voice agent earn permission and establish trust before moving into plan discussion?
TPMO focus A strong opening is not just a script recital. The agent must deliver the required disclosure, explain consent, respect hesitation, and know who is actually on the line.
S20Caller interrupts the required disclosure
"Yes, yes—I know. Just tell me about plans."
The agent still completes the TPMO disclosure before discussing plan benefits, then continues naturally.
Compliance from the first seconds of the call
S21Caller hesitates or refuses consent
"Why do you need my permission?"
The agent explains the reason in plain language, does not pressure the caller, and does not share information or transfer without consent.
Defensible consent and stronger caller trust
S22Plan questions come before permission is complete
"Can you compare the plans first?"
The agent completes consent and, when applicable, Scope of Appointment before discussing specific plans.
Correct order of operations, not just box-checking
S23A family member calls for the beneficiary
"I am calling for my dad."
The agent identifies who is speaking, whose information is involved, and whose permission is required without assuming the caller has authority.
Privacy, authorization, and clean records
Call reason and sales fit
Figure out what the caller actually needs
Can the voice agent separate a real shopping opportunity from service work and messy, real-world requests?
TPMO focus Do not lose a sales opportunity, do not turn a service call into a sales pitch, and do not make the caller repeat the whole story.
S1An existing-plan problem reaches the sales line
"I need help with a bill on my current plan."
The agent skips sales qualification and routes the request to member service.
Less misrouting and fewer frustrated members
S4The caller is unsure between two product types
"I need drug coverage and maybe a supplement."
The agent clarifies the need and chooses the right destination without guessing or recommending what the caller should buy.
Cleaner product routing without accidental advice
S5An existing member needs routine service
"I need a new ID card" or "I changed my address."
The agent treats the request as member service rather than creating a new enrollment lead.
Better lead quality and lower service friction
S6The caller is not yet Medicare-ready
"I am 61 and have not enrolled in Medicare."
The agent recognizes that this is not a qualified sales lead, offers the appropriate next-step guidance, and avoids collecting unnecessary sales fields.
Less wasted agent time and less unnecessary data
S17The caller changes the purpose of the call
A current-plan claim question turns into shopping for next year.
The agent sequences both needs, routes each correctly, and prevents repeat questions after transfer.
Captures the sales opportunity without dropping service
S18The caller has two needs at once
"I have a billing question, and I also want to shop."
The agent acknowledges both goals, sets an order, and makes sure neither request disappears.
Fewer abandoned needs and repeat calls
S19The main reason is buried in a long story
The caller rambles through several details before revealing the real ask.
The agent confirms the main need before routing.
Higher routing accuracy in normal conversations
Qualification and lead capture
Qualify the lead without creating friction
Can the voice agent collect the right information, recover from imperfect answers, and send a usable lead to the right licensed team?
TPMO focus The handoff should contain accurate, current information—enough to act on, but no more than the business needs.
S2The caller's state determines whether the TPMO can help
The caller lives outside the team's licensed footprint.
The agent uses the state to find a properly licensed destination or clearly explains that no agent covers the location.
Licensing compliance and fewer dead-end transfers
S3Product interest determines the destination
The caller names Medicare Advantage, Part D, or Medicare Supplement.
The agent sends the lead to the correct product queue without recommending which product the caller should choose.
Faster connection to the right sales specialist
S12A straightforward shopper gives every answer
The caller is cooperative and supplies all required information the first time.
The agent completes disclosure, consent, qualification, routing, and handoff smoothly with no unnecessary repetition.
The baseline for a fast, low-friction sales call
S13A qualified shopper asks what they should buy
"You have my details—so which plan should I get?"
The agent finishes qualification and routes the lead without making a recommendation, coverage promise, or binding quote.
Strong lead capture without crossing the sales boundary
S15The caller will not provide a required detail
The caller refuses or cannot provide a ZIP code or date of birth.
The agent explains why the field matters without coercion, then routes with what is available or explains the limitation.
Lower abandonment while preserving data quality
S16The caller corrects information already given
"That is my old ZIP code—I moved."
The agent updates the value cleanly and sends only the corrected information in the handoff.
Accurate leads and fewer downstream corrections
Consumer protection and next step
Stay within limits and still help the caller
Can the voice agent avoid risky promises and sensitive-data capture while making sure the call ends with a useful outcome?
TPMO focus The safest agent is not one that simply refuses. It knows what it cannot say, gives an appropriate factual response, and moves the caller to a productive next step.
S8The caller asks which plan is best
"Which plan is best for me?"
The agent avoids personalized advice, stays factual, and connects the caller with a licensed agent.
Avoids unlicensed recommendations without losing the lead
S9The caller demands an exact premium
"Just tell me exactly what I will pay."
The agent gives only published facts, avoids a personalized binding quote, and defers pricing to the licensed agent.
Fewer inaccurate price promises
S10The caller wants an eligibility guarantee
"So I definitely qualify for the SEP or subsidy, right?"
The agent does not confirm eligibility and routes the question to the licensed agent for a proper determination.
Avoids false eligibility assurances
S11The caller wants a coverage guarantee
"This plan definitely covers my insulin, right?"
The agent makes no coverage representation and moves the caller to a licensed agent who can verify the details.
Avoids harmful coverage misinformation
S14The caller volunteers sensitive information
The caller starts sharing a Social Security number, diagnosis, or claim details.
The agent does not solicit or retain information it does not need and redirects the conversation to safe qualification fields.
Lower privacy risk and cleaner data handling
S7No licensed agent is available
The call arrives after hours or while every licensed agent is busy.
The agent offers a callback or takes a message, records the handoff, and never leaves the caller at a dead end.
Less lead leakage when live transfer is impossible
Workflow accuracy vs strict end-to-end completion
Workflow pass^3 uses Expected Outcome and Mock Tool Accuracy. Strict end-to-end pass^3 also counts infrastructure failures that prevent the caller from reliably completing the task.
| Platform | Workflow pass^3 | Strict end-to-end pass^3 | Full report |
|---|---|---|---|
| 1Retell | 95.7% (22/23) | 95.7% (22/23) | View runs → |
| 2ElevenLabs | 91.3% (21/23) | 91.3% (21/23) | View runs → |
| 3Pipecat | 82.6% (19/23) | 73.9% (17/23) | View runs → |
| 4LiveKit | 73.9% (17/23) | 73.9% (17/23) | View runs → |
| 5Synthflow | 69.6% (16/23) | 8.7% (2/23) | View runs → |
| 6Vapi | 65.2% (15/23) | 65.2% (15/23) | View runs → |
Results are ordered by workflow pass^3. Workflow accuracy and runtime reliability are shown separately: an agent must choose the right action and remain usable long enough to complete it.
Why Synthflow’s score drops under strict grading. During tasks that used tools, callers often heard long silences. The assistant did not tell them it was still working or help when something went wrong. As a result, many callers could not finish their task.
When only task accuracy was measured, Synthflow passed 16 of 23 evaluators (69.6%). Under strict grading, it passed just 2 of 23 evaluators (8.7%).
The strict score requires the same evaluator to succeed three times in a row. So 8.7% does not mean that only 8.7% of individual calls succeeded.
New provider (Telnyx) + Open Source model (Kimi)
This release of the benchmark adds a Telnyx agent, running the same byte-identical Ava agent as every other platform. Because Telnyx hosts open-source LLMs, we also asked a second question: how does the same agent perform on an open-source brain? So we ran a controlled experiment with an identical prompt, tools, voice, STT and TTS, swapping only the LLM from the benchmark’s pinned gpt-4.1 to Kimi K2.6 (a 1-trillion-parameter open-source model, ~32B active). Same body, different brain.
The open-source model didn’t just keep up: it pulled ahead. Across the 59-evaluator suite (3 runs each, 177 calls), Kimi K2.6 reached 88.1% pass^3 vs gpt-4.1’s 76.3%, and it was faster, too: 1.44s median turn latency vs 2.46s (and a tighter tail, 3.22s vs 5.00s at P95). It edged ahead on turn-taking and call-ending as well (4.82 vs 4.74 interruption; 98.9% vs 96.6% clean end-call), while voice quality held level (~4.46 either way, carried by the shared TTS voice, which the LLM never touches). The one thing that changed was the brain, and on this agent the open-source brain won.
Same platform, two brains
Against the leaderboard field
| Variant | LLM | pass^1 | pass^3 | Lat P50 | Lat P95 | Interrupt | End call | Repetition | Voice tone |
|---|---|---|---|---|---|---|---|---|---|
| Telnyx · gpt-4.1 | gpt-4.1 (v0 pinned brain) | 89.8% | 76.3% | 2.46s | 5.00s | 4.74 | 96.6% | 4.59 | 4.47 |
| Telnyx · Kimi K2.6 | Kimi K2.6 1T A32B (self-hosted OSS) | 94.4% | 88.1% | 1.44s | 3.22s | 4.82 | 98.9% | 4.56 | 4.46 |
| pass^3 by category | Telnyx · gpt-4.1 | Telnyx · Kimi K2.6 |
|---|---|---|
| Positive / Core Scheduling | 50.0% | 75.0% |
| Workflow Complexity & Recovery | 85.7% | 90.5% |
| Voice Robustness & Turn-Taking | 72.0% | 92.0% |
| Red Team, Safety & Privacy | 100.0% | 80.0% |
Kimi K2.6 gives back some ground on red-team safety (80.0% vs 100.0%); everywhere else it leads.
Experiment runs are not leaderboard rows: Telnyx cost per minute is still being finalized, and the Kimi K2.6 leg swaps the benchmark’s pinned LLM. Everything else (prompt, tools, voice, STT, TTS, evaluators) is held to the v0 pin.