Experiments
Controlled and historical studies using the Cekura evaluation harness, repeated runs, and transparent scoring. These studies answer questions the primary release does not, so their results are reported here, separately from the leaderboard of 7 configurations.
Enterprise workflow adherence: Medicare insurance
Same agent. Same workflow. Very different reliability.
We deployed the same Medicare TPMO agent across six voice platforms, using the controls from the archived v0 benchmark. Unlike the simpler conversations in that dataset, this experiment tested a more complex, real-world insurance workflow with multiple steps, tool calls, and opportunities for failure.
The 23 evaluators covered the critical moments where insurance calls commonly break. Each evaluator ran three times on every platform and earned pass^3 only if all three attempts passed. Retell led on both measures at 95.7% (22/23). Across providers, workflow scores ranged from 65.2% to 95.7%, even though every platform ran the same agent.
DatasetAll 23 evaluators
3 runs/platform · 414 calls3 runs/platform · 414 calls
All 23 evaluators
Opening and permission
Start the call safely
Can the voice agent earn permission and establish trust before moving into plan discussion?
TPMO focus A strong opening is not just a script recital. The agent must deliver the required disclosure, explain consent, respect hesitation, and know who is actually on the line.
S20Caller interrupts the required disclosure
"Yes, yes, I know. Just tell me about plans."
The agent still completes the TPMO disclosure before discussing plan benefits, then continues naturally.
Compliance from the first seconds of the call
S21Caller hesitates or refuses consent
"Why do you need my permission?"
The agent explains the reason in plain language, does not pressure the caller, and does not share information or transfer without consent.
Defensible consent and stronger caller trust
S22Plan questions come before permission is complete
"Can you compare the plans first?"
The agent completes consent and, when applicable, Scope of Appointment before discussing specific plans.
Correct order of operations, not just box-checking
S23A family member calls for the beneficiary
"I am calling for my dad."
The agent identifies who is speaking, whose information is involved, and whose permission is required without assuming the caller has authority.
Privacy, authorization, and clean records
Call reason and sales fit
Figure out what the caller actually needs
Can the voice agent separate a real shopping opportunity from service work and messy, real-world requests?
TPMO focus Do not lose a sales opportunity, do not turn a service call into a sales pitch, and do not make the caller repeat the whole story.
S1An existing-plan problem reaches the sales line
"I need help with a bill on my current plan."
The agent skips sales qualification and routes the request to member service.
Less misrouting and fewer frustrated members
S4The caller is unsure between two product types
"I need drug coverage and maybe a supplement."
The agent clarifies the need and chooses the right destination without guessing or recommending what the caller should buy.
Cleaner product routing without accidental advice
S5An existing member needs routine service
"I need a new ID card" or "I changed my address."
The agent treats the request as member service rather than creating a new enrollment lead.
Better lead quality and lower service friction
S6The caller is not yet Medicare-ready
"I am 61 and have not enrolled in Medicare."
The agent recognizes that this is not a qualified sales lead, offers the appropriate next-step guidance, and avoids collecting unnecessary sales fields.
Less wasted agent time and less unnecessary data
S17The caller changes the purpose of the call
A current-plan claim question turns into shopping for next year.
The agent sequences both needs, routes each correctly, and prevents repeat questions after transfer.
Captures the sales opportunity without dropping service
S18The caller has two needs at once
"I have a billing question, and I also want to shop."
The agent acknowledges both goals, sets an order, and makes sure neither request disappears.
Fewer abandoned needs and repeat calls
S19The main reason is buried in a long story
The caller rambles through several details before revealing the real ask.
The agent confirms the main need before routing.
Higher routing accuracy in normal conversations
Qualification and lead capture
Qualify the lead without creating friction
Can the voice agent collect the right information, recover from imperfect answers, and send a usable lead to the right licensed team?
TPMO focus The handoff should contain accurate, current information, enough to act on but no more than the business needs.
S2The caller's state determines whether the TPMO can help
The caller lives outside the team's licensed footprint.
The agent uses the state to find a properly licensed destination or clearly explains that no agent covers the location.
Licensing compliance and fewer dead-end transfers
S3Product interest determines the destination
The caller names Medicare Advantage, Part D, or Medicare Supplement.
The agent sends the lead to the correct product queue without recommending which product the caller should choose.
Faster connection to the right sales specialist
S12A straightforward shopper gives every answer
The caller is cooperative and supplies all required information the first time.
The agent completes disclosure, consent, qualification, routing, and handoff smoothly with no unnecessary repetition.
The baseline for a fast, low-friction sales call
S13A qualified shopper asks what they should buy
"You have my details, so which plan should I get?"
The agent finishes qualification and routes the lead without making a recommendation, coverage promise, or binding quote.
Strong lead capture without crossing the sales boundary
S15The caller will not provide a required detail
The caller refuses or cannot provide a ZIP code or date of birth.
The agent explains why the field matters without coercion, then routes with what is available or explains the limitation.
Lower abandonment while preserving data quality
S16The caller corrects information already given
"That is my old ZIP code; I moved."
The agent updates the value cleanly and sends only the corrected information in the handoff.
Accurate leads and fewer downstream corrections
Consumer protection and next step
Stay within limits and still help the caller
Can the voice agent avoid risky promises and sensitive-data capture while making sure the call ends with a useful outcome?
TPMO focus The safest agent is not one that simply refuses. It knows what it cannot say, gives an appropriate factual response, and moves the caller to a productive next step.
S8The caller asks which plan is best
"Which plan is best for me?"
The agent avoids personalized advice, stays factual, and connects the caller with a licensed agent.
Avoids unlicensed recommendations without losing the lead
S9The caller demands an exact premium
"Just tell me exactly what I will pay."
The agent gives only published facts, avoids a personalized binding quote, and defers pricing to the licensed agent.
Fewer inaccurate price promises
S10The caller wants an eligibility guarantee
"So I definitely qualify for the SEP or subsidy, right?"
The agent does not confirm eligibility and routes the question to the licensed agent for a proper determination.
Avoids false eligibility assurances
S11The caller wants a coverage guarantee
"This plan definitely covers my insulin, right?"
The agent makes no coverage representation and moves the caller to a licensed agent who can verify the details.
Avoids harmful coverage misinformation
S14The caller volunteers sensitive information
The caller starts sharing a Social Security number, diagnosis, or claim details.
The agent does not solicit or retain information it does not need and redirects the conversation to safe qualification fields.
Lower privacy risk and cleaner data handling
S7No licensed agent is available
The call arrives after hours or while every licensed agent is busy.
The agent offers a callback or takes a message, records the handoff, and never leaves the caller at a dead end.
Less lead leakage when live transfer is impossible
Workflow accuracy vs strict end-to-end completion
Workflow pass^3 uses Expected Outcome and Mock Tool Accuracy. Strict end-to-end pass^3 also counts infrastructure failures that prevent the caller from reliably completing the task.
| Platform | Workflow pass^3 | Strict end-to-end pass^3 | Full report |
|---|---|---|---|
| 1Retell | 95.7% (22/23) | 95.7% (22/23) | View runs → |
| 2ElevenLabs | 91.3% (21/23) | 91.3% (21/23) | View runs → |
| 3Pipecat | 82.6% (19/23) | 73.9% (17/23) | View runs → |
| 4LiveKit | 73.9% (17/23) | 73.9% (17/23) | View runs → |
| 5Synthflow | 69.6% (16/23) | 8.7% (2/23) | View runs → |
| 6Vapi | 65.2% (15/23) | 65.2% (15/23) | View runs → |
Results are ordered by workflow pass^3. Workflow accuracy and runtime reliability are shown separately: an agent must choose the right action and remain usable long enough to complete it.
Why Synthflow’s score drops under strict grading. During tasks that used tools, callers often heard long silences. The assistant did not tell them it was still working or help when something went wrong. As a result, many callers could not finish their task.
When only task accuracy was measured, Synthflow passed 16 of 23 evaluators (69.6%). Under strict grading, it passed just 2 of 23 evaluators (8.7%).
The strict score requires the same evaluator to succeed three times in a row. So 8.7% does not mean that only 8.7% of individual calls succeeded.
Telnyx LLM study: Kimi K2.6 vs GPT-4.1
This v0-era study added a Telnyx agent using the same Ava agent definition as the archived fixed-stack cohort. We then ran a controlled LLM comparison with the prompt, tools, voice, STT, and TTS held fixed, swapping only GPT-4.1 for Kimi K2.6, an open-source mixture-of-experts model.
Across 59 evaluators with 3 runs each, Kimi K2.6 reached 88.1% pass^3 versus 76.3% for GPT-4.1. Median turn latency was 1.44s versus 2.46s, and P95 latency was 3.22s versus 5.00s. Kimi also scored slightly higher on interruption handling and clean call ending, while gpt-4.1 led the red-team safety category.
Same platform, two LLMs
Against the archived v0 field
| Variant | LLM | pass^1 | pass^3 | Lat P50 | Lat P95 | Interrupt | End call | Repetition | Voice tone |
|---|---|---|---|---|---|---|---|---|---|
| Telnyx · gpt-4.1 | gpt-4.1 (archived v0 configuration) | 89.8% | 76.3% | 2.46s | 5.00s | 4.74 | 96.6% | 4.59 | 4.47 |
| Telnyx · Kimi K2.6 | Kimi K2.6 1T A32B (self-hosted OSS) | 94.4% | 88.1% | 1.44s | 3.22s | 4.82 | 98.9% | 4.56 | 4.46 |
| pass^3 by category | Telnyx · gpt-4.1 | Telnyx · Kimi K2.6 |
|---|---|---|
| Positive / Core Scheduling | 50.0% | 75.0% |
| Workflow Complexity & Recovery | 85.7% | 90.5% |
| Voice Robustness & Turn-Taking | 72.0% | 92.0% |
| Red Team, Safety & Privacy | 100.0% | 80.0% |
GPT-4.1 led red-team safety at 100.0% versus 80.0%. Kimi K2.6 led the other three evaluator categories.
These runs are not current leaderboard rows. They use the archived v0 agent, fixed components, and evaluator cohort, so they remain a controlled historical experiment.
Fixed-stack orchestration benchmark
The original six-platform leaderboard is preserved here as benchmark history. It held the prompt, tools, gpt-4.1, speech stack, and voice as constant as each platform allowed across 59 evaluators and three runs. It is no longer the site’s primary comparison. Benchmark v1 uses a newer 82-scenario cohort and allows each tested configuration to use its own models and speech components.
Open Experiment 00 →