Cekura builds testing and monitoring tools for voice agents, and runs these benchmarks itself. Because we sell into the same market we measure, every benchmark is built so that you can check it rather than take our word for it.
Last updated
Every model or platform in a benchmark runs the same scenarios, recordings, prompts, tools and scoring. Only the thing being compared changes.
Places are read straight off the measured results. Ties are shown as ties, and figures that can't be compared fairly are marked rather than ranked.
Test cases, agent definitions and benchmark code are public on GitHub, and the speech-to-speech and workflow pages link to the recorded calls behind their numbers. The licensed Ocular recordings are the one exception.
If a figure looks wrong, open an issue on the repository. We rerun or correct it and the page's last updated date changes with the fix.
Founding Engineer, Cekura
Forward Deployed Engineer, Cekura
Co-founder & CTO, Cekura
Results are published under CC BY 4.0, so you can reuse them with credit. The data behind the speech benchmarks can be downloaded as JSON.
Cekura Bench (2026). Voice AI Benchmarks. Cekura. https://benchmarks.cekura.ai
@misc{cekurabench2026,
title = {Cekura Bench: Voice AI Benchmarks},
author = {Dileep Chagam and Luis Ojeda and Shashij Gupta},
year = {2026},
url = {https://benchmarks.cekura.ai},
note = {Cekura}
}