Cekura enables Conversational AI teams to automate QA across the entire agent lifecycle—from pre-production simulation and evaluation to monitoring of production calls. We also support seamless integration into CI/CD pipelines, ensuring consistent quality and reliability at every stage of development and deployment.
This is the 5th launch from Cekura. View more

Cekura Bench
Launching today
Cekura Bench publishes voice AI benchmarks you can verify. Our new speech-to-speech benchmark tests 9 realtime voice models, including GPT Realtime 2.1, Gemini Live, Grok and Phonic, as complete phone agents on live calls: 82 scenarios, three runs each. Models are ranked on reliability, data accuracy, stalled calls, response time and cost, and every call transcript is public. Cekura Bench also covers voice agent benchmarks and STT benchmarks, with TTS benchmarks coming soon.






Free
Launch Team





Publishing the call transcripts alongside the scores is the part I like here. Being able to look at what actually happened on a failed call makes the ranking much easier to judge. Congrats to the team!
Really interesting results 👀 The cascade baseline beating all the realtime models was definitely not what I expected.
Excited to see what the community thinks - and what models we should put through this next.
Every speech-to-speech model sounds great in a demo, so we put them all through the same real phone calls with real tool use and scored what actually happened. Excited to see how the rankings shift as new models drop 🚀
I like that they’re publishing the actual call transcripts with the scores. Way easier to see what went wrong instead of just taking the numbers at face value.
Most comprehensive benchmarks in voice and it's not even close. 👏
Glad to see this live, super useful..