Cekura enables Conversational AI teams to automate QA across the entire agent lifecycle—from pre-production simulation and evaluation to monitoring of production calls. We also support seamless integration into CI/CD pipelines, ensuring consistent quality and reliability at every stage of development and deployment.
This is the 5th launch from Cekura. View more

Cekura Bench
Launching today
Cekura Bench publishes voice AI benchmarks you can verify. Our new speech-to-speech benchmark tests 9 realtime voice models, including GPT Realtime 2.1, Gemini Live, Grok and Phonic, as complete phone agents on live calls: 82 scenarios, three runs each. Models are ranked on reliability, data accuracy, stalled calls, response time and cost, and every call transcript is public. Cekura Bench also covers voice agent benchmarks and STT benchmarks, with TTS benchmarks coming soon.






Free
Launch Team





Exciting to actually quantify the explosive growth of voice AI 🚀
Finally, a serious benchmark for speech-to-speech models. No cherry-picked demos. Same agent, same prompt, same tools, and real end-to-end scenarios.
Excited to see what the community does with it. 🚀
Choosing a speech-to-speech model today can be difficult, with new models launching constantly and limited ways to compare them fairly. These results provides transparent, verifiable results to help teams evaluate real-world performance and choose with confidence. 🎯
Cool stuff!
Really exciting to see this live! The fact that these models were tested on real phone calls with the same agent, prompts, tools, and scenarios makes the benchmark especially interesting. Looking forward to seeing how the results evolve as more models are added 🚀
Love it! Congrats on the launch team
Happy to see it live 🚀