Cekura enables Conversational AI teams to automate QA across the entire agent lifecycle—from pre-production simulation and evaluation to monitoring of production calls. We also support seamless integration into CI/CD pipelines, ensuring consistent quality and reliability at every stage of development and deployment.
This is the 5th launch from Cekura. View more

Cekura Bench
Launching today
Cekura Bench publishes voice AI benchmarks you can verify. Our new speech-to-speech benchmark tests 9 realtime voice models, including GPT Realtime 2.1, Gemini Live, Grok and Phonic, as complete phone agents on live calls: 82 scenarios, three runs each. Models are ranked on reliability, data accuracy, stalled calls, response time and cost, and every call transcript is public. Cekura Bench also covers voice agent benchmarks and STT benchmarks, with TTS benchmarks coming soon.






Free
Launch Team





Voice AI is moving fast, and having an independent way to compare different AI providers on real phone calls is incredibly useful. Really cool work, team!
Really comprehensive benchmark, Voice ai companies should definitely check this out 🚀🚀
Always glad to see more benchmarks for voice AI!