Coval helps developers build reliable voice and chat agents faster with seamless simulation and evals. Create custom metrics, run 1000s of scenarios, trace workflows and integrate with CI/CD pipelines for actionable insights and peak agent performance.
Hi Product Hunt Community 🐱👋
I’m Brooke, Founder of Coval!
Today, we’re excited to launch Coval, a platform that transforms how you test, debug, and monitor voice and chat agents. Simulate thousands of scenarios from a few test cases. You create the prompts, we simulate environments to test your agents from all directions.
👉 Why did I build Coval?
Before founding Coval, I led the evaluation job infrastructure team at Waymo, building simulation tools that tested every code change to ensure the Waymo Driver improved with every iteration. This shift from manual testing on racetracks to scalable, automated simulation transformed autonomous vehicles from early prototypes into reliable systems now navigating the streets of San Francisco.
Today, AI agents face similar challenges: promising prototypes often hit reliability roadblocks as they scale. Drawing on my Waymo experience, I built Coval to bring automated simulation and evaluation to AI agents, helping teams move faster and deliver reliable, real-world performance.
Coval’s mission? To ensure AI agents can be trusted with critical tasks, just as simulation helped unlock the potential of self-driving cars. It’s a tool built by developers, for developers—designed to save time, increase confidence, and eliminate the headaches of conversational AI development.
❓What Problems Does Coval Solve?
👉 Manual Testing Wastes Time
Manually calling or chatting with agents is inefficient. Coval integrates into your CI/CD pipeline, running 1000s of simulations automatically with each prompt change. This saves time, increases test coverage, and boosts confidence in production performance.
👉 Debugging is a Nightmare
Fixing one issue often breaks something else. Coval eliminates this frustration by providing actionable insights into agent workflows, tracking metrics for each simulation to help you pinpoint and resolve problems effectively.
👉 Production Monitoring is Hard
Identifying the root cause of agent mistakes in production can be a nightmare. Coval’s monitoring offers immediate, actionable insights into custom metrics like LLM-as-a-Judge or tool calls, making it easier to ensure reliable performance.
❓Why Us?
Our team brings deep experience in LLM evaluations at Berkeley & Stanford, building distributed systems for Fortune 500 companies, and crafting intuitive user interfaces.
🚀 Special Launch Offer
As part of our Product Hunt launch, enjoy a free 2-week trial with personalized onboarding. We’ll help you set up custom metrics, run your first evaluations, and get the most out of Coval.
👉 Start Your Free Trial: https://www.coval.dev
👉 Check out our Docs: https://docs.coval.dev/overview
👉 Book a Demo Call: https://cal.com/bnhopkins/demo
Excited to help you ship reliable AI agents faster!
P.S. Drop by the comments—we’d love your feedback!
@brooke_hopkins3 Coval feels like a step up for building smarter voice and chat agents. The seamless simulation and custom metrics are intriguing, how deeply can it analyze edge cases? I’m curious if it uncovers patterns that might otherwise go unnoticed. Sounds like a tool that pushes beyond the basics!
Report
@brooke_hopkins3 Seeing all the 100 hour weeks you put into this, I’m so proud of you and the team! Congrats!
@tonyabracadabra Great question Tony! Yes, we definitely support developers with spotting edge cases. We do this by offering - for example - tool call evaluations where we check for hard-to-spot function calls by your agent. It has definitely helped our customers with debugging incorrect tool calls with very precise identification of error spots.
Additionally, we offer topic analysis to catch new arising topics in conversations; as well as workflow monitoring where we tell you where and how often your agents go off the beaten path.
Make sure to check out our product, we're offering a free trial and you can get started on running your own evals very easily --> coval.dev
It's exciting to see this launch. Congratulations, Brooke and team!
When we all first started building voice agents on top of the new generation of LLMs, the hard problems were things like phrase endpointing, interruption handling, and squeezing latency out of all the steps in the processing pipeline.
Now we have really good frameworks -- both open source and proprietary -- that make it possible to build flexible, capable voice agents that perform well "most of the time." But most of the time isn't good enough for lots of use cases.
To get to "performs well all the time" we need great tooling for evals (testing), real-time observability, and performance metrics. Coval's tools are a huge step in this direction and we're all going to benefit from their work.
@kwindla Thank you! You’ve hit the nail on the head.
Getting AI agents to perform reliably every time is the next frontier. At Coval, we’re focused on helping teams close that gap with powerful eval tools and real-time monitoring. Excited to be part of this journey and looking forward to seeing how we can all push the boundaries of voice agent performance together with Daily!
Report
Congrats @brooke_hopkins3 & the team! I can see your platform becoming an integral part of the tech stack for developing and deploying AI agents. I've seen people outsourcing testing to cheaper countries, your solution by far beats the manual way - I'll be happy to spread the word about it. I am also just curious what's the word play behind the Coval name?
@katka_sabo thank you soooo much!
Coval draws inspiration from Sofya Kovalevskaya, the first woman to earn a Ph.D. in mathematics!
Report
@brooke_hopkins3 this is super helpful! I wanted to mention your startup to someone the other day but I couldn't recall the name - now I'll remember - best of lucj
@mwseibel@ashitvora we integrate with any tech stack! It is agnostic to tech stack because we call your API or phone number
Report
This is pretty cool! Seems like its an agent that chat with another agent through different use cases and use an agent to do eval? Would i have the capability to manually evaluate and see sessions?
@parodyyut exactly! You're able to check each session once the simulation is performend dive into each transcript.
Here's an example how this would look for a tool call evaluation: https://www.loom.com/share/d6f22...
Report
@brooke_hopkins3 congrats on the launch! Where are you seeing the most traction especially with voice agents? Any non-obvious uses you’re excited about?
@tjinsta thank you! Our voice agent customers have been loving our workflow analysis. It helps them track all the different paths their agents take during a conversation, revealing many scenarios where agents go off track.
Additionally, we’ve recently introduced sentiment analysis, which provides in-depth metrics for various sentiments like “Frustrated😠”, “Angry😡”, and “Confused😕”. This allows you to monitor users in production and easily debug any faulty agent responses, ensuring a smoother experience for everyone.
@brooke_hopkins3 Cool product! I am currently struggling to evaluate hours of calls at Focus Buddy, especially classify the calls into various buckets based on different criterias and analyze them together (essentially like SQL for databases but for agentic interactions). Do you have a feature for that?
Hi @adnan_sherif1! This is an interesting challenge – I think Coval might be just the right match for you. Do you wanna schedule a quick call to discuss how to approach your problem on our platform? here's the link: https://cal.com/bnhopkins/demo
Coval
ScaryStories Live
Coval
Coval
Daily.co
Coval
Coval
IntroJoy
Coval
Coval
Coval
Focus Buddy (YC S24)
Coval