Coval helps developers build reliable voice and chat agents faster with seamless simulation and evals. Create custom metrics, run 1000s of scenarios, trace workflows and integrate with CI/CD pipelines for actionable insights and peak agent performance.
Every time I take a Waymo, I'm struck by how safe and reliable it feels. As any AI engineer knows, that level of trust comes from rigorous evaluation frameworks implemented from day one. It's incredible that the same evaluation best practices that helped build these self-driving cars are now accessible to everyone, thanks to someone who actually built Waymo's evaluation infrastructure! This is exactly what we need.
Coval is tackling one of the biggest challenges in conversational AI: ensuring agents perform reliably in complex, real-world scenarios. The ability to simulate thousands of edge cases, integrate seamlessly with CI/CD, and provide actionable insights makes this a must-have tool for developers pushing the limits of voice and chat technology. Congratulations to Brooke and the team for bringing such an impactful solution to market!
@chase_martens thank you for sharing your thoughts! We're really pushing for the mission of having reliable AI agents across all industries and we're excited to share this with the world!
Report
As someone working extensively with AI agents, Coval is exactly what the developer community needs right now. The ability to run thousands of test scenarios and create custom metrics is game-changing for quality assurance. The CI/CD pipeline integration is particularly impressive - it makes continuous testing of AI agents as seamless as traditional software testing.
What really stands out is the comprehensive workflow tracing feature. It's invaluable for debugging and optimizing agent performance. This tool could significantly reduce the time from development to deployment while ensuring higher reliability of AI agents.
Excited to see how this will evolve and shape the future of AI agent development! 🚀
Hi Ethan, so great to hear! The workflow tracing feature has been also getting the most love from our customers so far, and it's been a game-changer for debugging conversational AI.
Thank you for sharing your perspective today! 🚀
Report
Congrats Brooke! This is a game-changer for teams building and scaling conversational AI. The ability to integrate thousands of simulations directly into the CI/CD pipeline is particularly exciting—it mirrors best practices from industries like autonomous vehicles and software development, but adapted to the unique challenges of AI agents.
I also love the focus on actionable insights. Debugging and monitoring are notoriously difficult in conversational AI, especially when dealing with LLMs. Coval's approach to custom metrics like LLM-as-a-Judge or tool call tracking feels incredibly forward-thinking.
As someone working on building reliable AI tools for research scientists, I can definitely see how Coval would fit into our stack to ensure our agents remain robust and trustworthy as we scale.
Unsloth
Coval
Coval
Coval
Coval
Coval