Agents in prod
by•
We have had a lot of fun since the PH launch and making it to number 1. Such a great experience and everyone who got involved and helped us I am insanely grateful to.
As we move on from PH into helping teams evaluate their agents in real time, I'm looking to understand how everyone is doing it. Are you using static evals and running a hit and hope approach? Are you using off the shelf agent evals from OSS? Are you using a tool?
I'd love to know so we can improve what we have to make it the world class solution you don't have to think twice about when building your agents.
Matt
42 views


Replies
The “hit and hope” approach feels painfully familiar 😂. It’s easy to test agent before launch and forget that its behavior can drift afterward.