
Most AI agent tools treat launch day as the finish line. You upload your knowledge, publish, and the agent just runs. Forever, as far as the system's concerned. Nothing ever checks back in and asks if it still holds up.
That's a real problem once an agent is actually making someone money. Knowledge goes stale, prices change, policies get updated, and an agent keeps answering with the same confidence whether it's right or not. Revenue coming in doesn't tell you anything about whether the knowledge behind it expired. If anything it hides the problem, because nobody goes looking while the money's still showing up.
Kopai
Hey Omri, thanks for actually building something and being specific about what worked, that's more useful to us than a star rating alone.
On the eval gate, I appreciate you noticing that, it's been our main focus. Every agent gets scored across dimensions (system prompt quality, scope adherence, safety, plus KB/tool accuracy where relevant) before it can publish, and any edit re-triggers the check, so it's not publish-once-and-hope. The passing threshold is 70/100 and it's visible in the builder next to the per-dimension breakdown, not a black box.
About live vs frozen snapshot, it pulls the current published version, not a snapshot frozen at embed time. So when you update the agent, anyone who's embedded it sees the update too. Fair callout that this wasn't clear up front, we'll get it stated explicitly wherever embed is set up.
On technical categories being thin; Fair, and it tracks who's found us first, mostly consulting, coaching, and finance backgrounds. Getting devops/infra/security folks publishing is next, if you know anyone in that world sitting on knowledge worth packaging, send them our way!