Ahead of Product Hunt we have open sourced Prefactor Evals, No LLM Required. Apache 2.0, public repo.
Most agent failure is behavioural: a silent loop, a step that failed, a run that never finished, a task that took four times the usual work. All of that is checkable in code, deterministically, for nothing. There is no model client anywhere in the project and there never will be, so it cannot cost you a token by construction.
I really like the idea of tracking how the agent is working, as it is performing actions. Being more proactive is better than finding out something went wrong several days ago. It's not to be overlooked that the interruption isn't just killing the agent run, but handing it to a human too. I am going to have to play around with this and my own agents.
Prefactor
@philnash Thanks for the comment Phil. Would be happy to share some of the examples Ive built too, I'm building an eval triaging tool. 3 phases.
First phase: Tokenless Evals/ technical reviews
Second phase: pattern management against prior activity
Third phase: LLM as judge
All live. Let me know where you get to.
Prefactor
@philnash Thanks Phil, exactly! Killing a bad run is only half of it, the handoff to a human is where the actual value sits, otherwise you have just built a more expensive way to fail.
Would love to hear what breaks when you try it on your own agents, that is the feedback we learn the most from.
What are you running them on at the moment?
Congratulations on the launch! Runtime enforcement is a really interesting direction. One challenge I've been thinking about is that agents are evolving quickly, so evaluation itself needs to evolve as well? Looking forward to seeing where the platform goes.
Voquill
I like that this is built around production instead of just benchmarks. That's where things usually get interesting. Congrats on the launch!
Prefactor
@henry_habib Thanks Henry, and thanks for the comment. Benchmarks are great in controlled envs. Agents don't operate in hermetically sealed containers. They exist in the wild. We needed to meet that challenge.
Voquill
@matt_doughty Exactly. That's where you find the issues you'd never catch in a controlled environment.
Prefactor
@henry_habib Thank you! Yes, production is where things tend to get real pretty quickly. Understanding how your benchmarks/evals translate to the real world is super important.
Prefactor
@henry_habib thanks for the comment, Henry! We agree, an agent demo on a developers laptop is one thing, but once it hits a production environment that where the fun really begins. Are you deploying anything into production currently?
Voquill
@joshgillies Couldn't agree more. Yeah, hoping to get our first deployment live soon.
Prefactor
@henry_habib great to hear! Let us know how you go getting it going, and any issues we're here to help!
Prefactor
@henry_habib Thanks for the support Henry and yes 100% agreed. Evals and testing shouldn't stop the moment it hits prod.
Are you building much agents?
Voquill
@ethan_lee8 Yeah, working on an app right now.
Prefactor
@henry_habib Nice whats the app?
Intent by Upflowy
Fantastic product and team, very excited to see them launch here, congratulations on the successful launch of the product!
Prefactor
@matthew_browne1 Thanks Matt. It means a lot coming from you! Thanks for all of the support. To the moon!!!!!!
FuseBase
we're all out here shipping agents with our eyes closed and calling it deployment 😅 congrats on a great launch @ethan_lee8 @matt_doughty
Prefactor
@matt_doughty @kate_ramakaieva thanks so much Kate for your help!
Prefactor
Novu
Great product, and very much in need these days!
Prefactor
@tomer_barnea1 thanks, Tomer! We believe it's an essential platform for anyone building production agents. How are you managing your agents at scale currently?
Prefactor
@tomer_barnea1 Thank you Tomer. I owe you a callback.
Prefactor
@tomer_barnea1 Thanks for the support Tomer!
Awesome product team!
Prefactor
@ethan_lee14 Thanks Ethan! Appreciate the support!
Prefactor