Launching today

Prefactor
Evaluate your AI Agents in real-time
410 followers
Evaluate your AI Agents in real-time
410 followers
Most agents pass their evals and fail in production. Prefactor is the evaluation layer that closes the gap. We score every agent run in real time, surface quality regressions and drift as they happen, and show engineering teams exactly how their agents are performing at scale. Built for the teams shipping agents to customers.















How does Prefactor decide what parts of the code need attention without creating unnecessary changes? Congrats @ethan_lee8 & team!
Prefactor
@ethan_lee8 @hamza_afzal_butt We dont claim to know every agent and how they work. We give you the tools to track things as they go wrong, whether thats through meta data, llm as judge evals etc.
We help the unnecessary change bit by version controlling each agent, enabling different environments so you can track the problems as they occur in dev and when you're ready push them to prod.
I spent months building internal scripts to catch exactly this kind of drift. Would've saved me a lot of late nights to just plug into something like this instead.
Prefactor
@malka_parveen many such cases unfortunately... Which is the reason we have built Prefactor, to save your engineering team the time and burden of building, and maintaining a platform like this. We should compare notes at some point, would love to know if we missed anything you'd consider essential for keeping your agents on the straight and narrow.
TimeToCoda
Working in high-trust environments, I don’t think the future is humans approving every AI action. It’s agents operating independently 95% of the time, with enough visibility and confidence that when they drift outside the guardrails, a human can intervene before it matters. Looks like you’re tackling that problem head on. Congrats @matt_doughty @simon_russell1 and team!
Send me my 1M spans to give it a solid rev up! ;)
Prefactor
@simon_russell1 @emotf Done. You've got to set up an agent first and then you get them in your account.
You were one of our first conversations wayyyy back. LFG!
TimeToCoda
@simon_russell1 @matt_doughty Big fan of your work. Following your journey all the way... I love seeing how this idea has progressed, and now PH launch 🤘
Intent by Upflowy
Fantastic product and team, very excited to see them launch here, congratulations on the successful launch of the product!
Prefactor
@matthew_browne1 Thanks Matt. It means a lot coming from you! Thanks for all of the support. To the moon!!!!!!
Voquill
I like that this is built around production instead of just benchmarks. That's where things usually get interesting. Congrats on the launch!
Prefactor
@henry_habib Thanks Henry, and thanks for the comment. Benchmarks are great in controlled envs. Agents don't operate in hermetically sealed containers. They exist in the wild. We needed to meet that challenge.
Prefactor
@henry_habib Thank you! Yes, production is where things tend to get real pretty quickly. Understanding how your benchmarks/evals translate to the real world is super important.
Prefactor
@henry_habib thanks for the comment, Henry! We agree, an agent demo on a developers laptop is one thing, but once it hits a production environment that where the fun really begins. Are you deploying anything into production currently?