Launching today

Prefactor
Evaluate your AI Agents in real-time
643 followers
Evaluate your AI Agents in real-time
643 followers
Most agents pass their evals and fail in production. Prefactor is the evaluation layer that closes the gap. We score every agent run in real time, surface quality regressions and drift as they happen, and show engineering teams exactly how their agents are performing at scale. Built for the teams shipping agents to customers.















Drift is the one nobody talks about until it's already cost them a bad week. Real detection here feels like the right instinct.
Prefactor
@ado_audu Absolutely does. When a model can change underneath you without a notification, how can you possibly contain drift. Being able to sit inside the agent is absolute gold dust.
Prefactor
@ado_audu this guy gets it! How are you monitoring this drift in production currently, and whats preventing you from adopting something like Prefactor?
Prefactor
Friends, Josh here. AI Engineer @ Prefactor. I built the agent instance tracking that scores quality and flags data risk while your agent runs, plus the SDKs and docs that get teams from zero to live without guessing.
Here's the thing I keep coming back to. You can't have confidence in something you can't see. An agent in production is a bucking bronco; it'll throw you the moment you stop paying attention. I've talked to too many teams who deployed, watched it work for a week, then realised they had no idea what it was actually doing. No visibility, no guardrails, just hope.
That's the piece Prefactor owns: the quality and data risk scoring that runs in-process, flagging misbehaviour mid-gallop rather than in a trace you read after the damage is done.
If you've got agents live right now, how do you actually know they're behaving? Or are you just holding on and hoping?
The line "most agents pass their evals and fail in production" is going to resonate with more teams than you probably realize. I've said some version of that sentence out loud in at least three postmortems this year.
Prefactor
@billy_boy Hey Billy, thanks for comment. You're right. I find that most companies don't even realise the difference.
I can tell someone what their evaluation approach is before they've even opened their mouth because they are so standardised. Sampled runs, LLM as judge and golden dataset. All are point in time.
They have their value in dev for sure. But in prod, you need a suite of tools which reduce the need for tokens and non-deterministic outcomes. Being in the agent run/ live, allows us to do that.
Would love to get your feedback on the tool and learn if it would have solved some of those PMs earlier in the year
Prefactor
@billy_boy It's such a big problem to solve -- so many projects failing, and it really doesn't have to be that way. We're trying to put an end to those post-mortems!
How does Prefactor decide what parts of the code need attention without creating unnecessary changes? Congrats @ethan_lee8 & team!
Prefactor
@ethan_lee8 @hamza_afzal_butt We dont claim to know every agent and how they work. We give you the tools to track things as they go wrong, whether thats through meta data, llm as judge evals etc.
We help the unnecessary change bit by version controlling each agent, enabling different environments so you can track the problems as they occur in dev and when you're ready push them to prod.
I spent months building internal scripts to catch exactly this kind of drift. Would've saved me a lot of late nights to just plug into something like this instead.
Prefactor
@malka_parveen many such cases unfortunately... Which is the reason we have built Prefactor, to save your engineering team the time and burden of building, and maintaining a platform like this. We should compare notes at some point, would love to know if we missed anything you'd consider essential for keeping your agents on the straight and narrow.
scoring the run live instead of charting it three days later is the part that actually matters. catching a bad agent after it already shipped the damage is just reporting.
Prefactor
@alex_watson2110 Totally agree -- evaluation needs to be part of the control loop for an agent to really be trustworthy.
TimeToCoda
Working in high-trust environments, I don’t think the future is humans approving every AI action. It’s agents operating independently 95% of the time, with enough visibility and confidence that when they drift outside the guardrails, a human can intervene before it matters. Looks like you’re tackling that problem head on. Congrats @matt_doughty @simon_russell1 and team!
Send me my 1M spans to give it a solid rev up! ;)
Prefactor
@simon_russell1 @emotf Done. You've got to set up an agent first and then you get them in your account.
You were one of our first conversations wayyyy back. LFG!
TimeToCoda
@simon_russell1 @matt_doughty Big fan of your work. Following your journey all the way... I love seeing how this idea has progressed, and now PH launch 🤘
Prefactor
@matt_doughty @emotf Thanks! Those 1M are yours if you sign up now :D