Launching today
AgentDiff
Trajectory regression testing for AI agents
0 followers
Trajectory regression testing for AI agents
0 followers
Traditional evals test outputs, missing when agents loop on tools or burn 50% more tokens. AgentDiff fixes this by recording agent runs as DAGs and diffing execution trajectories against golden baselines directly in CI. It catches tool loops, cost spikes, and latency regressions before merge evaluating how the model arrived at the answer, not just what it returned. Framework-agnostic with adapters for OpenAI Agents, Langfuse, LangSmith, and OpenInference.
AgentDiff Reviews
Pros
Cons
Reviews