Connect your production agent to Latitude and get a quality score across outcome, reliability, cost, speed, and safety.
It updates daily, so you always know whether your agent is improving or getting worse.
Hey Product Hunt 👋 I’m César, founder of Latitude.
Over the past year, we’ve spoken with many teams running AI agents in production. Across all those conversations, one question kept coming up:
Is your agent getting better over time?
Most teams could answer for a few regression tests or isolated metrics, but not for the whole agent. Current eval systems have a coverage problem.
That’s why we built Agent Score.
Connect your production traces and Latitude automatically evaluates your agent across:
• Outcome
• Reliability
• Cost
• Speed
• Safety
It brings everything together into one quality score that updates daily, so you can see whether your agent is improving, what’s lowering its score, and the production sessions behind every result.
From there, Latitude helps you investigate recurring failures, track them over time, and send the evidence to your coding agent to open a fix.
We’d love for you to try Agent Score and tell us what you think. Any feedback is more than welcome 🙏
Can this detect issues that traditional logs and traces usually miss ?
Report
@heycesr and @paula_cavero Congratulations on the launch of #AgentScore 🌟✌️. Quick Q, Is there any coding involved in the association of the production agent or is it more passive integration?
the "one score across outcome/reliability/cost/speed/safety" framing is useful but also where I'd worry about a single number hiding a tradeoff, an agent that gets faster and cheaper by being slightly less reliable could show a flat or improving score if the weighting favors cost. do you expose the five sub-scores individually as well, or is the daily score genuinely a blend with no visibility into which dimension moved
Replies
Latitude
Hey Product Hunt 👋 I’m César, founder of Latitude.
Over the past year, we’ve spoken with many teams running AI agents in production. Across all those conversations, one question kept coming up:
Is your agent getting better over time?
Most teams could answer for a few regression tests or isolated metrics, but not for the whole agent. Current eval systems have a coverage problem.
That’s why we built Agent Score.
Connect your production traces and Latitude automatically evaluates your agent across:
• Outcome
• Reliability
• Cost
• Speed
• Safety
It brings everything together into one quality score that updates daily, so you can see whether your agent is improving, what’s lowering its score, and the production sessions behind every result.
From there, Latitude helps you investigate recurring failures, track them over time, and send the evidence to your coding agent to open a fix.
We’d love for you to try Agent Score and tell us what you think. Any feedback is more than welcome 🙏
Mastra
oss ftw!
Can this detect issues that traditional logs and traces usually miss ?
@heycesr and @paula_cavero Congratulations on the launch of #AgentScore 🌟✌️. Quick Q, Is there any coding involved in the association of the production agent or is it more passive integration?
EverTutor AI
one daily score is nice but i'd worry it hides stuff. like if cost goes down but outcome quality drops a bit, the total could still look fine
can we set our own weights for each category? for us safety matters way more than speed
congrats on the 9th launch btw, that's wild
Dial
the "one score across outcome/reliability/cost/speed/safety" framing is useful but also where I'd worry about a single number hiding a tradeoff, an agent that gets faster and cheaper by being slightly less reliable could show a flat or improving score if the weighting favors cost. do you expose the five sub-scores individually as well, or is the daily score genuinely a blend with no visibility into which dimension moved