ARRM is a deterministic CI release gate for AI agents. It compares baseline and candidate releases, detects economically harmful regressions in measurable outcomes such as success, cost and latency, and returns a clear PASS, REVIEW or BLOCK decision before deployment. No LLM in the decision path.
Hi Product Hunt π
Iβm Antonio Orlando, founder of ZEODEC.
We built ARRM around a narrow problem in AI-agent development: a new release can still work technically while quietly becoming worse economically.
Success rate can fall. Cost per successful task can rise. Latency can increase. A business-critical outcome can regress β while the conventional test suite still passes.
ARRM compares a baseline release with a candidate and turns measurable outcomes into a deterministic release decision:
PASS β proceed
REVIEW β investigate before release
BLOCK β economically harmful regression detected
ARRM is deliberately not another general-purpose AI evaluation framework. There is no LLM in the decision path and no subjective judge.
Weβre opening a 14-day Early Access trial and looking for teams willing to run ARRM against one real baseline/candidate agent release.
The question Iβd most like Product Hunt users to help us answer is:
Would ARRM have stopped a real regression that your existing tests or eval stack would have allowed through?
Thanks for taking a look.
Antonio
Founder, ZEODEC