AgentVet.ai - AI agent reviews, benchmarks & live token pricing

by
Which AI Agent can I use to do my work today? AgentVet.ai combines AI agent discovery, community reviews, benchmark testing, and live token pricing into one platform. Compare agents side-by-side, verify performance claims, and discover which AI tools are truly worth using.

Add a comment

Replies

Best

Hey Product Hunt 👋 Saeed here from .

I kept seeing the same pattern: AI agent looks incredible in a demo, falls apart in production. No one was tracking real-world performance.

So we built AgentVet, where real users verdict AI agents, not marketing teams. Then we back it up with hard benchmark data.

This relaunch adds AgentVet Lab: independent benchmarks where we run agents through real tasks and score them across 5 dimensions. Vendors can get certified. Buyers can compare side-by-side.

Community reviews from real users

• Independent agent benchmarks & rankings

Head-to-head comparisons

Live token pricing

• ⚡ Token Efficiency Audit: independently verify your tool's actual token savings (not just claims)

106 agents listed. Growing daily.

A huge thank you to everyone who gave feedback and tested the platform — you shaped what this became. Keep it coming. We’re just getting started 🙏

How are you running the benchmark tests, like do you execute the same tasks across agents with a fixed prompt and dataset?

 Clever! Yes, completely standardized. Every agent gets identical instructions and identical input files (same CSVs, same JSONs, same logs). We don't give agents a pre-written prompt, they read the task spec and work autonomously, which is the point. We're testing the agent as a system, not just the underlying model.

Correctness is graded deterministically against a fixed answer key (exact field values, tolerances for floats). Qualitative dimensions , completeness and output quality go to two independent LLM judges (Claude + ChatGPT) who score blind, and we flag any gap > 1.5 points for manual review.

Speed is measured against a category-specific reference time, and first-try success is tracked separately so you can see if an agent needed hand-holding to get there.

congrats for your launch

 thanks Naim!