AgentVet.ai - AI agent reviews, benchmarks & live token pricing
by•
Which AI Agent can I use to do my work today?
AgentVet.ai combines AI agent discovery, community reviews, benchmark testing, and live token pricing into one platform. Compare agents side-by-side, verify performance claims, and discover which AI tools are truly worth using.

Replies
Hey Product Hunt 👋 Saeed here from AgentVet.ai.
I kept seeing the same pattern: AI agent looks incredible in a demo, falls apart in production. No one was tracking real-world performance.
So we built AgentVet, where real users verdict AI agents, not marketing teams. Then we back it up with hard benchmark data.
This relaunch adds AgentVet Lab: independent benchmarks where we run agents through real tasks and score them across 5 dimensions. Vendors can get certified. Buyers can compare side-by-side.
• Community reviews from real users
• Independent agent benchmarks & rankings
• Head-to-head comparisons
• Live token pricing
• ⚡ Token Efficiency Audit: independently verify your tool's actual token savings (not just claims)
106 agents listed. Growing daily.
A huge thank you to everyone who gave feedback and tested the platform — you shaped what this became. Keep it coming. We’re just getting started 🙏
mailX by mailwarm
How are you running the benchmark tests, like do you execute the same tasks across agents with a fixed prompt and dataset?
@manal_essalek1 Clever! Yes, completely standardized. Every agent gets identical task.md instructions and identical input files (same CSVs, same JSONs, same logs). We don't give agents a pre-written prompt, they read the task spec and work autonomously, which is the point. We're testing the agent as a system, not just the underlying model.
Correctness is graded deterministically against a fixed answer key (exact field values, tolerances for floats). Qualitative dimensions , completeness and output quality go to two independent LLM judges (Claude + ChatGPT) who score blind, and we flag any gap > 1.5 points for manual review.
Speed is measured against a category-specific reference time, and first-try success is tracked separately so you can see if an agent needed hand-holding to get there.
Mailwarm
congrats for your launch
@naimz thanks Naim!