Paste a website URL or API endpoint and Agent Torture Lab throws a messy, adversarial customer at your live chatbot — the kind that asks for a refund that doesn't exist or pushes for information it shouldn't get. You get a report with transcript evidence, severity, and fixes. Free preview, $29 one-time for the full report (launch price). No SDK, no integration, no account.
Hey — maker here.
The reason I built this: someone builds you a chatbot — a freelancer, an agency, a no-code tool, or you wired it up yourself over a weekend — and it demos fine. Then it goes live and a real customer asks it something you never tested: "what's your refund policy for orders over 90 days," or "can you just give me the account details for order #4521," or something that isn't even a real question, just a customer being a customer. How do you actually know your bot won't invent a policy, leak something it shouldn't, or say something embarrassing on day one?
Agent Torture Lab answers that before your customers find out for you. Paste a URL or API endpoint, it runs an adversarial simulation against the live bot, and hands you a report: what broke, the actual transcript proving it, how bad it is, and what to fix. Deterministic checks are the ground truth; AI judging is a bounded second opinion that never overrides them, because I didn't want "the AI thinks it's fine" to be the whole product.
Preview is free. Full report is $29 one-time — no subscription, and I'm treating that price as a launch experiment, not gospel. If you run an agency and test bots for multiple clients, there's a $10/client rate through workspace credits once you top up.
Happy to answer anything about how the evaluation works, what it catches, or what it deliberately doesn't do yet (voice bots are next, not now). And if you want to see it on your own bot, paste the URL — I'll be watching this thread.