Launched this week

BotCritic
Stress-test your AI chatbot before your customers do
30 followers
Stress-test your AI chatbot before your customers do
30 followers
Most chatbot testing tools check if your bot follows a script. BotCritic tests whether it survives a real customer — frustrated, confused, multilingual, or trying to break it. What makes it different: full conversation transcripts as evidence, a rewritten system prompt targeting the exact failures found, and the ability to re-test after fixes to prove they worked. We hold our own product to the same standard — we ran BotCritic on its own pitch and published the result, hallucination included.






honestly the idea of re-running tests after fixes to prove they actually worked is so underrated, most tools just give you a score and wave goodbye. one thing i'd love to see though is some kind of side-by-side diff view of the conversation transcripts before and after the prompt rewrite, basically so you can pinpoint exactly which turn changed and why. would make it way easier to trust the fix actually addressed the right thing and wasn't just a lucky side effect of rewording.
@demirkutlu12473 Really glad the re-test loop resonates — that was a deliberate bet, since “trust me it’s fixed” without proof felt hollow.
The diff view idea is excellent, and honestly a gap I hadn’t fully solved yet. Right now you can see before/after scores and the full transcripts separately, but you’re right that turn-by-turn diffing would make causality way clearer instead of just correlation. Adding this to the roadmap — appreciate you pointing at exactly the right thing to build next.
@demirkutlu12473 Also — thanks for the thoughtful feedback! First 10 people who comment, reply with your email and I’ll send you a free Pro-tier audit code.
love the idea of testing bots against real chaos instead of tidy scripts, and dogfooding it on your own pitch is a clever move. one thing i'd want to see is a way to share just the failing transcript snippets with my engineering team, maybe as a public permalink, so we can reproduce and discuss the exact failure without exporting whole logs.
@nturucal26502 Really like this — you’re right that exporting entire logs is friction nobody wants when all they need is the one broken exchange. A shareable permalink for a specific failing transcript snippet is a clean way to solve that.
Adding it to the roadmap alongside the diff-view idea from earlier in this thread — feels like both of these point at the same underlying need: making failures easy to isolate, share, and act on, not just score. Appreciate the sharp feedback.
@nturucal26502 Also — since you took the time for such sharp feedback, happy to send you a free Pro-tier audit code if you want to try it on your own bot. Just drop your email and I’ll get it right over.
The self-critique part genuinely impressed me, publishing the hallucination instead of hiding it shows real confidence in the product. Love that you ship proof alongside the tool itself.
@azat3yws Appreciate that — it felt important to hold ourselves to the same bar we’re asking bots to meet. If we’re going to tell people “don’t trust a demo, trust proof,” the tool itself has to walk that talk. Glad it landed that way.
@azat3yws Also — since you took the time for such sharp feedback, happy to send you a free Pro-tier audit code if you want to try it on your own bot. Just drop your email and I’ll get it right over.