We've been building chatbots for years, and the same gap keeps showing up: dashboards tell you how much your bot did, not how well it did.
Conversations handled. Containment rate. CSAT. Response time.
All green. Meanwhile the bot spent last Tuesday confidently quoting a refund policy that changed in March to eleven people, none of whom filled out a survey.
That kind of thing only surfaces when someone happens to read transcripts, which mostly nobody has time to do.
Inquio