Can we trust AI agents with real business responsibilities?
by•
For those building or deploying agents, what checks do you use before giving them responsibility for real business work?
My concern is that completing a task successfully doesn’t necessarily mean the agent respected the business’s rules while doing it.
Do your evals cover situations where an agent completes the task but also takes an unauthorized action or presents invented information as fact? A concrete example of a test that caught this would be useful.
And if something slips through, what happens next? Who gets alerted, who can stop the agent, and how do you investigate or reverse its actions?
I’d like to hear what has worked in practice, including failures that changed how you test or deploy agents.
26 views
Replies