We used to treat every prod failure like a scavenger hunt check Cloud Build, check logs, copy stack trace, write ticket, ping Slack, hope someone sees it.
The failure wasn't expensive. The queue that didn't exist was.
Production reports itself - your agent opens the PR.
Errors, metric breaches, and user reports land as one triaged queue. A scheduled coding agent picks up each item and opens the fix as a reviewable pull request.