In 2025, I was digging into attack vectors in the AI-for-AI space, and something didn't add up: every "solution" I found could be broken with a fragmented, multi-step attack. That sent me down a rabbit hole.
I dug into the internal guardrail architectures of open-source LLMs. For a while I assumed I just hadn't found the right product yet surely someone had solved this. So I started talking to governance and AI security experts. They showed me what enterprise vendors were shipping. The problem was always the same: risk detection stayed locked to the rules defined at setup. Nothing self-corrected. Nothing adapted across departments or use cases. The system either blocked everything (killing productivity) or someone had to manually retune the rule set constantly (not sustainable at scale).