trending

AI systems can appear trustworthy without understanding trust

One thing becoming increasingly obvious while testing AI systems:

AI can simulate trust surprisingly well.

AI systems can fail convincingly

One thing becoming increasingly clear while testing AI systems:

AI systems can follow instructions perfectly… and still fail

O

ne thing becoming increasingly clear while testing AI systems:

The weirdest AI attacks aren’t technical. They’re conversational.

One pattern we keep seeing while testing AI systems:

Many failures don t happen through traditional exploits.

AI confidence is becoming a security problem

One thing we ve noticed while testing AI systems:

Confidence is often mistaken for correctness.

A system can:

sound certain

What happens when AI gets conflicting instructions?

Tested an AI system with conflicting instructions.

It didn t fail.

Are you actually testing your AI… or just hoping it works?

Serious question for builders here:

Are you actively testing your AI systems for adversarial inputs?

Or mostly:

build test deploy

Real result: 217 vulnerabilities in 62 seconds

We ran a security test on an AI system.

How current AI security tools compare (and what’s missing)

We spent some time looking at existing AI security tools while building in this space.

There s a lot of strong work already out there.

But when you look closely, most tools fall into a few patterns:

Research-focused tools

AI agents are being deployed everywhere.

AI agents are being deployed everywhere.

But most of them are never tested for security.

Not because teams don t care.

Because the tools don t fit how developers actually build.