Launching today
Redline AI

Redline AI

Attack-test your AI agents and grade what they did

2 followers

Most agent evals grade the final answer. Redline grades the behaviour the server's own ledger of what the agent called, not its account of it. One task runs against Claude Code, Codex and the agent in your own repo, each in a clean container with identical tools. On top sit 16 security packs: 11,204 adversarial cases delivered the way real attacks arrive hidden in a support ticket, a document footer, a tool result. Your agent connects with redline dev and never leaves your machine.

Redline AI Reviews

Reviews