Playground - Earn $100K+ in weekly rewards for hacking AI agents.

Every challenge is a live and open-source AI agent guarding a secret - with its system prompt published for you to read. Talk it past its own defenses. Land the most approved breaks in a week and win $100K+ in rewards. Free to play, no account needed. New challenge every Monday.

Add a comment

Replies

Best

publishing the system prompt and still winning is the honest way to run this. who judges an "approved" break, a human review queue or another model scoring the transcript?

a human! 🫡

publishing the full system prompt and still making it hard is a genuinely good flex, most "jailbreak me" demos quietly rely on the prompt being secret. do the weekly challenges get harder over time as people share successful breaks publicly, or is each one designed independently so last week's winning technique doesn't just carry over?

On average, how many tokens does a test like this consume? I have an AI travel service, and we're in the final testing phase. I'm interested in using something like your platform, but I need to understand what it would cost us, including the token usage. Running 1,000+ tests could end up being quite expensive.

And one more question: does your service evaluate only the text responses? In our case, the responses include text, images, and an interactive map.

Is it simply going to notify me of the problem with this didn't work or is it going to be able to notify me of the exact input that caused the problem?