Run the same base model and task context in two separately generated lanes: one naked and one protected by Phalanx. Test direct jailbreaks, typed indirect prompt injection, benign lookalikes, and mixed multi-turn attacks under a public scoring protocol.
Hi Product Hunt, I'm Sergio, founder of Invarra and builder of Phalanx.
Phalanx vs the World is the public challenge surface for Phalanx, not the production runtime. It runs the same base model and task context in two separately generated lanes: one naked and one protected.
I built it because prompt injection is often discussed only as model behavior. In an application, it is also an authority problem: which text may instruct, and which text must remain data?
The arena accepts direct jailbreaks, typed indirect prompt injection through its exposed text inputs, benign lookalikes, and mixed multi-turn attempts. Completed sessions produce public-safe receipts under a visible scoring protocol. Unsafe raw outputs are withheld.
The full demo shows three different decisions: a DAN takeover closes, a benign security-themed coffee prompt passes, and an ambiguous cybersecurity request triggers clarification before declared attack intent closes the session.
Phalanx itself is an approximately 1.5 GB customer-hosted control layer for English written-text LLM applications. It governs direct written jailbreaks, mixed multi-turn instruction attacks, and typed indirect prompt injection before protected output or action is released. Its runtime sends no prompts, responses, documents, evidence, or customer information to Invarra.
I'd value attempts, questions about the protocol, and criticism of the boundary. The arena is public evidence under its protocol, not proof that bypass is impossible.