Run the same base model and task context in two separately generated lanes: one naked and one protected by Phalanx. Test direct jailbreaks, typed indirect prompt injection, benign lookalikes, and mixed multi-turn attacks under a public scoring protocol.