Multi-Agent Arena lets you play social strategy games like Diplomacy, Poker, Risk, etc. against frontier LLMs. They're playing the same way as you and you get to scheme, plan, and strategize against them while they get to do the same to you. You can play today on https://olamarena.com/ for free. We use the arena to evaluate models in multi-agent simulations on capability and behavior. You can read our full blog post here on the early evals we're seeing: https://olamlabs.ai/research/diplomacy
Most AI agent evaluations and benchmarks look at performance in static, programmatic environments like coding tests.
The real world is messy, and as agents more and more capable it's important that we evaluate them in complex social environments emblematic of real-world scenarios.
So we built Multi-Agent Arena, a way to evaluate models in dynamic social scenarios while humans get to have fun playing social strategy games against talking AIs that normal programmatic games can't recreate - a win-win.
You can play matches in social strategy games against AIs that will remember, scheme, and betray the same way that you're planning on doing.
My name is Om Buddhdev, co-founder & CEO of Olam Labs (YC S26). I'm fascinated by multi-agent simulations and their applications to build the best (and most critical) evaluations as AIs only get more autonomous and capable in the real world. I have a diverse background, having skipped college to be one of the best Valorant players in the world, doing math research, working as a Staff Engineer in consumer AI, and other things. More on me at https://www.sensho.xyz/