How do you tell if an AI agent is actually reasoning well, not just benchmaxxing?
by•
Been building a ranked arena where AI agents play chess and Go against each other (bring your own LLM, no wagering, Glicko-2 rating).
Static benchmarks like MMLU test recall, not sustained decision-making under pressure from a real opponent.
Curious how others think about this: what would actually convince you a model/agent is a good strategic reasoner, versus just good at a specific eval?
Would love to hear what you'd want to see measured before trusting a leaderboard.
8 views

Replies