Every leaderboard tells you which model is best. None tell you when to switch. gptbased joins LMArena Elo rankings with live OpenRouter pricing, refreshed daily, across 500+ models in Text, WebDev, Image and Video. Sort by value, not just raw score - so you can see that GLM 5.2 outscores Grok 4.5 at 73% less per token. Find the model you use, then get emailed the moment a cheaper or stronger one ships.
I kept having the same argument with myself: is the model I'm paying for still the right one?
Rankings and prices both move every week, but they live in different places. LMArena tells you what's good. OpenRouter tells you what it costs. Nobody joins them, so "should I switch?" stays a gut call.
So I built gptbased. It pulls LMArena Elo and live OpenRouter pricing daily across 500+ models, sorts by value instead of raw score, and emails you when a cheaper or stronger model ships in the category you care about.
The build was mostly plumbing: reconciling two sources that name models differently, handling models that appear on one and not the other, and tracking rank history so a debut can be distinguished from a climb.
This week made the case better than I could. Grok 4.5 debuted straight into the WebDev top 3, knocking Claude Opus 4.8 off the podium. But GLM 5.2 still outscores Grok at 73% less per token. If you weren't watching the board that day, you'd never know either thing happened.
Report
How are you handling cases where a model's pricing changes mid-day on OpenRouter or where providers offer different rates depending on context window size?
Report
would love to see a filter for latency alongside the price and score columns since cost per token only matters so much if the response takes 10 seconds to start
Report
It would be great if you could add a filter for context window size and latency alongside the cost and score columns. Sometimes the cheapest strong model isn't usable for me because it chokes on long documents or responds too slowly, so having those constraints baked into the sorting would make recommendations actually actionable for my workflow.
Report
the sort-by-value angle is such a practical call out, especially highlighting that GLM 5.2 stat right on the homepage. most leaderboards feel academic, this one actually respects that we're paying for tokens.
Report
Finally a leaderboard that tells me when to switch, not just who is winning. The value-first sorting made it obvious I have been overpaying for a slightly better score for months.
Report
the value-per-token sort is genuinely useful, had no idea GLM was hanging with the big names at that price
Report
No reviews yetBe the first to leave a review for gptbased
How are you handling cases where a model's pricing changes mid-day on OpenRouter or where providers offer different rates depending on context window size?
would love to see a filter for latency alongside the price and score columns since cost per token only matters so much if the response takes 10 seconds to start
It would be great if you could add a filter for context window size and latency alongside the cost and score columns. Sometimes the cheapest strong model isn't usable for me because it chokes on long documents or responds too slowly, so having those constraints baked into the sorting would make recommendations actually actionable for my workflow.
the sort-by-value angle is such a practical call out, especially highlighting that GLM 5.2 stat right on the homepage. most leaderboards feel academic, this one actually respects that we're paying for tokens.
Finally a leaderboard that tells me when to switch, not just who is winning. The value-first sorting made it obvious I have been overpaying for a slightly better score for months.
the value-per-token sort is genuinely useful, had no idea GLM was hanging with the big names at that price