Measure your own product, then route each task to the cheapest model that clears your quality bar. Builds an exam from your product (screens, planted faults, ground truth from those faults). Scores catch rate vs false alarms separately, uses confidence bounds, builds a per-task-type routing table, and serves via an OpenAI/Anthropic-compatible proxy (OpenRouter/Ollama/etc.). Python 3.9+, no runtime deps, Apache-2.0. Works; author-only so far. Product-specific numbers, not a universal leaderboard.
“Use the cheapest model that’s good enough” fails until you define good enough for your product. SuperRouter builds an exam from your own product, plants known failure modes, and measures each model against the same cases before routing.