Measure your own product, then route each task to the cheapest model that clears your quality bar. Builds an exam from your product (screens, planted faults, ground truth from those faults). Scores catch rate vs false alarms separately, uses confidence bounds, builds a per-task-type routing table, and serves via an OpenAI/Anthropic-compatible proxy (OpenRouter/Ollama/etc.). Python 3.9+, no runtime deps, Apache-2.0. Works; author-only so far. Product-specific numbers, not a universal leaderboard.
A desktop overlay that sees your screen, answers questions about it, and points at the control you need. Ask a question; Handrail sends your screenshot and question to your chosen vision model, then draws arrows at real controls. It watches multi-step checklists as you work. Electron, Apache-2.0; local threads, OS-keychain API key, no account/telemetry; web search off by default. BYO OpenRouter/OpenAI/Anthropic key. Working daily by maker; author-only so far.