Check whether a task rejects broken and incomplete candidates while accepting the reference fix. Broken baseline. Plausible shortcut. Reference fix. The web workbench replays committed reports. It does not run untrusted candidate code in the browser.
A test suite can accept the intended solution and still accept a plausible shortcut. That is why I built RepoGauntlet with three control candidates: the broken baseline, an incomplete fix and the reference fix.
Task packs contain frozen source, allowed edit paths, public and held-out tests, explicit commands and a canonical evidence report. The included packs cover Python, Java, Rust, C++ and TypeScript. A Python quickstart runs locally without an API key or Docker.
The web workbench lets you inspect committed reports. It does not pretend to run untrusted code in your browser. The runner also distinguishes candidate failures from infrastructure problems, which matters if scores become training rewards.
Free and MIT licensed. I would value feedback from benchmark authors on the incomplete-fix control: which shortcuts have your test suites accidentally accepted?
Try it: https://shi1720.github.io/repo-g...
Source: https://github.com/shi1720/repo-...
Built by Shivam Gupta. More work and contact: https://shivamgupta.web.app/
LinkedIn: https://www.linkedin.com/in/shiv...