I use Jev as a judge for whether my AI coding agent's tests really passed!

by•

Hi, I'm Rikin, maker of dotpals (a small open-source pal that shows what your AI coding agents did). Most people seem to use Jev for lead-gen and booking calls. I'm using it for something different: as a judge for test results. My coding agent says "All tests pass. Done!" far too often. dotpals reads the test runner's own output first, and when the output and the exit code disagree, it sends Jev a short, redacted snippet and asks one question: "did these tests pass?"

It answers in about 0.2 seconds with a probability. In the first day it caught a real false "pass" (a run where the tests never actually ran) at 92% certainty, and it rated our own 148-test suite as passed, 98% sure.

Two questions for the Jev team and other users: Is there a better way to phrase the question for better calibration? We use thresholds of 0.7 (passed) and 0.3 (failed). Is there a documented rate limit or a recommended max input size for the free tier?

Happy to share the integration; it's open source. Credit to TypeSafe for the model.

12 views

Add a comment

Replies

Be the first to comment