An independent, reproducible benchmark of TypeSafe's Jev against openai/gpt-6-luna (cheap LLM) and openai/gpt-6-astra (frontier LLM). Tests accuracy, calibration, latency, and cost. All code is open source — run it yourself and verify the results.
I built this because I wanted to know how TypeSafe's Jev actually performs compared to top LLMs — not from marketing claims, but from reproducible tests. The results surprised me in several ways (Jev beats frontier models on calibration but not latency). All code is open source so you can run the same benchmark yourself and check my numbers.