Launched this week
jev-bench
Open-source Jev benchmark: tested against frontier LLMs
3 followers
Open-source Jev benchmark: tested against frontier LLMs
3 followers
An independent, reproducible benchmark of TypeSafe's Jev against openai/gpt-6-luna (cheap LLM) and openai/gpt-6-astra (frontier LLM). Tests accuracy, calibration, latency, and cost. All code is open source — run it yourself and verify the results.
jev-bench Reviews
Reviews