trending
•

1d ago

jev-bench - Open-source Jev benchmark: tested against frontier LLMs

An independent, reproducible benchmark of TypeSafe's Jev against openai/gpt-6-luna (cheap LLM) and openai/gpt-6-astra (frontier LLM). Tests accuracy, calibration, latency, and cost. All code is open source — run it yourself and verify the results.