We publish our AI's exam scores, including the one we failed. Should this be standard?

by

We build GeoPard AI, a reasoning assistant that plans farm prescriptions (fertilizer, seeding) from field data. Launching here Friday (tomorrow).

Before launch, one decision we made that I'd love this community's take on: we test the assistant on 503 CCA-style agronomy questions (the exam human crop advisors take) and publish every category score. 100% precision ag, 95.8% soil and water, 93.8% crop management, 92.9% nutrient management, and 77.5% pest management, the one we haven't cracked yet.

The question set stays private so it can't leak into training data and inflate future scores. But the results, including the weak one, are public.

Our logic: if an AI recommends what to put on someone's field, they deserve to know exactly where it's strong and where it isn't. Black-box confidence is how you lose farmers forever.

Curious what you think: would published benchmark scores change how much you trust an AI product? And is publishing your failures brave or just bad marketing?

2-min demo if you want context:

4 views

Add a comment

Replies

Be the first to comment