Using AI models to judge AI-written replies: what worked and what didn't

by•

I'm launching here on Saturday. It's an iPhone keyboard that drafts replies from a screenshot of a chat. To check reply quality I use other AI models as judges, and I'd like to hear how others handle this.

My setup: one model writes the replies and models from other labs score them. A model never scores output from its own lab. When the judges disagree, I read that case myself.

Three things I didn't expect:

  • The first full run passed 38 of 52 cases against a 90% target. Part of that was my own test being wrong: one test screenshot had the two speakers swapped, so the model was being marked down for a chat that made no sense. I now read a sample of cases by eye before I trust any score.

  • Instructions I added to make the three drafts differ from each other made them worse. They read as over-written. Removing them improved the scores.

  • The judges are noisy. The same setup run twice can move the pass rate by about ten points, so small improvements mean nothing.

If you've used model judges for conversational output, do you trust them, or do you still read everything yourself?

9 views

Add a comment

Replies

Be the first to comment