Most annotation platforms already run honeypot tasks that part isn't new. What they don't do is model the trend across those checks per worker, per session, to tell "one bad guess" apart from a fatigue curve heading toward a bad batch. We sit on top of your existing gold-standard checks and read the slope, not just the score.