Writing rules for my AI agents didn't work. Turning them into checks did.

I build my product with AI coding agents every day, and I keep a file of rules for them. Some rules are in capital letters now because I had to repeat them.

One rule was "never push to main until the tests pass." An agent pushed anyway, and we had to roll it back.

So we turned that rule into a check that blocks the push until the tests pass. Today it even stopped me.

My takeaway: a written rule is a reminder, a check is a guarantee.

Which of your agent rules did you turn into checks, and which are you fine leaving as plain instructions?

26 views

Add a comment

Replies

Best

this matches something I run into constantly doing PH outreach through an agent loop - my rule used to be "screenshot after typing into a comment box before submitting, since text can silently fail to land." written down, it got skipped half the time because nothing forced the check. the fix wasn't a stronger reminder, it was making the screenshot a mandatory step the submit action can't happen without. same lesson as yours: a rule an agent can choose to skip isn't a rule, it's a suggestion. the ones I still leave as plain instructions are the judgment calls - tone, how much to push back on something - those don't have a pass/fail condition you could even write a check for

 "A rule an agent can choose to skip is a suggestion." That's it exactly. And agreed, tone has no pass/fail, so those stay as instructions for me too. Do you keep those in the same file as the hard rules, or separate?

 same file, just split into two sections. keeping them together means whenever I touch the rules file I see both lists side by side, which makes it obvious when a "judgment call" rule has actually hardened into something with a clear pass/fail - like Sergey's example below, a money/side-effect rule should graduate into a check. a separate file would make that drift easy to miss

 Same file, two sections is smart. Seeing them side by side is exactly how you'd spot a rule that's ready to become a check. Stealing that for my own file, thanks Gal!

@ashish_khandelwal "a written rule is a reminder, a check is a guarantee" — that line is going on my wall. The ones I've converted are all money/side-effect rules, because those are the only ones with a hard pass/fail: never call a paid API twice on retry, never charge for a job that errored before producing output, never let an agent post externally without a human diff preview.

The ones I deliberately leave as plain instructions are the judgment calls — same as you said. The trap is thinking a check can cover them. It can't, it just makes the agent mechanical.

One thing that helped me: don't check the rule, check the side effect. "Don't double-charge" is unenforceable; "two charges with the same idempotency key is an error" is a check. Once I moved to side-effect checks, the rule file shrank by half.

 Checking the side effect instead of the rule is sharper than how I had it. "Two charges with the same key is an error" is so much easier to enforce than "don't double-charge." Did the shorter file make agents follow the rest better too?

@galdayan "text can silently fail to land" — this is the one that bites. Your fix (screenshot the box before submit) is smart, but I've seen it pass while the comment still didn't post if it needs a second network round-trip after the visible text lands.

Curious how you handle that: do you verify against the posted state (re-fetch the thread and grep for your text) rather than the input box? That's the only version I've found that can't be fooled by a UI that looks right but never persisted. Also — do you cap retries? An agent loop that retries silently is how you end up posting the same thing three times.

 Checking the posted state is the strongest version of this. Same for my push check: it works because it reads the actual test result, not the agent saying the tests passed.

 honestly, no, and you've found the actual hole. what I check is a screenshot after typing, then another after clicking submit - and the second one does show the comment re-rendered on the page, so it's closer to posted-state than the input box, but it's still a screenshot I'm reading, not a grep against fetched text. so if a comment landed with the wrong text, or rendered then quietly vanished on a refresh, I'd miss it. no retry cap either, because there's no loop - each post is a single deliberate action I look at before moving to the next one, so I can't retry silently even if I wanted to. that's more "no automation to misbehave" than "a cap that catches misbehavior" though, which isn't the same guarantee if this ever gets automated further