Have you ever had AI write technically perfect code that broke real-world requirements?
by•
We had an AI agent refactor our database submission tables to fix load times. The pull request looked super clean and performance metrics improved instantly so everything seemed fine at first.
A few days later we realized it stripped out essential audit columns because nothing in the database constraints explicitly flagged them as required. It optimized strictly for raw execution speed over domain context.
What system or checks do you rely on to keep coding agents from making assumptions about administrative workflow rules?
19 views
Replies
Yes, and I think the pattern is that an agent only respects what's written down somewhere it can read. Audit columns are required by a regulator or a contract, not by the schema, so to the agent they just look like dead weight.
What helped us was writing those reasons into the repo — a short note next to the migration saying why a column exists, and a test that fails if it disappears. If the only place a rule lives is in someone's head, the agent will delete it eventually.
Agree with HongSJ about writing the reason next to the rule. We do exactly that, and the part that bit us next was the test itself.
Most of our rules of that kind are held by a test that scans the source: every signed-in route sits behind the onboarding gate, every error handler that shows a message to the user catches our refusal type and not the database failure, every heading font weight is preloaded. Those tests share a failure no agent will flag for you. If the scan matches nothing, it finds zero violations and passes. The first draft of the font test did exactly that, because of a typo in its pattern, and it reported a clean codebase while examining nothing.
So each of them now asserts how much it looked at before it asserts what it found: the error handler scan fails unless it examined more than forty catch blocks, the font scan unless it saw more than eight templates. A guard that cannot see what it guards is worse than none, because it is the reason nobody checks by hand.
The other habit is a real figure in the test instead of a tidy one. Our payment check compared amounts in minor units, which is right for dollars and wrong for yen, where Stripe takes whole yen. An invoice of 108250 minor units charges 1082, and 1082 converted back is 108200, fifty short, so the check would refuse a payment the buyer had already made. Every line was correct. The test that holds it uses those exact yen numbers now, because round dollar fixtures are where that assumption hides.
Toone
Yeah been there. I started making agents run a second pass just for the rules that live outside the code, because clean tests can still miss the real workflow. Did you add those audit columns to a test after?