Agents hand you five working versions. How do you pick the one that ships?

byโ€ข

Any coding agent gives you a working version of a feature in minutes. Ask again and you get another one, also working, slightly different. Working stopped being the filter.

My rule is simple: I ship the version with the fewest moving parts, because I'm the one debugging it at 2am. Speed of writing means nothing against speed of fixing.

What's your filter? Curious if anyone has a better rule than simplest one wins

335 views

Add a comment

Replies

Best

Great rule. Iโ€™d add one more strict filter to the 'simplest wins' approach: zero unnecessary dependencies. Often, an agent will give you a working version that quietly introduces a new package or deviates from the framework's native tools just to get the job done quickly. I always pick the version that strictly adheres to our core stack and doesn't bloat the architecture. Speed of fixing > speed of writing, 100%.

My deciding factor is usually the test coverage. a simple implementation without confidence checks can still hurt you later. The winner is the one I can safely change not just the one that works today.

ย Test coverage is a different kind of filter, it measures whether you can change the thing, not just whether it runs. Do you let the agent write those tests too, or is that the part you keep for yourself?

me choosing the simplest version has saved many future fixes. have you tried tracking bug counts after release to see whether simpler implementations actually reduce maintenance over time?

ย No real tracking, my scale is too small for bug stats to mean much. My signal is cruder, which versions I end up reopening. The simple ones mostly stay closed

Great question, this is exactly the kind of problem more teams are running into as agents get better at generating options. My honest answer: pick the version that's easiest to debug 6 months from now, not the one that's cleverest today.

Most of the "wow, this is beautiful" versions I've picked from AI-generated batches turned into maintenance nightmares. The ones that felt boring but had clear structure and readable logic aged way better. Ship-ability isn't just about now, it's about the version you'd still be willing to inherit later.

ย The agent optimizes for right now. Everyone in this thread is optimizing for whoever opens the file in six months. That second person is the one the agent doesn't have, which is why "inherit later" is the filter and the agent can't apply it for you.

ย Great point Alexander, honestly this reframes the whole conversation. Most reviewers are unknowingly acting on behalf of a future teammate the agent doesn't even model. That's why "inherit later" holds up as a filter even when the surface-level output looks impressive.

The agent is optimizing for local elegance while we're the ones responsible for long-term coherence. That mismatch is where a lot of AI-generated code quietly fails, not because it's wrong, but because it wasn't written with the future maintainer in mind. Really appreciate you sharpening this idea further.

Would stronger tests make a more complex version worth shipping instead?

Simplest-one-wins is close to my rule, but I sharpened it after enough 2am debugging of agent output: I ship the version whose failure I can predict. Fewest moving parts usually correlates, but not always. Sometimes the shortest version leans on clever implicit behavior, and that is exactly the one that bites. So my filter is: read each version and ask, when this breaks, will the error land near the cause? The version where failures surface loudly and locally wins, even if it is a few lines longer. Agents made writing code free, so the whole game moved to what happens six weeks later.

ย Predictable failure is a better word for what I meant by simple. The extra lines are cheap insurance against the six-weeks-later bill

Cheap insurance against the six-weeks-later bill is the phrase I was missing. Stealing it, with credit. Looking forward to the results post.

and what would your approach be if you are non-tech?
other agents in the loop?

ย If you can't read the code, judge behavior. Spend ten minutes trying to break each version and pick the one whose failures are loud and easy to describe. A reviewer agent helps a little, but it checks the code. What bites you later is behavior, and you still own that part.

Completely agree with the 'fewest moving parts' rule. 2 AM debugging sessions are brutal.

My secondary filter is 'Standardization over Cleverness.' Agents love to introduce overly complex frameworks, weird library dependencies, or 'clever' custom logic to solve simple problems. I always pick the version that sticks strictly to the standard, native features of the stack I'm using. If an agent introduces three new dependencies just to build a simple utility feature, it's an immediate reject. Fewer dependencies = faster fixing

great point i usually check if the code is easy to test if i cannot write unit tests for it quickly the simplest version is definitely the one i keep

simplest-one-wins is right and mine is close to it. i add one more filter: pick the version i can hand to one other person and have them say "yeah this did the thing" without me narrating what they're looking at.

when the agent hands me five, four of them require me standing next to the reader explaining what the buttons do. the fifth one lands solo. that one ships. the others are demos disguised as features.

123
Next