I stopped asking an AI agent to “fix the bug” and started asking it to reproduce it first

by•

I noticed a small problem in my debugging workflow with AI agents.

When something broke, my first instinct was:

“Find the bug and fix it.”

The agent would usually do exactly that.

Sometimes it worked.

But occasionally it would make a plausible change without actually understanding what triggered the failure.

So I changed the order.

Now my first request is:

“Reproduce the failure before changing anything.”

I ask the agent to:

  • identify the exact failing scenario

  • reproduce it locally

  • show the expected vs actual behavior

  • identify the smallest failing case

  • explain what it thinks is happening

  • only then propose a fix

That extra step has been useful because it separates:

“I found some code that looks suspicious”

from:

“I can actually reproduce the problem this code causes.”

It also catches a second problem: sometimes the reported bug isn't actually a bug in the code being inspected. The reproduction exposes that the real issue is somewhere else entirely.

My current flow is becoming:

Reproduce → Understand → Plan → Fix → Verify

rather than:

Error → Fix → Hope

I'm curious what other people are doing with AI debugging agents.

Do you make the agent reproduce a bug before allowing it to modify the code, or do you usually let it investigate and fix in the same pass?

70 views

Add a comment

Replies

Best

I have found reproduction especially useful when the error message points in the wrong direction. It gives the agent evidence to work from instead of assumptions.

 Exactly. I've run into that too.

An error message can point you toward the symptom rather than the actual cause. Reproduction gives the agent something much more concrete to reason from instead of anchoring on the message alone.

That's probably one of the biggest benefits I've noticed from adding that step.

i like the approach because reproduction forces the agent to earn confidence before touching the code. I would also make the reproduction test part of the final fix verification.

 I agree. That’s actually a useful addition to the workflow.

The reproduction shouldn't just be a debugging step — it can become part of the verification process too.

So the flow becomes:

Reproduce → Understand → Plan → Fix → Reproduce again → Verify

That second reproduction matters because otherwise it's easy for an agent to say "fixed" simply because the code changed without proving the original failure is actually gone.

I think that small distinction is especially important with AI agents: a plausible fix isn't the same as a verified fix.

I started doing something similar with tricky bugs. Reproduction gives me much more confidence in the fix, especially when the agent suggests changes across multiple files.

 Same here. The confidence difference is noticeable, especially when the proposed fix touches several files.

A reproduction gives me a much better checkpoint before accepting a multi-file change: I can see the actual failure first, then verify that the final change removes that specific failure.

Yes, but with one guardrail that took me a while to learn: make it write the failing case before it reads the buggy code, and make it show you the test failing. Otherwise you get a reproduction that was quietly written to match whatever the code does today. It looks like evidence, it passes review, and it locks the bug in as expected behavior. An agent that has already seen the implementation will write an assertion that agrees with it almost every time.

The other thing reproduction buys you is a decision you didn't know you had. Once you're down to the smallest failing case, a decent share of "bugs" turn out to be nobody's fault, the code does exactly what it was asked to do and the ask was wrong. That's not a fix, that's a product conversation, and Error to Fix to Hope skips straight past it. You only see it because the small case is small enough to argue about.

Where I don't do this: anything I can eyeball in ten seconds. Reproduce first is a real tax, and paying it on a typo or an off by one just teaches you to stop paying it on the ones that matter. Cheap and obvious goes straight to the fix, everything else earns the full loop.

 This is a really important guardrail. I hadn't been explicit enough about separating the failing case from the implementation.

Writing the failing case first and actually watching it fail makes the evidence much stronger. Otherwise the agent can accidentally turn the current behavior into the definition of "correct."

I also like your distinction between "bug" and "the code is doing exactly what it was asked to do." That can completely change the next step from debugging to a product decision.

And agreed on the cost: I wouldn't pay the full reproduction loop for every tiny typo either. It makes more sense for anything ambiguous, consequential, or difficult to reproduce mentally.

 The step I'd add to your second Reproduce is cheap: before you accept the fix, put the old code back and watch the new test go red. A stash andone command

Sounds redundant next to "reproduce again after the fix", but it catches adifferent failure. Reproducing after the fix proves the bug is gone. Reverting proves the test was ever capable of catching it. Those come apart more often than you'd expect, usually because the agent nudged the assertion while it was making things pass, and you end up with a green test that would stay green with the bug put back.

Did exactly this today on a timeout bug where two sequential calls could each claim the full budget and blow past a proxy cutoff. Two new tests, both green after the fix. Reverted the fix, both went red, and only then did I believe them. Thirty seconds of work.

Where it doesn't apply: fixes that remove a precondition instead of changing behaviour. I moved some work off a critical path this week so a later failure can't happen at all. Nothing goes red against the old version, because the old version was fine right up until something else went wrong first

 That distinction is really useful.

I hadn't separated "the fix removes the failure" from "the test actually detects the failure" before.

The revert check makes the second one explicit:

Fix → test passes → revert fix → test fails → restore fix → test passes

That gives the test its own evidence instead of assuming that a green result means the test is meaningful.

And your timeout example makes the point well. Thirty seconds of verification is a pretty cheap price for knowing the regression test can actually catch the bug it was written for.

I have started asking agents for evidence before solutions too. Reproduction helps seperate real root causes from guesses and makes the debugging process much more predictable.

 That's exactly what I'm finding too.

The useful part isn't just getting the fix. It's reducing the number of guesses between the error and the fix.

Once the agent has evidence for the failure, the debugging process becomes much easier to reason about and verify.