Vibe coding version one is basically solved. What's your month-two experience?
The tools keep getting better at the first week. Describe an app, get an app. Impressive and the demos aren't lying.
What I keep hearing about (and lived myself) is month two. You ask for one change and five files move. Styling drifts. The AI forgets decisions from last week because the only place your app's structure exists is the chat history and the chat history runs out. Every builder community has the same complaints in almost the same words.
So, honest questions for people actually shipping with these tools:
What's the longest you've kept a vibe-coded app alive and evolving and what finally made it painful?
What's your workaround: restore points, rebuilding cleanly from a mockup, exporting to Cursor, something else?
Has anyone gotten one through a real reorganisation, new roles, changed process without a rebuild?
Full disclosure: I'm building in this space but this thread isn't a pitch, the workarounds people have actually found are more interesting to me than my own answer.
Replies
I’ve had the same issue where the first version feels easy, then small changes start touching unexpected parts of the project.
@brody_vincent Yes and the counterintuitive part is that it's not the AI being careless. It literally can't prove a small edit is safe so regenerating wide is the rational move from where it sits. Thirty changes later, the graveyard.
What helps regardless of tool: one small change per request, name the exact file or component, snapshot before every change so a bad one is a revert and do your heavy iteration on a disposable copy then apply the final version once.
Wrote up the full mechanism and the workarounds here if useful: chromoly.io/why-ai-changes-unrelated-code the first half is tool-agnostic advice, no signup anywhere.
@david_brandt3
Yeah, that makes sense. I think the biggest shift is treating AI like a fast collaborator rather than expecting it to understand the whole project forever.
Keeping decisions documented and having a safe way to test changes seems like the missing workflow. The speed is already there, it’s just about keeping things manageable as the project grows.
@brody_vincent "Fast collaborator rather than understanding the whole project forever" is a really good frame. That's exactly the shift and it changes where the memory has to live. A collaborator doesn't hold the project in their head between sessions, so the project needs its own memory: the decisions, the structure, the safe way to try changes. Get that part right and the collaborator can be as forgetful as it wants.
Good talking through this, your comments sharpened it for me.
Agent patched edge case, added test. Looked clean. Reverted patch to check, test still passed. It wrote assertion that "tested" nothing relatd to bug. Month two realizing half safety net hollow
@konstantin_tikhaev The revert trick is the real lesson here. A test only counts if you've watched it fail. If you undo the fix and it stays green, you don't have a test, you have a decoration. And AI is unusually good at decorations, because it optimizes for looking done.
A hollow net is worse than no net since you take risks you'd never take bare. My rule has become: never accept the artifact, only the demonstration. A test proves itself by failing on the revert. A feature proves itself by a walkthrough, not a "done."
Month one you read what the AI wrote. Month two you find out what it did.
@david_brandt3 so we ended up automate that check, agent now has to run test agianst unpatched branch and log failure before PR opens
@david_brandt3 the chat history point is the real issue. Once your app's structure only exists in a conversation that eventually runs out, there is no durable record of the decisions that got you there. Calling it forgetting undersells it, the AI never had anywhere to keep that information to begin with. The workaround that tends to hold up is treating a written spec or architecture doc as the source of truth (we use Notion for that) and feeding the agent from that. Rebuilding clean from a mockup works too, it just costs more time than people admit upfront.
@alex_gidirim The AI never had anywhere to keep that information to begin with - that's a sharper way to put it than forgetting and I might steal it. Exactly, it's not memory loss, it's that no memory was ever designed.
The Notion-spec workaround is the correct instinct. You've hand-built the missing artifact. Its limit shows up around week six when the doc is advisory, so nothing stops the code and the spec from quietly disagreeing. You catch drift by reading, which works right up until you're busy.
That gap is basically why we built what we're building: the spec as the thing the system is actually compiled from, so code and spec can't disagree - disagreement isn't caught, it's impossible. Same idea as your Notion doc, just enforced by the machine instead of discipline.
And agreed on rebuild-from-mockup. It works and everyone underestimates the cost of the one clean pass.
Fourteen months on the same codebase, with the same agent, and what finally worked was not a huge spec. It was a short list of rules the agent has to read every session, with the important ones backed by tests.
The doc is short on purpose: one rule per line, no long explanations, plus a table showing which file to read before changing each part of the app. The details stay in those files, so the main doc stays short enough that the agent actually reads it.
Every time things went off track, it was because some rule existed only in my head. So now the rule is simple: if the code needs a new rule, I add that rule in the same change. Not later.
I would argue a bit on the idea that "code and spec can never disagree." The biggest problems I had were not really code problems. For example: text still promising a feature we removed, an old pricing claim in the FAQ, or the wrong feature being marked as paid. The compiler will not catch that.
A simple test can. It can search the app for old claims and fail the build if it finds them. Pretty dumb, but it has caught more issues than code review.
What it still does not solve is design drift and taste. A test cannot tell you "this looks weird." I still need to open the page and check it myself
And the rules doc can also become messy if the agent is allowed to change it freely. I only add to mine manually, one line at a time.
@siarheihamanovich Fourteen months of receipts beats theory and this is the right place to push. Let me split it, because the compiler covers more of your examples than you'd expect and less than everything.
Your examples - text promising a removed feature, the wrong thing marked as paid are mostly references, and references are exactly what a compiled system tracks. In our setup the app's UI isn't hand-maintained text, it's generated from the spec: remove a workflow from the spec and the screens referencing it regenerate without it. There's no orphaned button or stale label, because the label was never written by hand in the first place. And the dependency map flags anything still pointing at a removed thing before the change applies. So the app still mentions a feature we deleted is structurally hard to produce here.
What the compiler genuinely can't catch: A welcome message mentioning a retired workflow, an email template with an old price - that's meaning, not references, and your dumb grep test is the honest fix there. I'm stealing it.
Full disclosure: our own marketing site isn't built on Chromoly, so it gets exactly the manual discipline you describe and yes, a stale claim survived months there until I caught it by reading. The gap between compiled surfaces and hand-written ones is real, and it's the strongest argument for compiling more of them.
Same experience. What fixed it for me was moving the app's memory out of the chat and into the repo.
Every project gets its own folder on my desktop, opened in VS Code with Claude Code. In that folder I keep a file with the decisions - brand elements, structure, naming conventions, why certain things were built the way they were. Claude updates it as we go, and I point it back at that file at the start of every session. So when the chat dies, nothing important dies with it.
Longest running project is 8 months, still evolving. It still picks up where we left off at maybe 80-90% accuracy. The 10-20% gap is usually stuff that got decided in conversation and never made it into the file.
No restore points, no rebuild, no export. Just the folder plus the decisions file.
your line about the structure only existing in the chat history is the whole thing. month two is when that runs out and you find the app has no memory outside a transcript nobody can query. what fixed it for me was moving decisions out of the conversation and into files the agent reads at the start of every session. not documentation for humans, just the settled calls: this is the folder layout, this is why we went with x over y, dont touch z. when context resets it re-reads those instead of re-deriving them, and the re-deriving is exactly where the five-files-move problem was coming from. the honest cost is that its work nobody wants to do in week one, when everything still fits in the window and writing it down feels pointless. i only started because id already lost the same argument with an agent three separate times.