What's your "is it safe to ship?" checklist for AI-built apps?
AI can get me to a working app in a day, but "it works on my machine" and "it's safe for real users" are very different things.
The part that worries me is the boring stuff that's easy to miss when the code looks clean and everything runs: API keys sitting in the frontend, auth checks that exist on the page but not on the endpoint, database rules left wide open from the prototype stage, no rate limits on anything. The AI rarely warns you about these unless you ask, and by the time you think to ask, you've already moved on to the next feature.
Right now my pre-launch process is mostly manual and a bit random: I skim through the auth flow, search the repo for hardcoded secrets, and try to break things as a logged-out user. It catches some things, but I never feel fully confident.
I'd love to hear how others handle it:
What do you always check before letting real users in?
Do you use any tools or prompts to audit AI-written code, or do you rely on your own review?
What's the scariest thing you found after shipping?
Hoping we can put together a practical checklist from people who've been through it. Happy to share a summary in the comments once the replies come in.
Replies
Do you test AI built apps with real users before shipping or mostly rely on your own testing
@shivam_kushwaha16 Good question. Right now it's mostly my own testing, which is exactly the gap I'm trying to close. My plan is a small closed beta with a handful of people, but with limited access and spend caps in place first, since real users will find paths I'd never think to try. How do you handle it? Do you bring in testers before launch, or after?
@kailong_20 I’d bring a few testers in before launch so we can catch the obvious issues early and make the launch version more solid
The agent built a summary report with a metric nobody had measured. The value looked plausible, but it was just a hardcoded placeholder from the initial prompt. It took a day to realize the calculation was never wired up. Now every rendered number gets traced to raw data.
@konstantin_tikhaev That's a scary one because nothing errors out. The UI looks right, so it passes every "does it run" check. Tracing every rendered number back to raw data is a great rule. I'm adding it to the list: for anything the AI generates that displays a value, find where the value is computed and confirm it isn't a placeholder. Did you end up automating that trace, or is it still a manual pass?
Dial
the one that bit us wasn't a security hole exactly, it was cost. we run a voice API where every call costs real money per minute, and an early build had a demo endpoint with no rate limit and no auth check server-side, only client-side. someone (not maliciously, just poking around) hit it in a loop and it took us a few hours to notice usage spiking before we caught it. nothing was stolen, but it could have been an expensive night.
what actually changed our process: any endpoint that costs money or touches another user's data gets a rule "assume the frontend doesn't exist." if the AI wrote the check only in the component, we ask it explicitly to also enforce it in the handler/middleware and show us the diff. cheap to ask, easy to skip if you don't.
for tools: nothing fancy, just grep for API keys before every deploy and a checklist item for "does this route work if I call it directly with curl and no session." that curl test alone catches most of what the frontend hides.
@galdayan Thanks for sharing this, and good that it was only a few hours. I hadn't put cost abuse in the same category as security, but it clearly belongs there, since an unprotected endpoint can be as expensive as a breach. "Assume the frontend doesn't exist" is a great rule to remember, and asking the AI to show the diff of the server-side enforcement is a cheap habit. The curl test with no session is going straight into the checklist. One follow-up: do you have any spend alerts or hard usage caps now, so you'd notice sooner than a few hours?
Dial
@kailong_20 yeah, we do now. a daily spend cap per API key that just hard-stops the key if it's blown through, plus a slack alert at 50% of normal daily volume so someone sees it same-hour instead of same-day. the cap is the one that actually matters, the alert just tells you why. before that incident we had neither, we were just trusting the auth check that turned out not to be there.
In my Next.js apps every export from a "use server" file is a public POST endpoint. An agent adds a small exported helper next to an action, it takes a userId argument because that was convenient, and now anyone can call it with someone else's id.
An action never trusts a user id it's handed, it reads the user from the session itself. And anything that costs money runs the same fixed order of checks (auth, rate limit, balance, input) before it spends anything, so a rejected request can't burn a paid call.
@siarheihamanovich This is exactly the kind of thing I was worried about, and I wouldn't have thought to check it. A helper that looks like an internal function but is actually a public endpoint is easy to miss in review, especially when the diff is small. Two rules I'm taking from this: never trust a user id passed as an argument (read it from the session), and run a fixed order of checks (auth, rate limit, balance, input) before anything that spends money. Do you have a way to spot these, like grepping every export in "use server" files, or do you catch them in review?
@kailong_20 Partly a test: it walks every file that starts with "use server" and fails CI if any of them exports the 2 functions that take a user as an argument. Those are the ones that must never be reachable from outside, so they're pinned by name.
But, the weak spot is that it only knows the names I gave it. A brand new helper with a userId parameter goes straight past it. For those it's review: after any change to actions I run a separate agent pass whose only job is checking that the user comes from the session