đź§© We built the workflow. It took longer than doing the work.

I keep running into the same problem with GTM tools and AI workflows.

A task looks repetitive, so we start connecting tools, defining prompts, fixing edge cases and testing automations. A few hours later, the workflow still needs supervision, and the manual version would have been faster.

The tools also add up. Enrichment, outreach, CRM, automation and AI credits can quickly become a meaningful monthly cost, before accounting for the time spent building and maintaining the setup.

I’m curious how other makers approach this:

  • What is a task you tried to automate that you later moved back to a manual process?

  • Do you have a rule for when to automate, keep it manual, or hire for it?

34 views

Add a comment

Replies

Best

Yeah "how much do I let agents do" is where a lot of people get stuck. Same for me, I'm trying stuff every day. I do think the more you can hand off the better though. So I try not to let the LLM make calls at all and push as much as I can into scripts, that way nothing goes off the rules. When I really do need an LLM to judge something, I get other LLMs to check it. Like Claude makes the first call and Gemini double checks. Sometimes I throw Codex and Grok in too for a triple check lol. Jev just came out and looks good too.

 I agree. Scripts make more sense for tasks with fixed rules. LLMs can help with judgement, but they still need clear boundaries and a way to verify the outcome.

 For me the "way to verify" part is E2E tests and asking twice. If two runs don't agree, a human looks at it.

Mine is small and it is on this site: reordering the images in a launch gallery. I scripted it with browser automation, and the drag froze the tab twice, once while merely opening the section. By hand it is a ten second drag, so I stopped.

Next to it sits one automation that has paid for itself many times over, and the contrast is what gave me a rule. It is a script that records the demo GIF for that same gallery. It fills the invoice in before the recording starts, adds a second line on camera, and prints the total it reads back out of the live preview. Once that printed 24,000 dollars for an invoice meant to be 2,400, because the rate field already held a zero. The video looked perfectly convincing. Only the printed number gave it away.

So the rule is two questions. Will this run again? And can it tell me whether it worked without me watching it? The recording says yes to both: re-recording after a product change is one command, and it reports what it produced. The gallery drag said no to both. It happens once per launch, and when it failed it failed as a frozen tab rather than an error.

The second question is the one I underrated. The costly automations were not the ones that failed loudly. They were the ones that reported success while doing nothing. Typing into a text box that never got focus came back as done four times, with nothing in the box. Anything shaped like that needs a person to look at every result, and a person looking at every result could have done the task.

Applied to your stack, my guess is that enrichment passes both questions, since a missing field is visible in the row, and anything that writes to a prospect fails the second, since a wrong message looks exactly like a right one until somebody replies.

 Yeah, that makes sense. If someone needs to verify the result every time, it often ends up taking more work than doing it manually.

I had not really thought about it in terms of whether the automation can prove it worked, but that is especially relevant for anything that writes to prospects.

 Your line that an automation someone has to verify every time costs more than the work itself has stayed with me, and the product I was automating around launched on Product Hunt today. If you follow the product page you will see what comes next, and if you try it, I would like to hear what fails first.

 Congrats on the launch. I’ll take a look.

“What fails first” is probably the right test for any automation, especially when it can appear to work while doing nothing

 Thank you. Appearing to work while doing nothing is the case I watch hardest in mine: an invoice that looks sent while the email never left. So an invoice is marked sent only once the mail has actually gone out, and pressing Send only queues it. A failure lands on the invoice timeline with the reason, instead of a success that was never true.