We build agents all day, and the thing I keep noticing is that they almost never fail in the middle. They fail at the very last step, when something has to happen in the real world. Ours kept dying in the same spot: a signup form that texts a six-digit code. Everything up to that point ran fine, and then it just sat there waiting for a human to go and read a text message. I'm curious what everyone else's version of this is. Not the thing your agent does badly - the thing it does perfectly, right up until it has to stop and hand it back to you. Verification codes? A booking that only happens by phone? A supplier who answers email once a week? Something that needs an actual human voice on the other end? Mostly I want to know whether the wall sits in the same place for everyone, or whether it moves depending on what you're building.
Would love to hear thoughts. Here's my perspective. I'm against AI slop, but I'm not against AI. I believe AI is a great tool but it also needs greats pilots to use it well. This is what I'm working on. AI Fluency.
Biggest marker of AI fluency is AI judgement. Where and when to use AI and where and when not to.