What would you actually trust an agent to do while you're asleep?
I've spent seven months building one, so I've thought about this more than is healthy.
What surprised me: capability was never the hard part. Models got good enough a while ago. The hard part is finishing. Getting something to still be working at 4am, recover when a page changes shape or a login expires, and leave you something usable by breakfast is a completely different engineering problem from getting it to start well.
But the real blocker isn't technical. It's trust, and everyone draws the line somewhere different.
Some people will let it send emails but not book anything.
Some will let it spend money but won't let it near their inbox.
Almost nobody will let it reply to their boss.
Our users have run 260 sessions and pushed over 100 million tokens through it, and the pattern I keep seeing is that people hand over research and drafting instantly, and hand over anything irreversible almost never. Even when the irreversible thing is the one costing them the most time.
So, two questions for this crowd:
What's the one recurring task you'd hand over tonight if you trusted it completely?
And what would you never let it touch, no matter how good it got?
I'm trying to work out whether the line is really about reversibility, or about who sees the output when it goes wrong.
(Context, if you want it: fulmina.re. Generous free tier.)
ps. we're live on Product Hunt today. There's a code in the comments there for a free month of Lite, which works out to around 50 sessions. Free either way, no card.
Replies