Every number that says your AI product is working arrives before the one that says it isn't

by

AI apps convert trials 52% better and churn annual subscriptions 30% faster. The flattering metric lands in week one. The honest one lands in month twelve — and by then you've already spent money on the gap.

Here's a pairing I can't stop thinking about, and both halves come from the same dataset.

RevenueCat published its 2026 State of Subscription Apps report in March. It's built on the subscription apps running through their platform — more than 75,000 developers, over a billion in-app transactions, north of $11 billion in annual developer revenue. Big enough to be worth arguing with.

The good half: AI-powered apps convert trials to paid 52% better than non-AI apps, 8.5% against 5.6% at the median. They monetise downloads about 20% better. Realised lifetime value per paying user runs $18.92 a month against $13.59, and $30.16 annually against $21.37 — roughly 40% higher on both.

The other half, same report: annual retention for AI apps is 21.1%. For non-AI apps it's 30.7%. Monthly retention is 6.1% against 9.5%. Refund rates run about 20% higher, 4.2% against 3.5% at the median, with a much uglier upper bound — 15.6% against 12.5%.

So the AI label makes people say yes faster, pay more, and leave sooner.

Now notice the timing, because that's the part that actually costs you money. Trial conversion is visible in a week. Refund rate takes a month. Annual retention takes twelve months, which for most indie products means it doesn't exist yet as a number you've ever seen. Every signal that arrives early flatters you, and the one that would correct you shows up after you've made the decision it was supposed to inform.

I think that ordering explains a lot of behaviour I've seen in makers I respect, including me. You launch, conversion looks better than the benchmarks you read about, and you reasonably conclude the product is landing. So you put money and time into the top of the funnel, because that's the part that's responding. Twelve months later the cohort you were so pleased with is 21% intact and you're wondering what changed. Nothing changed. You just finally got the second number.

There's independent data pointing the same way with a sharper edge on it. Kyle Poyar, working with ChartMogul, classified roughly 3,500 software companies and compared retention across them, looking at businesses past $250k ARR. Median net revenue retention for B2B SaaS came in at 82%. For AI-native companies it was 48%, with gross revenue retention at 40% — worse than B2C.

The breakdown by price is the bit I'd tape to the wall. AI products priced above $250 a month: 70% gross revenue retention, essentially B2B SaaS territory. Priced $50 to $249: 45%. Priced under $50 a month: 23%. Being cheap and easy to buy is the same property as being cheap and easy to cancel, and at consumer price points that property dominates everything else you do.

One genuine counterweight, because the pessimistic read isn't automatically the correct one. a16z's growth team looked at retention across hundreds of AI companies and argued the problem is partly measurement, not durability. A large share of early churn is "AI tourists" — people who pay $20 to satisfy curiosity and were never going to stay. Anchoring your retention curve at signup bakes those people into your denominator forever and makes a healthy product look sick. Their suggestion is to rebase from month zero to month three, once the curve flattens, and read M12 divided by M3 as the real signal. On that basis the leading AI companies look fine, and some show a "smiling" curve where lapsed users return as the product improves.

I find that convincing and I still think it cuts both ways. Rebasing gives you a truer read of your committed users. It does not make the tourists free — you paid to acquire every one of them, and if you're spending on growth, cost per customer still retained at M3 is the number that decides whether you're building or leaking.

And there's a third complication that makes the dashboard worse rather than better. Mixpanel's 2026 benchmarks, drawn from something like 290 billion AI events, found devices using AI products up 26% year over year while total AI events dipped slightly. In North America, engagement depth fell 38% year over year even as adoption held. Their reading is that agents now do in one run what used to take five prompts — so falling engagement can mean the product got better. Which means the mid-funnel signal you'd normally use to catch a retention problem early has become genuinely ambiguous. Usage down can be efficiency or abandonment, and the line looks identical.

So what do you actually do with one product and no analytics team.

Pick your M3 number before you launch anything, and write down what you expect it to be. Not to be rigorous — to stop yourself retro-fitting a story onto whatever you get.

Treat trial conversion as a marketing result, not a product result. It tells you the pitch worked. It tells you nothing about whether the thing works.

If you're under $50 a month, assume the 23% until you have your own evidence otherwise, and price your acquisition spend as though most of it evaporates.

And stop reading engagement as a health metric without a second signal next to it. "Sessions down" is now a question, not an answer.

Murror sits awkwardly inside all of this, which is why I've been chewing on it. We're a journaling and reflection product. Someone using us less could mean they've drifted off, or it could mean they're doing better and need us less — which is, unhelpfully for my dashboard, the outcome I actually want. I don't have a clean way to tell those apart yet. What I've stopped doing is assuming the graph knows.

26 views

Add a comment

Replies

Be the first to comment