Is on-device inference the next big shift in consumer AI?

by

I’ve spent much of the past eight months exploring a question that started out fairly simple:

How much of an AI product actually needs the cloud?

I’m building The Change, a private AI coach for perimenopause. Early on, it worked the way most AI apps do: the phone was essentially the interface, while the intelligence lived somewhere else.

But that architecture felt increasingly strange for a product designed around deeply personal conversations.

So I started experimenting with moving the intelligence onto the device itself.

That turned into a much bigger rabbit hole than I expected.

Running useful conversational AI locally means accepting constraints that cloud developers can mostly ignore: memory, model size, thermal limits, latency, context budgets, battery life and the enormous differences between what looks good in a benchmark and what actually feels good in someone’s hand.

But there’s another side to those constraints that I find much more interesting.

Once inference happens on hardware the customer already owns, the economics and product design begin to change too.

There’s no cloud round trip for every conversation. Personal prompts don’t inherently need to leave the device. And inference stops being something you necessarily have to meter every time a user presses “send.”

For The Change, that led to a principle I’ve become rather attached to:

There shouldn’t be a meter on reassurance.

I suspect we’re still very early in figuring out what becomes possible when phones stop being interfaces to AI and start becoming the machines actually running it.

I’d love to hear from other builders experimenting with local models:

What do you think will ultimately drive on-device AI adoption—privacy, latency, cost, offline availability, or something we haven’t recognized yet?

11 views

Add a comment

Replies

Be the first to comment