The 4 costs I think AI products should track beyond API spend
I've been thinking about AI product economics, and I think "cost per API call" is only part of the problem.

An AI feature can look cheap from the infrastructure side and still become expensive to operate.
I like thinking about four different costs:
1. Inference cost
How much are the model calls actually costing?
That's the obvious one.
2. Retry cost
How often does the system call the model again because the first result failed validation, timed out, or produced something unusable?
A cheap model that needs three retries isn't necessarily cheap.
3. Verification cost
Who or what checks the result?
A response that costs $0.01 but requires a person to spend two minutes reviewing it can be more expensive than a $0.05 response that is reliable enough to use directly.
4. Context cost
How much information does the system need to repeatedly send or retrieve to produce a useful answer?
Long context, repeated retrieval, tool calls, and multi-step workflows can change the economics quickly.
So instead of tracking only:
AI cost = model tokens
I'm increasingly thinking about:
AI cost = inference + retries + verification + context
That also changes how I'd evaluate an AI feature.
A feature isn't really efficient because the model call is cheap.
It's efficient when the whole workflow is cheaper, faster, or better than the process it replaces.
I'd be interested in how other people building AI products measure this.
What's the AI cost that surprised you most after moving from a prototype to a real workflow?
Replies
Hi there! The number that surprised me was cost per accepted result, not cost per request.
In a flashcard tool Im building right now, a model call can succeed and still produce an empty or severely underfilled deck, or fail while the result is being saved. The user gets the credit back on those paths, but the API bill stays with us. That made retry and refund behavior part of the unit economics, not just error handling.
I now track model spend against accepted, persisted output. I also don’t retry every bad result. At some moment, Refunding is cheaper and more honest than quietly spending more tokens to make one request look successful.
@siarheihamanovich That’s a really useful distinction. “Cost per request” can hide a lot once the workflow has validation, persistence, refunds, or retries around it.
I especially like the idea of measuring spend against accepted, persisted output. It shifts the question from “what did the model cost?” to “what did it cost us to produce something the user could actually use?”
The refund vs retry tradeoff is interesting too. Sometimes making the system retry until it succeeds is actually worse economics than simply accepting the failed path and refunding it.