The unit economics of an AI app, written out
Most AI pricing pages are a guess with a dollar sign in front of it. Here is the arithmetic that decides whether a subscription survives its power users, and the three levers that actually move it.
An AI subscription has a failure mode that ordinary SaaS does not: the marginal cost of a user is not close to zero, it varies by two orders of magnitude between users, and the users who cost the most are the ones least likely to churn.
That combination is worth writing out properly, because it decides more about a product than almost any feature does.
The arithmetic#
Take a flat $20 a month plan. After payment processing you keep roughly $19.40. Say you want a 70% gross margin, which is low for software and high for anything with inference in it. That leaves about $5.80 a month, per user, to spend on compute, and that has to cover everything: the model calls, the retries, the tool loops, the embeddings, the run that failed halfway and got retried.
Now the distribution. In every agentic product I have seen numbers for, usage is not normal, it is a power law. Something like:
| Share of users | Runs a month | Cost to serve | |
|---|---|---|---|
| Light | 70% | under 20 | well under budget |
| Regular | 25% | 100–300 | around budget |
| Heavy | 5% | 2,000+ | many times budget |
The average is fine. The average is always fine, and that is what makes this trap so easy to walk into. You look at total spend over total users, see a comfortable number, and ship.
The problem is that the 5% is not a rounding error on your P&L; it is most of your P&L. And it is not random which users they are. They are the ones who integrated you into their daily work, who tell other people about you, and who would be most expensive to lose.
The pathology
Your best customers are your worst customers. Every pricing lever that protects the margin (caps, throttles, credit packs) is aimed directly at the people most committed to the product.
Why agents make it sharper#
A chat product costs you one round trip per user action. An agentic product does not.
One user request becomes a plan, then a tool call, then a file read, then another model call to decide what the file means, then an edit, then a check, then a retry because the check failed. A single "run" in an agent product can be fifty model calls, and the ratio between a cheap run and an expensive one is not 3×. It is 100×.
That does two things to the arithmetic above. It stretches the tail, so the heavy users get heavier. And it makes cost per user genuinely unpredictable at signup time, because it depends on what the user is trying to do, not on how often they log in.
Which is why the honest version of most AI pricing pages would read: we do not know what you will cost us, so we have priced for the worst case and hoped you are not it.
The three levers#
There are exactly three things you can do to that equation. Everything else is a variation.
1. Lower the price you pay per token#
Cheaper models, smaller context, aggressive caching, batching, a cheap model to route to an expensive one. This is real engineering and it works: you can often take 40–60% out of a naive implementation.
It also has a floor, and you will hit it. Prompt caching cannot be applied twice. And every competitor is running the same playbook against the same price list, so the advantage is temporary by construction.
2. Raise the price you charge#
Per-seat becomes per-seat-plus-usage. Flat becomes tiered. Unlimited becomes 500 messages.
This works too, and it is the standard answer. The cost is that you have moved the variance onto your customer, who now has to think about how much your product costs every time they use it. Products that make users ration themselves get used less, and products that get used less get cancelled.
3. Change who buys the compute#
The third lever is the one most teams never consider, and it is the only one that changes the shape of the curve rather than a coefficient in it.
Your heavy users, the 5% who dominate your costs, are overwhelmingly the same people who already pay for Claude Pro, Claude Max, ChatGPT Plus or Pro. That is what heavy AI users do. They are already holding subsidised inference capacity, sitting idle most of the day, bought at a price you cannot get.
If a run can execute on the plan that user already pays for, the cost of serving them goes to zero. Not lower. Zero. The tail that was eating your margin becomes the cheapest cohort you have.
What that does to the table#
Run the same distribution again, with the heavy users on their own compute:
| Share | Cost to serve | Before | After | |
|---|---|---|---|---|
| Light | 70% | trivial | trivial | trivial |
| Regular | 25% | around budget | around budget | around budget |
| Heavy | 5% | zero | most of your COGS | nothing |
The average barely moves. The variance collapses, and variance is the thing that was actually dangerous. You can now offer an unlimited tier without it being a bet, because the users most likely to take you up on it are the ones who cost you nothing.
And you can pay them for it. A larger allowance while connected, or credit, or simply
uncapping them. A plan that says unlimited while you are connected to your own AI is an
offer no competitor buying API tokens can match, and it costs you less than the metered tier
it replaces. In Soba that is one field on the plan object,
connectedReward.
The part that is not free#
Someone still has to do the work: decide per run which compute is allowed, reach the user's machine, meter what the run actually consumed and in which economics, bill the user at your rate rather than at cost, and fall back cleanly to a metered tier when the user has nothing connected, all without your call site knowing or caring which happened.
That is a real system, and it is the one we build. But the reason to build it is upstream of any of the engineering: the compute your most expensive users need has already been bought. Charging them again for it, at the highest price in the market, is a choice.
Start at the quickstart, or read how a run is routed if you want the mechanism before the pitch.