newSoba is in private beta. Request access
← All posts

The unit economics of an AI app, written out

Most AI pricing pages are a guess with a dollar sign in front of it. Here is the arithmetic that decides whether a subscription survives its power users, and the three levers that actually move it.

Andrea ChelloFounder, Soba
SOBA 002 5% OF USERS

An AI subscription has a failure mode that ordinary SaaS does not: the marginal cost of a user is not close to zero, it varies by two orders of magnitude between users, and the users who cost the most are the ones least likely to churn.

That combination is worth writing out properly, because it decides more about a product than almost any feature does.

The arithmetic#

Take a flat $20 a month plan. After payment processing you keep roughly $19.40. Say you want a 70% gross margin, which is low for software and high for anything with inference in it. That leaves about $5.80 a month, per user, to spend on compute, and that has to cover everything: the model calls, the retries, the tool loops, the embeddings, the run that failed halfway and got retried.

Now the distribution. In every agentic product I have seen numbers for, usage is not normal, it is a power law. Something like:

Share of users Runs a month Cost to serve
Light 70% under 20 well under budget
Regular 25% 100–300 around budget
Heavy 5% 2,000+ many times budget

The average is fine. The average is always fine, and that is what makes this trap so easy to walk into. You look at total spend over total users, see a comfortable number, and ship.

The problem is that the 5% is not a rounding error on your P&L; it is most of your P&L. And it is not random which users they are. They are the ones who integrated you into their daily work, who tell other people about you, and who would be most expensive to lose.

The pathology

Your best customers are your worst customers. Every pricing lever that protects the margin (caps, throttles, credit packs) is aimed directly at the people most committed to the product.

Why agents make it sharper#

A chat product costs you one round trip per user action. An agentic product does not.

One user request becomes a plan, then a tool call, then a file read, then another model call to decide what the file means, then an edit, then a check, then a retry because the check failed. A single "run" in an agent product can be fifty model calls, and the ratio between a cheap run and an expensive one is not 3×. It is 100×.

That does two things to the arithmetic above. It stretches the tail, so the heavy users get heavier. And it makes cost per user genuinely unpredictable at signup time, because it depends on what the user is trying to do, not on how often they log in.

Which is why the honest version of most AI pricing pages would read: we do not know what you will cost us, so we have priced for the worst case and hoped you are not it.

The three levers#

There are exactly three things you can do to that equation. Everything else is a variation.

1. Lower the price you pay per token#

Cheaper models, smaller context, aggressive caching, batching, a cheap model to route to an expensive one. This is real engineering and it works: you can often take 40–60% out of a naive implementation.

It also has a floor, and you will hit it. Prompt caching cannot be applied twice. And every competitor is running the same playbook against the same price list, so the advantage is temporary by construction.

2. Raise the price you charge#

Per-seat becomes per-seat-plus-usage. Flat becomes tiered. Unlimited becomes 500 messages.

This works too, and it is the standard answer. The cost is that you have moved the variance onto your customer, who now has to think about how much your product costs every time they use it. Products that make users ration themselves get used less, and products that get used less get cancelled.

3. Change who buys the compute#

The third lever is the one most teams never consider, and it is the only one that changes the shape of the curve rather than a coefficient in it.

Your heavy users, the 5% who dominate your costs, are overwhelmingly the same people who already pay for Claude Pro, Claude Max, ChatGPT Plus or Pro. That is what heavy AI users do. They are already holding subsidised inference capacity, sitting idle most of the day, bought at a price you cannot get.

If a run can execute on the plan that user already pays for, the cost of serving them goes to zero. Not lower. Zero. The tail that was eating your margin becomes the cheapest cohort you have.

What that does to the table#

Run the same distribution again, with the heavy users on their own compute:

Share Cost to serve Before After
Light 70% trivial trivial trivial
Regular 25% around budget around budget around budget
Heavy 5% zero most of your COGS nothing

The average barely moves. The variance collapses, and variance is the thing that was actually dangerous. You can now offer an unlimited tier without it being a bet, because the users most likely to take you up on it are the ones who cost you nothing.

And you can pay them for it. A larger allowance while connected, or credit, or simply uncapping them. A plan that says unlimited while you are connected to your own AI is an offer no competitor buying API tokens can match, and it costs you less than the metered tier it replaces. In Soba that is one field on the plan object, connectedReward.

The part that is not free#

Someone still has to do the work: decide per run which compute is allowed, reach the user's machine, meter what the run actually consumed and in which economics, bill the user at your rate rather than at cost, and fall back cleanly to a metered tier when the user has nothing connected, all without your call site knowing or caring which happened.

That is a real system, and it is the one we build. But the reason to build it is upstream of any of the engineering: the compute your most expensive users need has already been bought. Charging them again for it, at the highest price in the market, is a choice.

Start at the quickstart, or read how a run is routed if you want the mechanism before the pitch.

Keep reading

Stop buying tokens
your users already own.

Point your OpenAI client at one base URL. Soba prices the plan, routes the run, and bills the user at your rate.

Claude Code, Codex, Gemini CLI and Ollama are the products of their respective owners. Soba is an independent tool, not affiliated with or endorsed by any of them, and each run stays on a machine dedicated to one user, signed in with that user's own account and under that provider's terms.© 2026 Soba