For twenty years software had a marginal cost of approximately zero. One more user cost you a rounding error in storage and bandwidth, which is why the entire commercial architecture of SaaS, seats, flat tiers, unlimited usage, land and expand, works the way it does.
Every pricing instinct we have was built on that assumption. Generative features broke it, and most pricing has not caught up.
The inversion
Under a flat subscription with an AI feature inside it, your heaviest users are now your least profitable users. In a sufficiently lopsided distribution, the top decile is unprofitable in absolute terms.
Read that again in the context of everything you believe about customer success. Engagement has been a proxy for retention, expansion and advocacy for as long as any of us have been doing this. Your health scores are built on it. Your CS team is compensated on it. Your board deck celebrates it.
The customer using your product hardest is now the one you would quietly rather lose. Nothing in your operating model is set up to handle that, and most companies discover it a full year after it becomes true, in a gross margin conversation with a CFO who has been raising it for two quarters.
What makes this different from hosting costs
Two things.
The magnitude. Hosting was single-digit percentages of revenue. Inference on a heavy generative workload can be double digits, and on the wrong plan it is unbounded.
The direction of travel is not reliably down. Per-token prices have fallen enormously, and that is the argument for doing nothing. But usage per task has risen faster: reasoning models, agentic loops, longer context, multi-step chains. The cost of a unit of intelligence is collapsing while the number of units per job is climbing, and betting on the first line without modelling the second is how a plan becomes a write-off.
What I would actually do
Instrument cost per customer before you change anything. Most companies cannot answer “what does our tenth-percentile customer cost us” and every subsequent decision depends on it.
Separate the wedge from the layer. A single high-value capability can carry usage-based or value-based pricing, because the customer can see what they are buying. A layer of twelve small conveniences cannot, because none of them is individually worth metering and the pricing page becomes an accounting exercise.
Price the outcome where you can. “$X per qualified lead” or “$X per resolved ticket” tracks cost without exposing the customer to token mechanics they neither understand nor want to. Where you cannot, meter credits. Design the credit so the customer can predict their bill, because unpredictable bills churn accounts faster than expensive ones.
Put a ceiling on the flat plan now. Generous, explicit, and visible on the pricing page. Retrofitting a limit onto a plan customers experienced as unlimited is a trust event. Introducing one at signup is a footnote.
On lifetime plans, since somebody always asks
They work. They buy you time. And you will spend eighteen months unwinding them.
I have used them deliberately, to lift realized ARPU while tenure was still short, and it was the right call at the time. As retention improved they became a liability against future compute cost, and rebuilding tiering and add-ons around them took considerably longer than introducing them did.
If you are considering one, do it with that trade priced in rather than discovered.
The wider point
The question is never where to add AI. It is which single customer or economic problem has just become solvable, what that is worth, and what it costs to serve.
A layer is a cost with a press release attached. A wedge has a number on both sides of it.
