AI is a wedge, not a layer

Inference is a variable cost. Your pricing model is not.

GROSS MARGIN BY USAGE DECILE, FLAT-RATE AI PLANGROSS MARGIN BY USAGE DECILE, FLAT-RATE AI PLAN-50%0%50%100%D1D2D3D4D5D6D7D8D9D10
Your most engaged customers are now your least profitable. Every instinct your pricing was built on has inverted.

For twenty years software had a marginal cost of approximately zero. One more user cost you a rounding error in storage and bandwidth, which is why the entire commercial architecture of SaaS, seats, flat tiers, unlimited usage, land and expand, works the way it does.

Every pricing instinct we have was built on that assumption. Generative features broke it, and most pricing has not caught up.

The inversion

Under a flat subscription with an AI feature inside it, your heaviest users are now your least profitable users. In a sufficiently lopsided distribution, the top decile is unprofitable in absolute terms.

Read that again in the context of everything you believe about customer success. Engagement has been a proxy for retention, expansion and advocacy for as long as any of us have been doing this. Your health scores are built on it. Your CS team is compensated on it. Your board deck celebrates it.

The customer using your product hardest is now the one you would quietly rather lose. Nothing in your operating model is set up to handle that, and most companies discover it a full year after it becomes true, in a gross margin conversation with a CFO who has been raising it for two quarters.

What makes this different from hosting costs

Two things.

The magnitude. Hosting was single-digit percentages of revenue. Inference on a heavy generative workload can be double digits, and on the wrong plan it is unbounded.

The direction of travel is not reliably down. Per-token prices have fallen enormously, and that is the argument for doing nothing. But usage per task has risen faster: reasoning models, agentic loops, longer context, multi-step chains. The cost of a unit of intelligence is collapsing while the number of units per job is climbing, and betting on the first line without modelling the second is how a plan becomes a write-off.

What I would actually do

Instrument cost per customer before you change anything. Most companies cannot answer “what does our tenth-percentile customer cost us” and every subsequent decision depends on it.

Separate the wedge from the layer. A single high-value capability can carry usage-based or value-based pricing, because the customer can see what they are buying. A layer of twelve small conveniences cannot, because none of them is individually worth metering and the pricing page becomes an accounting exercise.

Price the outcome where you can. “$X per qualified lead” or “$X per resolved ticket” tracks cost without exposing the customer to token mechanics they neither understand nor want to. Where you cannot, meter credits. Design the credit so the customer can predict their bill, because unpredictable bills churn accounts faster than expensive ones.

Put a ceiling on the flat plan now. Generous, explicit, and visible on the pricing page. Retrofitting a limit onto a plan customers experienced as unlimited is a trust event. Introducing one at signup is a footnote.

On lifetime plans, since somebody always asks

They work. They buy you time. And you will spend eighteen months unwinding them.

I have used them deliberately, to lift realized ARPU while tenure was still short, and it was the right call at the time. As retention improved they became a liability against future compute cost, and rebuilding tiering and add-ons around them took considerably longer than introducing them did.

If you are considering one, do it with that trade priced in rather than discovered.

The wider point

The question is never where to add AI. It is which single customer or economic problem has just become solvable, what that is worth, and what it costs to serve.

A layer is a cost with a press release attached. A wedge has a number on both sides of it.

← All writing

Work with me

Something in the business is not working, and the explanations have stopped being convincing.

A first conversation is 45 minutes and costs nothing. If I am not the right person, I will say so and usually know who is.