Unit Economics of AI Features — Cost per Completion and Margin per User
AI features have real variable costs that compress traditional SaaS margins in unexpected ways.

SaaS pricing was built on a simple bet: once software is built, serving one more customer costs almost nothing. Traditional infrastructure costs run at a fraction of revenue, producing strong gross margins because adding one more user costs almost nothing at the margin, and the whole model depends on the marginal user being nearly free to serve. AI features quietly cancel that bet. Every interaction calls a model, and every model call has a real, variable cost attached to it, which destroys the near-zero marginal cost assumption traditional SaaS infrastructure depends on for its strong gross margins. This isn't a hypothetical risk sitting in a spreadsheet somewhere. Duolingo's Q4 2025 earnings filing projected meaningful gross margin compression across 2026, and the company pointed to one driver above the rest: expanding access to AI-powered features for all users. A public company, in a public filing, naming AI feature expansion as the reason its margins are set to shrink, is about as concrete a warning as this problem gets.
Consumption Behavior and Cost Distribution
That compression comes from straightforward math that compounds: cost per interaction equals tokens per interaction multiplied by cost per token, and cost per user equals cost per interaction multiplied by interactions per user. That second multiplication is where most teams get surprised, because interactions per user is not a fixed number. It depends entirely on how each person actually uses the product, and usage varies far more than most pricing models assume.
A product can look healthy on paper, with an average cost per user well under the price of a plan, while a small slice of heavy users quietly runs at a loss underneath that average. Profitability, in other words, is a distribution problem before it's a pricing problem. A flat-rate plan can be margin-positive for the vast majority of a user base and still bleed money overall if the top few percent of users are unbounded in how much they consume.
Fixing that starts with segmenting users by what they actually do, not by which plan they bought or which features they can access. Four metrics turn that distribution into something a team can actually act on: AI cost per monthly active user, AI cost per AI-active user, AI cost per paying user, and AI cost as a percentage of ARPU. None of these numbers replace the others. Each one answers a different question, from how much the AI feature costs across the whole base to how much it costs relative to what a paying customer brings in. Together they set up the next problem: none of these metrics mean anything unless the underlying cost is measured at the right level of granularity.
What cost per completion means at the request level
The unit that actually governs AI economics is cost per completion: the tokens consumed in a single request, multiplied by the per-token price of the model handling that request. That number has to be calculated at the level of the individual request. It can't be reconstructed after the fact from a monthly invoice.
End-of-month billing from a model provider rolls every call into one lump total. It can't tell a team which feature drove the spend, which user segment is expensive, or which prompt pattern is burning tokens for no real benefit, so the insight that would actually change a pricing or routing decision is visible only after the bill has already landed. By the time a finance team notices a spike, the quarter that caused it is over.
Getting request-level attribution means every call needs to carry metadata: a user identifier, a feature tag, which model handled it, and the token count on both the input and the output side. Without that metadata attached at the moment of the call, cost can't be traced back to the user, session, or feature that generated it.
Tools built for this exist and are already handling production volume. Langfuse instruments every call as a trace, attaching token counts, cost, latency, and quality to that trace, so attribution happens at the level of individual requests, users, and sessions. Another example in the category processes over 10 billion requests monthly across more than 650 organizations, with production safety and audit trails built directly into the API layer, and its core gateway went open-source under the MIT license in March 2026. These are examples of what the category does, not a claim that any one tool is the only answer.
The gap that keeps teams blind to this data is usually a choice, not a limitation. The OpenAI Usage API returns a null user for shared keys by default, so any team that hasn't set up virtual keys or per-user tagging simply has no per-developer or per-user visibility into cost. That's a configuration decision made or skipped early on, and skipping it means the cost data a team needs for any real pricing decision doesn't exist yet, no matter how good the analysis built on top of it looks.
Model choice is the largest single variable in cost per completion, and it changes the margin math dramatically
Once cost per completion is measurable, the single biggest lever for moving that number is which model handles the request. No other optimization available to a product team comes close to that kind of swing, given that the spread from model choice alone is roughly 81× between a premium frontier model and a small open model for an identical application.
The economics of model choice also keep shifting. Foundation model costs have dropped sharply year over year, to the point that a feature carrying a significant per-call cost in early 2023 can run for a fraction of that price in 2024.
The objection to any of this is obvious: a cheaper model produces worse output, and worse output drives churn. That's a fair concern, but it's an empirical question, not something to assume from the price tag alone. Teams need to measure the actual quality impact of a cheaper model on their specific task, because for a large share of tasks, a smaller model performs comparably on the dimensions users actually notice. Assuming quality loss without testing it means leaving the largest cost lever in the entire system untouched out of caution that may not be warranted.
Routing each request to the right model, not just any model, as a systematic cost and quality strategy
Model choice doesn't have to be a single decision made once at launch. It can be a decision made fresh for every request. Intelligent routing sends each request to the cheapest model capable of handling it well, cutting real LLM costs substantially with no visible drop in quality, and this pattern is already running in production at scale.
AT&T built a router on LiteLLM that sends coding tasks to the model that fits the job, cutting costs meaningfully while reporting roughly 2% quality loss. Databricks built its own Smart Router for the same purpose, and reports that it roughly matches the quality of the most expensive model available.
Routing matters more now because production AI architectures increasingly span several providers at once, drawing on OpenAI, Anthropic, Google, Mistral, and DeepSeek within a single application. Cloud-provider gateways weren't built for that kind of cross-provider routing, so teams working across providers end up needing a purpose-built routing layer instead.
Routing also doubles as a reliability tool. When the primary provider is rate-limited or goes down, a routing layer can redirect traffic to an equivalent model on another provider automatically, turning what would be a hard outage into a softer, slightly more expensive degradation rather than a failure the user actually sees.
None of this comes free. Quality-aware routers that run an ML classifier on every prompt add 50 to 100 milliseconds of classification latency, and for a latency-sensitive application, that's a real cost. The right routing setup depends on how much latency a given workload can tolerate against how much it needs to save on cost. Connecting cost data to quality evaluation directly is what closes this loop: when a particular step in an agent's process is found to consume a disproportionate share of the token budget for only a marginal gain in quality, a team can swap in a cheaper model for that step and check the quality impact before it ever reaches production.
The gateway as the infrastructure layer where cost measurement, routing, and governance converge
Everything described so far, request-level attribution, cross-provider routing, automatic failover, needs somewhere to live. Once an application depends on more than one AI provider, the problem shifts from model access to controlling that access at production scale, without every feature team rebuilding its own retry logic, budget enforcement, and observability from scratch.
The OpenAI Usage API returns a null user for shared keys, and without a layer enforcing virtual keys per developer, team, or project, the cost data that any unit economics analysis depends on simply doesn't exist at the granularity needed. It's a structural gap that has to be closed before the pricing work in the earlier sections is even possible, not a data problem to solve later.
A gateway built for this separates the control plane from the data plane, so authentication, authorization, and rate limiting run in-memory with consistent overhead no matter how complicated the governance rules get. That separation is what keeps governance from turning into a latency tax on every single request. One open-source example, written in Go, adds minimal overhead per request even at high throughput, and governs every consumer through virtual keys carrying budgets, rate limits, and model allow-lists, backed by published benchmarks from sustained runs on both t3.medium and t3.xlarge instances. Implementation language matters here in a way that's easy to underestimate at first. A Python proxy carries more overhead per request than a compiled Go binary, and that difference becomes visible once concurrency climbs and agentic workflows multiply the number of calls sitting behind a single user action.
Six capabilities separate a gateway built for production from one still running at proof-of-concept scale: exponential backoff on transient errors, fallback chains to another provider when one is degraded, a circuit breaker per provider so a dead endpoint isn't hammered with retries, token-level rate limits per consumer, cost attribution per virtual key and team, and audit-grade logs that satisfy compliance requirements. None of these six things are new ideas. Building each one once, centrally, is simply far more reliable than having every application team reinvent its own version.
The pressure to do this centrally is only growing. A Dataiku and Harris Poll survey of 600 enterprise CIOs found that 81% expect to rely on two or more LLM providers to stay competitive, and a majority have already switched providers at least once. Multi-provider architecture is the default condition most engineering organizations are already operating under, not an advanced configuration reserved for sophisticated teams.
Self-hosted gateway costs versus a managed layer
Once a team accepts that a gateway is necessary, the next question is who builds and runs it. Self-hosting shifts real operational weight onto engineering: standing up Redis and PostgreSQL, handling upgrades, owning incident response, and keeping up with every provider SDK change, all for a team whose leverage is really in shipping product features, not maintaining infrastructure.
LiteLLM, the most common self-hosted option, illustrates the shape of that cost. It requires Redis and PostgreSQL for production deployments, adds operational overhead, and depends entirely on third-party integrations for observability and evaluation, while guardrails and audit logs require an enterprise plan. None of that is a flaw specific to one project. It's the general cost profile of running gateway infrastructure yourself.
Self-hosting is still the right answer for a real subset of teams: those with data-residency requirements, in-VPC constraints, or air-gapped environments where traffic simply cannot cross a managed edge. That's a genuine constraint in regulated industries, not a reason to avoid managed infrastructure in general.
For teams without those hard constraints, a managed gateway can resolve the same routing, cost attribution, and governance requirements without adding a new system to operate and patch. What matters is who bears the operational cost of running it, the engineering team or the vendor.
Translating cost-per-completion data into pricing and packaging decisions that protect margin
None of the measurement work matters unless it changes how a product is priced. Accurate cost-per-completion data, broken down by user and by feature, is what any pricing decision has to be built on. Without it, flat-rate pricing is really just a bet that the average user's behavior represents everyone, and the power-user tail from earlier keeps eroding margin the average user generates without anyone noticing until the quarterly numbers come in.
Four responses exist, running roughly from lightest touch to firmest enforcement:
- Consumption-based pricing tiers: cap token consumption at each plan level so the cost of serving a user stays bounded by what that user pays. The power-user tail becomes a trigger to upgrade to a higher plan rather than a margin problem.
- Feature-level cost attribution: find which specific AI features drive disproportionate cost, then price or gate those features separately instead of spreading their cost across the entire user base.
- Model routing by tier: serve lower-tier users through cheaper models and reserve frontier models for premium plans, turning the model choice lever from earlier into a pricing tool rather than just an engineering decision.
- Usage budgets enforced at the gateway: set hard token limits per user or team that block further requests once a budget is hit, replacing passive dashboards with active enforcement before a cost spike ever reaches the monthly bill.
Each of these depends on the same foundation: knowing, at the request level, what a completion actually costs, and knowing which users and features are generating that cost. Everything from Duolingo's margin warning to AT&T's routing numbers traces back to that same starting point. Measure the completion. Everything else in pricing and packaging follows from getting that number right.
Sources
- 6 best LLM gateways for developers in 2026 - Articles - Braintrust
- Pricing AI Features: The Unit Economics Framework Engineering Teams Always Skip - TianPan.co
- Unit Economics of Consumer AI Apps 2026 - Inworld AI
- Subscription App Economics: The Hidden Cost of AI Features
- Pricing AI Features: The Unit Economics Framework Engineering Teams Always Skip - TianPan.co


