AI Spend Attribution in Multi-Tenant SaaS Products
Track AI costs per tenant at the gateway layer, not from billing exports.

AI spend attribution has to be solved at the gateway layer, in the request path, not stitched together after the fact from billing exports. That's the thesis, and the rest of this piece is about proving it, section by section, down to how the caching, routing, and reconciliation actually work.
Start with the money. AI has become the fastest-growing line item on most cloud bills, up 47% year over year by one 2026 estimate, and for most companies it appears as a single number on a vendor invoice Finout. One number. No breakdown by team, by feature, by customer.
A survey of 534 senior leaders found 82% concerned about token costs, but only 64% said they actively monitor token usage with clear budgets and guardrails. And in a multi-tenant SaaS product, the stakes are worse than a bloated internal cloud bill.
In practice, that looks like the following. A Series B vertical SaaS platform shipped an AI feature to 1,400 production customers on a Tuesday TrueFoundry Portkey. By Friday, one tenant stuck in a retry storm had eaten 84% of the shared OpenAI rate-limit pool TrueFoundry Portkey. The billing pipeline had under-reported AI usage by 9.4% against the actual OpenAI invoice TrueFoundry Portkey. A BYO-key enterprise tenant, whose contract guaranteed isolated key usage, had been silently falling back to the shared pool every time its own key throttled TrueFoundry Portkey. And when that same customer's security team needed a per-tenant audit log for a SOC 2 II review, the platform couldn't filter it to that tenant's traffic TrueFoundry Portkey.
None of that is a fluke. Shipping an AI feature without shipping an AI gateway to support it produces this default outcome. Three different audiences should be reading this scenario and seeing three different fires. Engineering leaders should see an infrastructure reliability failure. Finance should see a unit economics failure, because nobody can price a feature they can't cost. And security or compliance teams should see a SOC 2 and GDPR evidence failure, because auditors don't accept "we're pretty sure". Enterprise LLM API spending doubled in six months, rising from $3.5B in late 2024 to $8.4B by mid-2025 (per TrueFoundry's analysis), with Gartner forecasting $2.52 trillion in worldwide AI spending for 2026, a 44% year-over-year increase verticalapi.com.
The four structural failure modes that make attribution impossible without a gateway
Provider invoices, whether from OpenAI or Anthropic or anyone else, arrive as aggregate token counts. They tell a platform how much it spent in total. They say nothing about which tenant, which feature, or which team generated that spend, and no amount of clever spreadsheet work recovers information the invoice never contained in the first place.
That's failure mode one: no attribution. Failure mode two is that the platform is potentially under-billing customers, violating BYO-key isolation promises, and unable to produce per-tenant audit logs for SOC 2 II reviews. Cost anomalies in LLM workloads move fast, a retry storm or a runaway agent loop can burn through a month's budget in minutes, and without enforcement built into the infrastructure itself, the first anyone hears about it is the monthly bill. By then it's history, not an incident.
Failure mode three concerns the three audiences reading this: engineering leaders, for whom attribution is an infrastructure reliability problem; finance stakeholders, for whom it is a unit economics problem; and security and compliance teams, for whom it is a SOC 2 and GDPR evidence problem. The controls that actually move the needle on cost, semantic caching, model routing, fallback chains, only pay off when applied centrally across all traffic. Failure mode four is no model flexibility. If every service calls providers directly from its own code, switching models means touching every service, and no platform team can enforce a model-tier policy org-wide without one place where that policy actually lives.
Field audits of production LLM applications have found 40 to 60% of token budgets going to waste: redundant calls answering the identical prompt twice, flagship models doing work a cheaper model could do just as well, no rate limits on developer or CI pipelines eating budget nobody's watching TrueFoundry Portkey. Multi-tenant SaaS adds a fifth failure mode on top of all this. Noisy-neighbor effects, BYO-key traffic quietly commingling with the shared pool, audit logs that can't be scoped to one tenant, none of it visible without isolation built at the infrastructure layer itself. Trying to patch this after the fact, by reconstructing attribution from logs once the billing cycle closes, only ever approximates what a gateway would have captured exactly, at the moment the request was made.
What a gateway-layer attribution model looks like
Picture one gateway binary, sitting between every application service and every LLM provider, acting as a single control plane where each request gets tagged, metered, checked against a budget, and logged before it ever reaches a provider. That's the whole model. Every request carries its tenant context, tenant ID, virtual key, team, project, feature tag, decided the moment it authenticates, not guessed at later from scattered logs.
The mechanism for this is metadata tagging. Modern gateways support custom headers, TrueFoundry's implementation uses one called X-TFY-METADATA, that carry the application name, tenant ID, and feature dimension on every call, and cost dashboards segment automatically along those same lines. Get this metadata wrong and the whole attribution model falls apart downstream, no matter how good the reconciliation logic is later. The gateway becomes the single source of truth: a complete, queryable record of every model call, by user, team, virtual key, application, and model, at the moment it happens, rather than an approximation stitched together from provider invoices after the fact.
Caching stops redundant tenant queries from ever reaching a provider. Routing sends each request to the right model for the job, weighing cost against quality. Budget enforcement blocks a request before it goes over a cap, rather than flagging the overage after the fact. Observability tools display spend by model and by tenant in real time, in the platform's dashboards, not reconstructed weeks later at month-end.
Adding a layer like this obviously costs something in latency. Go-based implementations push that even lower, one benchmarked at 11 microseconds at 5,000 requests per second TrueFoundry verticalapi.com Maxim AI. And adoption friction turns out to be smaller than teams expect, because a gateway that accepts the same request format as the OpenAI SDK on the front end, then translates to whichever provider sits behind it, needs nothing more than a changed base URL in existing application code. No rewrite. No migration project.
Designing tenant isolation: virtual keys, budget hierarchies, and BYO-key separation
Virtual keys are the actual unit of tenant isolation. Each tenant, or team, or feature gets a scoped key carrying its own rate limit, its own budget cap, its own model permissions, and its own audit log stream, so that nothing under that key can ever reach into another tenant's allocation.
Budgets stack in a hierarchy: customer, then team, then virtual key, then provider config, with spend caps settable and enforceable at any of those four levels. Hitting the cap causes the gateway to block the request or trigger a fallback route. It does not let the overage through quietly.
BYO-key tenants need a separate kind of isolation. When an enterprise customer supplies its own provider API key, every one of its requests has to route through that key exclusively, never falling back to the shared pool the moment that key throttles. That means the gateway has to tag BYO-key traffic explicitly and keep a distinct data processing agreement path tied to it. The scenario from earlier, where a BYO-key tenant fell back silently to the shared pool, is precisely the compliance failure this isolation exists to prevent.
Caching needs the same tenant-level scoping, or it becomes a liability instead of a saving. But if cache entries aren't scoped per virtual key, one tenant's cached answer can get served to another tenant entirely, which is both a data leak and a billing error at the same time.
What actually needs to ride on every request? At minimum: tenant ID, environment, and a feature or endpoint name. For billing, add prompt tokens, completion tokens, and cached tokens, all broken out per tenant. For compliance, add a timestamp, which model and provider handled the call, latency, and any PII redaction event that fired. If any one of those fields is skipped, the omission is discovered later, usually at the worst possible moment, an audit or a disputed invoice.
Token counts alone undercount the real cost, too. Embedding calls, retry attempts, and the overhead of managing rate limits can add another 20 to 40% on top of raw API fees, and the gateway will miss all of that cost if it is only logging the headline completion TrueFoundry Portkey Maxim AI. Semantic caching mechanics rely on a dual-layer system combining exact hash matching with vector similarity search, with cache hits returning in roughly 5 milliseconds across vector store backends (Weaviate, Redis, Valkey, Qdrant, Pinecone), as Maxim AI's analysis of Bifrost found.
Routing strategy as a cost-attribution tool, not just a cost-reduction tool
Model pricing spans a 1,000× range, from $0.075/M tokens (Gemini Flash input) to $75/M tokens (Claude Opus 4 output), and routing decisions made without tenant context will produce attribution data that does not reflect actual per-tenant economics clawrouters.com. Routing is a governance mechanism here. Routing decides whether the cost record is even meaningful.
There are three audiences to consider: engineering leaders, finance stakeholders, and security and compliance teams. Rule-based routing uses declarative policies, mapping tenant tier or feature type or prompt complexity to a model tier. It's predictable and it's easy to explain to a finance team asking why one tenant's bill looks different from another's. Classifier-based routing goes further: RouteLLM, developed at UC Berkeley and presented at ICLR 2025, trains a classifier to predict whether a cheaper model can handle a given prompt as well as a frontier model would LMSYS/UC Berkeley ICLR 2025 LMSYS/UC Berkeley ICLR 2025. The paper's results are striking, routing simple queries away from frontier models cut overall spend by more than 85% on standard benchmarks while holding onto 95% of frontier-level quality, with one matrix-factorization approach sending only 14% of all queries to the expensive model LMSYS/UC Berkeley ICLR 2025 LMSYS/UC Berkeley ICLR 2025. A third approach, quality-led routing, optimizes for predicted output quality directly rather than cost; one case study reported a 39% accuracy improvement on SRE benchmarks, though that's a single vendor's result and shouldn't be read as a universal guarantee NotDiamond / Rootly.
A broader survey of dynamic routing and cascading approaches frames the design space along three questions: when the routing decision gets made, what information it uses, and how it's actually computed, whether by rules, classifiers, reinforcement learning, or cascading through models in sequence. Routing well can beat even the single best model available, because it leans on each model's particular strengths instead of asking one model to be good at everything. Cost control turns out to be a side effect of better routing.
None of this matters for attribution unless the gateway logs which model actually served each request. Routing the same feature to different models on different days makes the cost record for that feature look completely different depending on which one got picked, so the gateway has to record the actual model served, not the one that was requested ianlpaterson.com.
Agentic tools raise the stakes further. Without per-request model tagging in the gateway log, there's no putting that cost back together after the fact TrueFoundry. And fallback chains carry their own trap: when a primary provider goes down and the gateway reroutes to a backup, the logged cost has to reflect the model that actually answered, not the one that was originally intended, or the billing record quietly drifts away from reality.
Making attribution accurate enough to bill on: reconciliation with provider invoices
That 9.4% under-reporting figure from earlier is the number to measure against. Usage-based billing that drifts more than a few percent from what the provider actually invoiced isn't a rounding error, it's a liability sitting on the balance sheet.
Several things cause that drift. Cached responses need to be logged at the cached rate rather than skipped entirely or counted twice. Retries can show up once in an application log but twice on the provider's invoice if the gateway doesn't tag the retry explicitly. Tiered pricing structures, like a provider charging a different rate for prompts over 200,000 tokens, or region-specific rates on certain models, will get flattened incorrectly by any system using a single per-token rate instead of keeping its pricing catalog current. And enterprise tenants with negotiated custom rates need their own pricing configuration inside the gateway, not the public rate card everyone else pays.
Standard cloud billing tools weren't built for any of this. They can't correlate an OpenAI invoice with GPU usage on a cloud compute platform and trace both back to a single microservice, let alone a single tenant. Purpose-built systems close that gap by attributing external API spend and internal infrastructure spend to the same tenant record.
The reconciliation workflow itself is fairly mechanical, once the logging is right. At the close of a billing cycle, sum up those per-tenant records and check them against the provider's invoice total. A gap bigger than a few percent points to something specific: a retry storm that wasn't logged, a service calling a provider directly and skipping the gateway, or a BYO-key tenant that fell back into the shared pool without anyone noticing. Enforcing that every single provider call runs through the gateway is, itself, a control. Anything that bypasses it is invisible to the whole system, full stop.
None of this holds up without a pricing catalog that stays current. Providers change prices without much warning, sometimes running introductory rates for a limited window before reverting to list price, so any gateway doing this work needs to check its numbers against live provider rate cards rather than trusting a static table TrueFoundry verticalapi.com. The gateway emits a stable webhook per request to a downstream billing system (e.g., Stripe Meter) with tenant_id, prompt tokens, completion tokens, cached tokens, model, provider, and timestamp.
The compliance layer: per-tenant audit logs, PII redaction, and SOC 2 requirements for the gateway
A SOC 2 review doesn't ask whether a platform has good intentions about tenant isolation. It asks for the log that proves it. For compliance, the gateway captures the timestamp, model and provider used, latency, and any PII redaction events.
The failure earlier, where a customer's security team needed a per-tenant audit log filtered to their own traffic for a SOC 2 II review and the platform couldn't produce it, is the exact gap such a review exists to catch TrueFoundry Portkey. Fixing it after an auditor asks is far harder than building the field into the log from day one. A per-tenant audit stream tied to a virtual key is core infrastructure required by the compliance checklist. It's the same infrastructure that makes billing accurate in the first place, just pointed at a different question. Attribution built for billing and attribution built for governance turn out to be the same system, read two different ways.


