Est.

Allocating AI Infrastructure Costs in Engineering Chargebacks

Engineer spending on AI models reveals the need for token-level cost tracking.

Correspondent · · 10 min read
Cover illustration for “Allocating AI Infrastructure Costs in Engineering Chargebacks”
Cost Attribution · September 30, 2026 · 10 min read · 2,187 words

AI spend doesn't play by those rules. Uber found this out in 2026, when the company acknowledged burning through its entire annual AI budget by April, tracing part of the overrun to a single engineer whose token usage alone ran $40,000 a month. That's not a story about one careless engineer. It's what usage-based pricing does when nothing is watching it. Gartner forecasts worldwide AI spending will reach $2.59 trillion in 2026, with AI infrastructure alone accounting for more than 45% of that total and generative AI model spending climbing roughly 80% year over year.

The reason it happened at Uber, and the reason it's happening quietly at plenty of other companies right now, comes down to what's actually being billed. A cloud bill charges for a resource that belongs to somebody. Nobody's invoice says which department ran the calls or which feature quietly tripled its usage last week. That's the mismatch finance teams are up against, and it's the mismatch every other section of this piece works through. Research cited in the sources puts employees using AI tools IT hasn't sanctioned above the majority of the workforce, meaning shadow AI compounds cost attribution failures as real spend accrues outside any budget line. The FinOps Foundation named granular monitoring of AI spend (tokens, LLM requests, GPU utilization) as the top requested tooling capability in its 2026 report, which tells you what the field is missing.

The API Key as the Wrong Unit of Attribution

Sharing an API key across a whole engineering org isn't a bad habit teams picked up by accident. It's the default, because the provider's billing primitive is the key itself, and one key typically serves several teams, several projects, and several features at once. Different teams end up managing different API keys, provider-specific SDKs, rate limits, retries, and usage reports, each creating a separate, unreconciled ledger of consumption.

Most teams start by routing LLM traffic through whatever gateway they already have, often a cloud-provider default like AWS API Gateway or Azure API Management, though production AI workloads demand capabilities such traditional gateways weren't designed to provide (token-based rate limiting, multi-provider model routing, semantic caching, and cost attribution at the request level), as Zuplo's 2026 evaluative guide finds. It doesn't hold up. Request-per-minute rate limiting, the standard mechanism in traditional API gateways, can't control AI spending because a single LLM request can consume anywhere from a handful to tens of thousands of tokens, and two requests to the same endpoint can differ by orders of magnitude in cost. A gateway that counts requests instead of tokens is counting the wrong thing entirely.

So the fix has to happen at the request level, not after the fact in a spreadsheet. That's the actual job of an LLM gateway: it sits between the applications and the model providers, and every single call, whatever team it comes from, whatever model it targets, passes through one place that handles routing, authentication, failover, caching, cost attribution, and policy all at once. As agentic workflows spread (Gartner expects the share of enterprise applications wired into task-specific AI agents to jump sharply by the end of 2026, up from a small minority in 2025), the math only gets more urgent, since each agent fires off multiple model calls per task. At that point the gateway becomes the piece of production infrastructure the rest of the cost picture depends on, no longer just a nice-to-have for developers.

Gateway enforcement of the tagging discipline that makes chargeback possible

None of the chargeback math works without clean tagging, and a gateway is the only layer positioned to enforce it consistently, across every provider and every model an org happens to use. The mechanism is the virtual key: instead of one shared API key, the gateway issues a distinct key per team, per project, or per feature, and that key carries metadata (team name, cost center, application, environment) that rides along with every request it makes.

Because every call has to cross the gateway to reach a provider at all, there's no side door. Spend that used to hide inside the org's own stack, someone's personal API key, a forgotten side project, appears in the same ledger as everything else.

Once the tagging is clean, the arithmetic itself is almost boring: multiply input tokens by the input rate, output tokens by the output rate, cached tokens by the cached rate, using whatever per-million-token pricing the provider publishes, and sum it by team. The gateway's per-request logs are what supply the raw numbers that arithmetic runs on. None of it works without that data.

Tagging on its own is still just reporting, though. Turning it into an actual control means pairing it with budgets enforced at the virtual key, the team, and the customer level. Hard budget limits block requests once a team's threshold is exceeded, preventing the kind of runaway consumption the Uber case illustrates. Hierarchical budgets let a platform team set a ceiling for the whole org while handing each team its own slice underneath it, so one department's experiment can't quietly eat the shared quota. And rate limiting has to be counted in tokens, not requests, since token volume is the only unit that actually tracks cost when request sizes vary this much.

A 2025 report from Mavvrik and Benchmarkit on AI cost governance found that data platforms were the leading source of unexpected AI spend, with network access costs close behind, and that only a minority of companies were even including on-premises AI costs in their reporting at all. Tagging discipline that only covers managed API calls to OpenAI or Anthropic misses a real chunk of the picture.

The staged path from showback to defensible chargeback

Start by showing teams what they're spending, with zero financial consequence attached, and only move to actual allocation once the attribution behind those numbers has been checked and rechecked. That staging period reveals requests reaching providers with no team tag attached, shared services generating cost that belongs to more than one owner, and a feature that grew into a real consumer of tokens while nobody was watching, before the mess costs someone money.

Nearly all organizations now manage AI spend in some form, per the FinOps Foundation's 2026 findings, and the conversation has mostly moved into how to do it without breaking something. The FinOps Foundation's own State of FinOps 2026 press release puts the figure at 98% of the 1,192 survey respondents now managing AI spend, up from just 31% two years earlier. Showback is the bridge that gets you there.

Chargeback itself needs a written allocation policy before it goes live, not after the first disputed invoice. Shared infrastructure, a common embedding pipeline, a centrally managed agent framework, needs a rule for splitting the cost (by usage volume, by headcount, by project count) agreed on in advance. Allocation exceptions should be documented: a proof-of-concept that one team runs but multiple teams benefit from shouldn't be charged to the team that happened to own the API key. And finance needs the whole model to be auditable. When a team pushes back on a line item, the gateway's immutable logs and per-request telemetry are what settle the argument.

Gateway features for chargeback accuracy

Gateways vary a lot on the dimensions that actually matter for chargeback accuracy, including how deep the governance goes, how granular the tagging gets, how much control over deployment the org retains, and how well the logging holds up to an audit. A handful of options each have a different shape and merit evaluation.

One open-source option, Bifrost by Maxim AI, written in Go, benchmarks at very low latency overhead at high request-per-second throughput with a high success rate in sustained runs, and connects to 23–25+ providers, including OpenAI, Anthropic, AWS Bedrock, Google Vertex AI, Azure OpenAI, Mistral, Groq, Cohere, Ollama, and others, through a single OpenAI-compatible API. Bifrost is confirmed as an open-source AI gateway written in Go by Maxim AI, both by the company's own materials and by its GitHub repository. It supports hierarchical governance through virtual keys with per-key, per-team, and per-customer budgets and rate limits, and it can run self-hosted, in-VPC, or air-gapped under Apache 2.0. Bifrost's GitHub repository and license file confirm it is licensed under the Apache 2.0 License. It also ships a native MCP gateway, SSO, role-based access control, immutable audit logs, and vault integration. One September 2026 enterprise guide rated it the strongest fit for organizations that need compliance-grade governance and full deployment isolation.

LiteLLM is popular in Python ecosystems, supports self-hosting, and offers virtual keys with basic budgets, making it well-suited for Python-first teams at prototyping scale or moderate throughput. It falls behind on latency once load gets serious, though.

A managed edge option brings strong caching and solid analytics but offers no self-hosted or in-VPC path at all, which takes it off the table immediately for any organization with data residency requirements. It's a reasonable fit where data sovereignty simply isn't the constraint. Separately, Zuplo provides purpose-built AI Gateway and MCP Gateway project types alongside its API gateway, supporting token-based rate limiting, multi-provider routing, semantic caching, hierarchical budget controls, and prompt injection detection, and can deploy across AWS, Azure, GCP, Akamai, Equinix, or 300+ edge locations; Zuplo's own 2026 evaluative guide rated it the top AI-workload option among the gateways it evaluated, which includes API management platforms, cloud-provider gateways, and dedicated AI gateways. Zuplo's own learning center confirms it offers three purpose-built project types (API Gateway, AI Gateway, and MCP Gateway), with the MCP Gateway reaching general availability in August 2026, and its features page confirms single-tenant deployment on AWS, Azure, GCP, Akamai, or Equinix, alongside a broader claim of deployment across 300+ edge locations without cloud-provider lock-in, while the guide itself states that Zuplo is its pick for the best API gateway for AI and LLM workloads in 2026.

Whichever option a team lands on, a few criteria decide whether the numbers it produces can survive an audit. Metrics and traces need to come out in standard formats, Prometheus, OpenTelemetry, request-level logs, so they plug into the SRE and FinOps tooling teams already run, rather than locking data inside a proprietary dashboard. Logging needs to be genuinely immutable, tamper-evident through WORM policies or cryptographic signing, covering every config change and every model call, because that's what a SOC 2, HIPAA, or ISO 27001 review is going to ask for. And latency overhead needs to be checked at the team's actual request volume, not taken from a vendor's marketing page, since overhead compounds fast in agentic workflows where one user action triggers a chain of model calls. Per one enterprise gateway guide, the single criterion that knocks the most options out of enterprise deals before any other feature gets a look is whether the gateway can run inside the organization's own network at all. The gateways the sources evaluate in 2026 that are worth including in a serious shortlist are mapped to the chargeback-relevant dimensions. Kong AI Gateway fits teams that already run Kong for traditional API management, offers plugin-based AI capabilities including failover and policy, is self-hostable, and was the subject of a July 2025 benchmark authored by Kong reporting its data plane as faster than Portkey OSS and LiteLLM on requests-per-second using a mocked LLM backend (a test that measures raw gateway overhead, not production behavior, and should be treated accordingly).

Security and compliance requirements that chargeback infrastructure must satisfy alongside cost attribution

Cost attribution and data security run through the exact same control layer, and a gateway built for chargeback ends up handling both jobs whether that was the plan or not. Getting the cost numbers right and keeping sensitive prompts out of the wrong hands are the same project. They're the same project.

Know provider defaults on data retention before assuming a gateway handles this automatically. OpenAI retains API data for a period of weeks by default for abuse monitoring, and Anthropic briefly shortened its standard log retention window before reverting to its prior default. Anthropic's own Privacy Center shows that a shorter seven-day API retention window it rolled out in September 2025 is no longer current, with the commercial baseline having returned to 30 days. Zero Data Retention exists as an option, but it isn't something a pay-as-you-go customer gets by default. It requires a negotiated enterprise agreement. Microsoft's own documentation confirms Zero Data Retention is available only to customers on an Enterprise Agreement or Microsoft Customer Agreement and not to Pay-As-You-Go subscriptions, a restriction echoed elsewhere as requiring a negotiated enterprise agreement rather than a standard pay-as-you-go plan.

Zero Data Retention, where it's in place, is an architectural guarantee rather than a checkbox: prompts, context, and outputs get processed in memory only and are never written to logs, databases, or training pipelines. A gateway is positioned to enforce exactly that, since it's the one point in the stack every request already has to cross for tagging and budget purposes. The same interception point that makes chargeback numbers trustworthy is the one place an organization can actually guarantee what happens to the data riding inside those requests.

Sources

  1. Best API Gateways for AI and LLM Workloads (2026): Evaluative - Zuplo
  2. Chargeback vs. Showback: FinOps Cost Allocation Guide 2026
Filed underCost Attribution

More in Cost Attribution