Est.

Prompt Compression Techniques for High-Volume LLM Workloads

Cut token costs at scale while boosting reasoning by pruning redundant prompt content.

Staff Writer, Optimization and Governance · · 10 min read
Cover illustration for “Prompt Compression Techniques for High-Volume LLM Workloads”
Prompt Optimization · October 10, 2026 · 10 min read · 2,328 words

Prompt length is a line item, not a style choice. At high request volumes, every extra token in a prompt gets paid for repeatedly across thousands or millions of calls, and the bill compounds the way any fixed cost compounds when it's repeated at scale. Research behind the CompactPrompt paper, from BNY and Carnegie Mellon, documents the "lost-in-the-middle" effect: transformer attention leans on tokens at the start and end of a prompt and gives less weight to what sits in the middle, so a longer prompt can quietly produce worse answers even as it costs more to run. The same research found a measurable drop in reasoning performance once prompts get close to around 3,000 tokens, a threshold far below the context windows most current models support. Agentic workflows make this worse: a single user request often chains several model calls (retrieval, reasoning, classification, verification), and every excess token in the shared context gets billed at each step along the chain.

How the two principal compression families work

Two main approaches exist for cutting prompt size, and they sit at opposite ends of a tradeoff between how easy they are to deploy and how much they can shrink a prompt. Hard compression works on the text of the prompt itself: it strips, reorders, or summarizes content before it ever reaches the model, and it needs nothing special from the model receiving it. Soft compression goes further by changing how the model reads its input in the first place, usually through fine-tuning or specialized encoding that requires access to the model's internals.

That difference decides who can use which family. Hard compression works against any provider's API, open or closed, because it never touches the model itself, which makes it the default choice for most production teams. Soft compression can reach much higher compression ratios, but it only works where a team controls the model end to end, such as a closed, fine-tuned inference stack, whereas for teams calling any hosted third-party API, hard compression is the only option on the table. The choice isn't a matter of taste: it follows directly from how much control a team has over the model layer.

LLMLingua and LongLLMLingua: how coarse-to-fine pruning works in practice

LLMLingua and LongLLMLingua stand out as the most deployable hard compression methods available today, because they cut prompt size by a meaningful margin without asking anything of the target model, and because they were built with retrieval-augmented generation in mind. Both use a small language model, something like GPT-2 or LLaMA-7B, to score how predictable each token is given its surroundings. Tokens the small model finds easy to guess carry little information and get cut; tokens that surprise the small model tend to carry the meaning and get kept.

LongLLMLingua builds on this with a two-step process tuned to the question being asked. First comes a coarse pass: it ranks whole chunks of context by how relevant they are to the query, using a method called contrastive perplexity. Then comes a fine pass: within the chunks that survive, it prunes token by token using the same self-information scoring. Compression, done this way, doesn't just hold quality steady: in several documented cases, cutting out low-value tokens actually improved answer quality, because it strips away exactly the noise that triggers the lost-in-the-middle effect described earlier. A method built to save money ends up fixing a reasoning problem along the way.

CompactPrompt: extending compression to file-level data in agentic pipelines

Pruning the prompt text solves only part of the cost equation. Agentic pipelines routinely attach files, spreadsheets, and structured data alongside the natural-language prompt, and those attachments often account for more tokens than the prompt itself. Every hard compression method up to this point works on the prompt in isolation and leaves attachments untouched. Standard file compression tools like gzip or numeric quantizers shrink files for storage, but they produce output that token-based LLM APIs can't read.

CompactPrompt, built at BNY with a contributor from Carnegie Mellon, closes that gap. It's a training-free pipeline that works end to end on both the prompt and its attachments, combining three techniques: pruning low-information tokens from the prompt itself using self-information scores and sentence structure; building a reversible dictionary that swaps common multi-word phrases in attached documents for short stand-in tokens; and converting floating-point numbers in structured data into smaller fixed-bit integers. Tested on two financial reasoning benchmarks, TAT-QA and FinQA, CompactPrompt cut total token usage and inference cost by as much as 60%, with only a small drop in accuracy on Claude-3.5-Sonnet and GPT-4.1-Mini. The choice to build and test this inside a financial institution says something beyond the benchmark numbers: large banks were among the earliest enterprises to deploy generative models at scale, so the token and data constraints they run into tend to show up later across other regulated, data-heavy industries.

Context-Aware Sentence Encoding for Fast Inference Without Training Overhead

Token-level pruning has a cost of its own. Running a small language model to score every token in a prompt takes time and compute, and in a latency-sensitive pipeline, that added step can offset some of the savings it's meant to produce. Context-aware sentence encoding offers a cheaper alternative. Instead of scoring individual tokens, it works at the sentence level: it measures how close each sentence's embedding is to the embedding of the query, using cosine distance, and keeps only the sentences that score high enough.

The tradeoff is coarseness for speed. A sentence that gets kept might still carry a few tokens that add little value, so the final compression ratio tends to be lower than what token-level pruning achieves. In exchange, the method runs faster and slots more easily into a streaming inference path. For a pipeline where the compression step runs on every single request, the time that step adds is itself a cost to track. A compressor that adds a few hundred milliseconds to every call can erase the savings it was built to deliver, so for high-throughput, latency-sensitive workloads, sentence-level encoding often makes more sense than token-level pruning even at a lower compression ratio.

The benchmark-dependency problem: why a technique that works on one task may fail on another

No compression method is safe to deploy everywhere just because it worked well on one test. Research into benchmark-dependent behavior in prompt compression shows that the same method can hold up fine on one task type and lose meaningful quality on another, and the difference has less to do with the compression ratio applied than with what kind of task is being compressed for. A method tuned for retrieval-heavy RAG pipelines, for instance, doesn't automatically carry that performance over to summarization or multi-step reasoning tasks.

Agentic systems make this risk sharper. A coding agent might read a file, search a codebase, and reason over a growing session history all within one workflow, and the kind of content it's working with keeps shifting as the session goes on. A compression method calibrated against RAG retrieval tasks may behave very differently once it's applied to a debugging trace or a multi-step planning prompt partway through a long session. Compression ratio alone tells a team how much smaller a prompt got, not whether the answer it produces will still be right. Teams need to test methods against the actual task types running through their pipeline, retrieval, reasoning chains, summarization, classification, before trusting a single compression ratio number from somewhere else. A benchmark that "showed it works" only speaks to the prompts it was built on, which are rarely the mix of prompts a production system actually sees day to day.

Multi-resolution and adaptive compression for procedural and long-context workflows

A single, flat compression ratio applied across an entire prompt treats every part of that prompt as equally important, and that assumption rarely holds up in procedural or multi-step work. A prompt built from step-by-step instructions, background context, and retrieved evidence has regions that carry very different amounts of information, and some regions can tolerate heavy pruning while others can't lose a single token without breaking the task.

Adaptive, multi-resolution compression handles this by giving each region of a prompt its own token budget based on how much it matters to the task at hand, rather than cutting every region by the same percentage. Dense, high-value procedural steps get protected while low-density background material gets pruned hard. This builds directly on an idea already present in LLMLingua's design: its budget controller assigns different compression ratios to instructions, demonstrations, and questions as separate components. Multi-resolution compression takes that same logic further, dividing a prompt into finer and more dynamically determined regions than three fixed buckets. For long agentic sessions, where something read early on (a file, an error trace) becomes less relevant as the session moves forward, this approach lets context get rescored and recompressed continuously instead of being pruned once at the start and left alone.

Batched prompting as a compression strategy for high-volume evaluation workloads

When many requests share the same system prompt, the same instructions, or the same set of few-shot examples, combining those requests into a single call spreads the fixed cost of that shared content across every input at once. That's compression through consolidation rather than through pruning, and it attacks a different part of the cost equation than the methods discussed so far.

Research on BatchGEMBA, built for evaluating machine translation output efficiently, shows what this looks like in practice: pairing batched prompting with prompt compression produces token savings that neither technique achieves on its own. The mechanism is straightforward. Shared material, a rubric, a set of instructions, a handful of examples, gets paid for once per batch instead of once per item, so the marginal cost of each additional item in the batch comes down to only its own variable content. The same pattern applies well beyond machine translation: content moderation running at scale, bulk classification of support tickets, large summarization runs, nightly evaluation passes over model outputs. Batch size isn't free to push upward without limit, though. A batch that grows too large runs into context-window limits and can reintroduce the same lost-in-the-middle effect described earlier, this time across the variable content inside the batch. The right batch size depends on the task and the model, and it's worth finding through direct testing.

Why Compression Belongs in the Infrastructure Layer

Compression applied by individual engineers, prompt by prompt, inside application code tends to produce patchy coverage. Some prompts get compressed well, others not at all, and nobody outside the team that wrote the code can see which is which, let alone measure or improve it over time.

The same logic that pushes PII redaction, API key management, and cost metering into a shared gateway layer applies here. Compression touches every model call, so it shouldn't depend on each individual engineer remembering to add it to their code. Running compression as a step in shared infrastructure, before a request ever reaches the model provider, makes it auditable: original token count, compressed token count, and the resulting savings can get logged per request, per team, per feature, per model. That visibility turns compression into an ongoing practice that teams can continually monitor and refine. Teams can see which workloads benefit most, what compression ratios they're actually hitting, and where quality is taking a hit, then adjust the method or dial the aggressiveness up or down based on real data. An LLM gateway is the natural place to put this, since it already sits between the application and the model provider, already counts tokens, and already carries the routing logic that decides which model handles a given request. Adding compression to that same layer is a small extension of work the gateway is already doing.

Integrating compression with routing, caching, and spend controls in a production AI stack

Diagram: Four Controls, One Order: How Compression Fits the Production Stack. Visualizes: Visualize a four-step sequential pipeline showing the exact order cost controls must run in a production AI stack.

Compression doesn't replace the other cost controls available in a production AI stack. It works alongside intelligent routing, semantic caching, and spend enforcement, each one acting on a different part of the overall cost picture, and skipping any one of them leaves savings on the table. Compression cuts the token size of a request before it goes anywhere. Routing sends that smaller request to whichever model fits its complexity best. Semantic caching stops a request from being sent at all when an equivalent one has already been answered.

The order these steps run in changes the outcome. Compression should happen first, since a smaller prompt is what the router evaluates and what the cache checks against. The cache check comes next, because a compressed prompt matches a cached response more reliably than the original, more verbose version would, which raises the cache hit rate. If nothing in the cache matches, routing sends the compressed request to the right model. Budget enforcement runs last, checking a team's remaining allocation before the call ever reaches the provider, which blocks overspend outright instead of flagging it after the money is already spent. Logging each of these steps in real time, original token count, compressed count, which model handled the request, cache hit or miss, and cost, is what closes the loop between these controls. A team that can only see its spend after the fact, with no way to stop a call before it's billed, will always find out about overspend too late to do anything about it; the check needs to happen before the call goes out, not after. For teams running several workloads across more than one model provider, a shared view across all of them also makes it possible to compare how well a given compression method performs against different tokenizers, since the same method can save more against one model's tokenizer than another's, and that difference is worth feeding back into routing decisions. Savings that get measured, attributed to the right team or workload, and acted on build on each other call after call. That compounding is what separates a real operational discipline from a one-time fix.

Sources

  1. Compression Method Matters: Benchmark-Dependent Output Dynamics in LLM Prompt Compression
  2. Prompt Compression with Context-Aware Sentence Encoding for Fast and Improved LLM Inference
  3. CompactPrompt: A Unified Pipeline for Prompt and Data Compression in LLM Workflows
  4. BatchGEMBA: Token-Efficient Machine Translation Evaluation with Batched Prompting and Prompt Compression
  5. Characterizing Prompt Compression Methods for Long Context Inference
  6. Adaptive Multi-Resolution Procedural Knowledge Compression for Large Language Models

More in Prompt Optimization