Context Pruning & Prompt Compression (LLMLingua) Guide
As context windows scale to 1M+ tokens, sending massive uncompressed prompts (such as full repository file trees or RAG search results) dramatically inflates API costs and increases generation latency. Prompt Compression uses lightweight language models (like Microsoft LLMLingua-2) to prune low-perplexity tokens—compressing prompts by 5x to 20x while preserving key reasoning facts.
1. How Perplexity-Based Prompt Compression Works
LLMLingua calculates the surprisal / perplexity of each token in a long prompt using a small, fast model (e.g. LLaMA-3-8B). Repetitive boilerplate tokens, stop words, and filler syntax are pruned, leaving only high-information semantic tokens:
Compression Ratio Comparison:
- Original Prompt (10,000 tokens): Raw RAG context + system instructions ➔ Cost: $0.030 per call.
- LLMLingua Compressed (1,200 tokens): 8.3x compression ratio ➔ Cost: $0.0036 per call (88% cost savings).
2. Python Implementation with LLMLingua
from llmlingua import PromptCompressor
compressor = PromptCompressor(
model_name="microsoft/llmlingua-2-xlm-roberta-large",
use_fp16=True
)
long_prompt = "..." # 10,000 tokens of raw repo code context
results = compressor.compress_prompt(
context=[long_prompt],
instruction="Write a unit test for user registration.",
rate=0.2 # Retain top 20% highest information tokens
)
print(f"Compressed Prompt Tokens: {results['compressed_tokens']}")
compressed_text = results["compressed_prompt"]
3. Compression Is a Quality Trade, Not Free Context
Prompt compression can remove repeated or low-information tokens before a request reaches the target model. It can also delete a negation, identifier, number, exception, delimiter, citation, or security instruction. Do not apply the same rate to every task. Preserve system policy, schemas, code syntax, quoted evidence, user constraints, and the question itself unless a test demonstrates that changing them is safe.
| Content | Default treatment | Reason |
|---|---|---|
| System policy and tool contract | Do not compress | Small changes can alter authority or output shape |
| Retrieved prose | Candidate for selective compression | Often contains redundancy, but citations must survive |
| Source code and stack traces | Structure-aware pruning first | Identifiers and punctuation carry meaning |
| Conversation history | Summarize resolved turns with provenance | Old dialogue is often lower value than current state |
| Numbers, dates, IDs, and legal text | Preserve or verify exactly | Token removal can change the fact itself |
4. Reproducible Evaluation Protocol
- Freeze a task set with ordinary, long-context, adversarial, and no-answer cases.
- Run the uncompressed baseline and record accepted-task quality, input tokens, output tokens, latency, and provider cost.
- Test several compression rates with the same target model and decoding settings.
- Check exact preservation of citations, identifiers, constraints, and required output fields.
- Include the compressor's own model loading, compute, latency, and memory in total cost.
- Choose the lightest transformation that meets a predeclared quality floor; do not select on cost alone.
Report results by slice. A compressor may work well on meeting transcripts and fail on source code or multilingual policy text. Measure unsupported-answer rate and refusal changes in addition to task accuracy. If compression makes a prompt fit but produces an unreliable answer, the context problem has not been solved.
5. Prefer Deterministic Pruning Where Possible
Before using a learned compressor, remove duplicate passages, irrelevant log levels, generated files, superseded messages, and retrieval results below an evidence threshold. Select code by symbols and dependencies, trim tool results to typed fields, and summarize completed steps into a state record. These transformations are easier to audit and reproduce.
Version the compressor model, tokenizer, parameters, forced tokens, source prompt hash, compressed prompt hash, and target model. Sensitive content processed by a local compressor still requires retention, access-control, and deletion rules.
6. Security Regression Set
Include prompts where one word changes permission or meaning: negations, currency, dates, object IDs, code operators, policy exceptions, and instructions hidden inside retrieved text. The compressed output must preserve the trusted instruction hierarchy and must not turn untrusted data into an instruction. Fail the candidate when any prohibited action becomes easier after compression.
7. Primary References
- Microsoft LLMLingua repository and usage guide
- LLMLingua EMNLP paper
- AI Agent Hub: Prompt caching guide
8. Related Token Estimators & Cost Tools
What compression saves, in dollars
Compression is usually sold on ratio. The ratio only matters through the price it multiplies. Below: twenty thousand requests a month, comparing a raw 10,000-token prompt against an 8.3x compressed 1,200-token prompt, with 800 output tokens either way. Raw prompt: 10,000 input tokens with no cache benefit. Compressed: 1,200 input tokens, an 8.3x ratio. Output is 800 tokens either way.
| Model | Raw prompt | Compressed | Saved per month |
|---|---|---|---|
| Claude Sonnet 5 | $0.028 | $0.010 | $352 |
| Gemini 3.1 Pro | $0.030 | $0.012 | $352 |
| GPT-5.6 Sol | $0.056 | $0.021 | $704 |
Note that the saving is far larger in relative terms than in absolute dollars at this volume, because output dominates once input has been squeezed. Compression pays best on long-context flagship workloads, where the input line is still the largest one.
Rates verified against provider documentation on September 18, 2026. Promotional rates expire, so re-check before budgeting: LLM API cost planning · September 2026 pricing update. Run your own numbers in the cost calculator.