Context Pruning & Prompt Compression (LLMLingua) Guide

By AI Agent Hub Editorial Desk · Review method · Corrections

Tutorial · 5 min read · Reviewed September 18, 2026

As context windows scale to 1M+ tokens, sending massive uncompressed prompts (such as full repository file trees or RAG search results) dramatically inflates API costs and increases generation latency. Prompt Compression uses lightweight language models (like Microsoft LLMLingua-2) to prune low-perplexity tokens—compressing prompts by 5x to 20x while preserving key reasoning facts.

1. How Perplexity-Based Prompt Compression Works

LLMLingua calculates the surprisal / perplexity of each token in a long prompt using a small, fast model (e.g. LLaMA-3-8B). Repetitive boilerplate tokens, stop words, and filler syntax are pruned, leaving only high-information semantic tokens:

Compression Ratio Comparison:

2. Python Implementation with LLMLingua

from llmlingua import PromptCompressor

compressor = PromptCompressor(
    model_name="microsoft/llmlingua-2-xlm-roberta-large",
    use_fp16=True
)

long_prompt = "..." # 10,000 tokens of raw repo code context

results = compressor.compress_prompt(
    context=[long_prompt],
    instruction="Write a unit test for user registration.",
    rate=0.2 # Retain top 20% highest information tokens
)

print(f"Compressed Prompt Tokens: {results['compressed_tokens']}")
compressed_text = results["compressed_prompt"]

3. Compression Is a Quality Trade, Not Free Context

Prompt compression can remove repeated or low-information tokens before a request reaches the target model. It can also delete a negation, identifier, number, exception, delimiter, citation, or security instruction. Do not apply the same rate to every task. Preserve system policy, schemas, code syntax, quoted evidence, user constraints, and the question itself unless a test demonstrates that changing them is safe.

ContentDefault treatmentReason
System policy and tool contractDo not compressSmall changes can alter authority or output shape
Retrieved proseCandidate for selective compressionOften contains redundancy, but citations must survive
Source code and stack tracesStructure-aware pruning firstIdentifiers and punctuation carry meaning
Conversation historySummarize resolved turns with provenanceOld dialogue is often lower value than current state
Numbers, dates, IDs, and legal textPreserve or verify exactlyToken removal can change the fact itself

4. Reproducible Evaluation Protocol

  1. Freeze a task set with ordinary, long-context, adversarial, and no-answer cases.
  2. Run the uncompressed baseline and record accepted-task quality, input tokens, output tokens, latency, and provider cost.
  3. Test several compression rates with the same target model and decoding settings.
  4. Check exact preservation of citations, identifiers, constraints, and required output fields.
  5. Include the compressor's own model loading, compute, latency, and memory in total cost.
  6. Choose the lightest transformation that meets a predeclared quality floor; do not select on cost alone.

Report results by slice. A compressor may work well on meeting transcripts and fail on source code or multilingual policy text. Measure unsupported-answer rate and refusal changes in addition to task accuracy. If compression makes a prompt fit but produces an unreliable answer, the context problem has not been solved.

5. Prefer Deterministic Pruning Where Possible

Before using a learned compressor, remove duplicate passages, irrelevant log levels, generated files, superseded messages, and retrieval results below an evidence threshold. Select code by symbols and dependencies, trim tool results to typed fields, and summarize completed steps into a state record. These transformations are easier to audit and reproduce.

Version the compressor model, tokenizer, parameters, forced tokens, source prompt hash, compressed prompt hash, and target model. Sensitive content processed by a local compressor still requires retention, access-control, and deletion rules.

6. Security Regression Set

Include prompts where one word changes permission or meaning: negations, currency, dates, object IDs, code operators, policy exceptions, and instructions hidden inside retrieved text. The compressed output must preserve the trusted instruction hierarchy and must not turn untrusted data into an instruction. Fail the candidate when any prohibited action becomes easier after compression.

7. Primary References

8. Related Token Estimators & Cost Tools

What compression saves, in dollars

Compression is usually sold on ratio. The ratio only matters through the price it multiplies. Below: twenty thousand requests a month, comparing a raw 10,000-token prompt against an 8.3x compressed 1,200-token prompt, with 800 output tokens either way. Raw prompt: 10,000 input tokens with no cache benefit. Compressed: 1,200 input tokens, an 8.3x ratio. Output is 800 tokens either way.

Model Raw prompt Compressed Saved per month
Claude Sonnet 5$0.028$0.010$352
Gemini 3.1 Pro$0.030$0.012$352
GPT-5.6 Sol$0.056$0.021$704

Note that the saving is far larger in relative terms than in absolute dollars at this volume, because output dominates once input has been squeezed. Compression pays best on long-context flagship workloads, where the input line is still the largest one.

Rates verified against provider documentation on September 18, 2026. Promotional rates expire, so re-check before budgeting: LLM API cost planning · September 2026 pricing update. Run your own numbers in the cost calculator.