Prompt Compression with Microsoft LLMLingua: 4-20x Token Cuts (2026)
LLMLingua compresses prompts up to 20x with ~1.5pt accuracy drop. We dissect LLMLingua-2's BERT-classifier approach, where it dominates (long RAG, doc Q&A) and where it breaks (tool calling), and how CallSphere blends it with prompt caching for compounding savings.