Name: Towards AI Legal Name: Towards AI, Inc. Description: Towards AI is the world's leading artificial intelligence (AI) and technology publication. Read by thought-leaders and decision-makers around the world. Phone Number: +1-650-246-9381 Email: pub@towardsai.net
228 Park Avenue South New York, NY 10003 United States
Website: Publisher: https://towardsai.net/#publisher Diversity Policy: https://towardsai.net/about Ethics Policy: https://towardsai.net/about Masthead: https://towardsai.net/about
Name: Towards AI Legal Name: Towards AI, Inc. Description: Towards AI is the world's leading artificial intelligence (AI) and technology publication. Founders: Roberto Iriondo, , Job Title: Co-founder and Advisor Works for: Towards AI, Inc. Follow Roberto: X, LinkedIn, GitHub, Google Scholar, Towards AI Profile, Medium, ML@CMU, FreeCodeCamp, Crunchbase, Bloomberg, Roberto Iriondo, Generative AI Lab, Generative AI Lab VeloxTrend Ultrarix Capital Partners Denis Piffaretti, Job Title: Co-founder Works for: Towards AI, Inc. Louie Peters, Job Title: Co-founder Works for: Towards AI, Inc. Louis-François Bouchard, Job Title: Co-founder Works for: Towards AI, Inc. Cover:
Towards AI Cover
Logo:
Towards AI Logo
Areas Served: Worldwide Alternate Name: Towards AI, Inc. Alternate Name: Towards AI Co. Alternate Name: towards ai Alternate Name: towardsai Alternate Name: towards.ai Alternate Name: tai Alternate Name: toward ai Alternate Name: toward.ai Alternate Name: Towards AI, Inc. Alternate Name: towardsai.net Alternate Name: pub.towardsai.net
5 stars – based on 497 reviews

Frequently Used, Contextual References

TODO: Remember to copy unique IDs whenever it needs used. i.e., URL: 304b2e42315e

Resources

Free: 6-day Agentic AI Engineering Email Guide.
Learnings from Towards AI's hands-on work with real clients.
2 Numbers Out of 128 Destroy the Entire KV Cache — Why LLM Compression Keeps Failing
Artificial Intelligence   Latest   Machine Learning

2 Numbers Out of 128 Destroy the Entire KV Cache — Why LLM Compression Keeps Failing

Last Updated on September 22, 2026 by Editorial Team

Author(s): Ansh Saxena

Originally published on Towards AI.

2 Numbers Out of 128 Destroy the Entire KV Cache — Why LLM Compression Keeps Failing

If you haven’t read Part 1, here’s the one-paragraph setup: every time an LLM generates a token, the attention mechanism needs to “look back” at every previous token. To avoid recomputing them, the model stores a Key vector and a Value vector for each token, per layer, per attention head. This is the KV cache — a filing cabinet the model reads from on every generation step.

Part 1 : How LLMs Actually Predict the Next Word — 7 Steps, Real Math, Zero Hand-Waving

Part 1 ended with the engineering punchline. Let’s start this piece with the same number, and zoom in on why it’s a crisis that nobody has cleanly solved.

The Memory Math That Should Scare You

A production Llama 3.1–70B model serving a 128K-token context stores:

128,000 tokens × 80 layers × 128 dimensions × 2 vectors (Key + Value)
= 2,621,440,000 numbers

At float16 (2 bytes each): ~40 GB. Per user. Per session.

For context: the model weights themselves are ~140 GB at float16. Serve 4 concurrent users with long contexts and the KV cache alone eats 160 GB — more than an entire NVIDIA H100’s 80 GB of HBM.

At 128K+ context lengths, the KV cache becomes larger than the model weights. This isn’t a footnote in a research paper — it’s the primary cost driver for every inference provider shipping long-context features in 2026.

And the problem compounds. KV cache memory scales linearly with context length. Double the context window, double the memory. Push to 1M tokens — which is where the industry is heading — and you’re looking at hundreds of gigabytes per user.

The bottleneck has flipped. Short-context inference is compute-bound (GPU cores are busy doing matrix multiplications). Long-context inference is memory-bound (GPU cores sit idle, waiting for KV data to arrive from HBM). Researchers have measured GPU compute utilization below 20% during long-context decode — the hardware is starving for data, not crunching numbers.

The Obvious Fix: Use Fewer Bits

The KV cache stores every number as a 16-bit float. Each number gets 65,536 possible values — far more precision than the attention mechanism actually needs.

What if we stored each number using 3 bits instead? That gives 8 possible values. The compression ratio is 16/3 = 5.3×. Our 40 GB cache shrinks to ~7.5 GB. Problem solved?

Let’s test it. Here’s a Key vector from our Part 1 example:

Key("cat") = [1.0, 0.1, 0.9]

With 2-bit quantization (4 buckets), we find the range [0.1, 1.0], divide into 4 equally spaced values, and round:

1.0 → Bucket 3 → 1.00 (error: 0.00)
0.1 → Bucket 0 → 0.10 (error: 0.00)
0.9 → Bucket 3 → 1.00 (error: 0.10)

Compressed Key ≈ [1.00, 0.10, 1.00]

Tiny errors. Looks fine. But let’s check what this does to attention.

Small Errors, Exponential Damage

Remember from Part 1: the attention score is a dot product between the Query and the Key, followed by softmax. Let’s trace the damage.

Original dot product:

Query · Key("cat") = [0.8, 0.8, 0.5] · [1.0, 0.1, 0.9] = 1.33

Compressed dot product:

Query · Key("cat") = [0.8, 0.8, 0.5] · [1.00, 0.10, 1.00] = 1.38

Difference: 0.05. That seems harmless. But softmax amplifies small differences exponentially.

Original scores [0.69, 1.33, 1.44] produce attention weights [0.20, 0.38, 0.42].

Compressed scores [0.69, 1.38, 1.44] produce weights [0.19, 0.40, 0.41].

“cat” went from 38% to 40% attention. “sat” dropped from 42% to 41%. The Value blend shifts. The FFN gets slightly different input. The logits shift. Over thousands of generated tokens, these errors don’t cancel — they accumulate.

And we only tested 3-dimensional vectors with small, well-behaved numbers. Real Key vectors have 128 dimensions with a much nastier property.

The Real Problem: Outlier Channels

Here’s what Key vectors actually look like in a trained LLM. Let’s use 8 dimensions (still tiny, but realistic in shape):

Real Key = [0.2, 0.1, 47.3, 0.3, -52.1, 0.1, 0.4, 0.2]

Dimensions 3 and 5: 47.3 and -52.1. These are outlier channels — coordinates with values 100× larger than the rest. They exist in virtually every trained LLM. Researchers have confirmed them across GPT, Llama, Mistral, Qwen, and every other major model family.

Become a Medium member

Why do they exist? Nobody designed them. During training, certain dimensions of the Key/Value space end up carrying disproportionately important information. The model learns to concentrate signal in a few coordinates. It’s an emergent property — a side effect of billions of gradient updates optimising for next-token prediction.

This is not a bug. The model put that information there deliberately — those outlier dimensions carry the most important signal for attention computation. You can’t remove them without destroying the model’s ability to attend correctly.

Why Outliers Destroy Quantization

Now try to compress [0.2, 0.1, 47.3, 0.3, -52.1, 0.1, 0.4, 0.2] with 3-bit quantization. That gives us 8 buckets.

First, find the range: min = -52.1, max = 47.3. Total range = 99.4.

Divide into 8 equally spaced buckets:

Bucket 0 → -52.1
Bucket 1 → -37.9
Bucket 2 → -23.7
Bucket 3 → -9.5
Bucket 4 → 4.7
Bucket 5 → 18.9
Bucket 6 → 33.1
Bucket 7 → 47.3

Each bucket covers ~14.2 units. Now quantize every coordinate:

0.2 → nearest bucket: 4.7 error = 4.5 ❌
0.1 → nearest bucket: 4.7 error = 4.6 ❌
47.3 → nearest bucket: 47.3 error = 0.0 ✅
0.3 → nearest bucket: 4.7 error = 4.4 ❌
-52.1 → nearest bucket: -52.1 error = 0.0 ✅
0.1 → nearest bucket: 4.7 error = 4.6 ❌
0.4 → nearest bucket: 4.7 error = 4.3 ❌
0.2 → nearest bucket: 4.7 error = 4.5 ❌

The two outlier coordinates quantize perfectly. But the 6 normal coordinates — the majority of the vector — all get errors of ~4.5 units. That’s an error of 2,250% relative to original values of ~0.2.

The compressed vector looks like:

Original: [0.2, 0.1, 47.3, 0.3, -52.1, 0.1, 0.4, 0.2]
Compressed: [4.7, 4.7, 47.3, 4.7, -52.1, 4.7, 4.7, 4.7]

Almost every coordinate is wrong. The outliers hijacked the entire quantization range. All 8 buckets had to span from -52 to +47, leaving zero precision for the values that needed it most.

The Dot Product Damage — Numerically

Let’s compute the attention score using a Query of [1, 0, 1, 0, 1, 0, 1, 0].

With the original Key:

[1,0,1,0,1,0,1,0] · [0.2, 0.1, 47.3, 0.3, -52.1, 0.1, 0.4, 0.2]
= 0.2 + 47.3 + (-52.1) + 0.4
= -4.2

With the compressed Key:

[1,0,1,0,1,0,1,0] · [4.7, 4.7, 47.3, 4.7, -52.1, 4.7, 4.7, 4.7]
= 4.7 + 47.3 + (-52.1) + 4.7
= 4.6

The attention score went from -4.2 to +4.6. It didn’t just shift slightly — it changed sign. In softmax, a negative score means “ignore this token.” A positive score means “attend to this token.” The model will now pay attention to something it was supposed to ignore, and ignore something it was supposed to focus on.

The error chain from Part 1 fires in full: Corrupted Key → wrong dot product → wrong softmax weight → wrong Value blend → wrong FFN input → wrong logits → wrong token predicted. All from rounding a few numbers.

The Wall: Every Previous Method Tried and Failed

Researchers spent years trying to handle outliers. None of them solved the root cause.

Mixed Precision keeps outlier channels at 16-bit precision while quantizing the rest to low bits. It works, but the savings are limited — you’re still storing a significant chunk of the cache at full precision, capping compression at roughly 2–3×.

Value Clipping caps outlier values before quantizing. This destroys the signal in those channels. The model put that signal there for a reason — clipping it breaks attention patterns on long sequences.

Per-Channel Scaling uses a different scale factor for each dimension, so outliers get their own range. Better, but it requires running calibration data through the model to determine the right scales. It’s model-specific, dataset-specific, and fragile when the inference distribution shifts.

Dense-and-Sparse Decomposition stores outlier values separately from the main quantized cache. It works in theory, but requires complex custom CUDA kernels, adds overhead during attention computation, and is hard to deploy in production serving frameworks like vLLM or TensorRT-LLM.

All of these are workarounds. All require calibration data, model-specific tuning, or both. None of them eliminate the outlier problem — they just try to patch around it.

The Question Nobody Had Asked

Every previous approach accepted the outlier distribution as a given and tried to work around it. The vectors have outliers, so we need special handling for outlier channels.

Then a team at Google Research, writing what would become the TurboQuant paper (accepted at ICLR 2026), asked a fundamentally different question:

“What if we could make the outliers disappear entirely — before quantizing — without losing any information?”

Not clipping. Not mixed precision. Not special treatment for outlier channels.

What if there existed a mathematical operation that takes a vector with wildly uneven energy distribution — two dimensions at ±50, six dimensions near 0 — and spreads that energy uniformly across all dimensions? And what if that operation preserved every dot product perfectly, so the attention mechanism never knew anything happened?

Such an operation exists. It’s called a rotation.

The answer, the proof, and the full engineering implications are in Part 3.

Join thousands of data leaders on the AI newsletter. Join over 80,000 subscribers and keep up to date with the latest developments in AI. From research to projects and ideas. If you are building an AI startup, an AI-related product, or a service, we invite you to consider becoming a sponsor.

Published via Towards AI


Towards AI Academy

We Build Enterprise-Grade AI. We'll Teach You to Master It Too.

15 engineers. 100,000+ students. Towards AI Academy teaches what actually survives production.

Start free — no commitment:

→ 6-Day Agentic AI Engineering Email Guide — one practical lesson per day

→ Agents Architecture Cheatsheet — 3 years of architecture decisions in 6 pages

Our courses:

→ AI Engineering Certification — 90+ lessons from project selection to deployed product. The most comprehensive practical LLM course out there.

→ Agent Engineering Course — Hands on with production agent architectures, memory, routing, and eval frameworks — built from real enterprise engagements.

→ AI for Work — Understand, evaluate, and apply AI for complex work tasks.

Note: Article content contains the views of the contributing authors and not Towards AI.