2 Numbers Out of 128 Destroy the Entire KV Cache — Why LLM Compression Keeps Failing
Author(s): Ansh Saxena Originally published on Towards AI. If you haven’t read Part 1, here’s the one-paragraph setup: every time an LLM generates a token, the attention mechanism needs to “look back” at every previous token. To avoid recomputing them, the model …