LLM Systems & Inference

KV Cache: Understanding Prefill, Decode, and Attention Caching

A technical walkthrough of KV caching, explaining how K and V are stored across decoder blocks, why Q is computed only for the current token, and how prefill and decode differ.

Oct 06, 20267 pages · Technical Note

Download the original PDF