LLM Systems & Inference
KV Cache: Understanding Prefill, Decode, and Attention Caching
A technical walkthrough of KV caching, explaining how K and V are stored across decoder blocks, why Q is computed only for the current token, and how prefill and decode differ.
Oct 06, 20267 pages · Technical Note