Your question is KV Caching for Inference. Take a moment with it on the right.
Talk me through your thinking if you like. When you're confident, submit your answer and I'll grade it like a real screen (7/10 or better passes).
Discuss the necessity of KV caching during inference versus training. Explain how causal self-attention reuses previous keys and values during autoregressive decoding, and why training typically computes all token representations in parallel without a cache. Include the computational and memory tradeoffs, plus implementation considerations for prefill and decode.