Xiao Wang Asks, Da Wang Answers: From KV Cache to YOCO and Shared Memory
An 11-minute animated explainer. Xiao Wang asks and Da Wang answers as one token's KV cache works through the four questions it must settle before it can become shared memory: identity, reuse, state and reading.
Chapters
- Start Is “compute each one, store it, find it again by hash” enough? Four questions: identity, reuse, state, reading.
- 01 Compute Q2–Q3 Q is compared with every K, then V is mixed by weight; prefill and decode; a cache hit still needs computation.
- 02 Identity Q3–Q7 Check identity before computing; one changed token invalidates everything after it; the same “apple” diverges in higher layers; chained hashes; block granularity 999 → 992.
- 03 Reuse Q8–Q12 Claude cache breakpoints and TTL; DeepSeek’s A+B / A+C / A+D; cache admission; a sliding window cannot be cut back; four conditions for reuse.
- 04 State Q13–Q15 KV grows with layers × length; YOCO shares one global KV, 16.4 GB → 0.42 GB; prefill skips the second half for past positions.
- 05 Ledger Q16–Q21 40 layers concentrated into 4 sources; 890 B/token added up item by item; K=V and inverse RoPE; a counterexample of bounded replay; coarse filtering, then reranking; Engram.
- 06 Application Q22 How to read usage; a counterexample where a higher hit rate costs more; back to the four questions.
The top-right corner of each frame cites the matching question in the companion article (for example “Article Q06”). Figures specific to V4.1 are recomputed from the technical excerpt that accompanies the article, not measured.
Download MP4 (20.7 MiB)