How Big Is the Cache?
Serve real Llama models to many users at once: add up weights and KV cache against GPU memory, change the number of key/value heads, and see the memory-bandwidth limit on tokens per second.
From Chapter 11: Inside Modern LLMs
Try it
How Big Is the Cache?
Know well8 min