Skip to content
Road to Intelligence

How Big Is the Cache?

Serve real Llama models to many users at once: add up weights and KV cache against GPU memory, change the number of key/value heads, and see the memory-bandwidth limit on tokens per second.

From Chapter 11: Inside Modern LLMs

Try it

How Big Is the Cache?

Know well8 min