The whole article is designed to answer the question for people who wonder if VRAM is really essential for AI models, but if you need a quick answer, the answer is yes.
Weights
AI models need to keep their parameters in memory while they’re running, kind of like a game loading its files before you can play.
During inference, those parameters are usually stored in the GPU’s VRAM, which is basically the GPU’s own memory.
An 8-billion-parameter model can need around 16 GB of VRAM if each parameter uses 16 bits. That’s a lot, so models are often quantized, which stores the parameters using fewer bits and greatly reduces memory usage.
For example, Llama 3.1 8B has about 8.03 billion parameters. With the Q4_K_M quantization format, it uses roughly 4.89 bits per parameter, making the model much smaller.
Q4_K_M is not simply “4 bits everywhere.” Most weights use a compact format, while some important parts of the model get extra precision to reduce quality loss.
| Format | Bits/weight | Approx. weight size. |
| FP16 | 16.0 | 16.1 GB |
| Q8_0 | 8.5 | 8.5 GB |
| Q4_K_M | 4.89 | 4.9 GB |
Bits/weight come from the llama.cpp quantize README.
KV cache
While you are chatting with the model, the model stores keys and values for every earlier token so it doesn’t recompute them. It’s like remembering past messages so it doesn’t lose the context.
The more context the model has to remember from earlier messages, the more space its KV cache needs, and that memory usage also grows with the model’s layers, KV heads, head size, and the precision used to store the keys and values.
For example Llama 3.1 8B. Its config shows 32 layers, 32 attention heads, 8 key/value heads, and a head dimension of 128. At 16-bit precision, that’s 128 KiB per token.
| Context | KV cache |
| 8k tokens | 1 GiB |
| 32k tokens | 4 GiB |
| 128k tokens | 16 GiB |
That’s why a model that fits in your VRAM with a 4k context can suddenly run out of memory at 32k. The more tokens the model has to remember, the larger its KV cache becomes.
However, Llama 3.1 also uses an optimization called GQA to reduce this memory usage, and it lets multiple attention heads share the same stored keys and values instead of each one keeping its own copy.
In this example, the model has 32 attention heads but only 8 KV heads, giving about a 4× reduction in KV cache size. Without GQA, the 128k context cache would need around 64 GB of memory.
Speed of memory
We talked about the 2 big reasons why we need a large VRAM capacity to feed our AI models, but larger space is not the only important factor for local AI models.
The other essential feature is memory bandwidth. When generating text token by token, performance is often limited more by memory bandwidth than raw compute. That’s why VRAM matters not only for storage, but also for speed.
Additional note: some local AI models might even use RAM for their computing, but it won’t be as fast as VRAM for computing. That’s why RAM speed is as essential as its storage.
If each token needs to read roughly the full 4.9 GB of Q4_K_M weights, the theoretical ceilings are about 200 tokens/sec on the 4090 and 70 tokens/sec on the 3060.
Final verdict
At 16-bit precision, yes. But once it’s quantized, a capable 8B model can fit in around 5 GB. The bigger problem usually starts with context length rather than the model size itself. So while calculating how much VRAM you need, you also have to consider the KV cache, not only the model file size.
Sources
Meta Llama 3.1 8B model page and config
Ainslie, J., Lee-Thorp, J., de Jong, M., Zemlyanskiy, Y., Lebrón, F., & Sanghai, S. (2023). “GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints.” Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 4895–4901, Singapore. DOI: 10.18653/v1/2023.emnlp-main.298
