Why can a 16 GB model run on an 8 GB GPU? Quantization. This article shows how it works with a small Python script, then explains the GGUF file format and what the name Q4_K_M actually means.
One thing to clarify first: the demonstration doesn’t use the real Q4_K_M algorithm. It uses a tiny hand-picked list of weights and a simplified method, because the goal is to show the principle.
Quantization
A model is basically storage for a huge list of numbers, and those numbers are called weights. If you want to learn more about how weights impact model performance and hardware, you can read why local models need so much VRAM.
However, billions of numbers are not easy to store in VRAM, so we need a reliable method to reduce the size of the weights. This is where quantization comes into play. It reduces their precision.
By precision, we mean how accurately each number is stored. Higher precision uses more bits to represent a number, while lower precision uses fewer bits and takes up less memory.
For example, quantization can reduce weights from 16-bit floating-point numbers to 4-bit integers. This makes the model much smaller and can speed up inference, although it may cause some loss in accuracy.
Imagine a weight has the value 0.73 stored as a 16-bit floating-point number.
With 4-bit quantization, the model cannot store every possible decimal value, so 0.73 might be mapped to a nearby 4-bit value such as 4. Using a scale factor of about 0.2, that value represents approximately:
4 × 0.2 = 0.8
The scale factor here is the value that tells us how much each quantized step represents.
So instead of storing 0.73 exactly with 16 bits, the model stores a much smaller 4-bit value and reconstructs it as approximately 0.8 when needed.
However, in the end, the reconstructed value is close to the original number but not exactly the same. So we lose some numerical precision, which can lead to some quality loss, but can gain better memory usage and inference performance.
Demonstrating Scaling Factor
Before moving further, we need to understand the scaling factor because it plays a key role in how quantization works.
This example won’t use an industry-standard method, but it will make the basic idea much easier to understand.
First, we need to find the largest absolute weight. We use the absolute value because we only care about how far each number is from zero.
With these weights:
weights = [-0.4, 0.4, 0.73, 1.0, 0.2]
the largest absolute weight is 1.0.
Now let’s imagine our quantized values can range from -5 to 5. To calculate the scaling factor, we divide the largest absolute weight by the largest quantized value, so the scale is 1.0 / 5 = 0.2.
Next, we divide each weight by the scale and round the result to the nearest integer. To see what we lost, we then multiply those integers by the same scale to restore approximate weights and compare them with the originals.
weights = [-0.4, 0.4, 0.73, 1.0, 0.2]
largest_weight = max(abs(w) for w in weights)
largest_quantized_value = 5
scale = largest_weight / largest_quantized_value
quantized = [round(w / scale) for w in weights]
restored = [q * scale for q in quantized]
errors = [abs(w - r) for w, r in zip(weights, restored)]
print(scale)
print(quantized)
print([round(r, 2) for r in restored])
print([round(e, 2) for e in errors])
The output will be 0.2, then [-2, 2, 4, 5, 1], then the restored values [-0.4, 0.4, 0.8, 1.0, 0.2], then the errors [0.0, 0.0, 0.07, 0.0, 0.0]. The quantized values fit in 4 bits instead of 16, which is much better for memory usage. Four of the five weights came back exactly, while 0.73 became 0.8. That 0.07 gap is the quality loss quantization trades for size.
Understanding GGUF
GGUF is basically a container here. It’s a single file format designed for fast loading and saving of models. Of course, it doesn’t store model numbers using simple arrays like we did.
A GGUF file stores the model’s tensors, which contain the weights and other numerical data the model needs. If the model is quantized, those values are usually stored in a packed quantized format, often together with things like scales used to reconstruct approximate values during inference.
A tensor is basically a structured group of numbers used by the model to store data such as weights. You can think of it like an array in Python.
Real model tensors can contain millions of numbers and can have multiple dimensions, but they are still just organized collections of numerical values.
GGUF also stores extra information such as model metadata, tensor shapes, tokenizer data, and configuration details.
Basically, GGUF is a container that holds the model’s numerical weights, which may be quantized, plus the information needed to understand and load them.
Quantization, on the other hand, describes how the numbers inside the GGUF file are stored.
A GGUF model can use formats like F16, Q8_0, or Q4_K_M. That’s why model filenames often end with the quantization type, such as -F16.gguf.
Understanding Quantization Types
Quantization types are not so different from what we already talked about. Names such as F16 and Q8_0 might look ugly at first, but they basically show how the model’s weights are stored inside the file.
F16 stores the weights in a 16-bit floating-point format, which keeps more precision but also makes the model much larger. Q8_0 reduces them to an 8-bit quantized format, making the model smaller while still keeping most of its quality. Lastly, Q4_K_M is a more aggressive 4-bit quantization format that saves much more memory.
The letters after the number also mean something: Q4 means roughly 4-bit quantization, K means it belongs to the k-quant family, which groups weights into super-blocks with compressed scales, and M means the Medium variant, which uses a mix of quantization levels to balance size and quality.
What Does the K in Q4_K_M Mean?
The K in Q4_K_M means it belongs to the K-quant family. Older formats like Q4_0 and Q8_0 already split weights into small blocks of 32, each with its own scale. K-quants go further: they group blocks into larger super-blocks of 256 weights, split those into sub-blocks, and store the sub-block scales in a compressed low-bit form. That saves space while keeping better accuracy.
You can think of it as dividing a huge list of weights into smaller groups and quantizing each group separately. This gives the quantizer more control and usually keeps better accuracy compared to using one simple scale for everything.
So unlike our earlier example, where we used one scale for all the weights, real formats use many small groups, each with its own scale. K-quants just organize those groups more cleverly.
What Does the M in Q4_K_M Mean?
The M in Q4_K_M means the Medium variant. It uses a mix of quantization levels for different tensors to balance model size and quality. Instead of using exactly the same quantization type for every tensor, Q4_K_M keeps some more important tensors at a slightly higher precision.
This helps reduce quality loss while still keeping the model much smaller than higher-precision formats.
So yes, some tensors can use higher-precision quantization than others, depending on how important they are to model quality.
Deciding Quantization For Local AI
Memory is the biggest factor here: VRAM first, then RAM. Lower-bit formats fit in less memory but lose more quality, so the real choice is always size versus quality.
Here is how one model compares (approximate file sizes for Llama 3.1 8B):
| Format | Bits per weight | File size |
| F16 | 16 | ~16 GB |
| Q8_0 | 8.5 | ~8.5 GB |
| Q4_K_M | ~4.9 | ~4.9 GB |
Q4_K_M is a common choice for local AI because it cuts the model to roughly a third of its F16 size while generally keeping quality acceptable. Q8_0 and F16 keep more precision if you have VRAM to spare, but they need much more space and memory bandwidth.
