Half the Size, Better Performance, Same Accuracy: Understanding W8A8 INT8 LLM Quantization
Large language models are expensive to serve. A model like Llama 3.1 8B in BF16 precision occupies roughly 15GB of GPU memory. In BF16, each of the 8 billion parameters takes 2 bytes to store, which adds up to roughly 15GB just for the weights — and that’s not all. The GPU also needs memory for the KV cache (which stores context for every active request) and for the intermediate computations (“activations”) during inference. A single GPU with 16GB VRAM might technically fit the model weights, but with barely 1GB left for KV cache and activations, it would struggle to serve even a single request. These memory demands also directly limit how many requests you can serve concurrently and how fast each response is generated. ...