I've been reading about LLM quantization for a while, and almost every explanation I found stopped at the same sentence: "it makes the model smaller." Fine. But how ? And what do you give up? So I wrote my notes down, got stuck on a few things, and built a small playground to help me visualize this concept better. For this article, we only go as far as 8-bit quantization. The denser things like GPTQ, GGUF and BitNet are something I'm still reading and will cover in the next few weeks. ...