Tech
Fine-tuning a 7B model needs 112 GB. The model is only 14 GB of it.
Ask how much memory it takes to fine-tune a 7B model and the instinct is "the model's 14 GB in fp16, so a bit more than that". The real figure is about 112 GB, before you've stored a single activation. The model is 14 GB of it.
Once you see where the other 98 GB goes, LoRA and QLoRA stop looking like clever tricks and start looking obvious.
Where the memory actually goes
The accounting comes from the ZeRO paper, and it's worth reading in their words:
In total, this result...
Read the full discussion on Dev.to
This article was aggregated from Dev.to. Click to join the conversation.
View on Dev.to