Originally published at cchinchilla.dev . Part 5 of From code to weights, a 12-part series on ML fundamentals for engineers. Part 4 ended on one stage. A decoder is a stack of blocks, twelve in GPT-2 small, and attention is one of two stages in each. This post reads a whole block the way you'd read a pipeline: what flows in, what each stage writes, and where the parameters and the cost end up. They don't end up in the same place. I read Block in nanoGPT's model.py with the shape...