Tech
A transformer block is six tensors and a bus: read it like a pipeline
Originally published at cchinchilla.dev . Part 5 of From code to weights, a 12-part series on ML fundamentals for engineers.
Part 4 ended on one stage. A decoder is a stack of blocks, twelve in GPT-2 small, and attention is one of two stages in each. This post reads a whole block the way you'd read a pipeline: what flows in, what each stage writes, and where the parameters and the cost end up.
They don't end up in the same place.
I read Block in nanoGPT's model.py with the shape...
Read the full discussion on Dev.to
This article was aggregated from Dev.to. Click to join the conversation.
View on Dev.to