Inside the Block: FFN, Residuals, Normalization
The other half of the transformer block: feed-forward networks, the residual stream, and why normalization placement matters.
Content last verified 2026-09.
Lessons
- The Feed-Forward Network
- The Residual Stream
- LayerNorm, RMSNorm, and Where They Sit
- Assembling the Block
Sources
- Vaswani et al. (2017), Attention Is All You Need, arXiv:1706.03762
- Shazeer (2020), GLU Variants Improve Transformer, arXiv:2002.05202
- Ba, Kiros & Hinton (2016), Layer Normalization, arXiv:1607.06450
- Zhang & Sennrich (2019), Root Mean Square Layer Normalization, arXiv:1910.07467
- He et al. (2015), Deep Residual Learning for Image Recognition, arXiv:1512.03385