Skip to content

Chapter 12 — The Transformer Architecture, Deep Dive

Part IV — Transformers and LLMs · 4–6 weeks

What you will learn

This chapter describes the Transformer, the neural network architecture used by almost all current large language models (LLMs), at the level of detail needed to implement it. It covers scaled dot-product attention, causal masking, multi-head attention, positional encodings, residual connections, layer normalization, and the feed-forward layer, first in the encoder-decoder model of the original paper and then in the decoder-only stack used by GPT-style models. It also summarizes the changes that current open models make to the 2017 design — rotary position embeddings (RoPE), root mean square normalization (RMSNorm), gated feed-forward layers (SwiGLU), grouped-query attention (GQA), and mixture-of-experts (MoE) layers — because the model reports read from Chapter 15 onward assume them. The milestone implements multi-head attention; the tokenizer follows in Chapter 13, and the full model with its training loop in Chapter 14.

Topics

  • The limits of recurrence and the design goals of "Attention Is All You Need"
  • Scaled dot-product attention: queries, keys, values, and the scaling factor
  • Self-attention, cross-attention, and causal masking
  • Multi-head attention and its tensor shapes
  • Positional encodings: sinusoidal, learned absolute, and rotary (RoPE)
  • The Transformer block: residual connections, normalization (LayerNorm and RMSNorm, pre-norm and post-norm placement), and the feed-forward layer
  • Encoder-only (BERT), encoder-decoder (T5), and decoder-only (GPT) models, and the reasons LLMs use the decoder-only design
  • Parameter counts, and the quadratic time and memory cost of attention in the sequence length
  • Changes in current open LLMs: SwiGLU, grouped-query attention (GQA), and mixture-of-experts (MoE) layers (overview)

Resources

Suggested path. Start with the two 3Blue1Brown videos and The Illustrated Transformer, then watch CS224N 2024 lecture 8 for the derivation. Next, watch Karpathy's "Let's build GPT" while typing the code, and work through Raschka Chapter 3 (or Prince Chapter 12, which is free) for a second implementation of attention. Read "Attention Is All You Need" after that, with Formal Algorithms for Transformers open as the reference for shapes and pseudocode. CS336 lectures 3 and 4 cover the changes made by current models and fit at the end; when time is short, omit CS25, The Annotated Transformer, and the optional papers.

University courses

  • Stanford — CS224N: Natural Language Processing with Deep Learning by Diyi Yang and Yejin Choi (free; Winter 2026 slides, notes, and assignments + Spring 2024 lecture videos by Christopher Manning; start here: 2024 lecture 8 "Transformers" derives self-attention from the limits of the recurrent neural networks (RNNs) of Chapter 11, and lecture 9 "Pretraining" compares encoder, encoder-decoder, and decoder models; Assignment 3 covers self-attention and Transformers).
  • Stanford — CS336: Language Modeling from Scratch by Tatsunori Hashimoto and Percy Liang (free; Spring 2026 lecture videos and assignments; advanced: lecture 3 "Architectures, hyperparameters" and lecture 4 "Attention alternatives and mixture of experts" survey the design choices of current LLMs; Assignment 1 is scheduled in Chapters 13 and 14).
  • Stanford — CS25: Transformers United (free seminar; optional; sixth edition in Spring 2026; the recordings on YouTube include the overview lectures, one of them by Andrej Karpathy, followed by guest talks on current research).

Online courses (MOOCs)

  • Andrej Karpathy — Let's build GPT: from scratch, in code, spelled out (free; 1h56m; Lecture 7 of Neural Networks: Zero to Hero, after the six lectures used in Chapter 10; start here: builds a character-level, decoder-only Transformer on Tiny Shakespeare in about 200 lines of PyTorch; Chapter 14 reuses this code and does not require a second viewing).
  • DeepLearning.AI — How Transformer LLMs Work by Jay Alammar and Maarten Grootendorst (free; 1h44m; the lessons from "Architectural Overview" to "Mixture of Experts" cover the Transformer block, self-attention, the key-value cache, grouped-query attention, and MoE; the tokenizer lessons belong to Chapter 13).
  • Hugging Face — LLM Course, Chapter 1: Transformer models (free; about 2 hours; sections 4–6 describe the encoder-only, decoder-only, and encoder-decoder families and the tasks each is used for).

Books

  • Book: Sebastian Raschka, Build a Large Language Model (From Scratch) (Manning, 2024) — official page (paid; free code at rasbt/LLMs-from-scratch and free companion videos; start here: Chapter 3 "Coding Attention Mechanisms", from simplified self-attention to causal multi-head attention; Chapter 4 is used in Chapter 14).
  • Book: Simon J.D. Prince, Understanding Deep Learning (MIT Press, 2023) — official page (free PDF; Chapter 12 "Transformers", with notebooks on self-attention, multi-head attention, and tokenization).

Lectures, papers and articles

Milestone

Implement multi-head self-attention with a causal mask from scratch in PyTorch, copy its weights into nn.MultiheadAttention, and verify that both modules produce the same output to within 1e-5. Verify causality by changing the last input token and checking that all earlier output positions are unchanged. Measure the size of the attention matrix at sequence lengths 512 and 2,048, and explain in one paragraph why time and memory grow as O(n²).

Estimated time

4–6 weeks.