What do today’s models change inside each block, and why?
A short test to skip this level
Solve these 4 questions on your own. Answer all of them correctly and the level counts as cleared with three stars, and every part of the page opens. Showing an answer doesn’t count.
Number A post-norm stack has 12 layers, each x = LayerNorm(x + f(x)). How many LayerNorms sit on the straight residual path from the input to the output?
CodeWrite RMSNorm for the last axis (eps keeps it safe when x is all zeros).
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
CodeRotate a pair v by pos × theta radians (counter-clockwise).
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
CodeWrite the SwiGLU FFN of section 5 for a batch of word rows x of shape (L, d_model), with the weights W_gate, W_up and W_down. Return one row of d_model numbers per word.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
Warm-up2 questions from earlier levels
A quick review before you start. Optional. Nothing here locks the level.
The model you built in level 17 already has two habits of today’s large language models: it is one stack, and it
puts the LayerNorm in front of each sublayer (pre-norm). A model built today has the same overall design:
embeddings, attention, a feed-forward network, residuals, normalization. Inside each block, three parts are new.
Each one answers a question. Can the normalization do less work? RMSNorm does. Can a score depend only on how far apart
two tokens are, not on where they are? RoPE, positions by rotation, does. Can the feed-forward network reach a lower loss
with the same number of weights? SwiGLU, a gated FFN, does. Section 2 first reviews pre-norm, the change your model
already has.
1. Decoder-only: one stack
The 2017 design had two stacks: an encoder that reads the source and a decoder that writes the target, joined by
cross-attention. Side trip N7 builds it. It is where the cross-attention in that diagram
comes from. Every model from level 16 to level 21 is one stack. Nearly every large language model today is
one stack too. Every block uses masked self-attention, and the “source” is just the beginning of the sequence.
ChooseWhy would one stack beat an encoder plus a decoder for a general-purpose model?
2. Pre-norm, again: nothing rescales the residual path
In level 16, section 3, each sublayer (attention or FFN, written f) was wrapped as x = LayerNorm(x + f(x)): add, then normalize. That is post-norm, the 2017 placement.
Level 16’s decoder block, level 17’s model and nearly every model today use pre-norm: x = x + f(LayerNorm(x)).
Only the input of the sublayer is normalized. One more LayerNorm comes after the last block.
The difference is on the residual path: the straight path from the input to the output (level 16).
Count the LayerNorms on it.
🔒 Answer the question above to unlock
Where the LayerNorm sits: on the residual path, or only in front of the sublayer
Switch between the 2017 design and today’s, and add layers. Count the LayerNorms the input must pass through on the residual path.
LayerNorms on the residual path from input to output: 3.
Every layer rescales the whole sum x, so the input reaches the top only through all of them.
the residual pathLNa LayerNorm on that path
Number A post-norm stack has 12 layers, each x = LayerNorm(x + f(x)). How many LayerNorms sit on the straight residual path from the input to the output?
I got stuck here Pre-norm still normalizes. Where did the normalization go?
Into the sublayer’s input, on the sublayer side. Each attention or FFN still sees a normalized input, so it works with
numbers of a sensible size. What changes is the residual path: things are only added to it. In post-norm, every layer
rescales the whole sum x, including all that the earlier layers added.
Take x = [1, −1, 0.5, −0.5] and a sublayer output of [3, 3, −3, −3]. Post-norm keeps LN(x + f(x)) = [1.29, 0.64, −0.81, −1.13].
From these numbers, x is hard to find again. Pre-norm keeps x + f(x) = [4, 2, −2.5, −3.5]: x is still in there, untouched, just added.
Why does this help training? The gradient goes back along the same path. In a pre-norm stack, the gradient passes the
one final norm first. After that it can go back along the residual path to an early layer, multiplied only by 1:
there is no norm per layer. So early layers always get a gradient
that is not too small, even in a very deep stack.
Go deeper What post-norm needs instead, and what pre-norm costs
In a post-norm stack, that gradient passes through one normalization per layer, each dividing it by its input’s standard deviation.
With dozens of layers, post-norm models only train with a long, careful learning-rate warmup; pre-norm models train well with many more settings.
The cost: in pre-norm the residual path keeps growing as layers add to it, which is why there is one final LayerNorm before the output.
3. RMSNorm: skip the mean
LayerNorm subtracts the mean of each vector, then divides by its standard deviation (std). RMSNorm skips the mean.
It divides by the root mean square, the square root of the average of the squares:
RMSNorm(x)=d1∑ixi2x
🔒 Answer the question above to unlock
LayerNorm and RMSNorm on the same four numbers
Edit any number of x. LayerNorm shifts and scales; RMSNorm only scales.
x =
input xLayerNorm: shifted, then scaledRMSNorm: only scaled
LayerNorm
1. mean = 2.5000
2. subtract it: [-1.50, -1.50, 0.50, 2.50]
3. std = 1.6583
4. divide: [-0.9045, -0.9045, 0.3015, 1.5076]
output mean 0, std 1
RMSNorm
1. mean of squares = ?
2. rms = √ that = ?
3. divide: [?, ?, ?, ?]
no mean subtracted: the output keeps its sign pattern
Some RMSNorm numbers show “?” until you answer the questions below. Edit x to explore other vectors.
Number x = [1, 1, 3, 5]. The root mean square is √((1² + 1² + 3² + 5²) / 4). What is it?
🔒 Answer the question above to unlock
Number x = [1, 1, 3, 5] and its root mean square is 3. What is the last number of RMSNorm(x)? (2 decimals)
Experiments showed that subtracting the mean contributes little: models with RMSNorm train just as well. RMSNorm computes
one average instead of two and has no shift. The saving is small for one norm, but every layer makes it at every step.
In real models, both versions then multiply each column by a learned gain. Level 16’s examples did not use it, to keep the
numbers simple. RMSNorm also has no learned shift.
In the code, keepdims=True keeps the averaged axis with size 1, so a (B, L, d) input gives a (B, L, 1) result.
So each word is divided by its own number: np.mean(np.ones((2, 3)), axis=-1, keepdims=True).shape is (2, 1).
CodeWrite RMSNorm for the last axis (eps keeps it safe when x is all zeros).
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
4. RoPE: positions by rotation
Level 16 added a position vector to each word. Today’s models instead rotate q and k by an angle that grows with
the position. Take the numbers of q in pairs; rotate each pair by position × θ. Do the same for k. Then take the dot product as usual.
Rotating a pair [x, y] counter-clockwise by an angle a gives
[xcosa−ysina,xsina+ycosa]
Check it on a quarter turn, a = 90°, where cos a = 0 and sin a = 1: [1, 0] becomes [0, 1], and [0, 1] becomes [−1, 0].
The length never changes; only the direction does. Angles here are in radians: a full turn is 2π ≈ 6.28, so 1 radian is about 57°.
Here is a single pair with θ = 0.5 radians per position. q = [1, 0] and k = [0, 1]: without rotation they are at
right angles, so their score is 0 wherever they are in the sentence.
🔒 Answer the question above to unlock
RoPE turns q and k by their positions; the score depends only on m − n
Move the two positions. Then use “both +1” and watch the score stay the same.
The arc is the angle between q and k (in radians); the score is its cosine, so it only changes when m − n does.
Without RoPE the score is always [1, 0]·[0, 1] = 0, wherever the words are.
score q·k = 0.0000· distance m − n = 0
q, turned by m × 0.5k, turned by n × 0.5before turning
Number θ = 0.5 radians per position. q sits at position 2. By how many radians is q rotated?
🔒 Answer the question above to unlock
Chooseq sits at position 2 and k at position 0, and their score is some number s. Now move both 10 positions later (12 and 10). The score…
Number q = [1, 0] at position 2 and k = [0, 1] at position 0, θ = 0.5. After rotating, q = [cos 1, sin 1] ≈ [0.54, 0.84] and k stays [0, 1]. What is the score q·k? (2 decimals)
Rotating q by mθ and k by nθ changes the angle between them by exactly (m − n)θ. A dot product only depends on
the lengths and the angle, so the score only depends on the distance m − n, never on where the pair sits in the text.
A language model needs this: “the word two places back” means the same at position 5 and at position 5,000.
CodeRotate a pair v by pos × theta radians (counter-clockwise).
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
Go deeper Many pairs, many frequencies
A real head has d_k numbers, so d_k / 2 pairs. Each pair rotates at its own speed: θᵢ = 10000^(−2i/d_k).
With d_k = 4 the two pairs turn at 1 and 0.01 radians per position. Fast pairs distinguish neighboring words.
Slow pairs still change between words that are thousands of positions apart. A clock works the same way: the minute
hand is fast and the hour hand is slow.
demo.py checks it with a random 4-number q and k: the score at positions (7, 4) and (107, 104) is the same 1.058526.
RoPE does not make a model work on texts longer than the ones it trained on. A distance longer than any seen in
training gives the slow pairs angles the model has never seen, so the scores there are still new to it.
With a KV cache (level 18), k is rotated before it is stored. So the k rows in the cache already carry their
positions, and an old k is never rotated again. v is not rotated at all: RoPE changes only the scores, not the
numbers that the weights mix.
5. SwiGLU: a gate inside the feed-forward network
The 2017 FFN is ReLU(x @ W1) @ W2. Today’s FFN has a third matrix and multiplies two things: a gate that decides
how much passes, and a value that is passed:
The symbol ⊙ means “multiply number by number”, like * in NumPy. Each hidden unit computes one number from its gate
column and one from its value column, then multiplies them.
🔒 Answer the question above to unlock
SwiGLU: a SiLU gate decides how much of the value passes
Move the gate g and the value u. Watch the output bar: the gate scales the value.
value u2.0
× SiLU(g)-0.2689
output-0.5379
Unlike ReLU, a slightly negative gate keeps a small part of the value, with its sign flipped.
SiLU(g) = g × sigmoid(g) = -0.2689 · output = SiLU(g) × u = -0.5379
SiLU(z) = z · sigmoid(z)ReLU, for comparisonyour gate gpositivenegative
Number SiLU(z) = z × sigmoid(z) = z / (1 + e^−z). Use e^−1 ≈ 0.37. What is SiLU(1)? (2 decimals)
🔒 Answer the question above to unlock
Three matrices instead of two would mean 50% more weights. So the hidden size d_ff is shrunk to keep the total the same.
Number A plain FFN has 2 matrices of d_model × d_ff. A SwiGLU FFN has 3 matrices of d_model × d_ff′. For the same number of weights, what is d_ff′ / d_ff? (2 decimals)
I got stuck here If the total number of weights is the same, why is it better?
There is no simple proof. Measured on real training runs, gated FFNs reach a lower loss for the same number of weights.
Here is one way to see it. In ReLU, a unit is on or off, and only its own input decides. In SwiGLU, one product of x
(the gate) decides how much of another product of x (the value) passes. So the FFN can multiply two different numbers
made from the same word. A plain FFN cannot do that in one layer.
Now write it. Each word is one row of x.
x @ W_gate and x @ W_up are both (L, d_ff).
* multiplies them number by number (the ⊙).
@ W_down turns each row into d_model numbers again.
CodeWrite the SwiGLU FFN of section 5 for a batch of word rows x of shape (L, d_model), with the weights W_gate, W_up and W_down. Return one row of d_model numbers per word.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
6. A modern block, all together
Put the three new parts into level 16’s pre-norm decoder block. You get the block of nearly every large language model today:
h = RMSNorm(x)# causal mask; q and k rotated by RoPE, no position vector addedx = x + attention(h)h = RMSNorm(x)x = x + SwiGLU(h)
A stack of these blocks, one final RMSNorm, then the output layer. Every part on this page made something better:
pre-norm keeps deep stacks stable;
RMSNorm is cheaper;
RoPE makes scores depend on distance only;
SwiGLU reaches a lower loss with the same number of weights.
🔒 Answer the question above to unlock
Try it
In the RoPE lab of section 4, set q’s position to 2 and k’s to 0, then press “both +1” a few times.
The score stays 0.84: both vectors turn, but the distance stays 2. Now move only k’s slider and watch the score
follow the distance m − n. In a very long text, the same distance appears at position 5 and at position 5,000.
Which of the three new parts makes the scores agree there?
The GPT you write in level 21 uses pre-norm and one of the three new parts, RMSNorm. It keeps learned position vectors and a
ReLU FFN. At the end of level 21, an optional exercise lets you replace the learned positions with RoPE. Before that, level 20 asks what
running such a model costs: how much GPU memory it needs while it writes, and how many tokens per second it can produce.
Recap
a summary for when you finish the level
The key formulas and common mistakes appear here once you clear the level.
You can now
Explain why pre-norm, which levels 16 and 17 already use, trains deep stacks better than the 2017 post-norm.
Compute RMSNorm and a RoPE rotation by hand, and write both in NumPy.
Write a SwiGLU FFN, and size its hidden layer so it has as many weights as a plain FFN.
Keep in mind
Pre-norm: x = x + f(norm(x)), plus one final norm; post-norm: x = norm(x + f(x))
RMSNorm(x)=x/d1∑ixi2, no mean subtracted
RoPE: rotate each pair of q and k by position × θ; the score depends only on the distance
FFN(x)=(SiLU(xWgate)⊙(xWup))Wdown, with dff′=32dff
A modern block: x = x + attn(RMSNorm(x)) with RoPE on q and k, then x = x + SwiGLU(RMSNorm(x))
Common mistakes
Averaging over the wrong axis (or without keepdims=True), so each word is not normalized separately.
Putting SiLU on the value instead of the gate, or multiplying gate and value with @ instead of number by number.