Level 19 · Theory · runs in your browser

Modern blocks

What do today’s models change inside each block, and why?

The model you built in level 17 already has two habits of today’s large language models: it is one stack, and it puts the LayerNorm in front of each sublayer (pre-norm). A model built today has the same overall design: embeddings, attention, a feed-forward network, residuals, normalization. Inside each block, three parts are new. Each one answers a question. Can the normalization do less work? RMSNorm does. Can a score depend only on how far apart two tokens are, not on where they are? RoPE, positions by rotation, does. Can the feed-forward network reach a lower loss with the same number of weights? SwiGLU, a gated FFN, does. Section 2 first reviews pre-norm, the change your model already has.

1. Decoder-only: one stack

The 2017 design had two stacks: an encoder that reads the source and a decoder that writes the target, joined by cross-attention. Side trip N7 builds it. It is where the cross-attention in that diagram comes from. Every model from level 16 to level 21 is one stack. Nearly every large language model today is one stack too. Every block uses masked self-attention, and the “source” is just the beginning of the sequence.

ChooseWhy would one stack beat an encoder plus a decoder for a general-purpose model?

2. Pre-norm, again: nothing rescales the residual path

In level 16, section 3, each sublayer (attention or FFN, written f) was wrapped as x = LayerNorm(x + f(x)): add, then normalize. That is post-norm, the 2017 placement. Level 16’s decoder block, level 17’s model and nearly every model today use pre-norm: x = x + f(LayerNorm(x)). Only the input of the sublayer is normalized. One more LayerNorm comes after the last block.

The difference is on the residual path: the straight path from the input to the output (level 16). Count the LayerNorms on it.

🔒 Answer the question above to unlock

Where the LayerNorm sits: on the residual path, or only in front of the sublayer

Switch between the 2017 design and today’s, and add layers. Count the LayerNorms the input must pass through on the residual path.

input xoutputlayer 1attention / FFN+LNlayer 2attention / FFN+LNlayer 3attention / FFN+LN
LayerNorms on the residual path from input to output: 3. Every layer rescales the whole sum x, so the input reaches the top only through all of them.
the residual path LNa LayerNorm on that path
Number A post-norm stack has 12 layers, each x = LayerNorm(x + f(x)). How many LayerNorms sit on the straight residual path from the input to the output?
I got stuck here Pre-norm still normalizes. Where did the normalization go?

Into the sublayer’s input, on the sublayer side. Each attention or FFN still sees a normalized input, so it works with numbers of a sensible size. What changes is the residual path: things are only added to it. In post-norm, every layer rescales the whole sum x, including all that the earlier layers added.

Take x = [1, −1, 0.5, −0.5] and a sublayer output of [3, 3, −3, −3]. Post-norm keeps LN(x + f(x)) = [1.29, 0.64, −0.81, −1.13]. From these numbers, x is hard to find again. Pre-norm keeps x + f(x) = [4, 2, −2.5, −3.5]: x is still in there, untouched, just added.

Why does this help training? The gradient goes back along the same path. In a pre-norm stack, the gradient passes the one final norm first. After that it can go back along the residual path to an early layer, multiplied only by 1: there is no norm per layer. So early layers always get a gradient that is not too small, even in a very deep stack.

Go deeper What post-norm needs instead, and what pre-norm costs

In a post-norm stack, that gradient passes through one normalization per layer, each dividing it by its input’s standard deviation. With dozens of layers, post-norm models only train with a long, careful learning-rate warmup; pre-norm models train well with many more settings. The cost: in pre-norm the residual path keeps growing as layers add to it, which is why there is one final LayerNorm before the output.

3. RMSNorm: skip the mean

LayerNorm subtracts the mean of each vector, then divides by its standard deviation (std). RMSNorm skips the mean. It divides by the root mean square, the square root of the average of the squares:

RMSNorm(x)=x1d∑ixi2\text{RMSNorm}(x) = \frac{x}{\sqrt{\frac{1}{d}\sum_i x_i^2}}
🔒 Answer the question above to unlock

LayerNorm and RMSNorm on the same four numbers

Edit any number of x. LayerNorm shifts and scales; RMSNorm only scales.

x =
input x
LayerNorm: shifted, then scaled
?
RMSNorm: only scaled

LayerNorm

1. mean = 2.5000

2. subtract it: [-1.50, -1.50, 0.50, 2.50]

3. std = 1.6583

4. divide: [-0.9045, -0.9045, 0.3015, 1.5076]

output mean 0, std 1

RMSNorm

1. mean of squares = ?

2. rms = √ that = ?

3. divide: [?, ?, ?, ?]

no mean subtracted: the output keeps its sign pattern

Some RMSNorm numbers show “?” until you answer the questions below. Edit x to explore other vectors.
Number x = [1, 1, 3, 5]. The root mean square is √((1² + 1² + 3² + 5²) / 4). What is it?
🔒 Answer the question above to unlock
Number x = [1, 1, 3, 5] and its root mean square is 3. What is the last number of RMSNorm(x)? (2 decimals)

Experiments showed that subtracting the mean contributes little: models with RMSNorm train just as well. RMSNorm computes one average instead of two and has no shift. The saving is small for one norm, but every layer makes it at every step. In real models, both versions then multiply each column by a learned gain. Level 16’s examples did not use it, to keep the numbers simple. RMSNorm also has no learned shift.

In the code, keepdims=True keeps the averaged axis with size 1, so a (B, L, d) input gives a (B, L, 1) result. So each word is divided by its own number: np.mean(np.ones((2, 3)), axis=-1, keepdims=True).shape is (2, 1).

CodeWrite RMSNorm for the last axis (eps keeps it safe when x is all zeros).

Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs

4. RoPE: positions by rotation

Level 16 added a position vector to each word. Today’s models instead rotate q and k by an angle that grows with the position. Take the numbers of q in pairs; rotate each pair by position × θ. Do the same for k. Then take the dot product as usual.

Rotating a pair [x, y] counter-clockwise by an angle a gives

[ xcos⁡a−ysin⁡a,    xsin⁡a+ycos⁡a ][\,x\cos a - y\sin a,\;\; x\sin a + y\cos a\,]

Check it on a quarter turn, a = 90°, where cos a = 0 and sin a = 1: [1, 0] becomes [0, 1], and [0, 1] becomes [−1, 0]. The length never changes; only the direction does. Angles here are in radians: a full turn is 2π ≈ 6.28, so 1 radian is about 57°.

Here is a single pair with θ = 0.5 radians per position. q = [1, 0] and k = [0, 1]: without rotation they are at right angles, so their score is 0 wherever they are in the sentence.

🔒 Answer the question above to unlock

RoPE turns q and k by their positions; the score depends only on m − n

Move the two positions. Then use “both +1” and watch the score stay the same.

1.57qk
q [1, 0] turned by 0 × 0.5 = 0.0 rad → [1.0000, 0.0000]
k [0, 1] turned by 0 × 0.5 = 0.0 rad → [0.0000, 1.0000]

The arc is the angle between q and k (in radians); the score is its cosine, so it only changes when m − n does. Without RoPE the score is always [1, 0]·[0, 1] = 0, wherever the words are.

score q·k = 0.0000 · distance m − n = 0
q, turned by m × 0.5 k, turned by n × 0.5 before turning
Number θ = 0.5 radians per position. q sits at position 2. By how many radians is q rotated?
🔒 Answer the question above to unlock
Chooseq sits at position 2 and k at position 0, and their score is some number s. Now move both 10 positions later (12 and 10). The score…
Number q = [1, 0] at position 2 and k = [0, 1] at position 0, θ = 0.5. After rotating, q = [cos 1, sin 1] ≈ [0.54, 0.84] and k stays [0, 1]. What is the score q·k? (2 decimals)

Rotating q by mθ and k by nθ changes the angle between them by exactly (m − n)θ. A dot product only depends on the lengths and the angle, so the score only depends on the distance m − n, never on where the pair sits in the text. A language model needs this: “the word two places back” means the same at position 5 and at position 5,000.

CodeRotate a pair v by pos × theta radians (counter-clockwise).

Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs

Go deeper Many pairs, many frequencies

A real head has d_k numbers, so d_k / 2 pairs. Each pair rotates at its own speed: θᵢ = 10000^(−2i/d_k). With d_k = 4 the two pairs turn at 1 and 0.01 radians per position. Fast pairs distinguish neighboring words. Slow pairs still change between words that are thousands of positions apart. A clock works the same way: the minute hand is fast and the hour hand is slow. demo.py checks it with a random 4-number q and k: the score at positions (7, 4) and (107, 104) is the same 1.058526.

RoPE does not make a model work on texts longer than the ones it trained on. A distance longer than any seen in training gives the slow pairs angles the model has never seen, so the scores there are still new to it.

With a KV cache (level 18), k is rotated before it is stored. So the k rows in the cache already carry their positions, and an old k is never rotated again. v is not rotated at all: RoPE changes only the scores, not the numbers that the weights mix.

5. SwiGLU: a gate inside the feed-forward network

The 2017 FFN is ReLU(x @ W1) @ W2. Today’s FFN has a third matrix and multiplies two things: a gate that decides how much passes, and a value that is passed:

FFN(x)=(SiLU(x Wgate)⊙(x Wup)) Wdown\text{FFN}(x) = \big(\text{SiLU}(x\,W_{\text{gate}}) \odot (x\,W_{\text{up}})\big)\,W_{\text{down}} SiLU(z)=z⋅sigmoid(z)\text{SiLU}(z) = z \cdot \text{sigmoid}(z)

The symbol ⊙ means “multiply number by number”, like * in NumPy. Each hidden unit computes one number from its gate column and one from its value column, then multiplies them.

🔒 Answer the question above to unlock

SwiGLU: a SiLU gate decides how much of the value passes

Move the gate g and the value u. Watch the output bar: the gate scales the value.

-4-2024024gate g →ReLUSiLU

Unlike ReLU, a slightly negative gate keeps a small part of the value, with its sign flipped.

SiLU(g) = g × sigmoid(g) = -0.2689 · output = SiLU(g) × u = -0.5379
SiLU(z) = z · sigmoid(z) ReLU, for comparison your gate g positive negative
Number SiLU(z) = z × sigmoid(z) = z / (1 + e^−z). Use e^−1 ≈ 0.37. What is SiLU(1)? (2 decimals)
🔒 Answer the question above to unlock

Three matrices instead of two would mean 50% more weights. So the hidden size d_ff is shrunk to keep the total the same.

Number A plain FFN has 2 matrices of d_model × d_ff. A SwiGLU FFN has 3 matrices of d_model × d_ff′. For the same number of weights, what is d_ff′ / d_ff? (2 decimals)
I got stuck here If the total number of weights is the same, why is it better?

There is no simple proof. Measured on real training runs, gated FFNs reach a lower loss for the same number of weights. Here is one way to see it. In ReLU, a unit is on or off, and only its own input decides. In SwiGLU, one product of x (the gate) decides how much of another product of x (the value) passes. So the FFN can multiply two different numbers made from the same word. A plain FFN cannot do that in one layer.

Now write it. Each word is one row of x.

  1. x @ W_gate and x @ W_up are both (L, d_ff).
  2. * multiplies them number by number (the ⊙).
  3. @ W_down turns each row into d_model numbers again.
CodeWrite the SwiGLU FFN of section 5 for a batch of word rows x of shape (L, d_model), with the weights W_gate, W_up and W_down. Return one row of d_model numbers per word.

Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs

6. A modern block, all together

Put the three new parts into level 16’s pre-norm decoder block. You get the block of nearly every large language model today:

h = RMSNorm(x)
# causal mask; q and k rotated by RoPE, no position vector added
x = x + attention(h)
h = RMSNorm(x)
x = x + SwiGLU(h)

A stack of these blocks, one final RMSNorm, then the output layer. Every part on this page made something better:

🔒 Answer the question above to unlock
Try it

In the RoPE lab of section 4, set q’s position to 2 and k’s to 0, then press “both +1” a few times. The score stays 0.84: both vectors turn, but the distance stays 2. Now move only k’s slider and watch the score follow the distance m − n. In a very long text, the same distance appears at position 5 and at position 5,000. Which of the three new parts makes the scores agree there?

The GPT you write in level 21 uses pre-norm and one of the three new parts, RMSNorm. It keeps learned position vectors and a ReLU FFN. At the end of level 21, an optional exercise lets you replace the learned positions with RoPE. Before that, level 20 asks what running such a model costs: how much GPU memory it needs while it writes, and how many tokens per second it can produce.

Recap

a summary for when you finish the level

The key formulas and common mistakes appear here once you clear the level.

You can now

  • Explain why pre-norm, which levels 16 and 17 already use, trains deep stacks better than the 2017 post-norm.
  • Compute RMSNorm and a RoPE rotation by hand, and write both in NumPy.
  • Write a SwiGLU FFN, and size its hidden layer so it has as many weights as a plain FFN.

Keep in mind

  • Pre-norm: x = x + f(norm(x)), plus one final norm; post-norm: x = norm(x + f(x))
  • , no mean subtracted
  • RoPE: rotate each pair of q and k by position × θ; the score depends only on the distance
  • , with
  • A modern block: x = x + attn(RMSNorm(x)) with RoPE on q and k, then x = x + SwiGLU(RMSNorm(x))

Common mistakes

  • Averaging over the wrong axis (or without keepdims=True), so each word is not normalized separately.
  • Putting SiLU on the value instead of the gate, or multiplying gate and value with @ instead of number by number.

Press ? for keyboard shortcuts

Reading mode · every part open, no stars