Level N7 · Theory · Classic networks · runs in your browser

The 2017 encoder–decoder Transformer

How did the first Transformer turn one sequence into another, with two stacks?

Side trip · best after level 17 · The whole Transformer

This level prepares you to read the original 2017 Transformer design. Its diagram has two stacks, and this level builds both.

Level 17’s model is one stack: the digits, then =, then the words, all in one sequence, with masked self-attention everywhere. The first Transformer, from 2017, was built differently. It was made for translation: read a source sentence in one language, write a target sentence in another. So it has two stacks, an encoder that reads and a decoder that writes. (On this page, “target” means the sentence to write. It is not the training target y.)

This level builds that design and runs it on level 17’s task, reading a number aloud, with the same 900 training numbers and the same 100 held-out ones:

"427"  →  four hundred twenty seven

If you did side trip N5, you solved this kind of task there with two LSTMs and attention between them. This model keeps that plan (an encoder, a decoder, and attention from one to the other) and replaces the LSTMs with attention layers. N5 is optional; you don’t need it for this level.

demo.py prints every number on this page. To run it yourself, download it into a folder of its own and run python demo.py there (setup). It needs PyTorch and takes about 20 seconds.

1. Two stacks

The two stacks split the work:

An encoder layer is self-attention, then the FFN. A decoder layer has three sublayers: masked self-attention over the words written so far, cross-attention to read the source, and the FFN.

Cross-attention is ordinary attention, with one difference: its Q comes from one sequence and its K and V from another. The encoder’s final output is called memory: one vector per source token. It is what the decoder reads.

ChooseIn the decoder’s cross-attention, where do K and V come from?
🔒 Answer the question above to unlock

The 2017 design also puts the LayerNorm in a different place from level 17. Each sublayer is wrapped as LayerNorm(x + sublayer(x)), “add & norm”: add the residual first, then normalize. That is the wrapping of level 16, section 3. It is called post-norm. Level 17 normalizes what goes into each sublayer instead (pre-norm), and level 19 shows why that trains deep stacks better.

One encoder layer and one decoder layer

Select a block to see what it does and the shapes it works on. The sliders change the sizes.

encoder layer
memory (1, 5, 8) → to the decoder’s cross-attention
decoder layer
cross-attention

Every target word looks at the source sentence. This is how the decoder reads the input.

input
(B, L, d_model) = (1, 4, 8) and memory (1, 5, 8)
output
(B, L, d_model) = (1, 4, 8)
Q, K, V
Q comes from the decoder. K and V come from memory, the encoder’s output.
shape
weights per head (B, L_tgt, L_src) = (1, 4, 5): one row per target word, one column per source word
parameters
256

Remove it: Without it, the decoder never sees the source: it can only write a sentence, not translate one.

Inside “cross-attention”, tensor by tensor
tensorshapewhat happens
1Q per head (B, heads, L_tgt, d_k) = (1, 2, 4, 4) from the decoder, d_model split into heads
2K, V from memory (B, heads, L_src, d_k) = (1, 2, 5, 4) from the encoder’s output
3weights (B, heads, L_tgt, L_src) = (1, 2, 4, 5) one row per query word, one column per key word
4out (B, L_tgt, d_model) = (1, 4, 8) heads side by side again, then W_O
cross-attention: output (B, L, d_model) = (1, 4, 8) · 256 parameters
attention feed-forward (FFN) residual + LayerNorm input
I got stuck here What exactly is “memory”?

It is the output of the last encoder layer: one vector per source token, shape (B, L_src, d_model). It is computed once per source. Every decoder layer reads the same memory through its cross-attention, using it for K and V. Nothing is stored or trained in it. It is a name for “the encoder’s answer”. It is not the hidden state of an RNN (level N3). On this page, “memory” alone means the encoder’s output. The “GPU memory” of U6 and level 20 is something else.

Q decides the rows of the attention weights and K the columns, as in level 14. In cross-attention, Q comes from the decoder input and K from memory.

The decoder input always begins with a start token, written <bos> (“beginning of sequence”). It gives the decoder a first input before it has written any word. Section 4 shows how the decoder uses it.

Shape Batch 2. The decoder input has 3 tokens (<bos> plus 2 target words), and the source has 6 words. What shape are one head’s cross-attention weights?
🔒 Answer the question above to unlock

Now compute one cross-attention by hand. The steps are those of level 14: one score per key, softmax, then a weighted average of the values. Only the place each table comes from is new.

Number One decoder token reads a memory of 2 source tokens, with dₖ = 2. Its query is q = [√2 · ln 3, 0], about [1.55, 0]. From memory, the keys are k₀ = [1, 0] and k₁ = [0, 1], and the values are v₀ = [4, 0] and v₁ = [0, 8]. The output is a vector of 2 numbers. What is its first number?
🔒 Answer the question above to unlock

2. The whole machine

The encoder reads the digits. The decoder writes the words. While it writes, it reads the encoder’s output (memory) through cross-attention. Select a block to see its input, its output, and how many parameters it holds. The shapes are for one number (“427”) and a decoder input of 5 tokens: the start token <bos> and up to 4 words.

The whole model, block by block

Pick a block to see its input, its output and how many parameters it holds. Change the sizes below.

Encoder: reads “427”
Decoder: writes the words
Decoder layer × 2 (1, 5, 64) → (1, 5, 64)

Each layer has three parts: masked self-attention (each word sees only earlier words), cross-attention (Q from the words, K and V from memory), and the FFN. One layer has 49,856 parameters.

parameters: 99,712 (58.2% of all)

total parameters: 171,232 · d_k per head = 16
fraction of all parameters not a layer (no parameters)
Shape A model reads numbers aloud, with d_model = 64. A batch holds 32 numbers. The source digits are padded to 3 tokens, and the decoder input has 5 tokens. What shape is memory, the encoder’s output?
🔒 Answer the question above to unlock
ChooseAt the default size, a decoder layer has 49,856 parameters and an encoder layer 33,280. The difference is 16,576, the same as one FFN. What does the decoder layer have that the encoder layer doesn’t?
🔒 Answer the question above to unlock
Go deeper Where the 171,232 parameters are

At the default size (d_model 64, 4 heads, 2 layers, d_ff 128):

partparameters
one attention block: WQ, WK, WV (64 × 64 each) + WO (64 × 64 + 64)16,448
one FFN16,576
one encoder layer: attention + FFN + 2 LayerNorms (128 each)33,280
one decoder layer: 2 attention blocks + FFN + 3 LayerNorms49,856
both embedding tables: (13 + 32) × 642,880
output layer: 64 × 32 + 322,080
total: 2 encoder + 2 decoder layers + embeddings + output171,232

The position numbers are added, not learned, so they have no parameters. The number of heads does not appear anywhere in this table. It only decides how the 64 columns are cut into slices, so 8 heads instead of 4 still give 171,232. The encoder and the decoder have 2 layers each here, but the two numbers are independent and can differ. There are two embedding tables: one for the 13 source tokens (digits and 3 special tokens) and one for the 32 target words.

🔒 Answer the question above to unlock

3. Under the microscope

Below, a tiny copy of this model (1 layer, 1 head, d_model = 6, d_ff = 12) reads “89” and writes the next word. All the sizes are different, so each axis can be recognized by its size. Each step shows the tensor’s shape with named axes, its numbers, and the shape of the same tensor in the full model on this page.

The whole model under a microscope

Step through one input, from token ids to the next token. Point at or tap any number to read where it sits.

Loading the recorded run…

Point at or tap a number to see its full index. ← → also move between steps.
positive negative −∞ blocked by the mask batch positions widths heads vocabulary

Here is the encoder layer of the same tiny model as a picture, with the numbers from the microscope. The line under the blocks is the residual path. In this 2017 layer, a LayerNorm follows each add.

In the 2017 layer, the residual path is normalized after every add

Press Next to follow “89” through one encoder layer. Select a block for its numbers.

Drag to turn · click, then scroll to zoom
Press Next to start.
Layers (7)

Look at the residual path under the blocks: after each add, a LayerNorm rescales the sum. This is post-norm. Compare it with level 17’s block, where x is never normalized until the end.

forward (this step) Select a layer to see its numbers.

I got stuck here Why is the embedding multiplied by a factor before the position is added?

Look at steps 3 and 4 in the microscope. Every position number is between −1 and 1. The factor decides how strong the token is compared with its position.

This model’s token tables start with numbers of size about 1/√dmodel, as in level 17. Multiplying by √dmodel makes them about size 1, the same as the position numbers. So neither one is much bigger than the other. With d_model = 64, the factor is √64 = 8. As in level 17, a token’s row is then about 8 long, and a position’s row about 5.7.

The 2017 design is usually explained the same way. There, the output layer shares its weights with the embedding table. That is one reason for a table with small numbers.

It matters. Suppose the table starts at size 1 instead. A token’s row is then about 64 long, 11 times longer than its position’s row, as in level 17. Here is the fraction of unseen numbers each model gets right:

  • level 17’s one-stack model, after 60 epochs: 55% to 79% with the size-1 table, 99% to 100% with the small table;
  • this two-stack model, after 20 epochs: 71% with the size-1 table, 98% with the small table;
  • this two-stack model, after 60 epochs: 99% with the size-1 table, 100% with the small table.

So with the size-1 table, this model needs more epochs to reach the same result.

With several heads, attention scores have the shape (B, h, L_q, L_k): batch, heads, then the length of Q and the length of K.

Shape A model reads numbers aloud. A batch holds 64 numbers, the source digits are padded to 3 tokens, the decoder input has 5 tokens (words), and there are 4 heads. What shape are the cross-attention scores?
🔒 Answer the question above to unlock

4. One word at a time, with a source

The decoder writes the way level 17’s model does: run, take the last row, pick the top word, append it, run again. One thing is new. Level 17’s model starts from the digits and =, which are already in its sequence. This decoder’s input holds only words, so it starts from the start token <bos>. The digits reach it only through cross-attention, and the encoder reads them once, before the first word.

Press Next step and watch each stage. The bars under the digits show which digit cross-attention looked at for each word.

Greedy decoding, one word at a time

Pick a number for the trained model to read aloud, then press Next step to watch it choose each word.

Loading the trained model’s recording…

…
positive number negative number the word it picks the right word, when it picks wrong
Number Reading 742 aloud. The model has already written “seven hundred forty”. How many rows go into the decoder at this step?
🔒 Answer the question above to unlock
Number A decoder with 2 layers reads 742 aloud: 4 words, one per step. The model keeps (caches) every result that cannot change between steps, and reuses it. In one decoder layer, how many times is memory multiplied by W_K for the whole answer?
🔒 Answer the question above to unlock

This decoder uses two caches. The cross-attention cache holds memory’s keys and values: one row per source token, made once. The self-attention cache gets one new row at every step. Level 20 counts what such a cache costs in memory and time.

Try it

Switch the model to after 2 epochs and read 896 aloud. It says “six hundred eighty”. It has learned the shape of an answer (a digit, “hundred”, a tens word) but not which digit goes where. Look at the cross-attention bars. They are spread out, almost even.

Then switch back to after 60 epochs and read 896 again. Now each word looks at its own digit: “eight” puts all its weight on the 8, and “six” puts 0.92 on the 6. The trained model gets all 100 unseen numbers right.

🔒 Answer the question above to unlock

5. Three attentions, three masks

In training, the decoder gets the whole correct answer at once, shifted by one:

427:    tgt_in   <bos>  four     hundred  twenty  seven
        tgt_out  four   hundred  twenty   seven   <eos>
I got stuck here Why are the decoder’s input and its targets the same words, shifted by one?

Because position t has to guess the next word. Write them one under the other: under <bos> the answer is “four”; under “four” the answer is “hundred”. tgt_out is tgt_in moved one step to the left, with <eos> at the end. Level 17 does the same shift on its one sequence; here the shift only covers the words.

The model uses three attentions, and each one has its own mask. A mask always has the shape of the scores it blocks, (length of Q, length of K). L_src is the source length and L_tgt the decoder input length, <bos> included. Every head uses the same mask:

attentionQ fromK and V frommask
encoder self-attentionthe digitsthe digitsblocks PAD digits only
decoder self-attentiontgt_intgt_incausal, and blocks PAD words
cross-attentiontgt_inmemoryblocks PAD digits only

Their scores have these shapes:

  • encoder self-attention: (B, h, L_src, L_src);
  • decoder self-attention: (B, h, L_tgt, L_tgt);
  • cross-attention: (B, h, L_tgt, L_src).

Here are the three masks for one short example, side by side. Rows are the tokens that Q comes from; columns are the tokens that K comes from.

encoder self-attention
(3, 3), 3 blocked
15PAD
1001
5001
PAD001
rows (Q): the digits
columns (K): the digits
decoder self-attention
(4, 4), 9 blocked: 6 future, 3 PAD
<bos>fifteenPADPAD
<bos>0111
fifteen0011
PAD0011
PAD0011
rows (Q): the decoder input
columns (K): the decoder input
cross-attention
(4, 3), 4 blocked
15PAD
<bos>001
fifteen001
PAD001
PAD001
rows (Q): the decoder input
columns (K): memory (the digits)
blocked: futureblocked: PAD1 means True: blocked. A blocked score becomes −∞, so its weight is 0.
Reading 15 aloud: the source “15” is padded to 3 digits, and the decoder input is<bos> fifteen PAD PAD. A cell that is both future and PAD is shown as future.
Shape The source has 5 tokens. The decoder input is <bos> plus 5 target words. What shape is the causal mask in the decoder’s self-attention (for one sentence)?
🔒 Answer the question above to unlock
I got stuck here Which length decides the size of the causal mask?

The causal mask is used only in the decoder’s self-attention, where both Q and K come from the decoder input: <bos> plus the 5 target words, 6 tokens. The source length never enters a causal mask.

I got stuck here Why don’t the encoder and cross-attention use a causal mask?

The causal mask stops a position from seeing the word it has to guess. Only the decoder guesses words, and only the decoder’s own input contains them. The digits are the question, not the answer: every word may read the whole number, and every digit may read every other digit.

Number Reading 42 aloud, in a batch where longer answers need 5 words. The decoder input is <bos> forty two PAD PAD: 5 positions. The decoder’s self-attention mask blocks a cell if the causal mask blocks it or if its column is a PAD word. How many of the 25 cells are blocked?
🔒 Answer the question above to unlock

Now build all three masks for one example. True means blocked, as in level 15. Level 15’s padding mask had one entry per column; to give it the shape of the scores, stretch that row over every query row. np.broadcast_to(row, (n, len(row))) does it: np.broadcast_to(np.array([False, True]), (3, 2)) is a (3, 2) table whose three rows are all [False, True]. | combines two True/False tables cell by cell: a cell is True if either one is True. For the causal part, recall level 15’s np.triu.

CodeBuild the three masks of one example. src holds the digit ids, tgt_in the decoder input ids, and pad is the id of PAD. Return (enc, dec, cross), True where a cell is blocked: enc is (Ls, Ls), dec is (Lt, Lt), cross is (Lt, Ls). Several lines.

Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs

🔒 Answer the question above to unlock

Now use the cross mask inside cross-attention. Write one head, with the PAD source tokens blocked.

CodeWrite one cross-attention head. y holds the decoder’s rows and memory the encoder’s output; src_pad is True for a PAD source token. Return (out, A): the output and the weights. Several lines.

Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs

Go deeper Label smoothing, and why the loss never reaches 0

This model uses label_smoothing=0.1, as the 2017 design did. The target for each position is not “100% the right word”, but “90% the right word, and the remaining 10% spread over all 32 words”. It stops the model from becoming completely certain, which tends to help it on unseen inputs.

Because of this, even a perfect model can’t reach a loss of 0. The lowest possible loss is the entropy of the smoothed target. The right word gets 0.9 + 0.1 / 32 = 0.903125, and each of the other 31 words gets 0.1 / 32 = 0.003125. So the lowest loss is −(0.903125 ln 0.903125 + 31 × 0.003125 ln 0.003125) ≈ 0.651. This run ends at a loss of 0.652, only 0.001 above it, while getting 100% of the unseen numbers right. Unseen accuracy over the run: 0% after epoch 1, 89% after 10, 98% after 20, 98% after 40, 100% after 60.

🔒 Answer the question above to unlock

6. Reading the original 2017 design

You now know every part of the original 2017 Transformer. Its diagram and its text use these names. The right column says where this course teaches each one. You may not have done some of those levels yet.

name in the 2017 designwhat it iswhere this course teaches it
Input Embedding, Output Embeddingthe token tables: one for the source, one for the targetlevel 13; section 2
multiply the embeddings by √dmodelscale the token before the position is addedlevel 16, section 1; section 3
Positional Encoding (sine and cosine)a fixed table of position numbers, added to the tokenslevel 16, section 1
Scaled Dot-Product Attentionsoftmax(Q Kᵀ / √dₖ) Vlevel 14
Multi-Head Attention, h = 8the attention split into 8 heads, then joined with WOlevel 15, section 2; level 21
Masked Multi-Head Attentionthe decoder’s self-attention with the causal masklevel 15, section 1; level 17
encoder–decoder attentioncross-attention: Q from the decoder, K and V from the encoder outputsection 1
Add & Normresidual, then LayerNorm (post-norm)level 16, section 3; level 19
Position-wise Feed-Forward Networksthe FFN, applied to each position separatelylevel 16, section 2
N× (N = 6)the number of stacked layers in the encoder and in the decoderlevel 17; section 2
Outputs (shifted right)the decoder input tgt_in: the target moved one step right, starting with <bos>section 5
Linear, Softmaxthe output layer and the probabilities over the vocabularylevel 17
shared embedding and output weightsthe output layer reuses the embedding table (this course doesn’t)section 3
label smoothing, 0.190% on the right token, 10% spread over all tokenssection 5
warmup, then a decaying learning ratethe learning-rate schedulelevel 7, section 4
dropout, 0.1set some numbers to 0 at random while traininglevel 8, section 2
beam search, width 4keep the 4 best partial answerslevel 18, section 4

The original sizes: d_model = 512, h = 8 heads, so dₖ = 512 / 8 = 64; d_ff = 2048 = 4 × d_model; 6 layers on each side.

Today’s chat models keep only the decoder stack, as level 17 does: the “source” is simply the start of the sequence. The encoder–decoder design is still used where the input and the output are clearly two different things, for example translation or speech to text.

Recap

a summary for when you finish the level

The key formulas and common mistakes appear here once you clear the level.

You can now

  • Name the sublayers of an encoder layer and a decoder layer, and say where cross-attention takes Q, K and V from.
  • Give the shape of every attention’s scores and mask in an encoder–decoder, <bos> included.
  • Build the three masks of one example in NumPy.

Keep in mind

  • Cross-attention: Q from the decoder, K and V from memory; scores (B, h, L_tgt, L_src)
  • tgt_in = <bos> + answer, tgt_out = answer + <eos>: position t guesses word t + 1
  • Only the decoder’s self-attention is causal; encoder self-attention and cross-attention block PAD source tokens only
  • The 2017 placement wraps each sublayer as LayerNorm(x + sublayer(x)) (post-norm)

Common mistakes

  • Making the decoder’s causal mask from the source length: it follows the decoder input, <bos> included.
  • Swapping the axes of cross-attention: rows are the decoder’s tokens (Q), columns the source tokens (K).

Press ? for keyboard shortcuts

Reading mode · every part open, no stars