# 17. The whole Transformer

> How does a model produce text one word at a time?

LLM by Hand · Theory · runs in your browser · interactive page: https://llm.liko.page/learn/full-model/

You have every part now: embeddings, attention, masks, the FFN, residuals and LayerNorm, and level 16 put them together
into a decoder block. In this level you stack those blocks into a whole model and watch it work on a real task.

The task: **read a number aloud.** The input is a string of digits, the output is English words.
The model sees both as **one sequence**: the digits, then `=`, then the words, then `<eos>` (“end”).

```
4 2 7 = four hundred twenty seven <eos>
```

The model only ever does one thing: guess the next token. Give it `4 2 7 =` and it continues the sequence, one word at a
time, until it writes `<eos>`. The part you give it is the **prompt**; the part it writes is the answer.

There are 1,000 numbers, from 0 to 999. The model trains on 900 of them. The other 100 are held out (the model never trains on them),
so the only way to get them right is to learn the rule, not to remember the answers.

**Question.** Level 17’s model reads 905 as one sequence: the digits, then =, then the words, then <eos>. One token per digit and one per word. How many tokens is that?

*Answer it on the page to check your work.*

## 1. The whole machine

In order, from input to output: a token table turns each id into a vector. Then the position numbers of level 16 are added.
Then come a stack of decoder blocks, one final LayerNorm, and an output layer that gives one score per token.
Select a block to see its input, its output, and how many parameters it holds.
The shapes are for one sequence of 8 tokens, the longest input this task has.

*[Interactive lab: Full model — open the page to use it]*

**Question.** A batch holds 64 numbers. The model input x is (64, 8): 8 tokens per number, padded. There are 42 tokens in the vocabulary and `d_model` = 64. What shape are the logits?

*Answer it on the page to check your work.*

**Predict.** `d_model` stays 64 and you change heads from 4 to 8. The parameter count…

A. doubles
B. stays exactly the same
C. halves

*Answer it on the page to check your work.*

Most of the parameters are in the blocks, and inside each block most of them are in two places:
the attention matrices and the FFN. You can count one FFN by hand.

**Question.** `d_model` = 64 and `d_ff` = 128. How many parameters does one FFN have? It is Linear(64 → 128), ReLU, Linear(128 → 64). Count weights and biases.

*Answer it on the page to check your work.*

**Deeper: Where the 72,106 parameters are**

At the default size (`d_model` 64, 4 heads, 2 blocks, `d_ff` 128):

| part | parameters |
|---|---|
| token table: 42 × 64 | 2,688 |
| one attention block: `W_Q`, `W_K`, `W_V` (64 × 64 each) + `W_O` (64 × 64 + 64) | 16,448 |
| one FFN | 16,576 |
| one block: attention + FFN + 2 LayerNorms (128 each) | 33,280 |
| final LayerNorm | 128 |
| output layer: 64 × 42 + 42 | 2,730 |
| **total: 2 blocks + token table + final LayerNorm + output** | **72,106** |

The position numbers are added, not learned, so they have no parameters.
The number of heads does not appear anywhere in this table. It only decides how the 64 columns are cut into slices.
The 42 tokens are `<pad>`, `<eos>`, the ten digits, `=`, and 29 words: zero to nineteen, the eight tens words, and “hundred”.

## 2. Under the microscope

The diagram shows the blocks. Now look inside them. Below, a tiny copy of this model (1 block, 1 head,
`d_model` = 6, `d_ff` = 12) reads `8 9 =` and writes the next word. Every size is different from every other size,
so each number shows you which axis it is. Each step shows the tensor’s shape with named axes, its numbers,
and the shape of the same tensor in the full model on this page.

*[Interactive lab: Microscope — open the page to use it]*

Here is the block of the same tiny model as a picture, with the numbers from the microscope. The line underneath is
the residual path: each sublayer reads a normalized copy of x and adds its result to x.

*[Interactive lab: Arch — open the page to use it]*

**If you are stuck: Why is the embedding multiplied by a number before the position is added?**

Look at steps 3 and 4 in the microscope. Every position number is between −1 and 1, so the factor decides how strong
the token is compared with its position.

This model’s token table starts with numbers of size about 1/√`d_model`. Multiplying by √`d_model` makes them about size 1,
the same as the position numbers, so that one doesn’t hide the other. In the full model, a token’s row has 64 numbers
of size 1/8, so its length is about √64 × 1/8 = 1. After × √64 = 8 it is about 8 long, and a position’s row is about 5.7.

It matters. Suppose the table starts at size 1 instead. A token’s row is then about √64 × 1 × 8 = 64 long, 11 times
longer than its position’s row, and the model can hardly see where a token is. It reads 82 as “eight hundred twenty two”. Over three training runs it
gets 55% to 79% of the unseen numbers right, while the model on this page gets 99% to 100%.
In level 21 the position table is learned too, and both tables start at the same size, so that model needs no factor.

With several heads, attention scores have the shape `(B, h, L_q, L_k)`: batch, heads, then the length of Q and the length
of K. In a decoder-only model Q and K come from the same sequence.

**Question.** Same batch: x is (64, 8), and the model has 4 heads. What shape are one block’s attention scores?

*Answer it on the page to check your work.*

## 3. One word at a time

Level 14 called it **inference**: running a trained model to get an answer. It is the whole forward pass, from token ids to
logits, with every weight fixed. Training runs the same forward pass, then the backward pass, and then changes the weights.

**Predict.** A trained model reads a number aloud (inference). Which of these changes while it runs?

A. `W_Q`, the matrix inside the attention
B. The attention scores
C. The embedding table

*Answer it on the page to check your work.*

The model never writes the whole answer at once. It writes **one word, then runs again**:

1. Start with the prompt: the digits and `=`.
2. Run the model over everything so far. The output has one row per token.
3. Only the **last row** matters. The output layer turns its 64 numbers into 42 scores, one per token.
   These raw scores are usually called **logits**.
4. Pick the highest score (this is called greedy decoding), and add that token to the end.
5. Repeat until the model picks `<eos>`.

Below is a real model, trained on this page's task. Press **Next step** and watch each stage.
The bars above the tokens show where the last row’s attention looked while it chose the next word.

*[Interactive lab: Decode — open the page to use it]*

**Question.** Reading 742 aloud. The model has already written “seven hundred forty”. How many rows go into the model at this step?

*Answer it on the page to check your work.*

Step 4 is the only new code. NumPy’s `np.argmax` gives the position of the largest value, not the value itself:
`np.argmax(np.array([0.2, 1.5, -0.3]))` is `1`.

**Code question.** Write one step of greedy decoding. h is the last row of the model’s output (after the final LayerNorm), W and b are the output layer. Return the id of the next word.

Fill in the blank (`____`):

```python
def next_token(h, W, b):
    scores = h @ W + b
    print("scores:", scores)
    return ____

h = np.array([1.0, -1.0])
W = np.array([[1.0, 0.0, 2.0],
              [0.0, 1.0, 1.0]])
b = np.array([0.0, 0.0, -0.5])
print("next token:", next_token(h, W, b))
```

*Answer it on the page to check your work.*

Now the whole loop from the list above. The model returns one row of scores per token, so `scores[-1]` is the last row.
`ids.append(nxt)` adds the picked id to the end, and `break` leaves the loop once the model picks `<eos>`.
A fixed limit on new tokens stops a model that never picks `<eos>`.

**Code question.** Write the whole greedy loop. Start from the prompt (the digits and =). Each step runs the model on everything so far, takes the last row of the scores, appends the top id, and stops after the model picks eos (keep eos in the list). Stop after `max_new` new tokens even without eos.

Fill in the blank (`____`):

```python
# A small stand-in model with 6 tokens:
# 0 <pad>, 1 <eos>, 2 "4", 3 "=", 4 four, 5 hundred.
# Row i of T holds the scores for the token that follows token i.
T = np.array([[0., 0., 0., 0., 0., 0.],
              [0., 0., 0., 0., 0., 0.],
              [0., 0., 0., 2., 0., 0.],    # after "4": =
              [0., 0., 0., 0., 2., 1.],    # after "=": four
              [0., 1., 0., 0., 0., 3.],    # after four: hundred
              [0., 2., 0., 0., 1., 0.]])   # after hundred: <eos>

def model(ids):
    # (len(ids), 6): one row of scores per token
    return T[np.array(ids)]

def greedy_decode(model, prompt, eos=1, max_new=6):
    ids = list(prompt)
    for step in range(max_new):
        ____
    return ids

print("ids:", [int(i) for i in greedy_decode(model, [2, 3])])
```

*Answer it on the page to check your work.*

One **epoch** is one pass over all 900 training numbers. The lab has the same model saved twice: after 2 epochs and after 60.

**Try it**

Switch the model to **after 2 epochs** and read 14 aloud. It says “six hundred twenty nine”.
It has learned the shape of a long answer: a digit, “hundred”, a tens word, a digit.
It has not learned which digit goes where, or even how long the number is.

Then switch back to **after 60 epochs** and read 896. Watch the bars: for “eight” the last row looks at the 8,
for “ninety” mostly at the 9, and for “six” at the 6. Nobody told it which digit goes with which word. This model gets
all 100 unseen numbers right.

## 4. How it learned

While training, the model does **not** run one word at a time. It gets the whole sequence and makes every guess in one
pass. The input x is the sequence without its last token, and the target y is the sequence without its first token:

```
427:    x   4  2  7  =     four     hundred  twenty  seven
        y   2  7  =  four  hundred  twenty   seven   <eos>
```

The causal mask stops each position from seeing the token it has to guess. Training on the correct previous tokens
like this, instead of on the model’s own guesses, is called **teacher forcing**.

**If you are stuck: Why are x and y the same tokens, shifted by one?**

Because position t has to guess the next token. Write them one under the other: under `=` the answer is “four”;
under “four” the answer is “hundred”. y is x moved one step to the left, with `<eos>` at the end.

If you fed the same row as both input and target, position t could copy its own input,
and the loss would drop to zero while the model learns nothing.

Not every guess is worth learning. Under the 4 the target is 2, and under the 2 it is 7: the model would be guessing the
digits of the question, and those are random. Only the targets that are words or `<eos>` are the answer.
Every target before them, `=` included, is part of the question. The loss counts the answer targets and skips the rest, by setting the skipped targets to PAD.

**Question.** For 427, x = 4 2 7 = four hundred twenty seven and y = 2 7 = four hundred twenty seven <eos>: 8 targets. The loss counts only targets that are part of the answer (the words and <eos>). How many of the 8 count?

*Answer it on the page to check your work.*

Training repeats four steps on batches of 64 numbers:

```python
for x, y in batches:              # 1. a batch; skipped targets are PAD
    logits = model(x)             # 2. forward: (64, L, 42)
    # 3. loss, ignore_index=PAD
    loss = loss_fn(logits.reshape(-1, 42), y.reshape(-1))
    # 4. backward, then update
    opt.zero_grad(); loss.backward(); opt.step()
```

Short numbers have short sequences. In a batch with the numbers 0 to 63, the longest x is 5 tokens (for example
`2 3 = twenty three`), so the logits are (64, 5, 42) and there are 320 targets. Only 167 of them count: 118 are digits
or `=`, and 35 are PAD after a short sequence.

**Question.** For the loss, the (64, 5, 42) logits are flattened so that every position is one row. What shape do they become?

*Answer it on the page to check your work.*

**If you are stuck: Why reshape to (B × L, V) at all?**

The loss function wants a simple list of guesses: one row of scores per guess, and one correct id per guess.
It does not care which sequence or which position a guess came from.
Every position in every sequence is one independent guess, so (64, 5, 42) is stored as 320 rows of 42.

**Predict.** You delete the causal mask and train again. What happens?

A. The training loss stays high and the model learns nothing
B. The training loss gets very low, but reading new numbers aloud fails
C. Nothing changes

*Answer it on the page to check your work.*

**Deeper: A debugging habit: the copy task**

When a model learns nothing, the problem can be the task, or it can be a bug in the code. What would you check first?

**Predict.** You build a new model for a new task. After 10 epochs it gets 0% right, and the training loss stopped falling long ago. What do you try first?

A. Make the model bigger
B. Train it on the copy task: input 1 2 3, output 1 2 3
C. Collect more data

*Answer it on the page to check your work.*

The **copy task** is `1 2 3 = 1 2 3`. Any correct model learns it in about a minute of training.
So if yours can’t, the bug is in the mask, the shift by one or the loss.

Last step. Write the loss that skips the PAD targets. In NumPy, an array of True and False picks entries:
`np.array([5, 6, 7])[np.array([True, False, True])]` is `array([5, 7])`.
The browser gives you `softmax(x)`, and it works on each row of x separately.

**Code question.** Write the whole loss that skips PAD, starting from the scores. logits is (N, V), targets is (N,), and the argument pad is the id of PAD (0 by default). Return the mean loss over the rows whose target is not PAD. Several lines.

Fill in the blank (`____`):

```python
def masked_loss(logits, targets, pad=0):
    N = len(targets)                     # the number of rows
    # p[np.arange(N), targets] picks one entry per row:
    # from row i, column targets[i]
    ____

logits = np.array([[2.0, 1.0, 0.0],
                   [0.0, 3.0, 0.0],
                   [1.0, 1.0, 1.0]])
targets = np.array([0, 1, 0])        # rows 0 and 2 are PAD
print("loss:", masked_loss(logits, targets))
```

*Answer it on the page to check your work.*

**Recommended next: side trip [N7](/learn/encoder-decoder/).** If you want to read the original 2017 Transformer
design, now is a good time. N7 builds the two-stack encoder–decoder on this same task. Level 18 continues the main line.

## You can now

- Follow the shapes through a whole decoder-only model, from token ids to `(B, L, V)` logits.
- Write the greedy loop: run the model, take the last row, append the top id, stop at `<eos>`.
- Build x and y by a shift of one, and compute a loss that counts only the answer.
