The fix: before attention, add a different vector to each position. Word vector + position vector. Now “cat at position 0” and “cat at position 2” are different inputs, so they give different outputs.
The position vectors are made of sines and cosines at different frequencies (how fast each one repeats).
With 2 numbers per word, position pos gets [sin(pos), cos(pos)].
d_model = 8. What shape is its positional encoding?I got stuck here Why not just add 1, 2, 3, … to every number?
Two problems. The numbers grow without limit. Position 500 would add 500, but a word’s numbers are only a few units in size. So the word’s own numbers would be too small to matter. And a model trained on sentences of length 20 would see position 300 for the first time when it is used.
Sines and cosines stay between −1 and 1 forever, and every position still gets its own pattern. In the stripes, the left columns change quickly (they distinguish neighbors) and the right columns change slowly (they distinguish far-apart positions). A clock works the same way: the minute hand is fast and the hour hand is slow.
I got stuck here Why does the position table have 16 rows when this sentence uses only 3?
The position table is built once, for the longest sentence the model will see (16 here). A sentence of L tokens takes
the first L rows: PE[:3] for 3 tokens. The two tables are used differently:
| table | which row a token gets |
|---|---|
| embedding table (level 13) | chosen by which token it is (any row; “cat” is always row 5) |
| position table | chosen by where the token is: always rows 0, 1, 2, … in order |
I got stuck here Why is the word vector multiplied by a number before the position is added?
Level 17’s model does x = embedding * √d_model + PE. With d_model = 4, take the word [−0.77, −0.62, 0.91, −0.28]
and PE[0] = [0, 1, 0, 1]:
- word + PE[0] = [−0.77, 0.38, 0.91, 0.72]. The second and fourth numbers change sign: the position replaced part of the word.
- word × 2 + PE[0] = [−1.54, −0.24, 1.82, 0.44]. The word’s numbers are now bigger than the position’s.
Position numbers are fixed, between −1 and 1, so the factor decides how strong the word is compared with its position.
The usual explanation: if the embedding table starts with numbers of size about 1/√d_model, multiplying by √d_model
makes the word about the same size as the position. So “what the word is” is not hidden by “where it is”.
Level 17 shows the numbers in a real model, and level 21 shows a model that needs no factor.
d_model = 64. Before the position vector is added, each word vector is multiplied by √d_model. By what number?Go deeper The formula for any model width
For position pos, the pair of columns 2k and 2k + 1 gets:
Columns are grouped in pairs. Pair 0 turns at speed 1 per position; each later pair turns more slowly. The last pair turns at : close to 1/10000 when d is large, and 1/100 when d = 4.
With d = 2 there is only pair 0, so the vector is [sin(pos), cos(pos)]. Many newer models learn their position vectors instead,
or rotate Q and K by the position (level 19). The goal is the same: give attention something that differs by position.
Go deeper A causal mask alone shows a little of the order
With the causal mask of level 15, word i can see exactly i + 1 words. So even without positional encoding, the first word always averages over one word and the tenth over ten: the outputs do depend a little on position. That is far too weak to distinguish “dog bites man” from “man bites dog”, which is why decoders still add position vectors (or rotate Q and K).