From token to vector: a table lookup
Pick a token to see its one-hot row times the table. Drag a point (or focus it and use the arrow keys) to change its row.
Token ids (5 tokens). Pick one.
Embedding table E, shape (4, 2)
| id | char | one-hot | E row |
|---|---|---|---|
| 0 | e | 0 | [-1.4, 0.1] |
| 1 | h | 1 | [1.5, -1.2] |
| 2 | l | 0 | [-0.1, -1.5] |
| 3 | o | 0 | [-1.3, 1.5] |
“Take row 2” can also be written as a matrix multiplication. Make a one-hot row: all zeros, except a 1 at position 2. Multiply it by the table, and every row except row 2 is multiplied by 0. Select a token in the lab and look at the one-hot column.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
I got stuck here Why not give the ids to the network as numbers? h = 1, l = 2 are already numbers.
Because the network would treat them as amounts. It would “think” l (2) is twice h (1),
and that o (3) is somewhere past l. None of that is true. The ids are only names, in alphabetical order.
A vector of several numbers per token, which the network is free to learn, avoids this. Two tokens can become close or far apart for real reasons, and there are many directions to be close in.
Go deeper Where the numbers in the table come from
With the BPE tokens of level 12, the table has tens of thousands of rows, and each row has hundreds to thousands of numbers. The idea is the same as here: an id picks a row.
The rows start as random numbers. They are weights, exactly like the weights in level 3, and training moves them with gradient descent. Tokens used in similar ways get similar rows, because that helps the model predict the next token. Nobody writes the meaning in by hand. Section 3 shows why this happens.