This is all the math the course uses. Each idea has one small example with real numbers. Read it once now, or come back when a level sends you here. Each section ends with a short “Try it”: compute the answer on paper first, then open it to check.
is a fixed number, like . (in code np.exp(x)) is “e multiplied by itself x times”. It works for any x, not only whole numbers, and it is always positive.
| x | −2 | −1 | 0 | 0.5 | 1 | 2 |
|---|---|---|---|---|---|---|
| 0.14 | 0.37 | 1 | 1.65 | 2.72 | 7.39 |
“Multiplied by itself x times” only makes sense for whole numbers. For x = 0, x = 0.5 or a negative x, just read the table: and .
A negative power means “one over”: . And .
ln (the natural logarithm, np.log in code) undoes : it answers “e to what power gives this number?”. So , , and . ln only takes positive numbers, and ln of a number below 1 is negative.
Four rules cover everything in the course:
Level 2 draws its 3D loss surface with height . Adding 1 keeps the number positive, and ln pulls big values down a lot: , , . A bigger loss is still higher.
, so .
Used in: level 3 (cross-entropy), level 10 (softmax, perplexity), level 11.
and . A square is never negative: . The square root (np.sqrt) undoes a square: . It always means the positive root.
In Python a power is written **: x ** 2 is , and on a NumPy array a ** 2 squares every element. Don't write x ^ 2: in Python ^ means something else, and on floats it gives an error.
(-2) ** 2 + 3 ** 2, and what is its square root?(−2) × (−2) = 4 and 3 × 3 = 9, so the sum is 13 and its square root is √13 ≈ 3.61.
Used in: level 2 (squared error), level 6 (the slope of tanh), level 7 (Adam).
means “add for every i from 0 to n − 1”. The small number under a letter is its position, counted from 0 as in code: is the first x. (Many math books count from 1 instead; the sum is the same.)
x.sum().16 + 1 + 9 = 26. Square first, then add.
Used in: level 1 (each cell of a matrix product is a sum), level 10 (softmax divides by a sum).
The mean is the sum divided by how many numbers there are. The variance is the mean of the squared distances from the mean. The standard deviation (std) is the square root of the variance: a typical distance from the mean, in the same units as the numbers.
Mean = 20 / 4 = 5. Distances: −3, −1, 1, 3. Squares: 9, 1, 1, 9. Variance = 20 / 4 = 5, so std = √5 ≈ 2.24. (The same spacing as the example, so the same std.)
Used in: level 7 (starting weights), level 10 (Gaussian noise), level 16 and N2 (LayerNorm), D1.
The derivative of a function says how fast its output changes when its input changes a tiny bit. It is the slope of its graph at that point. You can measure it with two numbers: change the input a little, divide the change in output by the change in input.
When a function has several inputs, the partial derivative is the slope when only w moves and the other inputs stay fixed. All the partial derivatives together are the gradient.
For a function with one input, the derivative is often written , read “f prime of x”. For , , so . A layer’s in level 4 is the slope of its activation function.
f(5.01) = 25.1001, so the change is 0.1001. Divided by 0.01 that is about 10, and the rule gives 2 × 5 = 10.
Used in: level 2 onward. Training is “move each parameter against its slope”.
If y depends on u, and u depends on x, then the slope of y with respect to x is the two slopes multiplied:
u = 6. dy/du = 2u = 12, du/dx = 2, so dy/dx = 12 × 2 = 24. Check: y = 4x², slope 8x = 24 at x = 3.
Used in: level 2 (where the gradient formula comes from), level 6 (backpropagation is this rule, over and over).
sin (np.sin) is a wave: it goes up and down between −1 and 1 forever. , and it repeats every . Level 8 uses as a smooth curve to learn: for x from −1 to 1 it goes down to −1, back through 0, up to 1 and back to 0.
| x | −1 | −0.5 | 0 | 0.5 | 1 |
|---|---|---|---|---|---|
| 0 | −1 | 0 | 1 | 0 |
cos (np.cos) is the same wave, shifted by a quarter turn: . Angles in sin and cos are in radians, not degrees. A full turn is radians, and a quarter turn is . So 1 radian is about 57°.
Half a turn points the other way: cos π = −1 and sin π = 0.
A polynomial is a sum of powers of x, each times a number called a coefficient: Its degree is the biggest power. A degree-1 polynomial is a straight line; a higher degree can bend more times.
2 − 2 + 8 = 8. The biggest power is 3, so the degree is 3.
Used in: level 8 (fitting a curve with too many parameters), level 16 (position encodings use sin and cos), level 19 (RoPE turns vectors by angles in radians).
Computers write 1e-5 for and 3e4 for . The number after the “e” says how many places to move the decimal point: a minus sign moves it to the left (a small number), a plus or no sign moves it to the right (a big number). This “e” has nothing to do with .
2e-11 = 0.00000000002. 1.5e3 = 1500.1e-3 and 4e2 as plain numbers.1e-3 = 0.001 (three places to the left). 4e2 = 400 (two places to the right).
To multiply powers of ten, add the exponents: . To divide them, subtract the exponents: . With a number in front, divide that number separately.
Subtract the exponents: 13 − 14 = −1. So it is 2.8 × 10⁻¹ = 0.28.
Used in: level 6 (checking a gradient with a tiny step), U2, level 20 (memory and speed of a large model).
A vector is a list of numbers, like [2, 0]. A matrix is a table of numbers; its shape (rows, columns) says how big it is. Rows and columns count from 0. Matrix multiplication A @ B needs the inner sizes to match: . In formulas, two matrices written next to each other with nothing between them, like , mean the same product X @ W. Level 1 teaches all of this from scratch.
(4, 2): the inner sizes are both 3. The second is an error: the inner sizes 3 and 2 don't match.
Used in: level 1 and every level after it.