Every question you got wrong at least once, so you can try it again. Redoing a question here never takes away a level you cleared. Progress is saved in this browser only.
No mistakes yet. When you get a question wrong, it shows up here.
A = [[1, 2, 3], [0, 1, 0]] and B = [[1, 0], [0, 1], [1, 1]] What is C[0][1]?
Number A = [[1, 2, 3], [0, 1, 0]] and B = [[1, 0], [0, 1], [1, 1]] What is C[0][1]?
What is the shape of (3, 4) @ (4, 2)?
Shape What is the shape of (3, 4) @ (4, 2)?
A is (3, 2) and B is (3, 2). What happens with A @ B?
ChooseA is (3, 2) and B is (3, 2). What happens with A @ B?
Write matrix multiplication with three loops, without @ or np.dot (a and b are lists of lists). Loop over the rows of the output, its columns, and the inner size, and add one product to an output cell on each pass. Several lines.
CodeWrite matrix multiplication with three loops, without @ or np.dot (a and b are lists of lists). Loop over the rows of the output, its columns, and the inner size, and add one product to an output cell on each pass. Several lines.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
X is (2, 5, 4): 2 sentences, 5 words each, 4 numbers per word. W is (4, 8). What shape is X @ W?
Shape X is (2, 5, 4): 2 sentences, 5 words each, 4 numbers per word. W is (4, 8). What shape is X @ W?
A column of shape (4, 1) plus a row of shape (1, 5). What shape is the result?
Shape A column of shape (4, 1) plus a row of shape (1, 5). What shape is the result?
a = [[1], [2], [3], [4]] has shape (4, 1). b = [[10, 20, 30, 40, 50]] has shape (1, 5). What is (a + b)[2][3]?
Number a = [[1], [2], [3], [4]] has shape (4, 1). b = [[10, 20, 30, 40, 50]] has shape (1, 5). What is (a + b)[2][3]?
Only the shapes matter here. Attention scores for a batch have shape (2, 4, 4): 2 sentences, 4 words each. The padding mask has shape (2, 1, 4). What shape do they broadcast to when the mask is applied to the scores?
Shape Only the shapes matter here. Attention scores for a batch have shape (2, 4, 4): 2 sentences, 4 words each. The padding mask has shape (2, 1, 4). What shape do they broadcast to when the mask is applied to the scores?
X has shape (4, 3). What is the shape of X.sum(axis=0)?
Shape X has shape (4, 3). What is the shape of X.sum(axis=0)?
Subtract each column’s mean from X, so every column has mean 0. Use .mean(axis=0) and broadcasting.
CodeSubtract each column’s mean from X, so every column has mean 0. Use .mean(axis=0) and broadcasting.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
Start at w = 0.5, b = 0. The predictions are 0.5, 1, 1.5 and the targets are 2, 4, 6. What is the mean squared error loss?
Number Start at w = 0.5, b = 0. The predictions are 0.5, 1, 1.5 and the targets are 2, 4, 6. What is the mean squared error loss?
A line ŷ = w·x + b is fit to the points (1, 2), (2, 4), (3, 6) with the loss L = mean((w·x + b − y)²). At w = 0.5, b = 0, what is ∂L/∂w?
Number A line ŷ = w·x + b is fit to the points (1, 2), (2, 4), (3, 6) with the loss L = mean((w·x + b − y)²). At w = 0.5, b = 0, what is ∂L/∂w?
w = 0.5 and its gradient is −14. With learning rate 0.05, what is w after one step?
Number w = 0.5 and its gradient is −14. With learning rate 0.05, what is w after one step?
A line ŷ = w·x + b is fit to the points (1, 2), (2, 4), (3, 6). At w = 0.5, b = 0 the loss is 10.5, ∂L/∂w = −14 and ∂L/∂b = −6. The best line has w = 2, b = 0. You take one step on both w and b with learning rate 0.2 instead of 0.05. What happens to the loss?
ChooseA line ŷ = w·x + b is fit to the points (1, 2), (2, 4), (3, 6). At w = 0.5, b = 0 the loss is 10.5, ∂L/∂w = −14 and ∂L/∂b = −6. The best line has w = 2, b = 0. You take one step on both w and b with learning rate 0.2 instead of 0.05. What happens to the loss?
Write the gradient of w.
CodeWrite the gradient of w.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
Write the body of the training loop: one gradient-descent step on w and b. Several lines.
CodeWrite the body of the training loop: one gradient-descent step on w and b. Several lines.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
The points are now (1, 3), (2, 5), (3, 7). Start at w = 0.5, b = 0 with learning rate 0.05 and take one step of gradient descent (by hand, or with descend from the code box). What is b after that one step?
Number The points are now (1, 3), (2, 5), (3, 7). Start at w = 0.5, b = 0 with learning rate 0.05 and take one step of gradient descent (by hand, or with descend from the code box). What is b after that one step?
The AND neuron has w1 = 1, w2 = 1, b = −1.5, and z = w1·x1 + w2·x2 + b. What is z for the point (1, 1)?
Number The AND neuron has w1 = 1, w2 = 1, b = −1.5, and z = w1·x1 + w2·x2 + b. What is z for the point (1, 1)?
A neuron has z = 0.5. Its output is p = sigmoid(0.5) = 1 / (1 + e^(−0.5)). Use e^(−0.5) ≈ 0.6. What is p? (Three decimals.)
Number A neuron has z = 0.5. Its output is p = sigmoid(0.5) = 1 / (1 + e^(−0.5)). Use e^(−0.5) ≈ 0.6. What is p? (Three decimals.)
XOR data: the points (0, 0) and (1, 1) have target 0, and (0, 1) and (1, 0) have target 1. One neuron draws one straight line. Can you find w1, w2, b so that all 4 points are correct?
ChooseXOR data: the points (0, 0) and (1, 1) have target 0, and (0, 1) and (1, 0) have target 1. One neuron draws one straight line. Can you find w1, w2, b so that all 4 points are correct?
Write z for one neuron.
CodeWrite z for one neuron.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
On XOR, one neuron answers p = 0.5 for every point. Suppose a model answered p = 0.25 for the two points whose target is 0, and p = 0.75 for the two whose target is 1. What is its loss −mean(y·ln p + (1 − y)·ln(1 − p))? Use ln 0.75 ≈ −0.288 and ln 0.25 ≈ −1.386.
Number On XOR, one neuron answers p = 0.5 for every point. Suppose a model answered p = 0.25 for the two points whose target is 0, and p = 0.75 for the two whose target is 1. What is its loss −mean(y·ln p + (1 − y)·ln(1 − p))? Use ln 0.75 ≈ −0.288 and ln 0.25 ≈ −1.386.
Write the binary cross-entropy loss: −mean(y·ln p + (1 − y)·ln(1 − p)) for arrays p and y.
CodeWrite the binary cross-entropy loss: −mean(y·ln p + (1 − y)·ln(1 − p)) for arrays p and y.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
Write the body of `neuron_loss`: return the binary cross-entropy of the neuron’s predictions against y. Several lines.
CodeWrite the body of neuron_loss: return the binary cross-entropy of the neuron’s predictions against y. Several lines.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
x = [2, 1], W1 = [[1, 2], [0, 1]] and W2 = [[1], [1]]. First h = x @ W1, then out = h @ W2. What is out?
Number x = [2, 1], W1 = [[1, 2], [0, 1]] and W2 = [[1], [1]]. First h = x @ W1, then out = h @ W2. What is out?
W1 = [[1, 2], [0, 1]] and W2 = [[1], [1]]. W1 @ W2 has shape (2, 1). What is its top entry?
Number W1 = [[1, 2], [0, 1]] and W2 = [[1], [1]]. W1 @ W2 has shape (2, 1). What is its top entry?
σ(0) = 0.5. Using f′(x) = σ(x)·(1 − σ(x)), what is the slope of sigmoid at x = 0?
Number σ(0) = 0.5. Using f′(x) = σ(x)·(1 − σ(x)), what is the slope of sigmoid at x = 0?
ReLU(x) = max(0, x). What is its slope at x = −2?
Number ReLU(x) = max(0, x). What is its slope at x = −2?
A gradient flows back through 10 sigmoid layers. Each layer multiplies it by that layer’s σ′, which is at most 0.25. Start with a gradient of 1 and ignore the weights. What reaches the first layer, at most?
ChooseA gradient flows back through 10 sigmoid layers. Each layer multiplies it by that layer’s σ′, which is at most 0.25. Start with a gradient of 1 and ignore the weights. What reaches the first layer, at most?
Five sigmoid layers, ignoring the weights, best case: the gradient is multiplied by 0.25 five times. 0.25⁵ = 1 / N. What is N?
Number Five sigmoid layers, ignoring the weights, best case: the gradient is multiplied by 0.25 five times. 0.25⁵ = 1 / N. What is N?
GELU(x) = x · Φ(x), and Φ(−1) = 0.1587. What is GELU(−1)? (Four decimals.)
Number GELU(x) = x · Φ(x), and Φ(−1) = 0.1587. What is GELU(−1)? (Four decimals.)
Write the slope of ReLU for every element of an array x (1 where x > 0, else 0).
CodeWrite the slope of ReLU for every element of an array x (1 where x > 0, else 0).
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
Write the slope of sigmoid for every element of an array x: first σ(x), then σ(x)·(1 − σ(x)). Two lines.
CodeWrite the slope of sigmoid for every element of an array x: first σ(x), then σ(x)·(1 − σ(x)). Two lines.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
x = [2, 1], W1 = [[1, −1], [0, 1]], b1 = [0, 0], activation ReLU. The hidden layer is h = ReLU(x @ W1 + b1). What is h[1], the second hidden value?
Number x = [2, 1], W1 = [[1, −1], [0, 1]], b1 = [0, 0], activation ReLU. The hidden layer is h = ReLU(x @ W1 + b1). What is h[1], the second hidden value?
h = [2, 0], W2 = [[2], [−1]], b2 = 0.5. What is z = h @ W2 + b2, the number that goes into the output sigmoid?
Number h = [2, 0], W2 = [[2], [−1]], b2 = 0.5. What is z = h @ W2 + b2, the number that goes into the output sigmoid?
X holds 50 examples with 2 numbers each, (50, 2). W1 is (2, 16). What shape is H = tanh(X @ W1 + b1)?
Shape X holds 50 examples with 2 numbers each, (50, 2). W1 is (2, 16). What shape is H = tanh(X @ W1 + b1)?
A network 2 → 8 → 1: 2 inputs, one hidden layer of 8 units, 1 output. How many weights and biases does it have in total?
Number A network 2 → 8 → 1: 2 inputs, one hidden layer of 8 units, 1 output. How many weights and biases does it have in total?
A network 2 → 8 → 8 → 1: 2 inputs, two hidden layers of 8 units, 1 output. How many weights and biases?
Number A network 2 → 8 → 8 → 1: 2 inputs, two hidden layers of 8 units, 1 output. How many weights and biases?
Write the hidden layer: tanh of X @ W1 plus b1.
CodeWrite the hidden layer: tanh of X @ W1 plus b1.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
Write the forward pass of a network with two hidden ReLU layers, then return the output of the last layer (W3 and b3, no sigmoid). Several lines.
CodeWrite the forward pass of a network with two hidden ReLU layers, then return the output of the last layer (W3 and b3, no sigmoid). Several lines.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
a = 0.5 and y = 1. What is the local gradient ∂L/∂a = 2(a − y)?
Number a = 0.5 and y = 1. What is the local gradient ∂L/∂a = 2(a − y)?
A neuron computes a = sigmoid(z), and the loss is L = (a − y)². Going back, ∂L/∂a = −1 and ∂a/∂z = a(1 − a) = 0.5 × 0.5 = 0.25. Use the chain rule: what is ∂L/∂z?
Number A neuron computes a = sigmoid(z), and the loss is L = (a − y)². Going back, ∂L/∂a = −1 and ∂a/∂z = a(1 − a) = 0.5 × 0.5 = 0.25. Use the chain rule: what is ∂L/∂z?
A neuron computes z = w·x + b with x = 2. Going back, ∂L/∂z = −0.25, and the local gradient is ∂z/∂w = x = 2. What is ∂L/∂w?
Number A neuron computes z = w·x + b with x = 2. Going back, ∂L/∂z = −0.25, and the local gradient is ∂z/∂w = x = 2. What is ∂L/∂w?
Check on an easy function: L(w) = w². At w = 3 with ε = 0.1, what is (L(w + ε) − L(w − ε)) / (2ε)?
Number Check on an easy function: L(w) = w². At w = 3 with ε = 0.1, what is (L(w + ε) − L(w − ε)) / (2ε)?
Two examples: X = [[1, 2], [3, 1]] and dz = [[0.5], [−1]]. What is dW[0][0], the first entry of X.T @ dz?
Number Two examples: X = [[1, 2], [3, 1]] and dz = [[0.5], [−1]]. What is dW[0][0], the first entry of X.T @ dz?
X is (4, 2). The gradient at the hidden layer, dz1, is (4, 6). What is the shape of dW1 = X.T @ dz1?
Shape X is (4, 2). The gradient at the hidden layer, dz1, is (4, 6). What is the shape of dW1 = X.T @ dz1?
Rule 3 for the XOR network: dz2 is (4, 1), one number per example, and W2 is (6, 1). What is the shape of dz2 @ W2.T?
Shape Rule 3 for the XOR network: dz2 is (4, 1), one number per example, and W2 is (6, 1). What is the shape of dz2 @ W2.T?
One example, a hidden layer of 2 tanh neurons, one output. dz2 = [[0.5]], W2 = [[2], [−1]] (so W2.T = [[2, −1]]), and the hidden outputs are a1 = [[0.5, 0]]. The slope of tanh at a neuron is 1 − a², where a is that neuron’s output. Use rule 3 from section 3. What is the first number of dz1?
Number One example, a hidden layer of 2 tanh neurons, one output. dz2 = [[0.5]], W2 = [[2], [−1]] (so W2.T = [[2, −1]]), and the hidden outputs are a1 = [[0.5, 0]]. The slope of tanh at a neuron is 1 − a², where a is that neuron’s output. Use rule 3 from section 3. What is the first number of dz1?
Write the whole backward pass: every gradient the training step needs, from dz2 at the output down to dW1 and db1. Several lines.
CodeWrite the whole backward pass: every gradient the training step needs, from dz2 at the output down to dW1 and db1. Several lines.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
Write the update, then run 1000 steps of gradient descent on XOR.
CodeWrite the update, then run 1000 steps of gradient descent on XOR.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
Two examples have dz = [[0.5], [−1]]. The layer has one output neuron with one bias. What is db = dz.sum(axis=0)?
Number Two examples have dz = [[0.5], [−1]]. The layer has one output neuron with one bias. What is db = dz.sum(axis=0)?
An array of shape (3, 5) is contiguous, so its strides are (5, 1) and it starts at memory position 0. At which memory position is element [2, 3]?
Number An array of shape (3, 5) is contiguous, so its strides are (5, 1) and it starts at memory position 0. At which memory position is element [2, 3]?
An array of shape (2, 3, 4) is contiguous. What are its strides, counted in elements? Type them like a shape, for example (1, 2, 3).
Shape An array of shape (2, 3, 4) is contiguous. What are its strides, counted in elements? Type them like a shape, for example (1, 2, 3).
Write `contiguous_strides`(shape): the strides of a contiguous array, in elements. Walk the shape from the right.
CodeWrite contiguous_strides(shape): the strides of a contiguous array, in elements. Walk the shape from the right.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
x = np.arange(12.0).reshape(3, 4) has shape (3, 4) and strides (4, 1), in elements. What are the strides of x.T? Type them like a shape.
Shape x = np.arange(12.0).reshape(3, 4) has shape (3, 4) and strides (4, 1), in elements. What are the strides of x.T? Type them like a shape.
x = np.arange(12.0).reshape(3, 4), so x[1, 0] is 4. Then y = x.T and y[0, 1] = 100. What is x[1, 0] now?
Choosex = np.arange(12.0).reshape(3, 4), so x[1, 0] is 4. Then y = x.T and y[0, 1] = 100. What is x[1, 0] now?
t has shape (2, 3, 4) and strides (12, 4, 1). What are the strides of np.swapaxes(t, 1, 2), which has shape (2, 4, 3)? Type them like a shape.
Shape t has shape (2, 3, 4) and strides (12, 4, 1). What are the strides of np.swapaxes(t, 1, 2), which has shape (2, 4, 3)? Type them like a shape.
small = np.arange(6.0).reshape(2, 3) is [[0, 1, 2], [3, 4, 5]]. What is small.T.reshape(6)[1]?
Number small = np.arange(6.0).reshape(2, 3) is [[0, 1, 2], [3, 4, 5]]. What is small.T.reshape(6)[1]?
x = np.arange(12.0).reshape(3, 4). Which of these must be a copy, with its own memory?
Choosex = np.arange(12.0).reshape(3, 4). Which of these must be a copy, with its own memory?
Attention heads out has shape (2, 4, 10, 16) = (B, h, L, dₖ) and strides (640, 160, 16, 1). What are the strides of out.transpose(1, 2), with shape (2, 10, 4, 16)? Type them like a shape.
Shape Attention heads out has shape (2, 4, 10, 16) = (B, h, L, dₖ) and strides (640, 160, 16, 1). What are the strides of out.transpose(1, 2), with shape (2, 10, 4, 16)? Type them like a shape.
v is a float32 vector of shape (4,), and each float32 number takes 4 bytes. How many bytes of memory does `np.broadcast_to(v, (1000, 4))` store?
Number v is a float32 vector of shape (4,), and each float32 number takes 4 bytes. How many bytes of memory does np.broadcast_to(v, (1000, 4)) store?
Write the one line that reads element [i, j] from memory buf, given strides and a start offset. The tests use a transpose, a step, a slice and a broadcast.
CodeWrite the one line that reads element [i, j] from memory buf, given strides and a start offset. The tests use a transpose, a step, a slice and a broadcast.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
A float16 number has the bits 0 10001 1000000000: sign 0, exponent 10001 (that is 17), mantissa 1000000000 (that is 512). float16 has 10 mantissa bits and bias 15. What number is it?
Number A float16 number has the bits 0 10001 1000000000: sign 0, exponent 10001 (that is 17), mantissa 1000000000 (that is 512). float16 has 10 mantissa bits and bias 15. What number is it?
bfloat16 has 7 mantissa bits. How many different bfloat16 numbers are there from 1 up to 2 (1 included, 2 not included)?
Number bfloat16 has 7 mantissa bits. How many different bfloat16 numbers are there from 1 up to 2 (1 included, 2 not included)?
In bfloat16 there are 128 numbers from 1 up to 2, so neighbors are 1/128 apart. The numbers from 256 up to 512 use the same 128 mantissa patterns, scaled by 256. How far apart are neighbors there?
Number In bfloat16 there are 128 numbers from 1 up to 2, so neighbors are 1/128 apart. The numbers from 256 up to 512 use the same 128 mantissa patterns, scaled by 256. How far apart are neighbors there?
In bfloat16, neighbors from 256 to 512 are 2 apart. What is 256 + 0.75, stored in bfloat16?
ChooseIn bfloat16, neighbors from 256 to 512 are 2 apart. What is 256 + 0.75, stored in bfloat16?
The largest float16 is 65504. What is np.float16(300) * np.float16(300)?
ChooseThe largest float16 is 65504. What is np.float16(300) * np.float16(300)?
e¹¹ ≈ 59 874 and e¹² ≈ 162 755. The largest float16 is 65504. What is the largest whole number z for which np.exp(np.float16(z)) is not inf?
Number e¹¹ ≈ 59 874 and e¹² ≈ 162 755. The largest float16 is 65504. What is the largest whole number z for which np.exp(np.float16(z)) is not inf?
Scores z = [1000, 1001] in float32, where e⁸⁹ is already inf. What does the plain formula e^z / sum(e^z) give?
ChooseScores z = [1000, 1001] in float32, where e⁸⁹ is already inf. What does the plain formula e^z / sum(e^z) give?
Scores z = [1000, 1001]. Subtract the largest score first: [−1, 0]. Use e⁻¹ ≈ 0.37 and e⁰ = 1. What is the softmax probability of score 1 (the 1001), to two decimals?
Number Scores z = [1000, 1001]. Subtract the largest score first: [−1, 0]. Use e⁻¹ ≈ 0.37 and e⁰ = 1. What is the softmax probability of score 1 (the 1001), to two decimals?
Use ln 2 ≈ 0.69. What is ln(e¹⁰⁰⁰ + e¹⁰⁰⁰), to two decimals? (A computer gets inf when it computes this directly, but you can do it by hand.)
Number Use ln 2 ≈ 0.69. What is ln(e¹⁰⁰⁰ + e¹⁰⁰⁰), to two decimals? (A computer gets inf when it computes this directly, but you can do it by hand.)
Write the line that makes softmax safe for big scores: subtract each row’s largest score before exp.
CodeWrite the line that makes softmax safe for big scores: subtract each row’s largest score before exp.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
A weight w = 1 is stored in bfloat16, where the next number above 1 is 1 + 1/128 ≈ 1.0078. An update adds 0.001. What is w after the update, stored in bfloat16?
Number A weight w = 1 is stored in bfloat16, where the next number above 1 is 1 + 1/128 ≈ 1.0078. An update adds 0.001. What is w after the update, stored in bfloat16?
Write `cross_entropy` from raw scores with log-sum-exp, so it never sees inf. z has shape (B, V); t holds each row’s target index. Compute lse, one log-sum-exp per row, shape (B,).
CodeWrite cross_entropy from raw scores with log-sum-exp, so it never sees inf. z has shape (B, V); t holds each row’s target index. Compute lse, one log-sum-exp per row, shape (B,).
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
w = [0.5, nan, −1] and x = [1, 2, 3]. What is w @ x?
Choosew = [0.5, nan, −1] and x = [1, 2, 3]. What is w @ x?
a = 2, b = 3, c = 1, d = −2. e = a·b, f = e + c, L = f·d. What is ∂L/∂d?
Number a = 2, b = 3, c = 1, d = −2. e = a·b, f = e + c, L = f·d. What is ∂L/∂d?
a = 2, b = 3, c = 1, d = −2. e = a·b, f = e + c, L = f·d. What is ∂L/∂a?
Number a = 2, b = 3, c = 1, d = −2. e = a·b, f = e + c, L = f·d. What is ∂L/∂a?
Write the gradient that a product sends to its first input.
CodeWrite the gradient that a product sends to its first input.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
y = x · x with x = 3. What is dy/dx?
Number y = x · x with x = 3. What is dy/dx?
Someone writes self.grad = … and other.grad = … (with =, not +=). For y = x · x at x = 3, what does their backward give for x.grad?
ChooseSomeone writes self.grad = … and other.grad = … (with =, not +=). For y = x · x at x = 3, what does their backward give for x.grad?
Write the backward rule of subtraction, out = self − other: one line for self, one line for other.
CodeWrite the backward rule of subtraction, out = self − other: one line for self, one line for other.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
Write the end of backward, after build(self): set the starting gradient of the output node, then call _backward on every node, in reverse topological order. Several lines.
CodeWrite the end of backward, after build(self): set the starting gradient of the output node, then call _backward on every node, in reverse topological order. Several lines.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
A dataset has 10 examples. The batch size is 4, and the last batch keeps whatever is left. How many batches are in one epoch?
Number A dataset has 10 examples. The batch size is 4, and the last batch keeps whatever is left. How many batches are in one epoch?
10 examples, batch size 4. How many examples are in the last batch?
Number 10 examples, batch size 4. How many examples are in the last batch?
10 examples, batch size 4, short last batch kept. You train for 3 epochs. How many steps (weight updates) is that?
Number 10 examples, batch size 4, short last batch kept. You train for 3 epochs. How many steps (weight updates) is that?
10 examples stored sorted by class: examples 0–4 are class 0, 5–9 are class 1. No shuffling, batch size 5. What does each training step see?
Choose10 examples stored sorted by class: examples 0–4 are class 0, 5–9 are class 1. No shuffling, batch size 5. What does each training step see?
Cut a random order into batches. Write the slice that takes one batch of indices, starting at i.
CodeCut a random order into batches. Write the slice that takes one batch of indices, starting at i.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
50 examples. You hold out 20% of them, once, and train on the rest with batch size 8. How many training steps are in one epoch?
Number 50 examples. You hold out 20% of them, once, and train on the rest with batch size 8. How many training steps are in one epoch?
Instead of splitting once, the code draws a new random train / held-out split at the start of every epoch. What goes wrong?
ChooseInstead of splitting once, the code draws a new random train / held-out split at the start of every epoch. What goes wrong?
A script sets the seed at the top. You run it with seed 0, run it again with seed 0, then once with seed 1. Same code, same computer. Which runs print the same loss curve?
ChooseA script sets the seed at the top. You run it with seed 0, run it again with seed 0, then once with seed 1. Same code, same computer. Which runs print the same loss curve?
Three sequences, [5, 3, 8], [2, 9] and [7], are padded with PAD = 0 at the end into one batch. What is the shape of the batch?
Shape Three sequences, [5, 3, 8], [2, 9] and [7], are padded with PAD = 0 at the end into one batch. What is the shape of the batch?
[5, 3, 8], [2, 9] and [7] padded into a (3, 3) batch. How many cells are PAD?
Number[5, 3, 8], [2, 9] and [7] padded into a (3, 3) batch. How many cells are PAD?
Four sequences of lengths 2, 3, 6 and 1 go into ONE batch, padded to the longest. How many cells of the batch are PAD?
Number Four sequences of lengths 2, 3, 6 and 1 go into ONE batch, padded to the longest. How many cells of the batch are PAD?
Write the loop of the collate function: copy each sequence into its row of X, and mark its real cells as not blocked (False) in the mask. Replace ____ with as many lines as you need.
CodeWrite the loop of the collate function: copy each sequence into its row of X, and mark its real cells as not blocked (False) in the mask. Replace ____ with as many lines as you need.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
pred has shape (4, 1) and y has shape (4,). What is the shape of pred − y?
Shape pred has shape (4, 1) and y has shape (4,). What is the shape of pred − y?
A line fit on 4 points has a shape bug: pred − y has shape (4, 4), so the loss compares EVERY prediction with EVERY target 3, 3, 7, 7. The best the model can do is predict one number everywhere: the one with the smallest mean squared distance to 3, 3, 7, 7. Which number is it?
Number A line fit on 4 points has a shape bug: pred − y has shape (4, 4), so the loss compares EVERY prediction with EVERY target 3, 3, 7, 7. The best the model can do is predict one number everywhere: the one with the smallest mean squared distance to 3, 3, 7, 7. Which number is it?
You train a small MLP on just 4 examples, again and again. After 500 steps the loss is still 0.69, the same as at the start. What does that tell you?
ChooseYou train a small MLP on just 4 examples, again and again. After 500 steps the loss is still 0.69, the same as at the start. What does that tell you?
f(w) = w², w = 3, h = 0.01. f(w + h) = 9.0601 and f(w − h) = 8.9401. What is the numeric slope (f(w + h) − f(w − h)) / (2h)?
Number f(w) = w², w = 3, h = 0.01. f(w + h) = 9.0601 and f(w − h) = 8.9401. What is the numeric slope (f(w + h) − f(w − h)) / (2h)?
Fitting y = w·x + b to (1, 3), (2, 3), (3, 7), (4, 7) at w = 0.5, b = 0. The numeric gradient is dw = −21.5, db = −7.5. Someone’s formula gives dw = −10.75, db = −3.75. What is the most likely bug?
ChooseFitting y = w·x + b to (1, 3), (2, 3), (3, 7), (4, 7) at w = 0.5, b = 0. The numeric gradient is dw = −21.5, db = −7.5. Someone’s formula gives dw = −10.75, db = −3.75. What is the most likely bug?
Write the body of the loop: the numeric slope for entry i of w. Move w[i] up by h, then down by h, and put it back where it was. Replace ____ with as many lines as you need.
CodeWrite the body of the loop: the numeric slope for entry i of w. Move w[i] up by h, then down by h, and put it back where it was. Replace ____ with as many lines as you need.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
f(w) = w², so the gradient is 2w. Start at w = 1 and take gradient steps w ← w − lr · 2w with lr = 1.5. What is w after 3 steps?
Number f(w) = w², so the gradient is 2w. Start at w = 1 and take gradient steps w ← w − lr · 2w with lr = 1.5. What is w after 3 steps?
The loss goes 16, 86, 466, 2550, … then inf, and a few dozen steps later NaN. Where can NaN come from, if every number started out normal?
ChooseThe loss goes 16, 86, 466, 2550, … then inf, and a few dozen steps later NaN. Where can NaN come from, if every number started out normal?
A loop forgets `opt.zero_grad()`. For one weight, backward() gives a gradient of 2 at each of the first 3 steps. What is that weight’s .grad when opt.step() runs at step 3?
Number A loop forgets opt.zero_grad(). For one weight, backward() gives a gradient of 2 at each of the first 3 steps. What is that weight’s .grad when opt.step() runs at step 3?
A next-token model is trained on [<bos>, a, b, c]. The bug: the targets are the inputs themselves, [<bos>, a, b], instead of [a, b, c]. The training loss falls almost to 0 within a few steps. What happens when it writes text?
ChooseA next-token model is trained on [<bos>, a, b, c]. The bug: the targets are the inputs themselves, [<bos>, a, b], instead of [a, b, c]. The training loss falls almost to 0 within a few steps. What happens when it writes text?
Boss, part A. The teammate’s version, in the comments, has three bugs. Write a working body for `train_line` in place of ____. As many lines as you need.
CodeBoss, part A. The teammate’s version, in the comments, has three bugs. Write a working body for train_line in place of ____. As many lines as you need.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
Boss, part B. `lm_batch` builds one next-token batch: inputs X, targets Y, and keep (True where the loss should count). The teammate’s version, in the comments, has two bugs. Write a working body in place of ____.
CodeBoss, part B. lm_batch builds one next-token batch: inputs X, targets Y, and keep (True where the loss should count). The teammate’s version, in the comments, has two bugs. Write a working body in place of ____.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
Boss, part C. Paste the last line that python check.py `my_train.py` printed, between the quotes. If a check still fails, you can paste its FAIL line for a hint on the first failing check.
CodeBoss, part C. Paste the last line that python check.py my_train.py printed, between the quotes. If a check still fails, you can paste its FAIL line for a hint on the first failing check.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
Softmax gave p = [0.5, 0.25, 0.25]. Its Jacobian is J = diag(p) − pᵀp, so J[i][i] = pᵢ(1 − pᵢ). What is J[1][1]?
Number Softmax gave p = [0.5, 0.25, 0.25]. Its Jacobian is J = diag(p) − pᵀp, so J[i][i] = pᵢ(1 − pᵢ). What is J[1][1]?
Softmax gave p = [0.5, 0.25, 0.25]. Off the diagonal, the Jacobian J = diag(p) − pᵀp has J[i][j] = −pᵢ·pⱼ. What is J[0][1]?
Number Softmax gave p = [0.5, 0.25, 0.25]. Off the diagonal, the Jacobian J = diag(p) − pᵀp has J[i][j] = −pᵢ·pⱼ. What is J[0][1]?
For p = [0.5, 0.25, 0.25], J = diag(p) − pᵀp, and J[i][j] is how fast pᵢ moves when zⱼ moves. Adding the same number c to every score leaves p unchanged. What is the sum of each row of J?
ChooseFor p = [0.5, 0.25, 0.25], J = diag(p) − pᵀp, and J[i][j] is how fast pᵢ moves when zⱼ moves. Adding the same number c to every score leaves p unchanged. What is the sum of each row of J?
Softmax gave p = [0.5, 0.25, 0.25], and the gradient arriving at p is dp = [1, 0, 2]. Use the short form for dz from above, without building J. What is dz[2]?
Number Softmax gave p = [0.5, 0.25, 0.25], and the gradient arriving at p is dp = [1, 0, 2]. Use the short form for dz from above, without building J. What is dz[2]?
Write the backward pass of a softmax that ran along each row: return dz = p ⊙ (dp − dp·p), with dp·p computed row by row.
CodeWrite the backward pass of a softmax that ran along each row: return dz = p ⊙ (dp − dp·p), with dp·p computed row by row.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
One attention head with 2 tokens: the weights are A = [[0.5, 0.5], [0.5, 0.5]] and the gradient arriving at the output O = A @ V is dO = [[1, 0], [0, 2]]. Use the matrix-product rule from the list above. What is dV[0][1]?
Number One attention head with 2 tokens: the weights are A = [[0.5, 0.5], [0.5, 0.5]] and the gradient arriving at the output O = A @ V is dO = [[1, 0], [0, 2]]. Use the matrix-product rule from the list above. What is dV[0][1]?
Two tokens. The weights are A = [[0.5, 0.5], [0.5, 0.5]] and the gradient arriving at them is dA = [[2, 1], [0, 2]]. Softmax backward, row by row: dS = A ⊙ (dA − (dA·A for that row)). What is dS[1][0]?
Number Two tokens. The weights are A = [[0.5, 0.5], [0.5, 0.5]] and the gradient arriving at them is dA = [[2, 1], [0, 2]]. Softmax backward, row by row: dS = A ⊙ (dA − (dA·A for that row)). What is dS[1][0]?
Two tokens, dₖ = 4. The score gradient is dS = [[0.25, −0.25], [−0.5, 0.5]] and the keys are K = [[1, −1, 0, 0], [0, 0, 1, −1]]. The scores were S = Q @ K.T / √dₖ. What is dQ[1][0]?
Number Two tokens, dₖ = 4. The score gradient is dS = [[0.25, −0.25], [−0.5, 0.5]] and the keys are K = [[1, −1, 0, 0], [0, 0, 1, −1]]. The scores were S = Q @ K.T / √dₖ. What is dQ[1][0]?
The same example: Q = [[1, 1, 0, 0], [0, 0, 1, 1]], dₖ = 4, and the score gradient is dS = [[0.25, −0.25], [−0.5, 0.5]]. What is dK[0][2], the gradient at entry 2 of key 0?
Number The same example: Q = [[1, 1, 0, 0], [0, 0, 1, 1]], dₖ = 4, and the score gradient is dS = [[0.25, −0.25], [−0.5, 0.5]]. What is dK[0][2], the gradient at entry 2 of key 0?
One attention head: L = 5 tokens, dₖ = 3, dᵥ = 6. Q and K are (5, 3), V is (5, 6). What shape is dK?
Shape One attention head: L = 5 tokens, dₖ = 3, dᵥ = 6. Q and K are (5, 3), V is (5, 6). What shape is dK?
Write the backward pass of one attention head: from dO, compute dV and dA, then dS (softmax backward, row by row), then dQ and dK. Several lines, all at the same indent.
CodeWrite the backward pass of one attention head: from dO, compute dV and dA, then dS (softmax backward, row by row), then dQ and dK. Several lines, all at the same indent.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
Softmax backward computes dz = p ⊙ (dp − dp·p). To run it for A = softmax(S), which value from the forward pass must training keep?
ChooseSoftmax backward computes dz = p ⊙ (dp − dp·p). To run it for A = softmax(S), which value from the forward pass must training keep?
Training with mixed precision and Adam keeps, for every parameter: the weight in bf16 (2 bytes), its gradient in bf16 (2 bytes), an fp32 master copy (4 bytes), and Adam’s m and v in fp32 (4 bytes each). How many GB does a model with 7 billion parameters need for these alone? (1 GB = 10⁹ bytes)
Number Training with mixed precision and Adam keeps, for every parameter: the weight in bf16 (2 bytes), its gradient in bf16 (2 bytes), an fp32 master copy (4 bytes), and Adam’s m and v in fp32 (4 bytes each). How many GB does a model with 7 billion parameters need for these alone? (1 GB = 10⁹ bytes)
One attention layer keeps its weights A, of shape (B, h, L, L), for the backward pass. B = 4 sequences, h = 8 heads, L = 1000 tokens, 2 bytes per number. How many MB is that? (1 MB = 10⁶ bytes)
Number One attention layer keeps its weights A, of shape (B, h, L, L), for the backward pass. B = 4 sequences, h = 8 heads, L = 1000 tokens, 2 bytes per number. How many MB is that? (1 MB = 10⁶ bytes)
One pre-norm block keeps the tensors in the table above for its backward pass. Count them in units of one (B, L, `d_model`) tensor: a tensor of width 4 × `d_model` counts as 4. How many units?
Number One pre-norm block keeps the tensors in the table above for its backward pass. Count them in units of one (B, L, d_model) tensor: a tensor of width 4 × d_model counts as 4. How many units?
For every layer, training keeps activations of shape (B, L, width) and attention weights of shape (B, h, L, L). You double L from 1000 to 2000 and keep B, h and the width. The activations double. What happens to the saved attention weights?
ChooseFor every layer, training keeps activations of shape (B, L, width) and attention weights of shape (B, h, L, L). You double L from 1000 to 2000 and keep B, h and the width. The activations double. What happens to the saved attention weights?
Section 2’s example: the outputs are O = [[1.5, 0.5], [1.5, 0.5]] and the gradient arriving at them is dO = [[1, 0], [0, 2]]. A fused kernel needs dA·A for each row without keeping A, so it uses dO·O instead. What is dO·O for row 1?
Number Section 2’s example: the outputs are O = [[1.5, 0.5], [1.5, 0.5]] and the gradient arriving at them is dO = [[1, 0], [0, 2]]. A fused kernel needs dA·A for each row without keeping A, so it uses dO·O instead. What is dO·O for row 1?
Points (1, 3), (2, 3), (3, 7), (4, 7); w = 0.5, b = 0, so the errors are −2.5, −2, −5.5, −5. Using all four points, what is dL/dw = mean of 2·error·x?
Number Points (1, 3), (2, 3), (3, 7), (4, 7); w = 0.5, b = 0, so the errors are −2.5, −2, −5.5, −5. Using all four points, what is dL/dw = mean of 2·error·x?
A line ŷ = w·x + b with w = 0.5, b = 0. Use a batch of just one point, (3, 7). Its error is −5.5. What is dL/dw = 2·error·x for this batch?
Number A line ŷ = w·x + b with w = 0.5, b = 0. Use a batch of just one point, (3, 7). Its error is −5.5. What is dL/dw = 2·error·x for this batch?
A training set has 1,000 examples and you use mini-batches of 50. How many steps make one epoch?
Number A training set has 1,000 examples and you use mini-batches of 50. How many steps make one epoch?
f(w1, w2) = 0.5·(w1² + 25·w2²), start at (−4, 1), gradient (−4, 25). One SGD step with lr = 0.03: what is the new w2?
Number f(w1, w2) = 0.5·(w1² + 25·w2²), start at (−4, 1), gradient (−4, 25). One SGD step with lr = 0.03: what is the new w2?
The loss is f(w1, w2) = 0.5·(w1² + 25·w2²), with gradient (w1, 25·w2). Start at (−4, 1). SGD with lr = 0.09 instead of 0.03, to make w1 move faster. What happens to w2?
ChooseThe loss is f(w1, w2) = 0.5·(w1² + 25·w2²), with gradient (w1, 25·w2). Start at (−4, 1). SGD with lr = 0.09 instead of 0.03, to make w1 move faster. What happens to w2?
Adam with lr = 0.3 starts at (−4, 1), gradient (−4, 25). On step 1, m̂ = g and v̂ = g², so the step is 0.3·g/√(g²). By how much does w1 change? (Give the size of the move, not the new position.)
Number Adam with lr = 0.3 starts at (−4, 1), gradient (−4, 25). On step 1, m̂ = g and v̂ = g², so the step is 0.3·g/√(g²). By how much does w1 change? (Give the size of the move, not the new position.)
Write momentum’s step: first update the velocity u from the new gradient g, then move w by lr times u. Return both.
CodeWrite momentum’s step: first update the velocity u from the new gradient g, then move w by lr times u. Return both.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
Write Adam’s update. mₕₐₜ and vₕₐₜ are already computed.
CodeWrite Adam’s update. mₕₐₜ and vₕₐₜ are already computed.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
A layer has 100 inputs (fan-in n = 100). Using the rule s = 1/√n, what standard deviation should its weights have?
Number A layer has 100 inputs (fan-in n = 100). Using the rule s = 1/√n, what standard deviation should its weights have?
Inputs have std 1. They pass through 10 layers with no activation, width 100, weights with std 0.2. Each layer multiplies the std by √100 · 0.2. What is the std after 10 layers?
Number Inputs have std 1. They pass through 10 layers with no activation, width 100, weights with std 0.2. Each layer multiplies the std by √100 · 0.2. What is the std after 10 layers?
A network has 10 layers of width 100 with tanh after every layer. The weights have std 0.01, and the inputs have std 1. Without tanh, each layer would multiply the std by √100 · 0.01. After 10 layers the outputs are…
ChooseA network has 10 layers of width 100 with tanh after every layer. The weights have std 0.01, and the inputs have std 1. Without tanh, each layer would multiply the std by √100 · 0.01. After 10 layers the outputs are…
A schedule with peak = 0.004, warmup = 100 and total = 1000. Step 25 is still in warmup, so lr = peak · t / warmup. What is lr at step 25?
Number A schedule with peak = 0.004, warmup = 100 and total = 1000. Step 25 is still in warmup, so lr = peak · t / warmup. What is lr at step 25?
A schedule has peak = 0.004, warmup = 100, total = 1000. Step 550 is after warmup, so lr = peak · (total − t) / (total − warmup). What is lr at step 550?
Number A schedule has peak = 0.004, warmup = 100, total = 1000. Step 550 is after warmup, so lr = peak · (total − t) / (total − warmup). What is lr at step 550?
Write the body of a training loop for the loss (w − 3)², with momentum and a learning-rate schedule: each step gets its own learning rate from `lr_at`. Several lines.
CodeWrite the body of a training loop for the loss (w − 3)², with momentum and a learning-rate schedule: each step gets its own learning rate from lr_at. Several lines.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
A polynomial of degree 9, c0 + c1·x + c2·x² + … + c9·x⁹, has how many parameters?
Number A polynomial of degree 9, c0 + c1·x + c2·x² + … + c9·x⁹, has how many parameters?
Weight decay adds a penalty to the loss: total loss = data loss + 0.1·w². The data’s gradient for w is 0 right now, w = 2, lr = 0.5. After one gradient step, what is w?
Number Weight decay adds a penalty to the loss: total loss = data loss + 0.1·w². The data’s gradient for w is 0 right now, w = 2, lr = 0.5. After one gradient step, what is w?
Write one gradient step with weight decay: add the penalty’s gradient 2·lam·w to the data’s gradient `g_data`, then step downhill with lr. It must work for one weight and for an array of weights.
CodeWrite one gradient step with weight decay: add the penalty’s gradient 2·lam·w to the data’s gradient g_data, then step downhill with lr. It must work for one weight and for an array of weights.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
Hidden units h = [2, 4, 6, 8], dropout p = 0.5, keep mask [1, 0, 1, 0]. The kept units are divided by 1 − p. What does the third unit, 6, become?
Number Hidden units h = [2, 4, 6, 8], dropout p = 0.5, keep mask [1, 0, 1, 0]. The kept units are divided by 1 − p. What does the third unit, 6, become?
Write an inverted dropout layer. When train is False, return h unchanged. When train is True, drop each unit whose random number u is below p, and divide the kept ones by 1 − p. Several lines.
CodeWrite an inverted dropout layer. When train is False, return h unchanged. When train is True, drop each unit whose random number u is below p, and divide the kept ones by 1 − p. Several lines.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
model = nn.Sequential(nn.Linear(2, 6), nn.Tanh(), nn.Linear(6, 1)). A Linear layer has one weight for every input-output pair and one bias per output. How many numbers does the model learn in total?
Number model = nn.Sequential(nn.Linear(2, 6), nn.Tanh(), nn.Linear(6, 1)). A Linear layer has one weight for every input-output pair and one bias per output. How many numbers does the model learn in total?
In NumPy, W1 has shape (2, 6): 2 inputs, 6 outputs. What is the shape of nn.Linear(2, 6).weight?
Shape In NumPy, W1 has shape (2, 6): 2 inputs, 6 outputs. What is the shape of nn.Linear(2, 6).weight?
Your training loop has loss.backward() and opt.step(), but you forgot `opt.zero_grad()`. What happens?
ChooseYour training loop has loss.backward() and opt.step(), but you forgot opt.zero_grad(). What happens?
w = 1, x = 2, y = 6, loss = (w·x − y)², so each backward computes the gradient −16. You call loss.backward() twice, with no `zero_grad` in between. What is w.grad now?
Number w = 1, x = 2, y = 6, loss = (w·x − y)², so each backward computes the gradient −16. You call loss.backward() twice, with no zero_grad in between. What is w.grad now?
backward adds to w.grad, like PyTorch. Write `zero_grad` so that train(…, clear=True) reaches w = 3.
Codebackward adds to w.grad, like PyTorch. Write zero_grad so that train(…, clear=True) reaches w = 3.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
pred has shape (4, 1) and y has shape (4,). What is the shape of pred − y?
Shape pred has shape (4, 1) and y has shape (4,). What is the shape of pred − y?
Write the gradient of the last layer’s weights. It must match PyTorch’s W2 gradient above.
CodeWrite the gradient of the last layer’s weights. It must match PyTorch’s W2 gradient above.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
Write the hidden layer h inside forward. Use the box’s own weights (self.W1, self.b1) and tanh.
CodeWrite the hidden layer h inside forward. Use the box’s own weights (self.W1, self.b1) and tanh.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
Run `first_run.py` (section 8 code) on your computer. What loss does it print after 100 steps?
Number Run first_run.py (section 8 code) on your computer. What loss does it print after 100 steps?
Write the body of the inner loop: the three calls of one training step, in the right order.
CodeWrite the body of the inner loop: the three calls of one training step, in the right order.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
The scores are mat 2.0, sofa 1.0, roof 0.5, moon −1.0. Rounded, e^score is 7.4, 2.7, 1.6 and 0.4, which add up to 12.1. What probability does “mat” get? (2 decimals)
Number The scores are mat 2.0, sofa 1.0, roof 0.5, moon −1.0. Rounded, e^score is 7.4, 2.7, 1.6 and 0.4, which add up to 12.1. What probability does “mat” get? (2 decimals)
Scores [2, 0], with e² ≈ 7.39 and e⁰ = 1. What probability does the first word get? (2 decimals)
Number Scores [2, 0], with e² ≈ 7.39 and e⁰ = 1. What probability does the first word get? (2 decimals)
Scores [2, 0] give the probabilities 0.88 and 0.12. Add 10 to both scores: [12, 10]. What happens to the probabilities?
ChooseScores [2, 0] give the probabilities 0.88 and 0.12. Add 10 to both scores: [12, 10]. What happens to the probabilities?
The scores are mat 2.0, sofa 1.0, roof 0.5, moon −1.0. At T = 1, sampling 1000 words gives about 600 mat. Temperature divides every score by T. At temperature 0.1, what will sampling 1000 words give?
ChooseThe scores are mat 2.0, sofa 1.0, roof 0.5, moon −1.0. At T = 1, sampling 1000 words gives about 600 mat. Temperature divides every score by T. At temperature 0.1, what will sampling 1000 words give?
Scores [2, 0] at temperature T = 2, with e¹ ≈ 2.72. What probability does the first word get? (2 decimals)
Number Scores [2, 0] at temperature T = 2, with e¹ ≈ 2.72. What probability does the first word get? (2 decimals)
Write the line that applies the temperature. scores arrives as a Python list, and a list can’t be divided by a number, so make it a float array first.
CodeWrite the line that applies the temperature. scores arrives as a Python list, and a list can’t be divided by a number, so make it a float array first.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
The model gave the correct word p = 0.1. What is the loss −ln p? Use ln 10 ≈ 2.30. (2 decimals)
Number The model gave the correct word p = 0.1. What is the loss −ln p? Use ln 10 ≈ 2.30. (2 decimals)
The model gave the correct word p = 0.5. What is the loss? Use ln 2 ≈ 0.69. (2 decimals)
Number The model gave the correct word p = 0.5. What is the loss? Use ln 2 ≈ 0.69. (2 decimals)
Return the cross-entropy loss: softmax the scores, take the correct one, then −log.
CodeReturn the cross-entropy loss: softmax the scores, take the correct one, then −log.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
p = [0.7, 0.2, 0.1] for cat, dog, car, and the correct word is dog, so y = [0, 1, 0]. What is ∂L/∂s for dog, p − y?
Number p = [0.7, 0.2, 0.1] for cat, dog, car, and the correct word is dog, so y = [0, 1, 0]. What is ∂L/∂s for dog, p − y?
Gradient descent moves every score against its gradient: s ← s − lr × (p − y). With p − y = [0.7, −0.8, 0.1] for cat, dog, car, which score goes up?
ChooseGradient descent moves every score against its gradient: s ← s − lr × (p − y). With p − y = [0.7, −0.8, 0.1] for cat, dog, car, which score goes up?
Return the gradient of the loss with respect to the scores: p minus the one-hot target.
CodeReturn the gradient of the loss with respect to the scores: p minus the one-hot target.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
A model always splits its guess evenly over 4 words, p = 0.25 each, and the right word is always one of them. What is its perplexity?
Number A model always splits its guess evenly over 4 words, p = 0.25 each, and the right word is always one of them. What is its perplexity?
Ten words. Model A gives the right word p = 0.9 nine times and p = 0.001 once. Model B gives it p = 0.5 every time. Which has the lower perplexity?
ChooseTen words. Model A gives the right word p = 0.9 nine times and p = 0.001 once. Model B gives it p = 0.5 every time. Which has the lower perplexity?
For “the cat sat on the mat” the model gave the right next words p = 0.5, 0.25, 0.125, 0.25, 0.25. The surprises −ln p add up to 6.93, and ln 4 ≈ 1.386. What is its perplexity?
Number For “the cat sat on the mat” the model gave the right next words p = 0.5, 0.25, 0.125, 0.25, 0.25. The surprises −ln p add up to 6.93, and ln 4 ≈ 1.386. What is its perplexity?
Return the perplexity of a list of probabilities (the p the model gave each right word).
CodeReturn the perplexity of a list of probabilities (the p the model gave each right word).
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
One noisy number is x = mean + std × z, where z comes from a standard normal. mean = 3, std = 2, and z = −1.5. What is x?
Number One noisy number is x = mean + std × z, where z comes from a standard normal. mean = 3, std = 2, and z = −1.5. What is x?
What is the shape of `rng.standard_normal((1000, 2))` * 2 + 3?
Shape What is the shape of rng.standard_normal((1000, 2)) * 2 + 3?
Scores [2, 0]. Softmax gives [0.88, 0.12]. If you pass these probabilities to `F.cross_entropy` as if they were scores, it applies softmax again. What happens to training?
ChooseScores [2, 0]. Softmax gives [0.88, 0.12]. If you pass these probabilities to F.cross_entropy as if they were scores, it applies softmax again. What happens to training?
Write `mean_loss` for a batch: for each row of scores, divide by T, take a safe softmax (subtract the max before np.exp), and take −ln of the probability of that row’s right word. Return the mean of these losses over the rows. Several lines.
CodeWrite mean_loss for a batch: for each row of scores, divide by T, take a safe softmax (subtract the max before np.exp), and take −ln of the probability of that row’s right word. Return the mean of these losses over the rows. Several lines.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
The picture is 5×5 with 1s in column 2 and 0s everywhere else. The filter is [[−1, 0, 1], [−1, 0, 1], [−1, 0, 1]]. What is out[0][0], the output for the top-left 3×3 window?
Number The picture is 5×5 with 1s in column 2 and 0s everywhere else. The filter is [[−1, 0, 1], [−1, 0, 1], [−1, 0, 1]]. What is out[0][0], the output for the top-left 3×3 window?
The picture is 5×5 with 1s in column 1 and 0s everywhere else. The filter is [[−1, 0, 1], [−1, 0, 1], [−1, 0, 1]], no padding, stride 1. What is out[0][1]?
Number The picture is 5×5 with 1s in column 1 and 0s everywhere else. The filter is [[−1, 0, 1], [−1, 0, 1], [−1, 0, 1]], no padding, stride 1. What is out[0][1]?
A 5×5 picture has 1s in column 2 and 0s everywhere else. With the horizontal-edge filter [[−1, −1, −1], [0, 0, 0], [1, 1, 1]], no padding and stride 1, what does the 3×3 output look like?
ChooseA 5×5 picture has 1s in column 2 and 0s everywhere else. With the horizontal-edge filter [[−1, −1, −1], [0, 0, 0], [1, 1, 1]], no padding and stride 1, what does the 3×3 output look like?
A 28×28 picture, a 5×5 filter, no padding, stride 1. How many pixels wide is the output?
Number A 28×28 picture, a 5×5 filter, no padding, stride 1. How many pixels wide is the output?
A 32×32 picture, a 3×3 filter, padding 1, stride 2. How many pixels wide is the output?
Number A 32×32 picture, a 3×3 filter, padding 1, stride 2. How many pixels wide is the output?
A 6×6 picture and a 3×3 filter give a 4×4 output. A dense layer that makes the same 4×4 output connects every one of the 36 pixels to every one of the 16 output cells. How many weights does it have (no biases)?
Number A 6×6 picture and a 3×3 filter give a 4×4 output. A dense layer that makes the same 4×4 output connects every one of the 36 pixels to every one of the 16 output cells. How many weights does it have (no biases)?
A 3×3 convolution layer reads 8 input channels and makes 16 output channels. How many numbers does it learn, weights plus biases?
Number A 3×3 convolution layer reads 8 input channels and makes 16 output channels. How many numbers does it learn, weights plus biases?
Two 3×3 layers are stacked, stride 1, no pooling. One output pixel of the second layer depends on a square patch of the picture. How many pixels wide is that patch?
Number Two 3×3 layers are stacked, stride 1, no pooling. One output pixel of the second layer depends on a square patch of the picture. How many pixels wide is that patch?
2×2 max pooling of [[1, 3, 0, 2], [4, 2, 1, 1], [0, 0, 5, 6], [1, 2, 7, 0]] gives a 2×2 result. What is its bottom-right number?
Number 2×2 max pooling of [[1, 3, 0, 2], [4, 2, 1, 1], [0, 0, 5, 6], [1, 2, 7, 0]] gives a 2×2 result. What is its bottom-right number?
Track the patch width and the gap layer by layer. A k×k conv adds (k − 1) × gap to the width. A 2×2 pool adds 1 × gap to the width, then doubles the gap. After a 3×3 conv and a 2×2 pool, the width is 4 and the gap is 2. Add a second 3×3 conv, then a second 2×2 pool. How many pixels wide is the patch one final output sees?
Number Track the patch width and the gap layer by layer. A k×k conv adds (k − 1) × gap to the width. A 2×2 pool adds 1 × gap to the width, then doubles the gap. After a 3×3 conv and a 2×2 pool, the width is 4 and the gap is 2. Add a second 3×3 conv, then a second 2×2 pool. How many pixels wide is the patch one final output sees?
Write one line: the output at (i, j) is the 3×3 window starting at (i, j), times the filter, summed.
CodeWrite one line: the output at (i, j) is the 3×3 window starting at (i, j), times the filter, summed.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
Write a convolution for any square filter and any stride, no padding. Set the output side m, make out, and fill every out[i, j]. Replace ____ with as many lines as you need.
CodeWrite a convolution for any square filter and any stride, no padding. Set the output side m, make out, and fill every out[i, j]. Replace ____ with as many lines as you need.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
A plain stack has 4 layers. Going down, each one multiplies the gradient by its slope, 0.5. The gradient at the top is 1. What reaches the bottom?
Number A plain stack has 4 layers. Going down, each one multiplies the gradient by its slope, 0.5. The gradient at the top is 1. What reaches the bottom?
30 plain tanh layers. Going down, each one shrinks the gradient to about three quarters of its size. The gradient at the top has std 1. Roughly how big is it when it reaches layer 0?
Choose30 plain tanh layers. Going down, each one shrinks the gradient to about three quarters of its size. The gradient at the top has std 1. Roughly how big is it when it reaches layer 0?
A residual layer computes y = x + f(x), with f(x) = 0.5 · x. What is the slope dy/dx?
Number A residual layer computes y = x + f(x), with f(x) = 0.5 · x. What is the slope dy/dx?
x has shape (64, 8, 8): 64 channels, each 8×8. Both convolutions on the layer side are 3×3 with padding 1 and 64 filters. What shape is F(x)?
Shape x has shape (64, 8, 8): 64 channels, each 8×8. Both convolutions on the layer side are 3×3 with padding 1 and 64 filters. What shape is F(x)?
Feature 0 (column 0) is [1, 1, 5, 5]. Its mean is 3. Find its std, then the BatchNorm value of X[2][0] = (5 − mean) ÷ std. Ignore the tiny ε.
Number Feature 0 (column 0) is [1, 1, 5, 5]. Its mean is 3. Find its std, then the BatchNorm value of X[2][0] = (5 − mean) ÷ std. Ignore the tiny ε.
BatchNorm’s running mean for a feature is 2. The next batch has mean 4 for that feature. The update keeps 0.9 of the old value and adds 0.1 of the new one. What is the running mean now?
Number BatchNorm’s running mean for a feature is 2. The next batch has mean 4 for that feature. The update keeps 0.9 of the old value and adds 0.1 of the new one. What is the running mean now?
Training with a batch of 1 example, x = [3, 7, 1]. BatchNorm subtracts each feature’s mean over the batch. What is the output?
ChooseTraining with a batch of 1 example, x = [3, 7, 1]. BatchNorm subtracts each feature’s mean over the batch. What is the output?
LayerNorm uses each row’s own mean and std. Example 2 (row 2) is [5, 1, 5, 1]. What is the LayerNorm value of X[2][1]? Ignore ε.
Number LayerNorm uses each row’s own mean and std. Example 2 (row 2) is [5, 1, 5, 1]. What is the LayerNorm value of X[2][1]? Ignore ε.
A batch has 32 examples with 64 features each: shape (32, 64). LayerNorm computes 32 means, one per example. How many means does BatchNorm compute?
Number A batch has 32 examples with 64 features each: shape (32, 64). LayerNorm computes 32 means, one per example. How many means does BatchNorm compute?
A model writes text one token at a time, so at each step its input is a single sequence. Which normalization still works there?
ChooseA model writes text one token at a time, so at each step its input is a single sequence. Which normalization still works there?
Write `layer_norm`: normalize each row (one example) by its own mean and variance, with ε added to the variance before the square root, as `batch_norm` does for columns.
CodeWrite layer_norm: normalize each row (one example) by its own mean and variance, with ε added to the variance before the square root, as batch_norm does for columns.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
wₓ = 1, wₕ = 0.5, b = 0, h₀ = 0 and x₁ = 1. Use tanh(1) ≈ 0.76. What is h₁ = tanh(wₓ·x₁ + wₕ·h₀ + b)? (2 decimals)
Number wₓ = 1, wₕ = 0.5, b = 0, h₀ = 0 and x₁ = 1. Use tanh(1) ≈ 0.76. What is h₁ = tanh(wₓ·x₁ + wₕ·h₀ + b)? (2 decimals)
wₓ = 1, wₕ = 0.5, b = 0. h₁ = 0.76 and x₂ = 0. Use tanh(0.38) ≈ 0.36. What is h₂ = tanh(wₓ·x₂ + wₕ·h₁ + b)? (2 decimals)
Number wₓ = 1, wₕ = 0.5, b = 0. h₁ = 0.76 and x₂ = 0. Use tanh(0.38) ≈ 0.36. What is h₂ = tanh(wₓ·x₂ + wₕ·h₁ + b)? (2 decimals)
This RNN has a one-number hidden state: hₜ = tanh(wₓ·xₜ + wₕ·hₜ₋₁ + b). It reads a sequence of 1,000 inputs. How many learned numbers (weights and biases) does it have?
Number This RNN has a one-number hidden state: hₜ = tanh(wₓ·xₜ + wₕ·hₜ₋₁ + b). It reads a sequence of 1,000 inputs. How many learned numbers (weights and biases) does it have?
wₕ = 0.5. Going back one step multiplies the gradient by wₕ·(1 − h²). Ignoring tanh, three steps back would give 0.5³ = 0.125. With tanh, ∂h₄/∂h₁ is…
Choosewₕ = 0.5. Going back one step multiplies the gradient by wₕ·(1 − h²). Ignoring tanh, three steps back would give 0.5³ = 0.125. With tanh, ∂h₄/∂h₁ is…
No activation, wₕ = 1.5. After 30 steps the gradient is a product of 29 factors of 1.5. About how big is it?
ChooseNo activation, wₕ = 1.5. After 30 steps the gradient is a product of 29 factors of 1.5. About how big is it?
An RNN with a 32-number hidden state learns “remember the first symbol” perfectly for k = 20. Now k = 40: same rule, 1,500 training steps. What accuracy will it reach?
ChooseAn RNN with a 32-number hidden state learns “remember the first symbol” perfectly for k = 20. Now k = 40: same rule, 1,500 training steps. What accuracy will it reach?
Write the update inside the loop: h becomes tanh of (x times Wₓ, plus h times Wₕ, plus b).
CodeWrite the update inside the loop: h becomes tanh of (x times Wₓ, plus h times Wₕ, plus b).
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
Write `rnn_all`: run the same loop, but keep every hidden state, so the result has one row per step, shape (L, H). Replace ____ with as many lines as you need.
CodeWrite rnn_all: run the same loop, but keep every hidden state, so the result has one row per step, shape (L, H). Replace ____ with as many lines as you need.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
cₚᵣₑᵥ = 1.0, f = 0.9, i = 0.5, g = 0.8. What is c = f·cₚᵣₑᵥ + i·g?
Number cₚᵣₑᵥ = 1.0, f = 0.9, i = 0.5, g = 0.8. What is c = f·cₚᵣₑᵥ + i·g?
c = 1.3 and the output gate o = 0.6. Use tanh(1.3) ≈ 0.86. What is h = o·tanh(c)? (2 decimals)
Number c = 1.3 and the output gate o = 0.6. Use tanh(1.3) ≈ 0.86. What is h = o·tanh(c)? (2 decimals)
The forget gate is f = sigmoid(`w_f`·x + `u_f`·hₚᵣₑᵥ + `b_f`). With `w_f` = 2, x = 1, `u_f` = 0, hₚᵣₑᵥ = 0 and `b_f` = 1, what is f? Use e⁻³ ≈ 0.05. (2 decimals)
Number The forget gate is f = sigmoid(w_f·x + u_f·hₚᵣₑᵥ + b_f). With w_f = 2, x = 1, u_f = 0, hₚᵣₑᵥ = 0 and b_f = 1, what is f? Use e⁻³ ≈ 0.05. (2 decimals)
Nothing new is added (i = 0) and f = 0.9 at every step. c starts at 1. Use 0.9⁵ ≈ 0.59. What is c after 10 steps? (2 decimals)
Number Nothing new is added (i = 0) and f = 0.9 at every step. c starts at 1. Use 0.9⁵ ≈ 0.59. What is c after 10 steps? (2 decimals)
Nothing new is added. To keep half of c after 50 steps, the forget gate f must be about…
ChooseNothing new is added. To keep half of c after 50 steps, the forget gate f must be about…
Two LSTMs differ only in the starting bias of their forget gate: 0 or 2. The task is “remember the first symbol” with k = 80, 1,500 training steps and a 32-number state. A plain RNN only guesses (about 50%). Which LSTM reaches about 100%?
ChooseTwo LSTMs differ only in the starting bias of their forget gate: 0 or 2. The task is “remember the first symbol” with k = 80, 1,500 training steps and a 32-number state. A plain RNN only guesses (about 50%). Which LSTM reaches about 100%?
There are 6 colors. If a model gets the grammar right but picks the second color at random, what fraction of its lines follow every rule? (percent, 1 decimal)
Number There are 6 colors. If a model gets the grammar right but picks the second color at random, what fraction of its lines follow every rule? (percent, 1 decimal)
Write the cell update: keep f of the old c, and add i times the candidate g.
CodeWrite the cell update: keep f of the old c, and add i times the candidate g.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
A GRU step with hₚᵣₑᵥ = 0.6, update gate z = 0.25 and candidate h̃ = −0.2. What is the new h = (1 − z)·hₚᵣₑᵥ + z·h̃?
Number A GRU step with hₚᵣₑᵥ = 0.6, update gate z = 0.25 and candidate h̃ = −0.2. What is the new h = (1 − z)·hₚᵣₑᵥ + z·h̃?
Inputs have D = 10 numbers and the state has H = 20. Each block of a recurrent layer has a (D, H) input matrix, an (H, H) state matrix and H biases: 10×20 + 20×20 + 20 = 620 numbers. An LSTM has 4 blocks. How many numbers does a GRU have?
Number Inputs have D = 10 numbers and the state has H = 20. Each block of a recurrent layer has a (D, H) input matrix, an (H, H) state matrix and H biases: 10×20 + 20×20 + 20 = 620 numbers. An LSTM has 4 blocks. How many numbers does a GRU have?
The sentence is “he sat by the bank of the river” (8 words, counting from 0, “bank” is word 4). The backward RNN starts at “river”. When it reaches “bank”, how many words has it read, counting “bank”?
Number The sentence is “he sat by the bank of the river” (8 words, counting from 0, “bank” is word 4). The backward RNN starts at “river”. When it reaches “bank”, how many words has it read, counting “bank”?
Each direction has H = 32. For an 8-word sentence, the output is one row per word, each row the forward and backward states side by side. What shape is it?
Shape Each direction has H = 32. For an 8-word sentence, the output is one row per word, each row the forward and backward states side by side. What shape is it?
Two jobs. (1) Label every word of a finished sentence as noun, verb and so on. (2) Write a sentence one word at a time. Which job can a bidirectional RNN do?
ChooseTwo jobs. (1) Label every word of a finished sentence as noun, verb and so on. (2) Write a sentence one word at a time. Which job can a bidirectional RNN do?
The boss, part 1. Write one GRU step: the update gate z, the reset gate r, the candidate h̃ and the new h, all from section 3. Replace ____ with as many lines as you need.
CodeThe boss, part 1. Write one GRU step: the update gate z, the reset gate r, the candidate h̃ and the new h, all from section 3. Replace ____ with as many lines as you need.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
The boss, part 2. Paste the PASS line that boss.py printed for your own LSTM cell, between the quotes.
CodeThe boss, part 2. Paste the PASS line that boss.py printed for your own LSTM cell, between the quotes.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
An LSTM encoder with H = 32 reads a string of 24 digits. The decoder receives only the encoder’s last h and c. How many numbers is that?
Number An LSTM encoder with H = 32 reads a string of 24 digits. The decoder receives only the encoder’s last h and c. How many numbers is that?
E = [[1, 0], [0, 1], [0, −1]] (rows h₀, h₁, h₂) and the decoder state s = [2, 0]. What is the score of h₀, that is h₀ · s?
Number E = [[1, 0], [0, 1], [0, −1]] (rows h₀, h₁, h₂) and the decoder state s = [2, 0]. What is the score of h₀, that is h₀ · s?
The scores are [2, 0, 0]. Use e² ≈ 7.39 and e⁰ = 1. After softmax, what weight does h₀ get? (2 decimals)
Number The scores are [2, 0, 0]. Use e² ≈ 7.39 and e⁰ = 1. After softmax, what weight does h₀ get? (2 decimals)
The weights are [0.79, 0.11, 0.11] and E = [[1, 0], [0, 1], [0, −1]]. What is the second number of context = weights @ E?
ChooseThe weights are [0.79, 0.11, 0.11] and E = [[1, 0], [0, 1], [0, −1]]. What is the second number of context = weights @ E?
E holds 8 encoder states of 32 numbers each: E is (8, 32). The decoder state s is (32,). What shape is context = weights @ E? (Write one number, like (5).)
Shape E holds 8 encoder states of 32 numbers each: E is (8, 32). The decoder state s is (32,). What shape is context = weights @ E? (Write one number, like (5).)
Finish one attention step: the context is the weighted average of the encoder states.
CodeFinish one attention step: the context is the weighted average of the encoder states.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
The same reversing task and the same LSTM with H = 32, now with attention. At 20 digits the plain model got 0%. Does the attention model do better than 0%?
ChooseThe same reversing task and the same LSTM with H = 32, now with attention. At 20 digits the plain model got 0%. Does the attention model do better than 0%?
The model reads 427015 and writes “four hundred twenty seven thousand fifteen”. When it writes “seven”, which digit should get the most weight?
ChooseThe model reads 427015 and writes “four hundred twenty seven thousand fifteen”. When it writes “seven”, which digit should get the most weight?
The boss, part 1. Write attention for every decoder step at once: one row of weights per decoder state, each a softmax over the encoder states, and one context row per decoder state. Write the softmax yourself. Replace ____ with as many lines as you need.
CodeThe boss, part 1. Write attention for every decoder step at once: one row of weights per decoder state, each a softmax over the encoder states, and one context row per decoder state. Write the softmax yourself. Replace ____ with as many lines as you need.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
The boss, part 2. Paste the PASS line that boss.py printed for your own `attend_all()`, between the quotes.
CodeThe boss, part 2. Paste the PASS line that boss.py printed for your own attend_all(), between the quotes.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
A 28×28 picture has 784 numbers. The code has 2. How many times fewer numbers is the code?
Number A 28×28 picture has 784 numbers. The code has 2. How many times fewer numbers is the code?
A batch of 32 pictures has shape (32, 784). The encoder maps each picture to a 2-number code. What shape is the batch of codes?
Shape A batch of 32 pictures has shape (32, 784). The encoder maps each picture to a 2-number code. What shape is the batch of codes?
A 4-pixel picture x = [1, 0, 1, 1] is rebuilt as x̂ = [0.5, 0, 1, 0.5]. What is the mean squared error?
Number A 4-pixel picture x = [1, 0, 1, 1] is rebuilt as x̂ = [0.5, 0, 1, 0.5]. What is the mean squared error?
Write the mean squared error between a picture x and its rebuild xₕₐₜ (both NumPy arrays of the same shape).
CodeWrite the mean squared error between a picture x and its rebuild xₕₐₜ (both NumPy arrays of the same shape).
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
A VAE’s encoder gives μ = 1 and σ = 0.5 for one code number. The random draw is ε = −2. What is the code z = μ + σ·ε?
Number A VAE’s encoder gives μ = 1 and σ = 0.5 for one code number. The random draw is ε = −2. What is the code z = μ + σ·ε?
Use KL = ½(μ² + σ² − 1 − ln σ²). For one code number, μ = 2 and σ = 1 (so ln σ² = ln 1 = 0). What is the KL term?
Number Use KL = ½(μ² + σ² − 1 − ln σ²). For one code number, μ = 2 and σ = 1 (so ln σ² = ln 1 = 0). What is the KL term?
In demo.py, the plain autoencoder’s random codes landed far from every real code 11% of the time (about a third in the lab, which keeps fewer codes). A VAE draws random codes from N(0, 1). How often do they land far from every real code?
ChooseIn demo.py, the plain autoencoder’s random codes landed far from every real code 11% of the time (about a third in the lab, which keeps fewer codes). A VAE draws random codes from N(0, 1). How often do they land far from every real code?
Someone writes the VAE loss as the rebuild error AVERAGED over the 784 pixels (instead of summed), plus the KL term. What happens in training?
ChooseSomeone writes the VAE loss as the rebuild error AVERAGED over the 784 pixels (instead of summed), plus the KL term. What happens in training?
Write the VAE’s sampling step: given arrays mu, sigma and a standard normal draw eps (all the same shape), return the codes.
CodeWrite the VAE’s sampling step: given arrays mu, sigma and a standard normal draw eps (all the same shape), return the codes.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
Write the VAE loss for one picture: the binary cross-entropy of every pixel, summed, plus the KL term of every code number, summed.
CodeWrite the VAE loss for one picture: the binary cross-entropy of every pixel, summed, plus the KL term of every code number, summed.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
In the decoder’s cross-attention, where do K and V come from?
ChooseIn the decoder’s cross-attention, where do K and V come from?
Batch 2. The decoder input has 3 tokens (<bos> plus 2 target words), and the source has 6 words. What shape are one head’s cross-attention weights?
Shape Batch 2. The decoder input has 3 tokens (<bos> plus 2 target words), and the source has 6 words. What shape are one head’s cross-attention weights?
One decoder token reads a memory of 2 source tokens, with dₖ = 2. Its query is q = [√2 · ln 3, 0], about [1.55, 0]. From memory, the keys are k₀ = [1, 0] and k₁ = [0, 1], and the values are v₀ = [4, 0] and v₁ = [0, 8]. The output is a vector of 2 numbers. What is its first number?
Number One decoder token reads a memory of 2 source tokens, with dₖ = 2. Its query is q = [√2 · ln 3, 0], about [1.55, 0]. From memory, the keys are k₀ = [1, 0] and k₁ = [0, 1], and the values are v₀ = [4, 0] and v₁ = [0, 8]. The output is a vector of 2 numbers. What is its first number?
A model reads numbers aloud, with `d_model` = 64. A batch holds 32 numbers. The source digits are padded to 3 tokens, and the decoder input has 5 tokens. What shape is memory, the encoder’s output?
Shape A model reads numbers aloud, with d_model = 64. A batch holds 32 numbers. The source digits are padded to 3 tokens, and the decoder input has 5 tokens. What shape is memory, the encoder’s output?
At the default size, a decoder layer has 49,856 parameters and an encoder layer 33,280. The difference is 16,576, the same as one FFN. What does the decoder layer have that the encoder layer doesn’t?
ChooseAt the default size, a decoder layer has 49,856 parameters and an encoder layer 33,280. The difference is 16,576, the same as one FFN. What does the decoder layer have that the encoder layer doesn’t?
A model reads numbers aloud. A batch holds 64 numbers, the source digits are padded to 3 tokens, the decoder input has 5 tokens (words), and there are 4 heads. What shape are the cross-attention scores?
Shape A model reads numbers aloud. A batch holds 64 numbers, the source digits are padded to 3 tokens, the decoder input has 5 tokens (words), and there are 4 heads. What shape are the cross-attention scores?
Reading 742 aloud. The model has already written “seven hundred forty”. How many rows go into the decoder at this step?
Number Reading 742 aloud. The model has already written “seven hundred forty”. How many rows go into the decoder at this step?
A decoder with 2 layers reads 742 aloud: 4 words, one per step. The model keeps (caches) every result that cannot change between steps, and reuses it. In one decoder layer, how many times is memory multiplied by `W_K` for the whole answer?
Number A decoder with 2 layers reads 742 aloud: 4 words, one per step. The model keeps (caches) every result that cannot change between steps, and reuses it. In one decoder layer, how many times is memory multiplied by W_K for the whole answer?
The source has 5 tokens. The decoder input is <bos> plus 5 target words. What shape is the causal mask in the decoder’s self-attention (for one sentence)?
Shape The source has 5 tokens. The decoder input is <bos> plus 5 target words. What shape is the causal mask in the decoder’s self-attention (for one sentence)?
Reading 42 aloud, in a batch where longer answers need 5 words. The decoder input is <bos> forty two PAD PAD: 5 positions. The decoder’s self-attention mask blocks a cell if the causal mask blocks it or if its column is a PAD word. How many of the 25 cells are blocked?
Number Reading 42 aloud, in a batch where longer answers need 5 words. The decoder input is <bos> forty two PAD PAD: 5 positions. The decoder’s self-attention mask blocks a cell if the causal mask blocks it or if its column is a PAD word. How many of the 25 cells are blocked?
Build the three masks of one example. src holds the digit ids, `tgt_in` the decoder input ids, and pad is the id of PAD. Return (enc, dec, cross), True where a cell is blocked: enc is (Ls, Ls), dec is (Lt, Lt), cross is (Lt, Ls). Several lines.
CodeBuild the three masks of one example. src holds the digit ids, tgt_in the decoder input ids, and pad is the id of PAD. Return (enc, dec, cross), True where a cell is blocked: enc is (Ls, Ls), dec is (Lt, Lt), cross is (Lt, Ls). Several lines.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
Write one cross-attention head. y holds the decoder’s rows and memory the encoder’s output; src_pad is True for a PAD source token. Return (out, A): the output and the weights. Several lines.
CodeWrite one cross-attention head. y holds the decoder’s rows and memory the encoder’s output; src_pad is True for a PAD source token. Return (out, A): the output and the weights. Several lines.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
In “the cat sat . the dog sat . a cat ran .”, how many times is “sat” directly followed by “.”?
Number In “the cat sat . the dog sat . a cat ran .”, how many times is “sat” directly followed by “.”?
In “the cat sat . the dog sat . a cat ran .”, “cat” appears twice: once before “sat”, once before “ran”. What is P(sat | cat)?
Number In “the cat sat . the dog sat . a cat ran .”, “cat” appears twice: once before “sat”, once before “ran”. What is P(sat | cat)?
A counting model looks one word back. It counted only the text “the cat sat . the dog sat . a cat ran .”. What probability does it give the sentence “a dog sat .”?
ChooseA counting model looks one word back. It counted only the text “the cat sat . the dog sat . a cat ran .”. What probability does it give the sentence “a dog sat .”?
A counting model looks one word back. Its training stories contain “… on the mat .” many times, but no story contains “the mat .” without “on” before it. Can the model write “the mat .”?
ChooseA counting model looks one word back. Its training stories contain “… on the mat .” many times, but no story contains “the mat .” without “on” before it. Can the model write “the mat .”?
After a “.”, a model reads “the cat ran” and gives each next word P(the | .) = 0.5, P(cat | the) = 0.5, P(ran | cat) = 0.5. What is the perplexity?
Number After a “.”, a model reads “the cat ran” and gives each next word P(the | .) = 0.5, P(cat | the) = 0.5, P(ran | cat) = 0.5. What is the perplexity?
The stories have a vocabulary of 10 words. W holds one row of scores per previous word, one score per possible next word. What is the shape of W?
Shape The stories have a vocabulary of 10 words. W holds one row of scores per previous word, one score per possible next word. What is the shape of W?
The counting model scores perplexity 2.016 on the 40 held-out stories. The neural version starts with W all zeros (perplexity 10) and trains on the same 160 stories. Where is its perplexity after training?
ChooseThe counting model scores perplexity 2.016 on the 40 held-out stories. The neural version starts with W all zeros (perplexity 10) and trains on the same 160 stories. Where is its perplexity after training?
Compute the mean cross-entropy of the real next words. W is (2, 2), prev and nxt are arrays of word ids.
CodeCompute the mean cross-entropy of the real next words. W is (2, 2), prev and nxt are arrays of word ids.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
With a vocabulary of 10 words, how many rows does a count table need to look 2 words back (one row per possible pair of previous words)?
Number With a vocabulary of 10 words, how many rows does a count table need to look 2 words back (one row per possible pair of previous words)?
A real vocabulary has 50,000 tokens. A count table that looks 3 tokens back needs 50,000³ = 125,000,000,000,000 rows. What goes wrong?
ChooseA real vocabulary has 50,000 tokens. A count table that looks 3 tokens back needs 50,000³ = 125,000,000,000,000 rows. What goes wrong?
C is the count table of “the cat sat . the dog sat . a cat ran .” (row = this word, column = the next word). Write probs: the next-word probabilities with add-one smoothing, so every row adds up to 1. A few lines.
CodeC is the count table of “the cat sat . the dog sat . a cat ran .” (row = this word, column = the next word). Write probs: the next-word probabilities with add-one smoothing, so every row adds up to 1. A few lines.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
With the text “hello”, how many characters are in the vocabulary?
Number With the text “hello”, how many characters are in the vocabulary?
In “hello”, what is the id of “l”?
Number In “hello”, what is the id of “l”?
Build stoi: a dictionary from each character to its id (its position in chars).
CodeBuild stoi: a dictionary from each character to its id (its position in chars).
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
Write encode: turn a string into a list of ids.
CodeWrite encode: turn a string into a list of ids.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
A word-level tokenizer knows only: the, a, cat, dog, sat, ran, on, mat, rug, “.”. How does it encode “the cats sat .”?
ChooseA word-level tokenizer knows only: the, a, cat, dog, sat, ran, on, mat, rug, “.”. How does it encode “the cats sat .”?
Words: cat ×4, cats ×2, hat ×3, hats ×1. Cut into characters, how many times does the pair a + t appear in total (counting each word as often as it appears)?
Number Words: cat ×4, cats ×2, hat ×3, hats ×1. Cut into characters, how many times does the pair a + t appear in total (counting each word as often as it appears)?
After merging a + t, the words are c·at ×4, c·at·s ×2, h·at ×3, h·at·s ×1. How many times does the pair c + at appear now?
Number After merging a + t, the words are c·at ×4, c·at·s ×2, h·at ×3, h·at·s ×1. How many times does the pair c + at appear now?
The merges, in order, are: 1) a + t → at, 2) c + at → cat, 3) h + at → hat. How many tokens is “cats”?
Number The merges, in order, are: 1) a + t → at, 2) c + at → cat, 3) h + at → hat. How many tokens is “cats”?
Same three merges (a + t → at, c + at → cat, h + at → hat). “that” was never in the word list. How many tokens is it?
Number Same three merges (a + t → at, c + at → cat, h + at → hat). “that” was never in the word list. How many tokens is it?
Complete merge: replace every neighboring (a, b) in seg with the single token a + b.
CodeComplete merge: replace every neighboring (a, b) in seg with the single token a + b.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
BPE runs on a list of 40 words, and you keep merging. Each merge adds one token to the vocabulary. What happens to the total number of tokens across all the words?
ChooseBPE runs on a list of 40 words, and you keep merging. Each merge adds one token to the vocabulary. What happens to the total number of tokens across all the words?
A vocabulary of 50,000 tokens, with 512 numbers per token in the embedding table. How many numbers does the table hold, in millions?
Number A vocabulary of 50,000 tokens, with 512 numbers per token in the embedding table. How many numbers does the table hold, in millions?
million
Write the inside of the loop: for one word seg that appears n times, add n to the count of each of its neighboring pairs. The words are cat ×4, cats ×2, hat ×3, hats ×1. Two lines.
CodeWrite the inside of the loop: for one word seg that appears n times, add n to the count of each of its neighboring pairs. The words are cat ×4, cats ×2, hat ×3, hats ×1. Two lines.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
The table E has shape (4, 2): 4 ids, 2 numbers each. You look up the 5 ids of “hello”, [1, 0, 2, 2, 3]. What is the shape of the result?
Shape The table E has shape (4, 2): 4 ids, 2 numbers each. You look up the 5 ids of “hello”, [1, 0, 2, 2, 3]. What is the shape of the result?
“hello” has two l’s, at positions 2 and 3 (counting from 0). After the lookup, are their two vectors the same?
Choose“hello” has two l’s, at positions 2 and 3 (counting from 0). After the lookup, are their two vectors the same?
Embed by multiplying one-hot rows by the table.
CodeEmbed by multiplying one-hot rows by the table.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
What is [2, 1] · [1, 3]?
Number What is [2, 1] · [1, 3]?
cat = [2, 1]. You turn car so it points the opposite way from cat. What happens to cat·car?
Choosecat = [2, 1]. You turn car so it points the opposite way from cat. What happens to cat·car?
The word “apple” is looked up in the same table in both sentences. Is its vector the same in “sweet apple” and in “apple releases phone”?
ChooseThe word “apple” is looked up in the same table in both sentences. Is its vector the same in “sweet apple” and in “apple releases phone”?
Average the three vectors of “apple releases phone” (apple [1, 1, 0, 0], releases [0, 1, 0, 1], phone [0, 1, 0, 0]). What is the fruit number, the first entry, of the result? (3 decimals)
Number Average the three vectors of “apple releases phone” (apple [1, 1, 0, 0], releases [0, 1, 0, 1], phone [0, 1, 0, 0]). What is the fruit number, the first entry, of the result? (3 decimals)
Sentences: “cat eats fish”, “dog eats meat”, “cat drinks milk”, “car needs fuel”. How many times does one of eats, drinks, needs sit right next to “cat” (one word before or after)?
Number Sentences: “cat eats fish”, “dog eats meat”, “cat drinks milk”, “car needs fuel”. How many times does one of eats, drinks, needs sit right next to “cat” (one word before or after)?
Over (eats, drinks, needs): cat = [1, 1, 0], dog = [1, 0, 0], car = [0, 0, 1]. What is cat · dog?
Number Over (eats, drinks, needs): cat = [1, 1, 0], dog = [1, 0, 0], car = [0, 0, 1]. What is cat · dog?
Take thousands of sentences like “cat eats fish”, “dog eats meat” and “car needs fuel”. cat and dog keep appearing in the same places, and car in different ones. Whose count vector is finally closest to cat’s?
ChooseTake thousands of sentences like “cat eats fish”, “dog eats meat” and “car needs fuel”. cat and dog keep appearing in the same places, and car in different ones. Whose count vector is finally closest to cat’s?
Count neighbors: for every word, add 1 for each word right before and right after it.
CodeCount neighbors: for every word, add 1 for each word right before and right after it.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
Write mix: look up the rows of ids in the table E, then mix them with the weights w (one weight per id, adding up to 1) into one vector of 4 numbers. For example, ids [1, 0] with w = [0.75, 0.25] is 0.75 × sweet + 0.25 × apple. Two lines.
CodeWrite mix: look up the rows of ids in the table E, then mix them with the weights w (one weight per id, adding up to 1) into one vector of 4 numbers. For example, ids [1, 0] with w = [0.75, 0.25] is 0.75 × sweet + 0.25 × apple. Two lines.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
In an RNN with w = 0.5, ignoring tanh, each later word multiplies word 1’s effect on the hidden state by 0.5. After 3 more words, what is word 1’s effect multiplied by?
Number In an RNN with w = 0.5, ignoring tanh, each later word multiplies word 1’s effect on the hidden state by 0.5. After 3 more words, what is word 1’s effect multiplied by?
Before dividing by √2, what is the score of cat looking at dog? (cat = [2, 0], dog = [1, 1])
Number Before dividing by √2, what is the score of cat looking at dog? (cat = [2, 0], dog = [1, 1])
cat = [2, 0], dog = [1, 1], car = [0, 2]. Cat’s attention weights over (cat, dog, car) are [0.768, 0.187, 0.045]. What is the first number (x) of cat’s output? (2 decimals)
Number cat = [2, 0], dog = [1, 1], car = [0, 2]. Cat’s attention weights over (cat, dog, car) are [0.768, 0.187, 0.045]. What is the first number (x) of cat’s output? (2 decimals)
In an attention weight table for the words cat, dog, car (in that order), rows are words and columns are words. weights[0][2] = 0.045. What does that number say?
ChooseIn an attention weight table for the words cat, dog, car (in that order), rows are words and columns are words. weights[0][2] = 0.045. What does that number say?
Three words attend to each other with scores = X @ X.T / √2, then softmax: cat = [2, 0], dog = [1, 1], car = [0, 2]. If car becomes [5, 5] (cat and dog stay the same), what happens to car’s row of weights?
ChooseThree words attend to each other with scores = X @ X.T / √2, then softmax: cat = [2, 0], dog = [1, 1], car = [0, 2]. If car becomes [5, 5] (cat and dog stay the same), what happens to car’s row of weights?
Write the scores line of attention.
CodeWrite the scores line of attention.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
Take the “one-way keys” example: `W_Q` is the identity and `W_K` = [[0, 0], [1, 0]], so K = X @ `W_K` turns each word [x, y] into the key [y, 0]. cat = [2, 0] and car = [0, 2]. What is the score of car looking at cat, car’s query · cat’s key?
Number Take the “one-way keys” example: W_Q is the identity and W_K = [[0, 0], [1, 0]], so K = X @ W_K turns each word [x, y] into the key [y, 0]. cat = [2, 0] and car = [0, 2]. What is the score of car looking at cat, car’s query · cat’s key?
A sentence has 5 words. Q and K are both (5, 4). What shape is Q @ K.T?
Shape A sentence has 5 words. Q and K are both (5, 4). What shape is Q @ K.T?
At c = 0, sweet’s query is [0, 0]. Its keys are apple = [1, 0] and sweet = [0, 1]. How much weight does sweet put on apple?
Number At c = 0, sweet’s query is [0, 0]. Its keys are apple = [1, 0] and sweet = [0, 1]. How much weight does sweet put on apple?
The model is trained and saved. A user types a new sentence. Which numbers are different from the last time the model ran?
ChooseThe model is trained and saved. A user types a new sentence. Which numbers are different from the last time the model ran?
Without dividing by √2, cat’s scores are cat·cat, cat·dog, cat·car = [4, 2, 0]. Use e⁴ ≈ 54.6, e² ≈ 7.4, e⁰ = 1. What weight does cat put on itself after softmax? (2 decimals)
Number Without dividing by √2, cat’s scores are cat·cat, cat·dog, cat·car = [4, 2, 0]. Use e⁴ ≈ 54.6, e² ≈ 7.4, e⁰ = 1. What weight does cat put on itself after softmax? (2 decimals)
Write one attention head from start to end: the three projections, the scaled scores, softmax, the weighted average. Several lines.
CodeWrite one attention head from start to end: the three projections, the scaled scores, softmax, the weighted average. Several lines.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
Cat’s scores are [0.34, 0.05, blocked]; the third word is a PAD token. What if a blocked score were set to 0 instead of −∞? Use e^0.34 ≈ 1.40, e^0.05 ≈ 1.05, e^0 = 1. The PAD token would get…
ChooseCat’s scores are [0.34, 0.05, blocked]; the third word is a PAD token. What if a blocked score were set to 0 instead of −∞? Use e^0.34 ≈ 1.40, e^0.05 ≈ 1.05, e^0 = 1. The PAD token would get…
A sentence has 6 tokens. What shape is its causal mask?
Shape A sentence has 6 tokens. What shape is its causal mask?
Causal mask only, 4 tokens. How many of the 16 weights can be non-zero?
Number Causal mask only, 4 tokens. How many of the 16 weights can be non-zero?
A batch has 2 sentences of 5 tokens. The padding mask is (2, 1, 5) and the scores are (2, 5, 5). What shape is the mask after broadcasting?
Shape A batch has 2 sentences of 5 tokens. The padding mask is (2, 1, 5) and the scores are (2, 5, 5). What shape is the mask after broadcasting?
Build the causal mask: True (or 1) wherever a word would look ahead (column > row), False (or 0) elsewhere.
CodeBuild the causal mask: True (or 1) wherever a word would look ahead (column > row), False (or 0) elsewhere.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
3 words, 2 heads, each head outputs 2 numbers per word. The heads’ outputs are joined into one table called concat, one row per word. What shape is concat?
Shape 3 words, 2 heads, each head outputs 2 numbers per word. The heads’ outputs are joined into one table called concat, one row per word. What shape is concat?
2 heads each output 2 numbers per word. Their outputs are placed side by side in one table, concat, head 0 first. Heads and columns both count from 0. At which column of concat does head 1’s output start?
Number 2 heads each output 2 numbers per word. Their outputs are placed side by side in one table, concat, head 0 first. Heads and columns both count from 0. At which column of concat does head 1’s output start?
A model has `d_model` = 512 and 8 heads. How many numbers does each head get (dₖ)?
Number A model has d_model = 512 and 8 heads. How many numbers does each head get (dₖ)?
Write one attention head with a causal mask, from X and the three weight tables to the output out. Several lines.
CodeWrite one attention head with a causal mask, from X and the three weight tables to the output out. Several lines.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
Self-attention with no position information runs on “cat dog car”, then on “car dog cat”. Is cat’s output different?
ChooseSelf-attention with no position information runs on “cat dog car”, then on “car dog cat”. Is cat’s output different?
With 2 numbers per word, position pos gets [sin(pos), cos(pos)]. What is PE[0][1], the second number of position 0’s vector?
Number With 2 numbers per word, position pos gets [sin(pos), cos(pos)]. What is PE[0][1], the second number of position 0’s vector?
A model has `d_model` = 64. Before the position vector is added, each word vector is multiplied by √`d_model`. By what number?
Number A model has d_model = 64. Before the position vector is added, each word vector is multiplied by √d_model. By what number?
A sentence has 10 tokens and `d_model` = 8. What shape is its positional encoding?
Shape A sentence has 10 tokens and d_model = 8. What shape is its positional encoding?
X is (3, 2) and W1 is (2, 4). What shape is the hidden layer?
Shape X is (3, 2) and W1 is (2, 4). What shape is the hidden layer?
cat = [2, 0], W1 = [[1, −1, 0, 1], [0, 1, −1, 1]], b1 = [0, 0, 1, −2]. Before ReLU cat’s 4 hidden units are cat @ W1 + b1. How many are on (above 0) after ReLU?
Number cat = [2, 0], W1 = [[1, −1, 0, 1], [0, 1, −1, 1]], b1 = [0, 0, 1, −2]. Before ReLU cat’s 4 hidden units are cat @ W1 + b1. How many are on (above 0) after ReLU?
The FFN computes ReLU(X @ W1 + b1) @ W2 + b2. First X holds three words, cat, dog and car. Then you run it on cat alone. Compared with cat’s row when all three are processed together, the result is…
ChooseThe FFN computes ReLU(X @ W1 + b1) @ W2 + b2. First X holds three words, cat, dog and car. Then you run it on cat alone. Compared with cat’s row when all three are processed together, the result is…
30 layers, each multiplies by a random matrix that shrinks things a bit (gain 0.5). No residual, no LayerNorm. After 30 layers the numbers are…
Choose30 layers, each multiplies by a random matrix that shrinks things a bit (gain 0.5). No residual, no LayerNorm. After 30 layers the numbers are…
LayerNorm the word [1, 1, 5, 5]. Its mean is 3 and its std is 2. What does each 5 become?
Number LayerNorm the word [1, 1, 5, 5]. Its mean is 3 and its std is 2. What does each 5 become?
X is (2, 5, 8): 2 sentences, 5 words, 8 numbers per word. How many separate means does LayerNorm compute?
Number X is (2, 5, 8): 2 sentences, 5 words, 8 numbers per word. How many separate means does LayerNorm compute?
Finish LayerNorm over the last axis.
CodeFinish LayerNorm over the last axis.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
Write the FFN sublayer with its residual, X + ffn(X): widen with W1 and b1, ReLU, narrow back with W2 and b2, then add the input X. Two lines.
CodeWrite the FFN sublayer with its residual, X + ffn(X): widen with W1 and b1, ReLU, narrow back with W2 and b2, then add the input X. Two lines.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
A pre-norm block runs the line x = x + ffn(`layer_norm`(x)). For one token, x = [3, 1], and the FFN’s output is [0, 0]. What is x after this line?
ChooseA pre-norm block runs the line x = x + ffn(layer_norm(x)). For one token, x = [3, 1], and the FFN’s output is [0, 0]. What is x after this line?
A model has 3 pre-norm decoder blocks, each with a LayerNorm before attention and one before the FFN, and one final LayerNorm before the output layer. How many LayerNorms does it have?
Number A model has 3 pre-norm decoder blocks, each with a LayerNorm before attention and one before the FFN, and one final LayerNorm before the output layer. How many LayerNorms does it have?
Write the decoder stack: `n_blocks` pre-norm decoder blocks, then the final LayerNorm. In each block, masked self-attention and then the FFN read the LayerNorm of x, and each result is added to x.
CodeWrite the decoder stack: n_blocks pre-norm decoder blocks, then the final LayerNorm. In each block, masked self-attention and then the FFN read the LayerNorm of x, and each result is added to x.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
Level 17’s model reads 905 as one sequence: the digits, then =, then the words, then <eos>. One token per digit and one per word. How many tokens is that?
Number Level 17’s model reads 905 as one sequence: the digits, then =, then the words, then <eos>. One token per digit and one per word. How many tokens is that?
A batch holds 64 numbers. The model input x is (64, 8): 8 tokens per number, padded. There are 42 tokens in the vocabulary and `d_model` = 64. What shape are the logits?
Shape A batch holds 64 numbers. The model input x is (64, 8): 8 tokens per number, padded. There are 42 tokens in the vocabulary and d_model = 64. What shape are the logits?
`d_model` stays 64 and you change heads from 4 to 8. The parameter count…
Choosed_model stays 64 and you change heads from 4 to 8. The parameter count…
`d_model` = 64 and `d_ff` = 128. How many parameters does one FFN have? It is Linear(64 → 128), ReLU, Linear(128 → 64). Count weights and biases.
Numberd_model = 64 and d_ff = 128. How many parameters does one FFN have? It is Linear(64 → 128), ReLU, Linear(128 → 64). Count weights and biases.
Same batch: x is (64, 8), and the model has 4 heads. What shape are one block’s attention scores?
Shape Same batch: x is (64, 8), and the model has 4 heads. What shape are one block’s attention scores?
A trained model reads a number aloud (inference). Which of these changes while it runs?
ChooseA trained model reads a number aloud (inference). Which of these changes while it runs?
Reading 742 aloud. The model has already written “seven hundred forty”. How many rows go into the model at this step?
Number Reading 742 aloud. The model has already written “seven hundred forty”. How many rows go into the model at this step?
Write one step of greedy decoding. h is the last row of the model’s output (after the final LayerNorm), W and b are the output layer. Return the id of the next word.
CodeWrite one step of greedy decoding. h is the last row of the model’s output (after the final LayerNorm), W and b are the output layer. Return the id of the next word.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
Write the whole greedy loop. Start from the prompt (the digits and =). Each step runs the model on everything so far, takes the last row of the scores, appends the top id, and stops after the model picks eos (keep eos in the list). Stop after `max_new` new tokens even without eos.
CodeWrite the whole greedy loop. Start from the prompt (the digits and =). Each step runs the model on everything so far, takes the last row of the scores, appends the top id, and stops after the model picks eos (keep eos in the list). Stop after max_new new tokens even without eos.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
For 427, x = 4 2 7 = four hundred twenty seven and y = 2 7 = four hundred twenty seven <eos>: 8 targets. The loss counts only targets that are part of the answer (the words and <eos>). How many of the 8 count?
Number For 427, x = 4 2 7 = four hundred twenty seven and y = 2 7 = four hundred twenty seven <eos>: 8 targets. The loss counts only targets that are part of the answer (the words and <eos>). How many of the 8 count?
For the loss, the (64, 5, 42) logits are flattened so that every position is one row. What shape do they become?
Shape For the loss, the (64, 5, 42) logits are flattened so that every position is one row. What shape do they become?
You delete the causal mask and train again. What happens?
ChooseYou delete the causal mask and train again. What happens?
You build a new model for a new task. After 10 epochs it gets 0% right, and the training loss stopped falling long ago. What do you try first?
ChooseYou build a new model for a new task. After 10 epochs it gets 0% right, and the training loss stopped falling long ago. What do you try first?
Write the whole loss that skips PAD, starting from the scores. logits is (N, V), targets is (N,), and the argument pad is the id of PAD (0 by default). Return the mean loss over the rows whose target is not PAD. Several lines.
CodeWrite the whole loss that skips PAD, starting from the scores. logits is (N, V), targets is (N,), and the argument pad is the id of PAD (0 by default). Return the mean loss over the rows whose target is not PAD. Several lines.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
Scores: mat 2.0, sofa 1.5, rug 1.0, roof 0.0, moon −1.0. With T = 1, greedy picks mat. Use e² ≈ 7.4, e^1.5 ≈ 4.5, e¹ ≈ 2.7, e⁰ = 1, e^−1 ≈ 0.4. What probability does softmax give mat? (2 decimals)
Number Scores: mat 2.0, sofa 1.5, rug 1.0, roof 0.0, moon −1.0. With T = 1, greedy picks mat. Use e² ≈ 7.4, e^1.5 ≈ 4.5, e¹ ≈ 2.7, e⁰ = 1, e^−1 ≈ 0.4. What probability does softmax give mat? (2 decimals)
You sample (not greedy) with T = 0.01, so every score is multiplied by 100. What does sampling do now?
ChooseYou sample (not greedy) with T = 0.01, so every score is multiplied by 100. What does sampling do now?
Probabilities: mat 0.46, sofa 0.28, rug 0.17, roof 0.06, moon 0.03. Top-k with k = 2 keeps mat and sofa and divides by their total. What is sofa’s new probability? (2 decimals)
Number Probabilities: mat 0.46, sofa 0.28, rug 0.17, roof 0.06, moon 0.03. Top-k with k = 2 keeps mat and sofa and divides by their total. What is sofa’s new probability? (2 decimals)
Sorted probabilities: 0.4631, 0.2809, 0.1704, 0.0627, 0.0231. Top-p with p = 0.9 keeps words, starting from the largest, until the running total reaches 0.9. How many words does it keep?
Number Sorted probabilities: 0.4631, 0.2809, 0.1704, 0.0627, 0.0231. Top-p with p = 0.9 keeps words, starting from the largest, until the running total reaches 0.9. How many words does it keep?
Finish `top_p`: `keep_sorted` must be True for every word to keep, in sorted order.
CodeFinish top_p: keep_sorted must be True for every word to keep, in sorted order.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
First word: a 0.5, the 0.4, one 0.1. After a: dog 0.4, cat 0.3, fox 0.3. After the: dog 0.9, cat 0.05, fox 0.05. After one: dog 0.5, cat 0.25, fox 0.25. Greedy takes the best word at each step. What is the probability of the two-word sentence greedy writes? (2 decimals)
Number First word: a 0.5, the 0.4, one 0.1. After a: dog 0.4, cat 0.3, fox 0.3. After the: dog 0.9, cat 0.05, fox 0.05. After one: dog 0.5, cat 0.25, fox 0.25. Greedy takes the best word at each step. What is the probability of the two-word sentence greedy writes? (2 decimals)
First word: a 0.5, the 0.4, one 0.1. After a: dog 0.4, cat 0.3, fox 0.3. After the: dog 0.9, cat 0.05, fox 0.05. After one: dog 0.5, cat 0.25, fox 0.25. Greedy writes “a dog”, with probability 0.20. Beam search with width 2 keeps the two best first words (a and the) and then looks at all their second words. Will it find a better sentence than greedy’s 0.20?
ChooseFirst word: a 0.5, the 0.4, one 0.1. After a: dog 0.4, cat 0.3, fox 0.3. After the: dog 0.9, cat 0.05, fox 0.05. After one: dog 0.5, cat 0.25, fox 0.25. Greedy writes “a dog”, with probability 0.20. Beam search with width 2 keeps the two best first words (a and the) and then looks at all their second words. Will it find a better sentence than greedy’s 0.20?
First word: a 0.5, the 0.4, one 0.1. After a: dog 0.4, cat 0.3, fox 0.3. After the: dog 0.9, cat 0.05, fox 0.05. After one: dog 0.5, cat 0.25, fox 0.25. What is the probability of the best sentence that beam width 2 finds? (2 decimals)
Number First word: a 0.5, the 0.4, one 0.1. After a: dog 0.4, cat 0.3, fox 0.3. After the: dog 0.9, cat 0.05, fox 0.05. After one: dog 0.5, cat 0.25, fox 0.25. What is the probability of the best sentence that beam width 2 finds? (2 decimals)
A counting model looks at the last two words. It has seen “the cat sat on the mat .” far more than any other sentence. Greedy writes 14 words starting from “the”. What do you expect?
ChooseA counting model looks at the last two words. It has seen “the cat sat on the mat .” far more than any other sentence. Greedy writes 14 words starting from “the”. What do you expect?
After “. the”, ln p(cat) = −0.7752 and ln p(dog) = −1.1852. “cat” has already been written once, “dog” never. With penalty 1.0, score = ln p − 1.0 × (times written). What is cat’s score? (4 decimals)
Number After “. the”, ln p(cat) = −0.7752 and ln p(dog) = −1.1852. “cat” has already been written once, “dog” never. With penalty 1.0, score = ln p − 1.0 × (times written). What is cat’s score? (4 decimals)
After “on the” the model gives mat 0.3445, rug 0.2451, park 0.1984, roof 0.1835, and 0.0029 to each of the other 10 words. Top-p with p = 0.9 keeps how many words?
Number After “on the” the model gives mat 0.3445, rug 0.2451, park 0.1984, roof 0.1835, and 0.0029 to each of the other 10 words. Top-p with p = 0.9 keeps how many words?
Tokens 0 to 98 have already been processed by the model, and their k and v are in the KV cache. The model has just picked token 99. To predict token 100, it gives token 99 to the model. For how many tokens must it compute k and v now?
Number Tokens 0 to 98 have already been processed by the model, and their k and v are in the KV cache. The model has just picked token 99. To predict token 100, it gives token 99 to the model. For how many tokens must it compute k and v now?
With a KV cache: tokens 0, 1, 2 are in the cache, and token 3 is given to the model. How many rows does Q have at this step?
Number With a KV cache: tokens 0, 1, 2 are in the cache, and token 3 is given to the model. How many rows does Q have at this step?
No cache. Writing tokens 0 to 99 one at a time, the k of 1 token is computed at the first step, 2 at the second, …, 100 at the last. How many k computations is that in total?
Number No cache. Writing tokens 0 to 99 one at a time, the k of 1 token is computed at the first step, 2 at the second, …, 100 at the last. How many k computations is that in total?
Write the sampler’s filter: divide the scores by T (given), keep only the k largest, and return probabilities that add up to 1, with 0 for the dropped words.
CodeWrite the sampler’s filter: divide the scores by T (given), keep only the k largest, and return probabilities that add up to 1, with 0 for the dropped words.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
Why would one stack beat an encoder plus a decoder for a general-purpose model?
ChooseWhy would one stack beat an encoder plus a decoder for a general-purpose model?
A post-norm stack has 12 layers, each x = LayerNorm(x + f(x)). How many LayerNorms sit on the straight residual path from the input to the output?
Number A post-norm stack has 12 layers, each x = LayerNorm(x + f(x)). How many LayerNorms sit on the straight residual path from the input to the output?
x = [1, 1, 3, 5]. The root mean square is √((1² + 1² + 3² + 5²) / 4). What is it?
Number x = [1, 1, 3, 5]. The root mean square is √((1² + 1² + 3² + 5²) / 4). What is it?
x = [1, 1, 3, 5] and its root mean square is 3. What is the last number of RMSNorm(x)? (2 decimals)
Number x = [1, 1, 3, 5] and its root mean square is 3. What is the last number of RMSNorm(x)? (2 decimals)
Write RMSNorm for the last axis (eps keeps it safe when x is all zeros).
CodeWrite RMSNorm for the last axis (eps keeps it safe when x is all zeros).
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
θ = 0.5 radians per position. q sits at position 2. By how many radians is q rotated?
Number θ = 0.5 radians per position. q sits at position 2. By how many radians is q rotated?
q sits at position 2 and k at position 0, and their score is some number s. Now move both 10 positions later (12 and 10). The score…
Chooseq sits at position 2 and k at position 0, and their score is some number s. Now move both 10 positions later (12 and 10). The score…
q = [1, 0] at position 2 and k = [0, 1] at position 0, θ = 0.5. After rotating, q = [cos 1, sin 1] ≈ [0.54, 0.84] and k stays [0, 1]. What is the score q·k? (2 decimals)
Number q = [1, 0] at position 2 and k = [0, 1] at position 0, θ = 0.5. After rotating, q = [cos 1, sin 1] ≈ [0.54, 0.84] and k stays [0, 1]. What is the score q·k? (2 decimals)
Rotate a pair v by pos × theta radians (counter-clockwise).
CodeRotate a pair v by pos × theta radians (counter-clockwise).
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
SiLU(z) = z × sigmoid(z) = z / (1 + e^−z). Use e^−1 ≈ 0.37. What is SiLU(1)? (2 decimals)
Number SiLU(z) = z × sigmoid(z) = z / (1 + e^−z). Use e^−1 ≈ 0.37. What is SiLU(1)? (2 decimals)
A plain FFN has 2 matrices of `d_model` × `d_ff`. A SwiGLU FFN has 3 matrices of `d_model` × `d_ff`′. For the same number of weights, what is `d_ff`′ / `d_ff`? (2 decimals)
Number A plain FFN has 2 matrices of d_model × d_ff. A SwiGLU FFN has 3 matrices of d_model × d_ff′. For the same number of weights, what is d_ff′ / d_ff? (2 decimals)
Write the SwiGLU FFN of section 5 for a batch of word rows x of shape (L, `d_model`), with the weights `W_gate`, `W_up` and `W_down`. Return one row of `d_model` numbers per word.
CodeWrite the SwiGLU FFN of section 5 for a batch of word rows x of shape (L, d_model), with the weights W_gate, W_up and W_down. Return one row of d_model numbers per word.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
A model has 32 layers and 32 attention heads with dₖ = 128, and every head keeps its own K and V. Each number takes 2 bytes. How many KiB of KV cache does one token need? (1 KiB = 1,024 bytes)
Number A model has 32 layers and 32 attention heads with dₖ = 128, and every head keeps its own K and V. Each number takes 2 bytes. How many KiB of KV cache does one token need? (1 KiB = 1,024 bytes)
A model has 32 query heads and 8 K/V heads, numbered from 0; consecutive query heads share a K/V head. Which K/V head does query head 13 use?
Number A model has 32 query heads and 8 K/V heads, numbered from 0; consecutive query heads share a K/V head. Which K/V head does query head 13 use?
A GPU has 24 GiB of memory, and the model’s weights take 16 GiB of it. The rest holds the KV cache. The model uses grouped-query attention with 4 K/V heads, dₖ = 128 and 32 layers, and stores each number in 2 bytes. How many tokens of KV cache fit? (1 KiB = 1,024 bytes, 1 GiB = 1,024 × 1,024 KiB)
Number A GPU has 24 GiB of memory, and the model’s weights take 16 GiB of it. The rest holds the KV cache. The model uses grouped-query attention with 4 K/V heads, dₖ = 128 and 32 layers, and stores each number in 2 bytes. How many tokens of KV cache fit? (1 KiB = 1,024 bytes, 1 GiB = 1,024 × 1,024 KiB)
A model has 7 billion weights. A GPU does 10¹⁴ FLOPs per second (100 trillion). Using the 2N rule for the forward pass, at most how many tokens per second can it produce, counting only the arithmetic?
Number A model has 7 billion weights. A GPU does 10¹⁴ FLOPs per second (100 trillion). Using the 2N rule for the forward pass, at most how many tokens per second can it produce, counting only the arithmetic?
One block has model width d = 64 and FFN width 256. A new token attends to L = 128 tokens. Its attention needs two products: the scores, (1, 64) @ (64, 128), and the weighted sum, (1, 128) @ (128, 64). How many FLOPs does this token cost in the whole block, weights and attention together?
Number One block has model width d = 64 and FFN width 256. A new token attends to L = 128 tokens. Its attention needs two products: the scores, (1, 64) @ (64, 128), and the weighted sum, (1, 128) @ (128, 64). How many FLOPs does this token cost in the whole block, weights and attention together?
One block has model width 64 and FFN width 256, so its weights part is 98,304 FLOPs per token. Its attention part is 4 × L × 64 FLOPs. At what context length L are the two parts equal?
Number One block has model width 64 and FFN width 256, so its weights part is 98,304 FLOPs per token. Its attention part is 4 × L × 64 FLOPs. At what context length L are the two parts equal?
A model has 7 billion weights of 2 bytes each. The GPU reads 10¹² bytes per second from its memory and must read every weight once for each new token. For one user, at most how many tokens per second can it write? (1 decimal)
Number A model has 7 billion weights of 2 bytes each. The GPU reads 10¹² bytes per second from its memory and must read every weight once for each new token. For one user, at most how many tokens per second can it write? (1 decimal)
A model has 7 billion weights of 2 bytes each. One decode step reads the 1.4 × 10¹⁰ bytes of weights at 10¹² bytes per second. Each user’s token costs 1.4 × 10¹⁰ FLOPs at 10¹⁴ FLOPs per second. One step serves B users at once, with one read of the weights. For which B does the arithmetic take as long as the read?
Number A model has 7 billion weights of 2 bytes each. One decode step reads the 1.4 × 10¹⁰ bytes of weights at 10¹² bytes per second. Each user’s token costs 1.4 × 10¹⁰ FLOPs at 10¹⁴ FLOPs per second. One step serves B users at once, with one read of the weights. For which B does the arithmetic take as long as the read?
Store each of the 7 billion weights in 1 byte instead of 2. The GPU still reads 10¹² bytes per second, and each user’s token still costs 1.4 × 10¹⁰ FLOPs at 10¹⁴ FLOPs per second. For which B does the arithmetic now take as long as the read?
Number Store each of the 7 billion weights in 1 byte instead of 2. The GPU still reads 10¹² bytes per second, and each user’s token still costs 1.4 × 10¹⁰ FLOPs at 10¹⁴ FLOPs per second. For which B does the arithmetic now take as long as the read?
A model has 7 billion weights of 2 bytes each. The GPU reads 10¹² bytes per second and does 10¹⁴ FLOPs per second. It reads a prompt of 2,000 tokens in one pass. The weights are read once (14 ms), and each token costs 1.4 × 10¹⁰ FLOPs. How many ms does the pass take?
Number A model has 7 billion weights of 2 bytes each. The GPU reads 10¹² bytes per second and does 10¹⁴ FLOPs per second. It reads a prompt of 2,000 tokens in one pass. The weights are read once (14 ms), and each token costs 1.4 × 10¹⁰ FLOPs. How many ms does the pass take?
The same model and GPU: 7 billion weights of 2 bytes, 10¹² bytes per second, 10¹⁴ FLOPs per second. Now the prompt has only 50 tokens. How many ms does the prefill pass take?
Number The same model and GPU: 7 billion weights of 2 bytes, 10¹² bytes per second, 10¹⁴ FLOPs per second. Now the prompt has only 50 tokens. How many ms does the prefill pass take?
A GPU has 80 GB of memory. It holds the 14 GB of weights of the 7-billion-weight model, and the KV cache of every user. Each user has 4,000 tokens of cache, at 128 KiB per token. At most how many users fit? (1 GB = 10⁹ bytes, 1 KiB = 1,024 bytes)
Number A GPU has 80 GB of memory. It holds the 14 GB of weights of the 7-billion-weight model, and the KV cache of every user. Each user has 4,000 tokens of cache, at 128 KiB per token. At most how many users fit? (1 GB = 10⁹ bytes, 1 KiB = 1,024 bytes)
Write serve for one decode step. B users each have a KV cache of the given number of tokens. The step reads the weights once and every user’s cache. It also does 2N FLOPs for each user’s token. `bytes_per` is the bytes of every stored number: the weights and the KV cache. The step takes the longer of the read and the arithmetic. Return the step time in ms, the tokens per second for each user, and the tokens per second in total.
CodeWrite serve for one decode step. B users each have a KV cache of the given number of tokens. The step reads the weights once and every user’s cache. It also does 2N FLOPs for each user’s token. bytes_per is the bytes of every stored number: the weights and the KV cache. The step takes the longer of the read and the arithmetic. Return the step time in ms, the tokens per second for each user, and the tokens per second in total.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
Sequences in a batch must all have the same length. The longest problem is 99 + 99. How many tokens are in <bos>99+99=198<eos>?
Number Sequences in a batch must all have the same length. The longest problem is 99 + 99. How many tokens are in <bos>99+99=198<eos>?
Positions count from 0: x = <bos> 2 3 + 5 8 = 8 1 … At which position t does the model make its guess for the first digit of the answer?
Number Positions count from 0: x = <bos> 2 3 + 5 8 = 8 1 … At which position t does the model make its guess for the first digit of the answer?
y for 23 + 58 is 2 3 + 5 8 = 8 1 <eos> <pad>. With the loss on the answer digits only, every target before the answer and every <pad> becomes IGNORE. How many targets still count?
Number y for 23 + 58 is 2 3 + 5 8 = 8 1 <eos> <pad>. With the loss on the answer digits only, every target before the answer and every <pad> becomes IGNORE. How many targets still count?
Split one sequence of ids into the input x and the target y.
CodeSplit one sequence of ids into the input x and the target y.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
A batch of 64 examples, each padded to 11 tokens, so each x has 10. The vocabulary has 15 tokens. What shape are the logits?
Shape A batch of 64 examples, each padded to 11 tokens, so each x has 10. The vocabulary has 15 tokens. What shape are the logits?
q has shape (32, 10, 64): batch 32, length 10, `d_model` 64. With 4 heads, dₖ = 16. After view(32, 10, 4, 16) and transpose(1, 2), what shape is q?
Shape q has shape (32, 10, 64): batch 32, length 10, d_model 64. With 4 heads, dₖ = 16. After view(32, 10, 4, 16) and transpose(1, 2), what shape is q?
Split the last axis into h heads and move the heads in front of the length: (B, L, d) → (B, h, L, d // h).
CodeSplit the last axis into h heads and move the heads in front of the length: (B, L, d) → (B, h, L, d // h).
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
After the attention, out has shape (32, 4, 10, 16): batch 32, 4 heads, length 10, dₖ 16. What shape is out.transpose(1, 2).contiguous().view(32, 10, -1)?
Shape After the attention, out has shape (32, 4, 10, 16): batch 32, 4 heads, length 10, dₖ 16. What shape is out.transpose(1, 2).contiguous().view(32, 10, -1)?
Finish masked multi-head attention. q, k, v are already split into heads, shape (B, h, L, dₖ).
CodeFinish masked multi-head attention. q, k, v are already split into heads, shape (B, h, L, dₖ).
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
With the loss on every position, the training loss stays around 1.0 forever, while the answers become almost all correct. Why can’t it go lower?
ChooseWith the loss on every position, the training loss stays around 1.0 forever, while the answers become almost all correct. Why can’t it go lower?
The vocabulary has 15 tokens, and ln 15 ≈ 2.71. Your untrained model’s first loss on one batch is 9.4. What does that tell you?
ChooseThe vocabulary has 15 tokens, and ln 15 ≈ 2.71. Your untrained model’s first loss on one batch is 9.4. What does that tell you?
Return every problem the model got wrong, written like ‘2+7=1’. problems is a list of (a, b) pairs, answers the model’s answers as strings, in the same order.
CodeReturn every problem the model got wrong, written like ‘2+7=1’. problems is a list of (a, b) pairs, answers the model’s answers as strings, in the same order.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
python check.py `my_gpt.py` printed a line like forward = “0a1b2c3d4e5f”. Paste that whole line over the first line below, or put only the 12 letters and digits in place of ____.
Codepython check.py my_gpt.py printed a line like forward = “0a1b2c3d4e5f”. Paste that whole line over the first line below, or put only the 12 letters and digits in place of ____.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
Paste the report line that python check.py `my_gpt.py` model.pt printed: everything after “report =”.
CodePaste the report line that python check.py my_gpt.py model.pt printed: everything after “report =”.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
With β = 0.1, 0.2, 0.5, what is alpha-bar after 3 steps?
Number With β = 0.1, 0.2, 0.5, what is alpha-bar after 3 steps?
The signal is multiplied by 0.6 and the noise by 0.8. What is 0.6² + 0.8²?
Number The signal is multiplied by 0.6 and the noise by 0.8. What is 0.6² + 0.8²?
x0 = [1, 2], eps = [0.5, −1], alpha-bar = 0.36. What is the first number of x₃?
Number x0 = [1, 2], eps = [0.5, −1], alpha-bar = 0.36. What is the first number of x₃?
Noise is mixed into a spiral of 2D points with xₜ = √ᾱₜ x0 + √(1 − ᾱₜ) eps. On a 50-step schedule, alpha-bar at t = 50 is 0.0045. What does the spiral look like there?
ChooseNoise is mixed into a spiral of 2D points with xₜ = √ᾱₜ x0 + √(1 − ᾱₜ) eps. On a 50-step schedule, alpha-bar at t = 50 is 0.0045. What does the spiral look like there?
Write `add_noise`. ab is alpha-bar. It should work for one point or for an array of points.
CodeWrite add_noise. ab is alpha-bar. It should work for one point or for an array of points.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
Write `noisy_at`: from the clean point x0, the noise eps, the list of betas and a step t (counting from 1), compute alpha-bar for step t, then return the noisy point xₜ.
CodeWrite noisy_at: from the clean point x0, the noise eps, the list of betas and a step t (counting from 1), compute alpha-bar for step t, then return the noisy point xₜ.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
At t = 3 (alpha-bar 0.36), x₃ = [0.7, 1.0] and the true noise was eps = [0.5, 0.5]. What is the second number of x0?
Number At t = 3 (alpha-bar 0.36), x₃ = [0.7, 1.0] and the true noise was eps = [0.5, 0.5]. What is the second number of x0?
The simplest network always guesses that the noise is [0, 0]. For one point the true noise is [0.5, −1]. What is the loss, mean((guess − eps)²)?
Number The simplest network always guesses that the noise is [0, 0]. For one point the true noise is [0.5, −1]. What is the loss, mean((guess − eps)²)?
One training batch has 128 noisy points. Each input row is the point’s x and y, plus 12 numbers that encode the step t. What shape goes into the network?
Shape One training batch has 128 noisy points. Each input row is the point’s x and y, plus 12 numbers that encode the step t. What shape goes into the network?
Write `guess_clean`: given a noisy point, the network’s noise guess, and alpha-bar, return the guess of x0.
CodeWrite guess_clean: given a noisy point, the network’s noise guess, and alpha-bar, return the guess of x0.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
β = 0.5 and alpha-bar = 0.36. What is β / √(1 − alpha-bar)?
Number β = 0.5 and alpha-bar = 0.36. What is β / √(1 − alpha-bar)?
One reverse step: mean = (xₜ − 0.625 × ε̂) / √α. Here x₃ = [1.0, 0.4], the noise guess ε̂ = [0.5, −1] and α = 0.5. Use 1 / √0.5 ≈ 1.41. What is the first number of the mean? (2 decimals)
Number One reverse step: mean = (xₜ − 0.625 × ε̂) / √α. Here x₃ = [1.0, 0.4], the noise guess ε̂ = [0.5, −1] and α = 0.5. Use 1 / √0.5 ≈ 1.41. What is the first number of the mean? (2 decimals)
At the very last step, from t = 1 to t = 0, do we add fresh noise z?
ChooseAt the very last step, from t = 1 to t = 0, do we add fresh noise z?
Write one reverse step. z is the fresh noise.
CodeWrite one reverse step. z is the fresh noise.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
Guidance mixes two noise guesses: ε̂ = `ε_free` + w (`ε_shape` − `ε_free`). `ε_free` = [0.2, 0.0], `ε_shape` = [0.6, −0.4], w = 3. What is the first number of the guided guess ε̂?
Number Guidance mixes two noise guesses: ε̂ = ε_free + w (ε_shape − ε_free). ε_free = [0.2, 0.0], ε_shape = [0.6, −0.4], w = 3. What is the first number of the guided guess ε̂?
Sampling with guidance, ε̂ = `ε_free` + w (`ε_shape` − `ε_free`), takes 50 steps. How many times do we run the network in total? (One run handles all 300 points at once.)
Number Sampling with guidance, ε̂ = ε_free + w (ε_shape − ε_free), takes 50 steps. How many times do we run the network in total? (One run handles all 300 points at once.)
Write the guidance mix.
CodeWrite the guidance mix.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
Write one guided step back, from x to the next x, with guidance weight w. z is the fresh noise (zeros on the last step).
CodeWrite one guided step back, from x to the next x, with guidance weight w. z is the fresh noise (zeros on the last step).
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
A 6 × 6 picture is cut into patches of size 2 (each patch 2 × 2 pixels). What shape is the token matrix (tokens, numbers per token)?
Shape A 6 × 6 picture is cut into patches of size 2 (each patch 2 × 2 pixels). What shape is the token matrix (tokens, numbers per token)?
A 28 × 28 digit is cut into patches of size 7 (each patch 7 × 7 pixels). How many tokens is that?
Number A 28 × 28 digit is cut into patches of size 7 (each patch 7 × 7 pixels). How many tokens is that?
On a 28 × 28 picture, go from patch size 4 to patch size 2. What happens to the number of attention scores in each layer?
ChooseOn a 28 × 28 picture, go from patch size 4 to patch size 2. What happens to the number of attention scores in each layer?
Split each axis of an (H, W) picture into (which block, which pixel in the block), for patch size p.
CodeSplit each axis of an (H, W) picture into (which block, which pixel in the block), for patch size p.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
Write patchify: an (H, W) picture → (number of patches, p·p), patches left to right, then top to bottom.
CodeWrite patchify: an (H, W) picture → (number of patches, p·p), patches left to right, then top to bottom.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
Patch size 4 on a 28 × 28 digit gives 49 tokens of 16 numbers. Wᵢₙ is (16, 96). What is the shape of tokens @ Wᵢₙ?
Shape Patch size 4 on a 28 × 28 digit gives 49 tokens of 16 numbers. Wᵢₙ is (16, 96). What is the shape of tokens @ Wᵢₙ?
With 49 tokens, how many attention scores does one head compute in one layer?
Number With 49 tokens, how many attention scores does one head compute in one layer?
The autoencoder squeezes 784 pixels into 16 numbers. How many times fewer numbers does diffusion have to work on?
Number The autoencoder squeezes 784 pixels into 16 numbers. How many times fewer numbers does diffusion have to work on?
The 16 code numbers have a std of about 4. Level D1’s formula needs clean data with variance 1, so std 1. You divide every code number by the same number first. Which number?
Number The 16 code numbers have a std of about 4. Level D1’s formula needs clean data with variance 1, so std 1. You divide every code number by the same number first. Which number?
An autoencoder turns each 28 × 28 digit into a 16-number code, and its decoder turns a code back into a picture. Take a 3 and a 7 and go halfway from one to the other. What does mixing the pixels give, compared with mixing the two codes and decoding the result?
ChooseAn autoencoder turns each 28 × 28 digit into a 16-number code, and its decoder turns a code back into a picture. Take a 3 and a 7 and go halfway from one to the other. What does mixing the pixels give, compared with mixing the two codes and decoding the result?
Write the sampling loop. `eps_exact` is the perfect noise guess for this ring of 8 points. Every sample must land on one of the 8 points, and all 8 must appear.
CodeWrite the sampling loop. eps_exact is the perfect noise guess for this ring of 8 points. Every sample must land on one of the 8 points, and all 8 must appear.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
The boss, part 2. Paste the PASS line that boss.py printed for your own `add_noise`, guide and `step_back`, between the quotes.
CodeThe boss, part 2. Paste the PASS line that boss.py printed for your own add_noise, guide and step_back, between the quotes.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs