What does PyTorch do for you, and what is still your job?
A short test to skip this level
Solve these 5 questions on your own. Answer all of them correctly and the level counts as cleared with three stars, and every part of the page opens. Showing an answer doesn’t count.
Shape In NumPy, W1 has shape (2, 6): 2 inputs, 6 outputs. What is the shape of nn.Linear(2, 6).weight?
Number w = 1, x = 2, y = 6, loss = (w·x − y)², so each backward computes the gradient −16. You call loss.backward() twice, with no zero_grad in between. What is w.grad now?
Shape pred has shape (4, 1) and y has shape (4,). What is the shape of pred − y?
CodeWrite the hidden layer h inside forward. Use the box’s own weights (self.W1, self.b1) and tanh.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
CodeWrite the body of the inner loop: the three calls of one training step, in the right order.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
Warm-up2 questions from earlier levels
A quick review before you start. Optional. Nothing here locks the level.
In level 6 you trained a network to learn XOR with NumPy. You wrote the forward pass, the loss,
every line of the backward pass, and the update. It worked, but the backward pass was the hard part,
and it gets longer with every layer you add.
PyTorch is NumPy plus two things: tensors that can run on a GPU (a graphics card, which does many
multiplications at once; you don’t need one for this course), and gradients computed for you. This level is the
translation: the same network, written both ways, line by line. Then come the few things PyTorch does not do
for you, because that is where most bugs come from.
Levels 10 to 19 stay in NumPy in your browser, so you can see every number. PyTorch comes back when you train real
models on your computer: the GPT you write in level 21, the LSTM boss, and the diffusion boss.
Every question on this page runs in your browser. Section 8 has one optional exercise that runs PyTorch on your
own computer. When you want to do it, Run it on your computer shows how to install Python and PyTorch.
1. The same network, written twice
One network, written twice
Point at or tap a line, or pick a part below. The matching lines are highlighted on both sides.
NumPy: you write everything
PyTorch
BackwardPyTorch does this
This is the part PyTorch does for you: four lines of chain rule become loss.backward(). It walks the recorded operations in reverse and fills .grad on every parameter, with the same numbers your hand-written lines give.
Backward: NumPy 4 lines → PyTorch 1 line
dashed: still your job in PyTorchthe part you picked
Point at or tap the backward lines in the NumPy code. Four lines of chain rule turn into loss.backward().
Everything else maps one to one. Two names in the code may be new. logits are the raw scores before the
sigmoid. binary_cross_entropy_with_logits is level 3’s yes/no cross-entropy with the sigmoid built in
(it is more accurate than doing the two steps separately).
Number model = nn.Sequential(nn.Linear(2, 6), nn.Tanh(), nn.Linear(6, 1)). A Linear layer has one weight for every input-output pair and one bias per output. How many numbers does the model learn in total?
🔒 Answer the question above to unlock
2. Tensors and parameters
A tensor is PyTorch’s array. Most of what you know from NumPy works the same way: @, .shape, .sum(0),
slicing, broadcasting. Three things are new:
NumPy array
PyTorch tensor
number type
float64 by default (about 16 digits)
float32 by default for weights (about 7 digits, half the memory)
where it lives
main memory
"cpu", "cuda" (NVIDIA GPU) or "mps" (Apple GPU)
gradients
none
if requires_grad=True, backward() fills .grad
The weights inside nn.Linear are created with requires_grad=True, so they get a .grad after every
loss.backward(). Your data does not need one.
Almost everyone makes one mistake the first time: how each side stores a weight matrix.
In NumPy you wrote X @ W with W of shape (inputs, outputs). nn.Linear(in, out) stores its weight in the opposite
layout, one row per output, and computes x @ W.T + b. For example, nn.Linear(3, 5).weight has shape (5, 3).
Shape In NumPy, W1 has shape (2, 6): 2 inputs, 6 outputs. What is the shape of nn.Linear(2, 6).weight?
🔒 Answer the question above to unlock
Go deeper Why nn.Linear stores (out, in) and multiplies by its transpose
Why one row per output? Each row is the 2 numbers that make one hidden unit, so weight[3] is everything
about unit 3. To compute the layer it does x @ W.T + b, which is (4, 2) @ (2, 6) → (4, 6),
exactly the NumPy shape.
Same numbers, stored in the opposite layout. It only matters when you copy weights between the two,
or when you print .weight.shape and expect NumPy’s layout. The demo for this level copies NumPy weights
into PyTorch and has to transpose them for exactly this reason.
3. A model is a class
On the PyTorch side of the lab, the model is one object, model, and you call it like a function: model(X).
That object is an instance of a class. A class is a recipe for a box that keeps its own numbers and its own
functions. Compare two ways to write the first layer:
def layer(x, W, b): # a function: you pass the weights in every time return x @ W + bclass Layer: # a class: the box keeps its weights def __init__(self, W, b): # runs once, when you make the box: layer = Layer(W, b) self.W = W # self means "this box" self.b = b def forward(self, x): # the computation, using the box's own weights return x @ self.W + self.b
A PyTorch model is a class built the same way, with four rules:
__init__ builds the parts: self.l1 = nn.Linear(2, 6). Its first line is always super().__init__() (see the box below).
forward(self, x) computes the output from the parts.
model(x) calls forward(x) for you. You never call forward by its name.
model.parameters() collects every weight of every part you stored as self.something. The optimizer only
updates what parameters() returns.
I got stuck here What are super().__init__() and __call__?
super().__init__() runs the setup of nn.Module itself, which parameters() needs later. Without it, PyTorch stops
with an error as soon as you store the first layer.
__call__ is the Python name for “what happens when you write model(x)”. nn.Module writes it for you, and it
calls forward. The NumPy version below writes it by hand.
Here is the same idea in NumPy, so it runs in your browser.
CodeWrite the hidden layer h inside forward. Use the box’s own weights (self.W1, self.b1) and tanh.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
🔒 Answer the question above to unlock
Rule 4 has a trap. Level 21 stacks several identical blocks. The natural Python way to keep them is a list:
self.blocks = [Block(), Block(), Block()].
Predict firstGuess before you continue: a model keeps its 3 blocks in a plain Python list, self.blocks = [Block(), Block(), Block()], and calls them in forward. What happens when you train it?
I got stuck here My model runs, but its blocks never change. Why?
A plain Python list hides its contents from model.parameters(). The blocks still run in forward, so nothing
crashes and the loss is a real number. But the optimizer never sees their weights, so they stay at their
random starting values. Use nn.ModuleList([Block(), Block(), Block()]) instead: it is a list that PyTorch looks
inside. A quick check: sum(p.numel() for p in model.parameters()) should match the count you expect.
4. backward() adds, it doesn’t replace
This one rule causes most beginner bugs. loss.backward() computes the gradients and
adds them to whatever is already in .grad. It never clears .grad. Clearing is your job,
with opt.zero_grad().
Why add? A parameter can be used in several places, and its total gradient is the sum of the pieces
(you build this yourself in level U3). Inside one backward pass, adding is right. Across training steps,
it means old gradients stay until you remove them.
ChooseYour training loop has loss.backward() and opt.step(), but you forgot opt.zero_grad(). What happens?
Take one weight, w = 1, and one example, x = 2, y = 6, with loss (w⋅x−y)2.
Its gradient is 2(wx−y)x=2×(2−6)×2=−16.
Number w = 1, x = 2, y = 6, loss = (w·x − y)², so each backward computes the gradient −16. You call loss.backward() twice, with no zero_grad in between. What is w.grad now?
🔒 Answer the question above to unlock
Now try it. The buttons do what the PyTorch calls do, on this one weight.
Gradients are added together until you clear them
Press the three calls in order, then press 1 twice in a row and see w.grad grow.
w = 1 · x = 2 · y = 6 · loss = (w·x − y)² = 16 · lr = 0.05
gradient of this loss-16
w.grad (what step() uses)0
one backward
w.grad
or run 10 steps:
w.grad starts at 0. Press 1, 2, 3 for one clean training step.
w.grad, the sum step() usesthe gradient of one backwardw after each stepbest w = 3
Try it
Press “10 steps with zero_grad”: w walks to 3, where w·x = y and the loss is 0. Reset, then press
“10 steps without”. Each step uses the sum of every gradient so far, so w jumps past 5 and swings back.
Now do it by hand: backward, step, backward, step, and watch w.grad grow.
Copy how PyTorch keeps track of gradients, in NumPy. Param below is a tiny class (section 3): a box that holds two numbers,
w.data and w.grad (inside the class, self means “this box”). backward adds to w.grad, as PyTorch does.
Write zero_grad.
Codebackward adds to w.grad, like PyTorch. Write zero_grad so that train(…, clear=True) reaches w = 3.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
🔒 Answer the question above to unlock
5. Shapes that broadcast without asking
PyTorch broadcasts exactly like NumPy. That is convenient, and it hides a classic bug. The model outputs
one score per example as a column, shape (4, 1). Labels are often stored flat, shape (4,).
Shape pred has shape (4, 1) and y has shape (4,). What is the shape of pred − y?
I got stuck here My loss looks fine but the model learns nothing. Why?
Check the shapes going into the loss. (4, 1) − (4,) does not fail: broadcasting stretches both into a
(4, 4) grid and compares every prediction with every label. The mean of that grid is a number,
so nothing crashes, but it is the wrong number and its gradient points the wrong way.
The fix is one line: make the labels a column with y.reshape(-1, 1) (or y.unsqueeze(1) in PyTorch),
or flatten the prediction with pred.squeeze(1). A good habit is to print pred.shape and y.shape
once, right before the loss.
I got stuck here Why does print(loss) show tensor(0.6958, grad_fn=...) instead of a number?
loss is still a tensor that remembers how it was computed (that is the grad_fn), so backward() can use it.
To get a plain Python number, write loss.item(). Use it for printing and logging.
Don’t keep the tensor itself in a list across steps, like history.append(loss). Each one keeps its
whole graph and memory grows every step. Append loss.item() instead.
🔒 Answer the question above to unlock
6. Same numbers, both ways
PyTorch’s gradients are not an approximation. They are exactly the chain rule you wrote by hand.
The demo for this level copies the same weights into NumPy and into PyTorch, runs one backward pass in each,
and prints both: they match to the last digit. (The demo uses float64 on both sides. With PyTorch’s default
float32 they agree to about 7 digits.) Check one of them yourself.
For a 2→3→1 version of the network, with the weights below, PyTorch reports this gradient for the last layer
(shown as a column):
∂W2∂L=[0.014675,−0.005854,0.004884]T
Write the NumPy line that produces it.
CodeWrite the gradient of the last layer’s weights. It must match PyTorch’s W2 gradient above.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
Go deeper What loss.backward() actually does
While the forward pass runs, PyTorch writes down every operation and which tensors went into it.
loss.backward() walks that record from the loss back to the weights, and at each step multiplies by the
local gradient of that one operation: the chain rule from level 6, applied by a program.
Level U3, “Build your own autograd”, builds that program in about 50 lines of Python.
After it, you will know exactly what loss.backward() does.
🔒 Answer the question above to unlock
7. A translation table to keep
When you write PyTorch later, most lines are a NumPy line you already know. This table is what you have used on
this page.
Used on this page
NumPy
PyTorch
example
np.array(x, dtype=float)
torch.tensor(x, dtype=torch.float32)
torch.tensor([[0., 1.]]) has shape (1, 2)
A @ B, A.T, x.sum(0), x.mean()
the same
x.sum(0) goes down the rows: one total per column
np.exp, np.log, np.tanh
torch.exp, torch.log, torch.tanh
torch.tanh(torch.tensor(0.)) is 0.
X @ W + b with W of shape (in, out)
nn.Linear(in, out)
stores (out, in), computes x @ W.T + b
a class with forward
a subclass of nn.Module
model(x) runs forward(x)
hand-written backward and update
loss.backward(), opt.step(), opt.zero_grad()
zero_grad before each backward
a number from an array
loss.item()
torch.tensor(0.5).item() is 0.5
Go deeper Reference for later levels: more PyTorch, Python you will read, one epoch
Skip this box now. Come back when a later level sends you here.
For later levels (the level that needs it is in the last column)
NumPy
PyTorch
example
level
softmax, then -np.log(p[correct])
F.cross_entropy(logits, target)
takes raw scores, not probabilities
10, 20
skip some targets in the loss
F.cross_entropy(..., ignore_index=0)
targets equal to 0 add nothing
17, 20
E[ids], a lookup in a table of shape (V, d)
nn.Embedding(V, d)(ids)
ids (2, 5) → vectors (2, 5, d)
13, 20
a weight you make yourself
nn.Parameter(torch.ones(d))
stored as self.g, it appears in parameters()
16, 20
a list of layers
nn.ModuleList([Block(), Block()])
a Python list hides them from parameters()
20
np.triu(np.ones((L, L)), k=1)
torch.triu(torch.ones(L, L), diagonal=1)
ones above the diagonal; PyTorch’s name is diagonal, not k
15, 20
np.where(mask, -np.inf, s)
s.masked_fill(mask, float('-inf'))
mask is True where blocked
15, 20
x.reshape(B, L, h, d_k)
x.view(B, L, h, d_k)
view needs the memory in order
U1, 20
x.transpose(0, 2, 1, 3)
x.transpose(1, 2)
PyTorch swaps exactly two axes
U1, 20
after a transpose, reshape
x.transpose(1, 2).contiguous().view(...) or .reshape(...)
view right after transpose raises an error
U1, 20
np.argmax(p, axis=-1)
logits.argmax(-1)
the index of the largest score
17, 18, 20
—
model.train() / model.eval()
switch training-only behavior on or off
20
—
with torch.no_grad():
no gradient record: faster, for testing
20
Python you will read in the code of later levels:
Python
example
result
list comprehension
[x * 2 for x in [1, 2, 3]]
[2, 4, 6]
enumerate
list(enumerate("ab"))
[(0, 'a'), (1, 'b')]
dict comprehension
{c: i for i, c in enumerate("ab")}
{'a': 0, 'b': 1}
slices
s = [5, 6, 7]: s[:-1], s[1:]
[5, 6], [6, 7]
range(a, b) stops before b
list(range(1, 4))
[1, 2, 3]
zip
list(zip([1, 2], "ab"))
[(1, 'a'), (2, 'b')]
a random order
torch.randperm(3)
for example tensor([2, 0, 1])
add an axis
torch.tensor([1, 2]).unsqueeze(1)
shape (2, 1)
One epoch, the loop you will write again and again (one epoch = one pass over all the training data):
for epoch in range(n_epochs): for xb, yb in batches(train, size=64): # shuffled mini-batches (level U4 writes batches) logits = model(xb) # (64, V): one score per word in the vocabulary loss = F.cross_entropy(logits, yb) opt.zero_grad() loss.backward() opt.step()
8. Your first PyTorch run (optional, on your computer)
Optional. Skip this section if you have not installed PyTorch; nothing on this page depends on it, and the
level clears without it. It is the only exercise on this page that needs PyTorch (Run it on your computer).
Copy these lines into a file first_run.py (or download first_run.py) and run python first_run.py. The seed makes the random starting
weights the same on every computer, so everyone gets the same loss.
Number Run first_run.py (section 8 code) on your computer. What loss does it print after 100 steps?
9. The loop, by yourself
Last, the skill this level is about: the four steps of training, in the right order. Below, Param is the box from
section 4, and the model is a line, w · x + b. backward adds to .grad, like PyTorch. Write the body of the
inner loop: three calls, one per line.
CodeWrite the body of the inner loop: the three calls of one training step, in the right order.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
Recap
a summary for when you finish the level
The key formulas and common mistakes appear here once you clear the level.
You can now
Translate a NumPy network into a PyTorch nn.Module with __init__ and forward.
Write the training step in the right order: zero_grad, backward, step.
Spot the bugs that give no error message: gradients that add up, labels that broadcast, layers hidden in a plain list.
Keep in mind
nn.Linear(in, out).weight has shape (out, in) and computes x @ W.T + b
loss.backward() adds to .grad; opt.zero_grad() clears it
A column (n, 1) minus a flat (n,) gives an (n, n) square: print both shapes before the loss
loss.item() turns a one-number tensor into a Python number
Common mistakes
Forgetting opt.zero_grad(): each step then uses the sum of all the old gradients, and w jumps past the lowest point.
Keeping layers in a plain Python list: parameters() cannot see them, so they never train. Use nn.ModuleList.