Solve these 4 questions on your own. Answer all of them correctly and the level counts as cleared with three stars, and every part of the page opens. Showing an answer doesn’t count.
Number wₓ = 1, wₕ = 0.5, b = 0, h₀ = 0 and x₁ = 1. Use tanh(1) ≈ 0.76. What is h₁ = tanh(wₓ·x₁ + wₕ·h₀ + b)? (2 decimals)
Number wₓ = 1, wₕ = 0.5, b = 0. h₁ = 0.76 and x₂ = 0. Use tanh(0.38) ≈ 0.36. What is h₂ = tanh(wₓ·x₂ + wₕ·h₁ + b)? (2 decimals)
Number This RNN has a one-number hidden state: hₜ = tanh(wₓ·xₜ + wₕ·hₜ₋₁ + b). It reads a sequence of 1,000 inputs. How many learned numbers (weights and biases) does it have?
CodeWrite rnn_all: run the same loop, but keep every hidden state, so the result has one row per step, shape (L, H). Replace ____ with as many lines as you need.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
Warm-up2 questions from earlier levels
A quick review before you start. Optional. Nothing here locks the level.
An MLP takes a fixed number of inputs. A sentence does not have a fixed length, and the order of its words matters.
A recurrent network (RNN) reads a sequence one item at a time. It carries a short list of numbers, the hidden state h, from one step to the next.
The hidden state is how the network remembers what it has read so far.
1. One step at a time
At every step the network mixes the new input with the hidden state, then passes the result through tanh, which keeps it between −1 and 1:
ht=tanh(wxxt+whht−1+b)
Below, the hidden state is a single number. The sequence is 1, 0, 0, 0, and h starts at 0. Press Next step to run it.
A recurrent network, one step at a time
Press Next step. Change the weights or the inputs x and watch h change at every step.
z = w_x·x_t + w_h·h_(t−1) + b, then h_t = tanh(z)
t
x_t
z
h_t
size of h
1
1.00×1 + 0.50×0.0000 + 0.00 = 1.0000
?
2
…
…
3
…
…
4
…
…
step 1 of 4: h1 = ?. The same w_x, w_h and b are used at every step; only h carries anything forward.
h above 0h below 0deeper color: bigger |h|
Number wₓ = 1, wₕ = 0.5, b = 0, h₀ = 0 and x₁ = 1. Use tanh(1) ≈ 0.76. What is h₁ = tanh(wₓ·x₁ + wₕ·h₀ + b)? (2 decimals)
🔒 Answer the question above to unlock
Starting at step 2, the input is 0. The only thing that carries the 1 forward is whht−1.
Number wₓ = 1, wₕ = 0.5, b = 0. h₁ = 0.76 and x₂ = 0. Use tanh(0.38) ≈ 0.36. What is h₂ = tanh(wₓ·x₂ + wₕ·h₁ + b)? (2 decimals)
🔒 Answer the question above to unlock
Each step keeps about half of what was there, so the 1 from step 1 gets smaller and smaller: 0.76, 0.36, 0.18, 0.09.
Number This RNN has a one-number hidden state: hₜ = tanh(wₓ·xₜ + wₕ·hₜ₋₁ + b). It reads a sequence of 1,000 inputs. How many learned numbers (weights and biases) does it have?
I got stuck here Is there a separate network for every step?
No. Drawings often show the network copied once per step, side by side. That drawing is called unrolling,
and it helps to see the flow. But every copy uses the same wx, wh and b. A sequence of 4 steps and a sequence of 4,000 steps
use the same 3 weights in this lab. With vectors, it is the same 3 matrices.
2. Training through time
🔒 Answer the question above to unlock
To train the network, the loss at the last step has to tell step 1 what to change. Backpropagation walks back through the unrolled steps.
At every step it multiplies by the same kind of factor: wh times the slope of tanh there, which is 1−h2.
∂h1∂h4=(wh(1−h22))(wh(1−h32))(wh(1−h42))
Choosewₕ = 0.5. Going back one step multiplies the gradient by wₕ·(1 − h²). Ignoring tanh, three steps back would give 0.5³ = 0.125. With tanh, ∂h₄/∂h₁ is…
The next question leaves tanh out, so every factor is just wh.
ChooseNo activation, wₕ = 1.5. After 30 steps the gradient is a product of 29 factors of 1.5. About how big is it?
How much does step 1 still matter at step t?
Drag w_h and switch the activation. The line is the gradient d h_t / d h_1 over 30 steps.
at t = 30, d h_30 / d h_1 = 1.55e-9: the first input has almost no say, so training does not see that it matters
input 1 at step 1, then 0; w_x = 1, b = 0each step multiplies by w_h × (slope of the activation); tanh’s slope is at most 1vanishing or exploding
Each factor is wh times a tanh slope, and the slope 1−h2 is always between 0 and 1.
If ∣wh∣ is below 1, every factor is below 1, and a long product shrinks toward 0.
This is the vanishing gradient: the first steps get almost no training signal.
If ∣wh∣ is well above 1, a factor can stay above 1 and the product grows very large: an exploding gradient.
tanh does not prevent that. It only helps where h is near ±1, because there the slope is close to 0.
Either way, a product of many factors almost never stays near 1. The explode question above left tanh out to make the
arithmetic easy, but the gradient grows the same way with tanh whenever wh(1−h2) stays above 1.
Go deeper An exploding gradient is easy to fix, a vanishing one is not
For an exploding gradient there is a simple fix: if its length is above some limit, scale it down to the limit.
That is gradient clipping, and demo.py uses it (clip_grad_norm_(…, 1.0)).
A vanishing gradient has no such fix. Making a number like 1e-9 bigger also makes the noise from every other step.
The information about step 1 is simply lost on the way back. The fix has to change the network itself. That is level N4.
Here is the same chain drawn as a network: the sequence 1, 0, 0, 0 with wx=1, wh=0.5, b=0, as in the lab of section 1.
One cell, used four times
Press Next for the forward steps, then keep going to send the gradient back.
Press Next to start. · 3 parameters in all
Layers (8)
Look at the color: every step is the same cell with the same three weights. Going back, the bright area shrinks at every step.
positive valuenegative valueA disc’s fill is its value: the stronger the color, the further from 0.the same weights, used againforward (this step)backward: gradient“∂ −0.5” = the gradient ∂L/∂(that value).Select a layer to see its numbers.
3. Remember the first symbol
🔒 Answer the question above to unlock
Here is a task where the network must remember. A sequence starts with A or B, then k − 1 random extra symbols (c, d or e).
At the end the network must say which symbol came first. Guessing gives 50%.
ChooseAn RNN with a 32-number hidden state learns “remember the first symbol” perfectly for k = 20. Now k = 40: same rule, 1,500 training steps. What accuracy will it reach?
Remember the first symbol: real training runs
Pick a sequence length k. Try the longer ones.
Loading runs…
…
RNN1,500 steps of 64 sequences, a 32-number hidden state, accuracy on 1,000 new sequences
Up to k = 20, the RNN learns the task perfectly. At k = 40 it stays at guessing, even though the rule is just as simple.
I got stuck here Why not train longer, or with a bigger learning rate?
The training signal that should reach step 1 is a product of 39 factors, each well below 1. It is far smaller than the
noise coming from the random symbols near the end. More steps or a bigger learning rate amplify that noise just as much.
The network never “hears” that step 1 mattered.
Try it
In the step lab, set wh to 1 and run all 4 steps. The 1 now shrinks much more slowly. Then set wh to −1. What happens to the sign of h at each step?
4. Write it yourself
With vectors, x is a list of D numbers and h is a list of H numbers. As in every level since level 1, vectors are rows
and a layer is “row @ matrix”:
ht=tanh(xtWx+ht−1Wh+b)
So Wx is (D, H): it turns D input numbers into H. Wh is (H, H), and b has H numbers. In code that is
x @ Wx + h @ Wh + b. With a one-number hidden state, xtWx is just wxxt, the formula from section 1.
Write the one line inside the loop.
CodeWrite the update inside the loop: h becomes tanh of (x times Wₓ, plus h times Wₕ, plus b).
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
🔒 Answer the question above to unlock
rnn keeps only the last h. Often you need the hidden state after every step, for example to read the whole sequence again
later (level N5 does that). Collect them in a list: start with hs = [], and after each update call hs.append(h).
At the end, np.stack(hs) turns a list of L rows of H numbers into one array of shape (L, H).
CodeWrite rnn_all: run the same loop, but keep every hidden state, so the result has one row per step, shape (L, H). Replace ____ with as many lines as you need.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
Recap
a summary for when you finish the level
The key formulas and common mistakes appear here once you clear the level.
You can now
Run an RNN by hand: mix the input and the old hidden state, then take tanh.
Explain why the gradient vanishes or explodes over many steps: it is a product of one factor per step.
Write an RNN loop in NumPy that keeps every hidden state, shape (L, H).
Keep in mind
ht=tanh(xtWx+ht−1Wh+b), with Wx (D, H), Wh (H, H), b (H,)
Every step uses the same Wx, Wh and b: the number of weights does not depend on the length
One step back multiplies the gradient by wh(1−h2); many such factors shrink to 0 or grow very large
Common mistakes
Counting new weights for every step: the steps share one set of weights.
Adding the factors instead of multiplying them: 29 factors of 1.5 give about 128,000, not 45.