Matrix multiplication lets you group the two weight matrices first: .
So the two layers are really one layer with weights . However many linear layers you stack, the result is still a single line. Extra layers add nothing unless something nonlinear sits between them.
Go deeper What the nonlinearity actually does
Put a nonlinear function between the layers, for example . tanh is sigmoid stretched to the range −1 to 1: . Now can’t be rewritten as one matrix, because tanh does not distribute over addition: .
Each hidden neuron still draws one line, then tanh bends its output into “this side” (about +1) or “that side” (about −1). The output neuron combines those answers. Two lines plus a combination can make the XOR pattern: “between the two lines” is 1, “outside” is 0. That is the pattern you’ll see the network find in the lab below.
Any nonlinear function works in principle. The last part of this page compares the four you will meet most.
Here is a network with 6 hidden neurons, learning XOR for real in your browser. Its hidden neurons use tanh (sigmoid stretched to −1…1, see the box above). Press Train.
A 2 → 6 → 1 network learns XOR
Press Train and watch the background bend around the four points. Then turn the activation off and try again.
| input | target | output |
|---|---|---|
| [0, 0] | 0 | 0.000 |
| [0, 1] | 1 | 0.000 |
| [1, 0] | 1 | 0.000 |
| [1, 1] | 0 | 0.000 |