3. How wrong is a yes/no answer?
The lab shows a loss. Level 2 measured a mistake with the squared error. For a yes/no answer there is a better
measure, the cross-entropy loss. It uses ln, the natural logarithm (NumPy’s np.log). You only need a few
values of −lnp (the math page explains ln):
| p | 1 | 0.75 | 0.5 | 0.25 | 0.1 |
|---|
| −lnp | 0 | 0.29 | 0.69 | 1.39 | 2.30 |
The smaller p is, the bigger the cost, and p=1 costs nothing. Take one point at a time:
- target 1: the cost is −lnp. If the neuron is sure and right (p=1) the cost is 0.
The closer p gets to 0, the bigger the cost.
- target 0: the cost is −ln(1−p). This is the same idea with the neuron’s confidence in 0.
One formula covers both, because y is 1 or 0, which makes one of the two terms 0:
loss=−mean(ylnp+(1−y)ln(1−p))This loss is also called binary cross-entropy (“binary” because there are two answers).
A neuron that answers p=0.5 for everything costs −ln0.5≈0.69 per point. This is the cost of a
coin flip. On XOR one neuron can do no better than that, so its loss stops at 0.6931.
Level 10 extends the loss to more than two answers and explains where it comes from.
How much better is a model that is only partly sure, but right on every point?