So far, every weight was chosen by hand. How do we choose them from data instead? We need a measure of how wrong a network is. That measure is the loss. We give each output a target, the answer we want, and the loss compares the two. Then we ask which way each output should move to make the loss smaller. That direction is its gradient.
d=1, h=1, seq=5, L2 loss
Let's start with the d=1, h=1, seq=5 network of the Recurrence article. Its weights and inputs are the same, so every yₜ is too. Under each output, we write its target, y\*ₜ, in cream, because it is given, like an input. The loss at step t is lₜ = ½(yₜ − y\*ₜ)². It is half the squared difference, called the L2 loss. It is zero when the output hits its target, and it grows the further away it is. The loss of the whole sequence, L, is the mean of the five, in purple, under l₁. In practice, a sequence runs hundreds or thousands of steps. The mean keeps L on the same scale however long it gets. The largest lₜ is at step 2, where the output is furthest from its target. In the download, type another target, or change a weight, and watch L move.


