In the last three articles, our network had one hidden layer. How do we make it deeper? We can stack layers, each reading the hidden state of the layer below. Or we can read the sequence in both directions at once. Each example builds on one we have already seen, so we keep an eye on what stays the same.
d=1, h=1, h′=1, seq=5
We stack a second layer on top of the first. We mark it with a prime. Its hidden state is h′, and its weights are W′x, W′h and b′. At each step, layer 2 reads the hidden state hₜ from layer 1. It also reads its own state from the step before, h′ₜ₋₁. The output now comes from layer 2. Why does the h row match the last example of the first article? Because layer 1 is exactly that network.


