Recurrent networks
Process sequences with a hidden state carried through time, and see why LSTMs and GRUs were invented.
- Run a recurrent cell over a sequence, carrying a hidden state
- Explain backpropagation through time and why gradients vanish or explode over long sequences
- Describe how LSTM and GRU gates help, and why transformers replaced RNNs
Text, speech, sensor readings and music are sequences, where order matters and lengths vary. A recurrent neural network (RNN) reads a sequence one step at a time and carries a hidden state - a running summary of everything so far:
The same weights , and are used at every step - weight sharing again, this time across time - so one small network handles sequences of any length. To predict, put a dense layer on the final hidden state (to classify a whole sequence) or on every state (to label each step, or to predict the next one).
1import numpy as np
2
3W_x = np.array([[0.5], [-1.0]]) # input (1) -> hidden (2)
4W_h = np.array([[0.8, 0.0], [0.3, 0.5]]) # hidden -> hidden
5b = np.zeros(2)
6
7h = np.zeros(2)
8for t, x in enumerate([1.0, 0.0, 0.0, 2.0]):
9 h = np.tanh(W_x @ [x] + W_h @ h + b)
10 print(f"t={t} x={x} h={np.round(h, 3)}")t=0 x=1.0 h=[ 0.462 -0.762] t=1 x=0.0 h=[ 0.354 -0.238] t=2 x=0.0 h=[ 0.276 -0.013] t=3 x=2.0 h=[ 0.84 -0.958]
Notice the memory fade: the input at t=0 still echoes through the hidden state at t=2, a little weaker each step.
Backpropagation through time
To train an RNN, unroll it: one copy of the cell per time step, all sharing weights, then backpropagate through the whole chain - backpropagation through time (BPTT). Gradients for the shared weights add up over every step.
Here’s the catch. The gradient that reaches step 1 from step 50 has been multiplied by the recurrent weights and tanh’s slope 49 times. Multiply by 0.5 forty-nine times and you get about - the gradient vanishes, and the network can’t learn long-range dependencies. Multiply by 1.5 and it explodes. (Gradient clipping tames explosions; vanishing needs a new design.)
LSTMs and GRUs
The LSTM (long short-term memory, 1997) adds a separate cell state that runs through time like a conveyor belt, changed only by additions and gated multiplications. Three gates - small sigmoid layers outputting values between 0 and 1 - control it:
- the forget gate decides what to erase from the cell state,
- the input gate decides what new information to write,
- the output gate decides what to reveal as the hidden state.
When the forget gate stays near 1, gradients flow back along the cell state almost unchanged - the same trick as a residual connection. The GRU (2014) is a simpler two-gate variant that often performs as well.
LSTMs ruled translation and speech until 2017. Their weakness: step t can’t start until step t−1 is done, so they don’t parallelize across a GPU. Transformers process every position at once with attention - next lesson.
Key takeaways
RNNs carry a hidden state through a sequence, using the same weights at every step.
Training unrolls them through time (BPTT); gradients get multiplied once per step.
Over long sequences gradients vanish or explode; clipping handles explosions.
LSTMs and GRUs use gates and an additive cell state to remember longer; transformers replaced them by parallelizing with attention.
Lesson quiz
7 questions · pass with 5 correct · up to 50 XP
Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.
Practice: write Python
Write Python in the editor and run it against sample inputs. Python runs locally in your browser using a WebAssembly runtime.
Run an RNN
The starter has a tiny RNN (1 input, 2 hidden units). The input is a sequence of numbers on one line. Run the RNN over it from a zero hidden state and print each step’s hidden state as t=0 h=0.462 -0.762 (3 decimals), then final: positive if the first hidden unit ends above 0, otherwise final: negative.
- Four steps
- Negative inputs
Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.
Gradients through time
In a simple linear RNN, the gradient reaching step 1 from step T is scaled by the recurrent weight w once per step: . Each input line is a weight w. For T = 10, 50 and 100 print the scale in scientific notation with 2 decimals, and a verdict for T = 100: vanishing (below 1e-3), exploding (above 1e3) or stable:
w=0.9: T=10 3.87e-01 T=50 5.73e-03 T=100 2.95e-05 -> vanishing- Three weights
Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.
Questions about this lesson
Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.
Loading posts…