Loading
0xD0Lesson 14 of 17

Recurrent networks

Process sequences with a hidden state carried through time, and see why LSTMs and GRUs were invented.

28 min 7-question quiz 2 code exercises
By the end of this lesson you can
  • Run a recurrent cell over a sequence, carrying a hidden state
  • Explain backpropagation through time and why gradients vanish or explode over long sequences
  • Describe how LSTM and GRU gates help, and why transformers replaced RNNs

Text, speech, sensor readings and music are sequences, where order matters and lengths vary. A recurrent neural network (RNN) reads a sequence one step at a time and carries a hidden state hh - a running summary of everything so far:

ht=tanh⁡(Wxxt+Whht−1+b)h_t = \tanh(W_x x_t + W_h h_{t-1} + b)

The same weights WxW_x, WhW_h and bb are used at every step - weight sharing again, this time across time - so one small network handles sequences of any length. To predict, put a dense layer on the final hidden state (to classify a whole sequence) or on every state (to label each step, or to predict the next one).

rnn.py
1import numpy as np
2
3W_x = np.array([[0.5], [-1.0]])          # input (1) -> hidden (2)
4W_h = np.array([[0.8, 0.0], [0.3, 0.5]])  # hidden -> hidden
5b = np.zeros(2)
6
7h = np.zeros(2)
8for t, x in enumerate([1.0, 0.0, 0.0, 2.0]):
9    h = np.tanh(W_x @ [x] + W_h @ h + b)
10    print(f"t={t} x={x} h={np.round(h, 3)}")
Output
t=0 x=1.0 h=[ 0.462 -0.762]
t=1 x=0.0 h=[ 0.354 -0.238]
t=2 x=0.0 h=[ 0.276 -0.013]
t=3 x=2.0 h=[ 0.84  -0.958]

Notice the memory fade: the input at t=0 still echoes through the hidden state at t=2, a little weaker each step.

Backpropagation through time

To train an RNN, unroll it: one copy of the cell per time step, all sharing weights, then backpropagate through the whole chain - backpropagation through time (BPTT). Gradients for the shared weights add up over every step.

Here’s the catch. The gradient that reaches step 1 from step 50 has been multiplied by the recurrent weights and tanh’s slope 49 times. Multiply by 0.5 forty-nine times and you get about 10−1510^{-15} - the gradient vanishes, and the network can’t learn long-range dependencies. Multiply by 1.5 and it explodes. (Gradient clipping tames explosions; vanishing needs a new design.)

LSTMs and GRUs

The LSTM (long short-term memory, 1997) adds a separate cell state that runs through time like a conveyor belt, changed only by additions and gated multiplications. Three gates - small sigmoid layers outputting values between 0 and 1 - control it:

  • the forget gate decides what to erase from the cell state,
  • the input gate decides what new information to write,
  • the output gate decides what to reveal as the hidden state.

When the forget gate stays near 1, gradients flow back along the cell state almost unchanged - the same trick as a residual connection. The GRU (2014) is a simpler two-gate variant that often performs as well.

LSTMs ruled translation and speech until 2017. Their weakness: step t can’t start until step t−1 is done, so they don’t parallelize across a GPU. Transformers process every position at once with attention - next lesson.

Key takeaways

  • RNNs carry a hidden state through a sequence, using the same weights at every step.

  • Training unrolls them through time (BPTT); gradients get multiplied once per step.

  • Over long sequences gradients vanish or explode; clipping handles explosions.

  • LSTMs and GRUs use gates and an additive cell state to remember longer; transformers replaced them by parallelizing with attention.

Lesson quiz

7 questions · pass with 5 correct · up to 50 XP

Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.

Practice: write Python

Write Python in the editor and run it against sample inputs. Python runs locally in your browser using a WebAssembly runtime.

Exercise 1

Run an RNN

+25 XP

The starter has a tiny RNN (1 input, 2 hidden units). The input is a sequence of numbers on one line. Run the RNN over it from a zero hidden state and print each step’s hidden state as t=0 h=0.462 -0.762 (3 decimals), then final: positive if the first hidden unit ends above 0, otherwise final: negative.

  • Four steps
  • Negative inputs
main.py
Loading editor…

Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.

Exercise 2

Gradients through time

+25 XP

In a simple linear RNN, the gradient reaching step 1 from step T is scaled by the recurrent weight w once per step: wT−1w^{T-1}. Each input line is a weight w. For T = 10, 50 and 100 print the scale in scientific notation with 2 decimals, and a verdict for T = 100: vanishing (below 1e-3), exploding (above 1e3) or stable:

w=0.9: T=10 3.87e-01  T=50 5.73e-03  T=100 2.95e-05  -> vanishing
  • Three weights
main.py
Loading editor…

Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.

Questions about this lesson

Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.

Loading posts…

Did you like the lesson? 😆👍
Consider a donation to support our work: