Loading
0x50Lesson 6 of 17

Backpropagation

Compute every gradient in a network with one backward sweep of the chain rule.

32 min 7-question quiz 2 code exercises
By the end of this lesson you can
  • Apply the chain rule along a computation graph
  • Compute gradients backward as “upstream gradient × local derivative”, summing over branches
  • Derive the gradients of a linear layer and check them numerically

To improve a weight, training needs its gradient: how much the loss changes when that weight changes a little, ∂L∂w\frac{\partial L}{\partial w}. A network can have billions of weights. Backpropagation gets every one of those gradients in a single backward pass that costs about the same as the forward pass. It’s the algorithm that makes deep learning possible.

The idea is the chain rule. If LL depends on dd, and dd depends on ww, then:

∂L∂w=∂L∂d⋅∂d∂w\frac{\partial L}{\partial w} = \frac{\partial L}{\partial d} \cdot \frac{\partial d}{\partial w}

Break any computation into simple steps - a computation graph - and every node only needs its local derivative. Going backward from the loss, each node’s gradient is the gradient flowing in from above (upstream) times its local derivative. When a value feeds into several places, the gradients from each place add up.

The local derivatives you need are few:

NodeLocal derivativeSo the gradient...
c = a + b∂c/∂a = 1, ∂c/∂b = 1is copied to both inputs
c = a − b1 and −1is copied, negated for b
c = a × b∂c/∂a = b, ∂c/∂b = ais multiplied by the other input
c = a²2ais scaled by 2a
c = relu(a)1 if a > 0 else 0passes through or is blocked
c = σ(a)σ(a)(1 − σ(a))shrinks (at most ×0.25)

Try it

Backprop stepper

Run each graph forward one node at a time, then backward. Before some gradients are revealed, predict them: take the gradient of the node above and multiply by the local derivative.

In the last graph, watch what a ReLU that’s switched off does to everything below it.

Predicted 0/0

A one-weight model predicts p=w⋅x+bp = w \cdot x + b and is scored with L=(p−y)2L = (p - y)^2, for w=2,x=3,b=1w = 2, x = 3, b = 1 and target y=10y = 10.

NodeValueGradient dL/d·How
w (input)?
x (input)?
b (input)?
y (input)?
z = w × x?
p = z + b?
d = p − y?
L = d²?

Backprop through a whole layer

Real layers work on matrices, but the rules are the same. For a linear layer Y=XW+bY = XW + b with input batch XX (N × in), and the upstream gradient G=∂L∂YG = \frac{\partial L}{\partial Y} (N × out):

∂L∂W=X⊤G,∂L∂b=∑rowsG,∂L∂X=GW⊤\frac{\partial L}{\partial W} = X^\top G, \qquad \frac{\partial L}{\partial b} = \sum_{\text{rows}} G, \qquad \frac{\partial L}{\partial X} = G W^\top

A good sanity check: every gradient has the same shape as the thing it’s the gradient of. And dX is what you pass on, upstream, to the layer before.

gradient_check.py
1import numpy as np
2
3rng = np.random.default_rng(1)
4X, W, b = rng.normal(size=(4, 3)), rng.normal(size=(3, 2)), rng.normal(size=2)
5target = rng.normal(size=(4, 2))
6
7def loss(W):
8    return np.mean((X @ W + b - target) ** 2)
9
10# Backprop: L = mean(D²) with D = XW + b - target, so dL/dD = 2D / D.size
11G = 2 * (X @ W + b - target) / target.size
12dW = X.T @ G
13
14# Numerical check: nudge each weight by h and watch the loss
15h = 1e-6
16numeric = np.zeros_like(W)
17for i in range(W.shape[0]):
18    for j in range(W.shape[1]):
19        nudged = W.copy()
20        nudged[i, j] += h
21        numeric[i, j] = (loss(nudged) - loss(W)) / h
22print(dW.shape, np.abs(dW - numeric).max() < 1e-5)
Output
(3, 2) True

Key takeaways

  • Backprop applies the chain rule backward through a computation graph, getting every gradient in one sweep.

  • Each node’s gradient = upstream gradient × local derivative; branches add up.

  • For Y = XW + b: dW = Xᵀ G, db = sum of G’s rows, dX = G Wᵀ.

  • Check gradients numerically when you write anything custom.

Lesson quiz

7 questions · pass with 5 correct · up to 50 XP

Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.

Practice: write Python

Write Python in the editor and run it against sample inputs. Python runs locally in your browser using a WebAssembly runtime.

Exercise 1

Gradients of a sigmoid neuron

+25 XP

A neuron computes p=σ(w⋅x+b)p = \sigma(w \cdot x + b) and is scored with L=(p−y)2L = (p - y)^2. Each input line is w x b y. Backpropagate by hand (chain rule!) and print the loss and the gradients dL/dw and dL/db to 4 decimals:

L=0.2500 dw=-0.5000 db=-0.2500

The starter checks your answer numerically, so you can see whether you got it right.

  • The stepper’s neuron
  • Two more
main.py
Loading editor…

Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.

Exercise 2

Backward pass of a linear layer

+25 XP

The starter builds a batch X (from the input lines), weights W, bias b, and an upstream gradient G. Compute dW, db and dX for Y=XW+bY = XW + b and print their shapes, then each one’s values to 2 decimals (row by row, space-separated). Finally print whether a numerical check of dW agrees (check: True).

  • Two examples
main.py
Loading editor…

Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.

Questions about this lesson

Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.

Loading posts…

Did you like the lesson? 😆👍
Consider a donation to support our work: