Loading
0xC0Lesson 13 of 17

Convolutional networks

Slide small learned filters over images, pool the results, and build the CNNs that taught computers to see.

32 min 7-question quiz 2 code exercises
By the end of this lesson you can
  • Explain why convolutions suit images: local patterns, shared weights
  • Compute convolution outputs and output sizes with padding and stride
  • Describe how conv, activation and pooling layers stack into a CNN, and count their parameters

Feed a 224×224 color photo to a fully connected layer with 1,000 neurons and you need 150 million weights - for one layer. Worse, a cat in the top-left corner and a cat in the bottom-right would be learned separately.

A convolutional layer fixes both problems with a small kernel (filter), say 3×3, that slides across the image. At each position it computes a weighted sum of the pixels under it. The same kernel - the same 9 weights - is used everywhere (weight sharing), and each output only looks at a small neighborhood (locality). The output is a feature map showing where the kernel’s pattern appears.

The kernels aren’t designed by hand: the network learns them by backprop, just like any other weights. Early layers end up with edge and color detectors much like the classic ones below.

Try it

Slide a kernel

Choose a kernel and click any output pixel to see the multiplication behind it. Edge kernels light up where brightness changes; notice that a vertical-edge detector ignores horizontal edges completely.

Input (click a pixel in the output)
Kernel
Edit any weight to make your own filter.
Output (orange = negative)

Output at row 6, column 6 = sum of (pixel × weight) over the highlighted 3×3 window (pixels outside the image count as 0):

220×-1 + 220×0 + 220×1 + 20×-2 + 20×0 + 20×2 + 20×-1 + 20×0 + 220×1 = 200

conv2d.py
1import numpy as np
2
3def conv2d(image, kernel, stride=1, padding=0):
4    image = np.pad(image, padding)
5    k = kernel.shape[0]
6    out_size = (image.shape[0] - k) // stride + 1
7    out = np.zeros((out_size, out_size))
8    for i in range(out_size):
9        for j in range(out_size):
10            patch = image[i * stride:i * stride + k, j * stride:j * stride + k]
11            out[i, j] = np.sum(patch * kernel)
12    return out
13
14doodle = np.array([[0, 0, 1, 1, 0],
15                   [0, 0, 1, 1, 0],
16                   [0, 0, 1, 1, 0],
17                   [0, 0, 1, 1, 0],
18                   [0, 0, 1, 1, 0]], dtype=float)
19vertical_edges = np.array([[-1, 0, 1], [-2, 0, 2], [-1, 0, 1]])
20print(conv2d(doodle, vertical_edges))
21print(conv2d(doodle, vertical_edges, padding=1).shape, conv2d(doodle, vertical_edges, stride=2).shape)
Output
[[ 4.  4. -4.]
 [ 4.  4. -4.]
 [ 4.  4. -4.]]
(5, 5) (2, 2)

Positive values mark where the doodle turns from dark to bright (left to right), negative where it turns back. (Strictly, deep learning “convolution” doesn’t flip the kernel - mathematicians call it cross-correlation - but since kernels are learned, it doesn’t matter.)

Output size. For input width WW, kernel KK, padding PP and stride SS:

Wout=⌊W−K+2PS⌋+1W_{out} = \left\lfloor \frac{W - K + 2P}{S} \right\rfloor + 1

  • Padding adds a border of zeros; with P=(K−1)/2P = (K-1)/2 (“same” padding) the size is preserved.
  • Stride skips positions: stride 2 roughly halves each dimension.
  • Channels: real images have 3 channels (RGB), and layers produce many feature maps. A kernel spans all input channels, so a layer with CinC_{in} inputs, CoutC_{out} outputs and K×K kernels has K⋅K⋅Cin⋅Cout+CoutK \cdot K \cdot C_{in} \cdot C_{out} + C_{out} parameters - tiny compared with a dense layer.

Pooling and the CNN recipe

Pooling shrinks feature maps by summarizing small windows - usually taking the maximum of each 2×2 block (max pooling). It makes the representation smaller and a little tolerant to small shifts: the feature was somewhere in that window.

Try it

Max pooling

Step the 2×2 window across the feature map with stride 2. Each output keeps only the strongest response in its window.

Feature map: 6 × 6
Pooled (2×2, stride 2): 3 × 3 - click a cell

max(1, 3, 4, 2) = 4

A classic CNN repeats conv → ReLU → pool a few times, so feature maps get smaller but more numerous (more channels), and each neuron sees a larger region of the original image - its receptive field grows layer by layer. Then it flattens the result and finishes with dense layers to classify. LeNet (1998) read handwritten digits this way; AlexNet, VGG and ResNet scaled the same recipe up.

Input image32×32×3Convolutionedge / texture mapsPoolingshrink, keep the maxConvolutionshape / part mapsFlatten + Densecombine everythingOutput"cat" 92%
A convolutional neural network passes an image through filters that gradually turn pixels into a decision.

Key takeaways

  • Convolutions slide small, shared, learned kernels over the input; each output sees only a local patch.

  • Output width = ⌊(W − K + 2P)/S⌋ + 1; “same” padding preserves size, stride shrinks it.

  • A conv layer has K·K·C_in·C_out + C_out parameters - far fewer than a dense layer.

  • CNNs stack conv → ReLU → pool, growing channels and receptive fields, then classify.

Lesson quiz

7 questions · pass with 5 correct · up to 50 XP

Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.

Practice: write Python

Write Python in the editor and run it against sample inputs. Python runs locally in your browser using a WebAssembly runtime.

Exercise 1

Write a convolution

+25 XP

The first input line is stride padding; the remaining lines are a square image. Implement conv2d(image, kernel, stride, padding) (no kernel flip, zero padding) and apply the horizontal-edge kernel [[-1, -2, -1], [0, 0, 0], [1, 2, 1]]. Print the output shape, then its rows with values as integers separated by spaces.

  • A horizontal bar
  • Padded and strided
main.py
Loading editor…

Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.

Exercise 2

Trace a CNN’s shapes

+25 XP

The first input line is the input shape channels height width. Each following line is a layer: conv out_channels kernel stride padding, pool size (max pooling with stride = size), or flatten. Print the shape after each layer and that layer’s parameter count, then the total:

1conv 16 3 1 1 -> (16, 32, 32) params 448
2pool 2 -> (16, 16, 16) params 0
3flatten -> (4096,) params 0
4total params 448
  • A small CNN
  • LeNet-style
main.py
Loading editor…

Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.

Questions about this lesson

Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.

Loading posts…

Did you like the lesson? 😆👍
Consider a donation to support our work: