Computer vision: how machines see
Treat images as grids of numbers, slide kernels across them to find edges, and stack convolutions into a CNN.
- Describe a grayscale image as a matrix of brightness values
- Apply a convolution kernel by hand
- Explain how a convolutional network builds from edges to objects
To a computer, a photo is just a grid of numbers. In a grayscale image each pixel holds a brightness from 0 (black) to 255 (white). A color image stores three such grids - red, green and blue. A phone photo has about 12 million pixels, so 36 million numbers.
Everything computer vision does - recognizing faces, reading signs, spotting tumors in scans - starts from those numbers.
Try it
Paint with numbers
Each cell is one pixel, and its number is its brightness. Pick a brush and click cells to repaint the face: give it a nose, or erase its smile. Notice that the Python list below the image changes with every click - that list is the image as far as a computer is concerned.
image = [
[0, 0, 255, 255, 255, 255, 0, 0],
[0, 255, 0, 0, 0, 0, 255, 0],
[255, 0, 192, 0, 0, 192, 0, 255],
[255, 0, 0, 0, 0, 0, 0, 255],
[255, 0, 128, 0, 0, 128, 0, 255],
[255, 0, 0, 128, 128, 0, 0, 255],
[0, 255, 0, 0, 0, 0, 255, 0],
[0, 0, 255, 255, 255, 255, 0, 0],
]Convolution: sliding a small window
How do you find an edge in a grid of numbers? An edge is where brightness changes suddenly, so compare each pixel with its neighbors. A kernel (or filter) is a small grid of weights, often 3 × 3. Convolution slides it over every position of the image; at each one, multiply the kernel’s weights by the pixels underneath and add them up:
It’s a dot product again - between the kernel and one patch of the image. Different kernels find different things: a blur kernel averages neighbors, an edge kernel subtracts one side from the other, so flat regions give 0 and edges give big numbers.
Try it
Try different kernels
Pick a kernel and click any pixel of the output to see the multiply-and-add that produced it. Which kernel lights up only the left and right sides of the square? Which one catches the diagonal line best?
Output at row 6, column 6 = sum of (pixel × weight) over the highlighted 3×3 window (pixels outside the image count as 0):
220×.11 + 220×.11 + 220×.11 + 20×.11 + 20×.11 + 20×.11 + 20×.11 + 20×.11 + 220×.11 = 108.89
1image = [
2 [10, 10, 10, 200, 200],
3 [10, 10, 10, 200, 200],
4 [10, 10, 10, 200, 200],
5]
6kernel = [[-1, 0, 1], [-1, 0, 1], [-1, 0, 1]] # left-to-right change
7
8row = 0 # one row of output: the kernel's top-left corner moves along row 0
9outputs = []
10for column in range(len(image[0]) - 2):
11 total = 0
12 for i in range(3):
13 for j in range(3):
14 total += kernel[i][j] * image[row + i][column + j]
15 outputs.append(total)
16print(outputs)[0, 570, 570]
Convolutional neural networks
In the widget you chose kernels made by people. A convolutional neural network (CNN) learns its kernel weights by gradient descent, discovering whatever filters help with the task. Because the same small kernel slides over the whole image, it needs only 9 weights rather than one per pixel (shared weights), and it finds a pattern wherever it appears.
A CNN stacks layers:
- Convolution layers apply many learned kernels, each producing a feature map.
- Pooling layers shrink each map, for example by keeping the maximum of every 2 × 2 block, so later layers see a wider area.
- Fully connected layers at the end turn the features into a final answer, such as “cat: 94%”.
Key takeaways
An image is a matrix of pixel values (three matrices for color).
Convolution slides a small kernel over the image, taking a weighted sum at each position.
A CNN learns its kernels, pools to shrink feature maps, and builds edges → shapes → objects.
Lesson quiz
6 questions · pass with 5 correct · up to 50 XP
Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.
Practice: write Python
Write Python in the editor and run it against sample inputs. Python runs locally in your browser using a WebAssembly runtime.
Convolve an image
The first 3 lines are a 3 × 3 kernel. The remaining lines are an image (space-separated numbers). Slide the kernel over every position where it fits completely (no padding) and print the output image, one row per line with values separated by spaces. Format numbers with :g.
- Vertical edge
- Box blur
Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.
Max-pool a feature map
Each input line is a row of a feature map with an even number of rows and columns. Apply 2 × 2 max pooling: split it into non-overlapping 2 × 2 blocks and keep the largest value of each. Print the pooled map, one row per line, values separated by spaces.
- Four by four
- Two by six
Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.
Questions about this lesson
Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.
Loading posts…