Loading
0x30Lesson 4 of 17

Layers and the forward pass

Stack layers into a multilayer perceptron, run a batch through it, and count its parameters.

24 min 7-question quiz 2 code exercises
By the end of this lesson you can
  • Describe a multilayer perceptron as a chain of x @ W + b and activations
  • Track tensor shapes through every layer of a forward pass
  • Count a network’s parameters and estimate its memory

A multilayer perceptron (MLP) - also called a fully connected or dense network - chains layers. For Neuro’s doodles (5×5 = 25 pixels) and 4 kinds of doodle, a small MLP looks like:

input  (batch, 25)
  → Linear 25→16  → ReLU    hidden  (batch, 16)
  → Linear 16→4   → softmax output  (batch, 4)  probabilities

Running inputs through all the layers to get predictions is the forward pass. In code it’s strikingly short:

Input layerHidden layerHidden layerOutput layer
A deep neural network is just neurons like the one above, arranged in connected layers.
mlp_forward.py
1import numpy as np
2
3rng = np.random.default_rng(42)
4W1, b1 = rng.normal(0, 0.3, size=(25, 16)), np.zeros(16)
5W2, b2 = rng.normal(0, 0.3, size=(16, 4)), np.zeros(4)
6
7def softmax(z):
8    exps = np.exp(z - z.max(axis=1, keepdims=True))
9    return exps / exps.sum(axis=1, keepdims=True)
10
11X = rng.integers(0, 2, size=(3, 25)).astype(float)   # 3 random doodles
12hidden = np.maximum(0, X @ W1 + b1)                    # (3, 16)
13probabilities = softmax(hidden @ W2 + b2)              # (3, 4)
14print(X.shape, hidden.shape, probabilities.shape)
15print(np.round(probabilities.sum(axis=1), 6))
16print(probabilities.argmax(axis=1))
Output
(3, 25) (3, 16) (3, 4)
[1. 1. 1.]
[3 2 3]

Notice the softmax works row by row (axis=1, keepdims=True), so every example gets its own distribution. argmax picks the most likely class. With random weights, Neuro’s guesses are, of course, nonsense - training is how they get good.

Counting parameters

A linear layer from ninn_{in} to noutn_{out} features has nin×noutn_{in} \times n_{out} weights plus noutn_{out} biases. Activations have no parameters. For the doodle MLP: 25·16 + 16 + 16·4 + 4 = 484 parameters.

Real networks are bigger: a classic MNIST MLP, 784→128→64→10, has about 109,000 parameters; GPT-3 has 175 billion. Each 32-bit float takes 4 bytes, so parameters × 4 is the memory just to store the weights - training needs several times more for gradients and optimizer state.

Key takeaways

  • An MLP chains linear layers and activations: h = relu(x @ W1 + b1), out = softmax(h @ W2 + b2).

  • The batch dimension rides along untouched; each layer changes only the feature dimension.

  • A linear layer has in × out weights plus out biases; activations have none.

  • Parameters × 4 bytes is the storage for float32 weights; training needs more.

Lesson quiz

7 questions · pass with 5 correct · up to 50 XP

Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.

Practice: write Python

Write Python in the editor and run it against sample inputs. Python runs locally in your browser using a WebAssembly runtime.

Exercise 1

Run Neuro’s forward pass

+25 XP

The starter has a tiny trained network (3 inputs → 4 hidden → 2 outputs, ReLU then softmax). Each input line is one example. Run the whole batch through it in one go, and for each example print the two probabilities to 3 decimals and the predicted class (argmax):

0.905 0.095 -> 0
  • Three examples
main.py
Loading editor…

Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.

Exercise 2

Parameter budget

+25 XP

Each input line describes an MLP as layer sizes, like 784 128 64 10. Print the number of weights, biases, the total, and the memory for float32 parameters in megabytes (1 MB = 1,000,000 bytes) to 2 decimals:

784 128 64 10: weights=109184 biases=202 total=109386 memory=0.44 MB
  • Two networks
  • A wide one
main.py
Loading editor…

Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.

Questions about this lesson

Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.

Loading posts…

Did you like the lesson? 😆👍
Consider a donation to support our work: