Loading
0x80Lesson 9 of 11

Neural networks from the ground up

Stack linear layers with nonlinear activations, run a forward pass, and see why a hidden layer solves XOR.

24 min 5-question quiz 2 code exercises
By the end of this lesson you can
  • Describe neurons, layers and activation functions
  • Run a forward pass with numpy
  • Explain backpropagation as gradient descent through layers

A neuron computes a weighted sum of its inputs plus a bias, then applies an activation function. A layer is many neurons side by side, so a whole layer is a matrix multiplication:

h=ReLU(Wx+b),ReLU(z)=max⁡(0,z)h = \mathrm{ReLU}(W x + b), \qquad \mathrm{ReLU}(z) = \max(0, z)

A neural network stacks layers: the output of one is the input of the next. The nonlinear activation is essential - without it, any stack of linear layers collapses into a single linear layer and can only draw straight boundaries.

The classic proof is XOR: output 1 when exactly one input is 1. No single straight line separates those points, but a network with one hidden layer of two neurons can.

Try it

Linear or not?

Which problems can a single linear model (one straight boundary) solve, and which need a hidden layer?

0 of 5 sortedScore 0/0
  • “AND: output 1 only when both inputs are 1”

  • “XOR: output 1 when exactly one input is 1”

  • “Points inside a circle versus outside it”

  • “Price grows steadily with floor area”

  • “Recognizing a cat in a photo”

How networks learn: backpropagation

Training is the same gradient descent as before - just with millions of weights. Backpropagation computes every gradient efficiently by applying the chain rule backwards from the loss through each layer. Frameworks such as PyTorch and JAX do this automatically (“autograd”), so in practice you define the forward pass and the loss, and the library handles the rest.

relu.py
import numpy as np
z = np.array([-2.0, -0.5, 0.0, 1.5])
print(np.maximum(0, z))
Output
[0.  0.  0.  1.5]

Key takeaways

  • A layer computes activation(Wx+b)\mathrm{activation}(Wx + b); a network stacks layers.

  • Nonlinear activations let networks learn curved boundaries like XOR.

  • Backpropagation applies the chain rule to get every gradient for gradient descent.

Lesson quiz

5 questions · pass with 4 correct · up to 50 XP

Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.

Practice: write Python

Write Python in the editor and run it against sample inputs. Python runs locally in your browser using a WebAssembly runtime.

Exercise 1

A forward pass with numpy

+25 XP

Line 1 is a JSON object with W1 (2×2), b1, W2 (1×2) and b2. Each following line is an input x1 x2. Compute

h=ReLU(W1x+b1),y=W2h+b2h = \mathrm{ReLU}(W_1 x + b_1), \qquad y = W_2 h + b_2

and print x1 x2 -> y with y formatted :g.

  • A network that computes XOR
main.py
Loading editor…

Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.

Exercise 2

Show that a line can’t do XOR

+25 XP

Try every linear rule predict 1 if w1x1+w2x2+b>0\text{predict } 1 \text{ if } w_1 x_1 + w_2 x_2 + b > 0 with w1,w2,bw_1, w_2, b each in −2, −1.5, …, 2 (steps of 0.5). Count how many of the four XOR points each rule gets right. Print best linear accuracy: N/4 and rules tried: R.

  • XOR
main.py
Loading editor…

Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.

Questions about this lesson

Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.

Loading posts…

Did you like the lesson? 😆👍
Consider a donation to support our work: