Loading
0x20Lesson 3 of 17

Neurons and activation functions

Build an artificial neuron, explore activation functions, and see why nonlinearity is essential.

26 min 7-question quiz 2 code exercises
By the end of this lesson you can
  • Compute a neuron’s output as an activation of a weighted sum plus bias
  • Compare sigmoid, tanh, ReLU, leaky ReLU and GELU, including their slopes
  • Explain why stacked layers need nonlinear activations, and compute a stable softmax

An artificial neuron does two things:

  1. a weighted sum of its inputs plus a bias: z=w1x1+w2x2+⋯+wnxn+b=w⋅x+bz = w_1 x_1 + w_2 x_2 + \dots + w_n x_n + b = \mathbf{w} \cdot \mathbf{x} + b
  2. an activation function applied to that sum: a=f(z)a = f(z)

The weights say how much each input matters (and in which direction); the bias shifts the threshold. A layer is many neurons sharing the same inputs, so its weights form a matrix and the whole layer is one x @ W + b.

x₁w₁x₂w₂x₃w₃Σ + bactivationy
A single artificial neuron: multiply each input by a weight, add a bias, then apply an activation function.

Activation functions

ActivationFormulaRangeNotes
Sigmoidσ(z)=11+e−z\sigma(z) = \frac{1}{1 + e^{-z}}(0, 1)squashes to a probability; flat tails
Tanhtanh⁡(z)\tanh(z)(−1, 1)zero-centered sigmoid; still flat tails
ReLUmax⁡(0,z)\max(0, z)[0, ∞)the default for hidden layers: cheap, gradient 1 when active
Leaky ReLUmax⁡(0.1z,z)\max(0.1z, z)(−∞, ∞)keeps a small gradient for negative inputs
GELUz⋅Φ(z)z \cdot \Phi(z)about [−0.17, ∞)a smooth ReLU used in transformers

What matters most for learning is the slope (derivative): during training, gradients are multiplied by it at every layer. Where the slope is near zero, learning stalls.

Try it

Activation explorer

Pick an activation and slide the input. The dashed curve is its slope.

  • Slide sigmoid to x = 5: the output is nearly 1, but the slope is almost 0. Stack ten such layers and the gradient all but vanishes.
  • Compare ReLU: its slope is exactly 1 for any positive input - which is why it made deep networks trainable. But for negative inputs its slope is 0 (a “dead” neuron); leaky ReLU fixes that.

σ(x) = 1 / (1 + e^-x)

Solid: the function. Dashed: its derivative (the slope backprop multiplies by).

output
0.731
slope
0.197
activations.py
1import numpy as np
2
3z = np.array([-3.0, -0.5, 0.0, 0.5, 3.0])
4sigmoid = 1 / (1 + np.exp(-z))
5relu = np.maximum(0, z)
6print("sigmoid:", np.round(sigmoid, 3))
7print("slope:  ", np.round(sigmoid * (1 - sigmoid), 3))
8print("tanh:   ", np.round(np.tanh(z), 3))
9print("relu:   ", relu)
10print("leaky:  ", np.where(z > 0, z, 0.1 * z))
Output
sigmoid: [0.047 0.378 0.5   0.622 0.953]
slope:   [0.045 0.235 0.25  0.235 0.045]
tanh:    [-0.995 -0.462  0.     0.462  0.995]
relu:    [0.  0.  0.  0.5 3. ]
leaky:   [-0.3  -0.05  0.    0.5   3.  ]

Why nonlinearity is essential

Without activations, a stack of layers is pointless: two linear layers, (x @ W1) @ W2, equal one linear layer with weights W1 @ W2. A hundred linear layers still only draw straight lines. The nonlinearity between layers is what lets depth add power:

collapse.py
1import numpy as np
2
3rng = np.random.default_rng(0)
4x = rng.normal(size=(4, 3))
5W1, W2 = rng.normal(size=(3, 5)), rng.normal(size=(5, 2))
6two_layers = (x @ W1) @ W2
7one_layer = x @ (W1 @ W2)
8print(np.allclose(two_layers, one_layer))
9with_relu = np.maximum(0, x @ W1) @ W2
10print(np.allclose(with_relu, one_layer))
Output
True
False

Softmax: scores into probabilities

For classification, the last layer outputs one raw score per class - the logits. Softmax turns them into probabilities that are positive and sum to 1:

softmax(z)i=ezi∑jezj\text{softmax}(z)_i = \frac{e^{z_i}}{\sum_j e^{z_j}}

Computed naively, np.exp(1000) overflows to infinity. Since softmax doesn’t change when you subtract the same number from every logit, subtract the maximum first - the standard stable softmax.

Key takeaways

  • A neuron computes f(w·x + b); a layer does it for many neurons at once with x @ W + b.

  • ReLU is the default hidden activation; sigmoid and tanh have flat tails that shrink gradients.

  • Without nonlinear activations, any stack of layers collapses into a single linear layer.

  • Softmax turns logits into probabilities; subtract the max first for numerical stability.

Lesson quiz

7 questions · pass with 5 correct · up to 50 XP

Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.

Practice: write Python

Write Python in the editor and run it against sample inputs. Python runs locally in your browser using a WebAssembly runtime.

Exercise 1

Activation table

+25 XP

Each input line is a number z. Print a row with z and its sigmoid, tanh, ReLU and leaky ReLU (slope 0.1) values, all to 3 decimals, separated by spaces. Write each activation as a function that works on numpy arrays.

  • Five inputs
main.py
Loading editor…

Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.

Exercise 2

A softmax that never overflows

+25 XP

Each input line holds the logits for one example. Print the softmax probabilities to 3 decimals, then sum=1.000. Logits can be huge (like 1000) - your softmax must subtract each row’s maximum so nothing overflows.

  • Huge logits
  • Ordinary logits
main.py
Loading editor…

Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.

Questions about this lesson

Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.

Loading posts…

Did you like the lesson? 😆👍
Consider a donation to support our work: