Loading
0x30Lesson 4 of 11

Classify with logistic regression

Turn a linear score into a probability with the sigmoid, train with log loss, and pick a decision threshold.

24 min 5-question quiz 2 code exercises
By the end of this lesson you can
  • Explain the sigmoid and why it outputs probabilities
  • Compute log loss and understand why it punishes confident mistakes
  • Turn probabilities into decisions with a threshold

For classification we want a probability, not an unbounded number. Logistic regression computes a linear score z=w⋅x+bz = w \cdot x + b and squashes it with the sigmoid:

σ(z)=11+e−z\sigma(z) = \frac{1}{1 + e^{-z}}

which maps any score into (0,1)(0, 1): large positive scores approach 1, large negative scores approach 0, and z=0z = 0 gives exactly 0.5. The result p=σ(z)p = \sigma(z) is read as the probability of the positive class.

Training minimizes log loss (binary cross-entropy):

L=−1n∑i=1n[yiln⁡pi+(1−yi)ln⁡(1−pi)]L = -\frac{1}{n} \sum_{i=1}^{n} \left[ y_i \ln p_i + (1 - y_i) \ln (1 - p_i) \right]

A confident wrong prediction - p=0.99p = 0.99 for an example that is really negative - costs −ln⁡0.01≈4.6-\ln 0.01 \approx 4.6, while a hesitant one costs much less. That pushes the model to be confident only when it’s right.

From probability to decision

A probability isn’t a decision. You choose a threshold: predict positive when p≥tp \ge t. The default 0.5 is rarely special - for a disease screening you may flag patients at p≥0.2p \ge 0.2 because missing a case is worse than an extra test; for blocking a payment you may require p≥0.9p \ge 0.9.

Try it

Move the threshold

These are a fraud model’s probabilities for 12 transactions. Slide the threshold and watch precision and recall trade off.

60%
Precision
Of the messages flagged, how many really are fraud
60%
Recall
Of the real fraud messages, how many were flagged
60%
F1
One number that balances the two
Confusion matrix
Flagged fraudCalled legitimate
Really fraud3 true positives2 missed (false negatives)
Really legitimate2 false alarms (false positives)5 true negatives
  • 0.95Card used in two countries within an hour (fraud)flagged fraud
  • 0.88Large electronics purchase at 3 a.m. (fraud)flagged fraud
  • 0.72New device, new shipping address (fraud)flagged fraud
  • 0.64Frequent small gift card purchases (legitimate)flagged fraud
  • 0.55First purchase at a new store (legitimate)flagged fraud
  • 0.48Many failed PIN attempts (fraud)legitimate
  • 0.41Travel booking abroad (legitimate)legitimate
  • 0.33Small purchase right after a password reset (fraud)legitimate
  • 0.18Online order to home address (legitimate)legitimate
  • 0.08Fuel at the usual station (legitimate)legitimate
  • 0.05Usual grocery store, usual amount (legitimate)legitimate
  • 0.02Monthly streaming subscription (legitimate)legitimate
sigmoid.py
import numpy as np
z = np.array([-4.0, -1.0, 0.0, 1.0, 4.0])
print(np.round(1 / (1 + np.exp(-z)), 3))
Output
[0.018 0.269 0.5   0.731 0.982]

Key takeaways

  • Logistic regression: p=σ(w⋅x+b)p = \sigma(w \cdot x + b), a probability between 0 and 1.

  • Log loss punishes confident mistakes heavily.

  • Choose the decision threshold from the costs of each kind of error.

Lesson quiz

5 questions · pass with 4 correct · up to 50 XP

Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.

Practice: write Python

Write Python in the editor and run it against sample inputs. Python runs locally in your browser using a WebAssembly runtime.

Exercise 1

Compute log loss

+25 XP

Line 1 holds true labels (0 or 1); line 2 the predicted probabilities. Clip probabilities to [10−15,1−10−15][10^{-15}, 1 - 10^{-15}] and print the log loss with 4 decimals.

  • Mostly right
  • One confident mistake
main.py
Loading editor…

Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.

Exercise 2

Train logistic regression

+25 XP

Line 1 is LEARNING_RATE STEPS. Each following line is x1 x2 label. Start with weights and bias at 0 and run gradient descent on log loss, with p=σ(Xw+b)p = \sigma(Xw + b):

∇wL=1nX⊤(p−y),∂L∂b=1n∑(pi−yi)\nabla_w L = \frac{1}{n} X^\top (p - y), \qquad \frac{\partial L}{\partial b} = \frac{1}{n} \sum (p_i - y_i)

Print w=[W1, W2] b=B (3 decimals), log loss: L (4 decimals) and training accuracy: A% using a 0.5 threshold.

  • Pass or fail from hours studied and slept
main.py
Loading editor…

Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.

Questions about this lesson

Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.

Loading posts…

Did you like the lesson? 😆👍
Consider a donation to support our work: