Loading
0xA0Lesson 11 of 18

Reasoning with uncertainty: Bayes’ rule

Discover why a positive test can still mean you’re probably fine, update beliefs as evidence arrives, and build a tiny spam filter.

22 min 7-question quiz 2 code exercises
By the end of this lesson you can
  • Compute a posterior probability with natural frequencies
  • Explain why base rates matter so much
  • Describe how a naive Bayes spam filter combines evidence

The real world is uncertain: sensors are noisy, tests make mistakes, and emails don’t come labelled “spam”. AI systems need a principled way to answer “given what I’ve just seen, how sure should I be?” That way is Bayes’ rule, worked out by Reverend Thomas Bayes in the 1760s.

Try this puzzle before reading on. A disease affects 1% of people. A test catches 90% of real cases, but also wrongly flags 5% of healthy people. You test positive. What’s the chance you’re actually sick?

Try it

The surprising positive test

First make your gut-feeling guess, then explore. Each dot is one of 1,000 people. Drag the sliders: what happens to the chance a positive is real when the disease gets rarer? When the false-alarm rate drops to almost zero? Switch scenarios to see the same maths in spam filters and airport scanners.

Gut check first: someone gets a positive test result. How likely is it they really have the disease?

  • 9 sick, test positive
  • 1 sick, test negative (missed)
  • 50 healthy, test positive (false alarm)
  • 940 healthy, test negative
Chance a positive is real
15.4%
9 real out of 59 positives

Out of 1,000 people, about 10 have the disease. The test result flags 9 of them - and also 50 healthy people. False alarms outnumber real cases, so a positive is more likely wrong than right!

Most people guess around 90%. The answer is about 15%. The trick is to stop thinking in percentages and count people instead - doctors call these natural frequencies:

1% sick99% healthy90% caught10% missed5% false alarm95% cleared1,000people tested10sick990healthy9positive ✔1negative50positive ✘940negativeP(sick | positive) = 9 ÷ (9 + 50) ≈ 15%
Natural frequencies: picture 1,000 people instead of juggling percentages. Only 9 of the 59 positive tests are real - about 15%.
natural_frequencies.py
1people = 1000
2sick = people * 0.01                      # 1% have the disease
3healthy = people - sick
4true_positives = sick * 0.90              # the test catches 90% of real cases
5false_positives = healthy * 0.05          # and wrongly flags 5% of healthy people
6
7positives = true_positives + false_positives
8print(f"{positives:.1f} people test positive")
9print(f"{true_positives:.1f} of them are really sick")
10print(f"P(sick | positive) = {true_positives / positives:.1%}")
Output
58.5 people test positive
9.0 of them are really sick
P(sick | positive) = 15.4%

The disease is so rare that the 5% false alarms among 990 healthy people (about 50) swamp the 9 real cases. Forgetting how rare something is to begin with is called base rate neglect, and it trips up doctors, juries and AI designers alike.

Written as a formula, Bayes’ rule says:

P(sick∣+)=P(+∣sick)⋅P(sick)P(+)P(\text{sick} \mid +) = \frac{P(+ \mid \text{sick}) \cdot P(\text{sick})}{P(+)}

  • P(sick)P(\text{sick}) is the prior - what you believed before the evidence (1%).
  • P(+∣sick)P(+ \mid \text{sick}) is the likelihood - how expected the evidence is if it’s true (90%).
  • P(sick∣+)P(\text{sick} \mid +) is the posterior - your updated belief (15%).

Updating beliefs, one clue at a time

Today’s posterior becomes tomorrow’s prior. If the first test is positive and you take a second, independent test, you start from 15.4% instead of 1%:

update_beliefs.py
1def update(prior, sensitivity, false_alarm, positive):
2    if positive:
3        hit, miss = sensitivity, false_alarm
4    else:
5        hit, miss = 1 - sensitivity, 1 - false_alarm
6    return prior * hit / (prior * hit + (1 - prior) * miss)
7
8belief = 0.01
9print(f"before any test: {belief:.1%}")
10for result in [True, True, False]:
11    belief = update(belief, 0.90, 0.05, result)
12    print(f"after a {'positive' if result else 'negative'} test: {belief:.1%}")
Output
before any test: 1.0%
after a positive test: 15.4%
after a positive test: 76.6%
after a negative test: 25.6%

This is how a robot vacuum works out where it is, how a self-driving car fuses camera and radar, and how a search-and-rescue team narrows down where a lost boat might be: start with a belief, and nudge it with every new clue.

A tiny spam filter

Early spam filters (and many still today) use naive Bayes. They learn how often each word appears in spam and in normal email, then multiply the evidence from every word together. It’s “naive” because it pretends words are independent - “free” and “winner” obviously aren’t - yet it works remarkably well.

naive_bayes_spam.py
1# How often each word shows up in spam and in normal email ("ham").
2p_word_given_spam = {"free": 0.30, "winner": 0.20, "meeting": 0.01}
3p_word_given_ham = {"free": 0.02, "winner": 0.001, "meeting": 0.10}
4p_spam = 0.4
5
6def spam_probability(words):
7    spam, ham = p_spam, 1 - p_spam
8    for word in words:                     # "naive": treat words as independent
9        spam *= p_word_given_spam[word]
10        ham *= p_word_given_ham[word]
11    return spam / (spam + ham)
12
13for email in [["meeting"], ["free"], ["free", "winner"], ["free", "meeting"]]:
14    print(f"{' + '.join(email):15} -> {spam_probability(email):.2%} spam")
Output
meeting         -> 6.25% spam
free            -> 90.91% spam
free + winner   -> 99.95% spam
free + meeting  -> 50.00% spam

Key takeaways

  • Bayes’ rule turns a prior and a likelihood into a posterior: an updated belief after seeing evidence.

  • Counting people (natural frequencies) makes it intuitive: positives = true positives + false positives.

  • Rare conditions + imperfect tests = most positives are false alarms (base rate neglect).

  • Beliefs update step by step; naive Bayes combines many clues by multiplying them, which works well but tends to be overconfident.

Lesson quiz

7 questions · pass with 5 correct · up to 50 XP

Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.

Practice: write Python

Write Python in the editor and run it against sample inputs. Python runs locally in your browser using a WebAssembly runtime.

Exercise 1

Natural frequencies calculator

+25 XP

The input is one line with three percentages: how common the condition is, the test’s sensitivity, and its false-alarm rate. Imagine 10,000 people and use round() for each count:

  • sick: N - how many have the condition
  • true positives: N - sick people the test catches
  • false positives: N - healthy people the test wrongly flags
  • P(sick | positive) = X% - true positives ÷ all positives, formatted with :.1%
  • The lesson example
  • More common
  • Very rare, great test
  • Common condition
main.py
Loading editor…

Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.

Exercise 2

Update beliefs test by test

+25 XP

Line 1: the prior, the sensitivity and the false-alarm rate, as percentages. Line 2: a sequence of test results, + or -, separated by spaces.

After each result, update the belief with Bayes’ rule and print the result symbol and the new belief, like + 15.4% (use :.1%). For a negative result, use the chances of a negative: 1 - sensitivity if sick, 1 - false_alarm if healthy.

  • The lesson sequence
  • Coin-flip prior
  • Mounting evidence
main.py
Loading editor…

Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.

Questions about this lesson

Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.

Loading posts…

Did you like the lesson? 😆👍
Consider a donation to support our work: