Learn a classifier with Naive Bayes
Let labeled examples, not hand-written word lists, decide which words matter.
- Explain how Naive Bayes turns word counts into class probabilities
- Apply add-one smoothing and work in log probabilities
- Train and run a spam filter from labeled messages
Hand-written word lists don’t scale: who decides that “prize” is spammy? A learned classifier works it out from labeled examples - messages people already marked as spam or not (“ham”). Naive Bayes is the classic learned text classifier: it powered the first effective spam filters, is fast to train, and is still a strong baseline.
Try it
Think like a spam filter
Before any math, sort these messages yourself. As you go, notice which individual words made you decide - that is exactly the evidence Naive Bayes collects.
“WIN a FREE prize now, click here!”
“Can we move the meeting to 3pm?”
“Your account has been selected for a cash reward”
“Thanks for lunch, see you tomorrow”
“Free tickets for the team dinner on Friday”
Counting evidence
Naive Bayes asks, for each class: how likely is this message, if it were spam? if it were ham? It multiplies:
- the prior: how common the class is (say, 40% of messages are spam), and
- for every word in the message,
P(word | class): how often that word appears in that class’s training messages.
It is naive because it pretends words are independent of each other - “free” and “prize” are scored separately even though they travel together. The assumption is wrong, yet the classifier works surprisingly well.
Two practical fixes make it work:
- Add-one smoothing: a word never seen in spam would get
P = 0and wipe out the whole product. Instead use(count + 1) / (total words in class + vocabulary size). - Log probabilities: multiplying hundreds of small numbers underflows to 0.0 on a computer. Adding their logarithms instead (
log(a × b) = log a + log b) keeps the numbers healthy, and the biggest sum still wins.
1import math
2from collections import Counter
3
4spam_words = Counter("win money win prize".split())
5vocabulary_size = 6
6total = sum(spam_words.values())
7for word in ["win", "meeting"]:
8 probability = (spam_words[word] + 1) / (total + vocabulary_size)
9 print(word, round(probability, 2), round(math.log(probability), 2))win 0.3 -1.2 meeting 0.1 -2.3
Key takeaways
Naive Bayes learns which words signal each class from labeled examples.
Score = log prior + sum of log word likelihoods; the highest score wins.
Smoothing handles unseen words; logs avoid numbers too small to store.
Lesson quiz
6 questions · pass with 5 correct · up to 50 XP
Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.
Practice: apply NLP with Python
Try each text-processing idea in Python, run it against sample inputs, and use the results to see where the method works or falls short.
Train a spam filter
The first line is N, the number of training messages. Each of the next N lines is label|message. The last line is a new message to classify.
Lowercase and split messages on whitespace. For each label, compute log(prior) + Σ log((count(word, label) + 1) / (total words in label + V)) over the new message’s words, where V is the number of distinct words in all training messages. Print the label with the highest score (ties: the alphabetically first label).
- Spammy words
- Work message
Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.
Questions about this lesson
Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.
Loading posts…