Loading
0xD0Lesson 14 of 18

Report cards for AI: measuring models

Why 99% accuracy can be useless, how a confusion matrix exposes a model’s mistakes, and how to trade precision against recall.

22 min 7-question quiz 2 code exercises
By the end of this lesson you can
  • Build a confusion matrix and compute accuracy, precision and recall
  • Explain the accuracy paradox and why baselines matter
  • Choose between precision and recall based on the cost of each mistake

A company proudly announces: “Our fraud detector is 99% accurate!” Impressive? Not necessarily. If only 1% of payments are fraud, a “model” that answers “not fraud” every single time is also 99% accurate - and catches nothing.

One number can hide a lot. To really understand a classifier, split its predictions into four boxes - the confusion matrix:

the model predicted…spamnot spamreally was…spamnot spamTrue positivespam caught ✔False negativespam slipped in ✘False positivereal email binned ✘True negativereal email kept ✔precision = TP ÷ this columnrecall =TP ÷ this rowaccuracy = (TP + TN) ÷ everything
A confusion matrix sorts every prediction into one of four boxes. Precision asks “of what I flagged, how much was right?” Recall asks “of what was really there, how much did I catch?”
  • Accuracy = (TP + TN) ÷ everything - how often it’s right overall.
  • Precision = TP ÷ (TP + FP) - when it raises the alarm, how often is it right?
  • Recall = TP ÷ (TP + FN) - of all the real cases, how many did it catch?
confusion_matrix.py
1actual    = [1, 1, 1, 1, 0, 0, 0, 0, 0, 0]    # 1 = spam
2predicted = [1, 1, 1, 0, 1, 0, 0, 0, 0, 0]
3
4pairs = list(zip(actual, predicted))
5tp = pairs.count((1, 1))
6fn = pairs.count((1, 0))
7fp = pairs.count((0, 1))
8tn = pairs.count((0, 0))
9print(f"TP={tp} FN={fn} FP={fp} TN={tn}")
10print(f"accuracy  {(tp + tn) / len(pairs):.0%}")
11print(f"precision {tp / (tp + fp):.0%}  (of emails flagged as spam, how many were spam)")
12print(f"recall    {tp / (tp + fn):.0%}  (of real spam, how much was caught)")
Output
TP=3 FN=1 FP=1 TN=5
accuracy  80%
precision 75%  (of emails flagged as spam, how many were spam)
recall    75%  (of real spam, how much was caught)

The accuracy paradox

accuracy_paradox.py
1# 1,000 card payments; 10 are fraud.
2actual = [1] * 10 + [0] * 990
3
4lazy_model = [0] * 1000                         # always says "not fraud"
5print(f"lazy model accuracy:  {sum(a == p for a, p in zip(actual, lazy_model)) / 1000:.1%}")
6print(f"lazy model recall:    {sum(a == p == 1 for a, p in zip(actual, lazy_model)) / 10:.0%}")
7
8# A real model: catches 8 of the 10 frauds, but also flags 30 honest payments.
9real_model = [1] * 8 + [0] * 2 + [1] * 30 + [0] * 960
10print(f"real model accuracy:  {sum(a == p for a, p in zip(actual, real_model)) / 1000:.1%}")
11print(f"real model recall:    {sum(a == p == 1 for a, p in zip(actual, real_model)) / 10:.0%}")
Output
lazy model accuracy:  99.0%
lazy model recall:    0%
real model accuracy:  96.8%
real model recall:    80%

The useless model has higher accuracy than the useful one! That’s the accuracy paradox, and it shows up whenever one class is rare - fraud, disease, defects, earthquakes. Two lessons: always compare against a baseline (like “always guess the most common answer”), and look at precision and recall, not just accuracy.

Turning the dial: the precision-recall trade-off

Most classifiers output a score, and a threshold turns it into a yes/no. Lower the threshold and you catch more real cases (recall ↑) but raise more false alarms (precision ↓). Raise it and the opposite happens. There is no free lunch - you choose the balance based on which mistake costs more.

Try it

Tune a cat detector

A model scored each photo for “is this a cat?”. Slide the threshold and watch the confusion matrix and the metrics change. Find the threshold that catches every cat. What did it cost you in precision? Then find the one with no false alarms at all.

57%
Precision
Of the messages flagged, how many really are cat
80%
Recall
Of the real cat messages, how many were flagged
67%
F1
One number that balances the two
Confusion matrix
Flagged catCalled not a cat
Really cat4 true positives1 missed (false negatives)
Really not a cat3 false alarms (false positives)3 true negatives
  • 0.97Tabby asleep on a sofa (cat)flagged cat
  • 0.92Kitten chasing a string (cat)flagged cat
  • 0.81Fluffy dog with pointy ears (not a cat)flagged cat
  • 0.74Black cat in a dark room (cat)flagged cat
  • 0.69Lion at the zoo (not a cat)flagged cat
  • 0.58Cat half-hidden behind a curtain (cat)flagged cat
  • 0.51Cat-shaped cushion (not a cat)flagged cat
  • 0.37Blurry cat running past (cat)not a cat
  • 0.29Fox in a garden (not a cat)not a cat
  • 0.18Bowl of cat food (not a cat)not a cat
  • 0.04Empty armchair (not a cat)not a cat

Try it

Which mistake hurts more?

For each system, decide what matters most: catching every real case (recall) or making sure alarms are trustworthy (precision).

0 of 6 sortedScore 0/0
  • “Screening for a serious but treatable disease, with a follow-up test”

  • “Moving emails to the spam folder”

  • “Detecting smoke in a building”

  • “Automatically blocking a customer’s bank card”

  • “Flagging possible weapons in airport bags for a human to check”

  • “A search engine’s first page of results”

Key takeaways

  • The confusion matrix splits predictions into TP, FP, FN and TN.

  • Precision = TP ÷ (TP + FP); recall = TP ÷ (TP + FN); accuracy = (TP + TN) ÷ all.

  • With rare classes, accuracy misleads (the accuracy paradox) - always compare with a baseline.

  • Moving the threshold trades precision against recall; choose based on which mistake costs more.

Lesson quiz

7 questions · pass with 5 correct · up to 50 XP

Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.

Practice: write Python

Write Python in the editor and run it against sample inputs. Python runs locally in your browser using a WebAssembly runtime.

Exercise 1

Build a confusion matrix

+25 XP

Line 1: the actual labels, 1 (positive) or 0. Line 2: the model’s predictions, in the same order.

Print TP=a FN=b FP=c TN=d, then accuracy: X, precision: X and recall: X, each to 2 decimal places. If precision or recall would divide by zero, print 0.00 for it.

  • The lesson example
  • Flags everything
  • Flags nothing
  • Mixed
main.py
Loading editor…

Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.

Exercise 2

Does the model beat the baseline?

+25 XP

Line 1: the actual labels (words like ok and fraud). Line 2: a model’s predictions.

The baseline always predicts the most common actual label (if there’s a tie, the alphabetically first one). Print baseline (LABEL): X% and model: Y% - both accuracies with :.1% - and then beats the baseline if the model’s accuracy is strictly higher, otherwise no better than the baseline.

  • The lazy fraud model
  • A useful model
  • Three classes
  • Worse than guessing
main.py
Loading editor…

Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.

Questions about this lesson

Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.

Loading posts…

Did you like the lesson? 😆👍
Consider a donation to support our work: