Loading
0x30Lesson 4 of 12

The three types of machine learning

Supervised learning from answers, unsupervised learning from structure, and reinforcement learning from rewards.

20 min 6-question quiz 2 code exercises
By the end of this lesson you can
  • Tell supervised, unsupervised and reinforcement learning apart
  • Split supervised learning into classification and regression
  • Describe how clustering and reward-based learning work

Machine learning comes in three flavors, depending on what kind of feedback the system gets while it learns.

Machine LearningSupervised
Learns from labeled examples (input → correct answer)
e.g. spam filtering
Unsupervised
Finds structure in unlabeled data on its own
e.g. customer segments
Reinforcement
Learns by trial and error, guided by rewards
e.g. game-playing agents
The three broad families of machine learning.
  • Supervised learning - every example comes with the right answer, like a student with an answer key. Predicting a category (spam / not spam, cat / dog) is classification; predicting a number (price, temperature) is regression.
  • Unsupervised learning - no answers at all. The system finds structure on its own: clustering groups similar items, dimensionality reduction squeezes many features into a few, anomaly detection flags the odd one out.
  • Reinforcement learning - an agent acts in an environment and gets rewards or penalties. Nobody shows it the right move; it discovers what works by trial and error, like a dog learning tricks for treats.

Try it

Which kind of learning?

Sort each task by the kind of feedback the system learns from.

0 of 7 sortedScore 0/0
  • “Predict a house’s price from past sales”

  • “Group shoppers into segments nobody defined in advance”

  • “Teach a robot arm to stack blocks by rewarding each success”

  • “Recognize handwritten digits from labeled scans”

  • “Flag a server whose traffic looks unlike anything seen before”

  • “Learn to play a video game from the score alone”

  • “Predict tomorrow’s temperature from years of weather records”

Unsupervised: finding clusters

The classic clustering algorithm is k-means. Drop k centers anywhere, then repeat two steps: assign every point to its nearest center, and move each center to the middle of its points. Nobody told it what the groups are; they emerge from the data.

Try it

Step through k-means

Press the step buttons and watch the three centers (crosses) travel. How many rounds until nothing changes any more? The “inertia” number measures how tightly packed the clusters are - it can only go down.

Diamonds are centroids. Next step: assign each point to its nearest centroid.

    Reinforcement: learning from rewards

    An agent facing three slot machines doesn’t know which pays best. It has to explore (try them) and exploit (keep playing the best one so far). This agent tries each machine twice, then sticks with the one whose average reward is highest:

    slot_machines.py
    1# What each machine pays out on successive pulls (unknown to the agent).
    2payouts = {"A": [1, 0, 1, 0, 1, 0], "B": [0, 1, 1, 1, 0, 1], "C": [0, 0, 1, 0, 0, 0]}
    3pulls = {name: 0 for name in payouts}
    4total_reward = {name: 0 for name in payouts}
    5
    6def pull(name):
    7    reward = payouts[name][pulls[name] % 6]
    8    pulls[name] += 1
    9    total_reward[name] += reward
    10
    11for name in payouts:        # explore: try each machine twice
    12    pull(name)
    13    pull(name)
    14averages = {name: total_reward[name] / pulls[name] for name in payouts}
    15best = max(averages, key=averages.get)
    16print("averages after exploring:", averages)
    17print("exploit:", best)
    Output
    averages after exploring: {'A': 0.5, 'B': 0.5, 'C': 0.0}
    exploit: A

    Look closely: after two pulls each, A and B are tied, so the agent settles on A - yet over six pulls B pays 4 times and A only 3. Two tries weren’t enough evidence. That is the exploration-exploitation trade-off: explore too little and you lock in a worse choice; explore too much and you waste pulls on machines you already know are bad. Change the exploring loop to pull each machine four times and see what it picks.

    Key takeaways

    • Supervised: learn from labeled examples - classification (categories) or regression (numbers).

    • Unsupervised: find structure without labels - clustering, dimensionality reduction, anomaly detection.

    • Reinforcement: an agent learns which actions earn rewards, balancing exploration and exploitation.

    Lesson quiz

    6 questions · pass with 5 correct · up to 50 XP

    Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.

    Practice: write Python

    Write Python in the editor and run it against sample inputs. Python runs locally in your browser using a WebAssembly runtime.

    Exercise 1

    Assign points to clusters

    +25 XP

    The first line holds cluster centers as x,y pairs separated by spaces. Every other line is a point x,y. Print the index (starting at 0) of each point’s nearest center by straight-line distance - the assign step of k-means. Then print how many points each center got, as center 0: N and so on.

    • Two clusters
    main.py
    Loading editor…

    Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.

    Exercise 2

    Exploit the best machine

    +25 XP

    Each line is machine reward - the result of one pull during exploration. Print each machine’s average reward to 2 decimal places in alphabetical order as A: 0.50, then play: NAME for the machine with the highest average (alphabetically first on ties).

    • Three machines
    • A tie
    main.py
    Loading editor…

    Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.

    Questions about this lesson

    Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.

    Loading posts…

    Did you like the lesson? 😆👍
    Consider a donation to support our work: