Where to go next
Review the core vocabulary, build a complete tiny classifier end to end, and pick your next track.
- Use the core AI vocabulary with confidence
- Combine training, prediction and evaluation in one program
- Choose a next step that fits your goals
You’ve gone from “what is AI?” to vectors, models, neural networks, vision, language and ethics. Here’s the vocabulary you’ve picked up - keep it handy.
| Term | Meaning |
|---|---|
| Feature | A measurable input a model uses to make a prediction. |
| Label | The correct answer attached to a training example. |
| Vector | An ordered list of numbers representing one thing. |
| Dot product | Multiply matching entries of two vectors and add them up. |
| Regression | Predicting a number. |
| Classification | Predicting a category. |
| k-nearest neighbors | Predict by a majority vote of the k closest training examples. |
| Loss function | One number measuring how wrong the predictions are. |
| Overfitting | Memorizing the training data instead of learning a general pattern. |
| Gradient descent | Repeatedly nudging parameters downhill on the loss. |
| Learning rate | The size of each gradient descent step. |
| Epoch | One full pass through the training data. |
| Neural network | Layers of connected artificial neurons. |
| Activation function | The nonlinear function applied to a neuron’s weighted sum. |
| Bias (parameter) | The learned constant added to a weighted sum. |
| Backpropagation | Computing every weight’s gradient by working backward from the output. |
| Deep learning | Machine learning with many-layered neural networks. |
| Convolution | Sliding a small kernel over a grid, taking a weighted sum at each spot. |
| Token | A unit of text (often a subword) that a language model processes. |
| Embedding | A learned vector that places similar items near each other. |
| Hallucination | A fluent, confident, false statement from a generative model. |
| Reinforcement learning | Learning by trial and error from rewards. |
Try it
Quick vocabulary sort
Which part of building a model does each term belong to?
“Features and labels”
“Activation function”
“Learning rate”
“Held-out test set”
“Embeddings”
“Epochs and batches”
“Per-group error rates”
“Train/test split”
Everything together
Here is a complete, tiny machine learning pipeline: split the data, train (compute each class’s average point, its centroid), predict (nearest centroid), and evaluate on held-out data. It’s a close cousin of both k-NN and k-means. The exercise asks you to write it yourself.
1import math
2
3data = [((1.0, 1.2), "small"), ((1.4, 0.9), "small"), ((0.8, 1.1), "small"), ((1.2, 1.5), "small"),
4 ((4.1, 3.8), "large"), ((3.7, 4.2), "large"), ((4.4, 4.0), "large"), ((3.2, 2.6), "large")]
5train, test = data[:3] + data[4:7], [data[3], data[7]]
6
7centroids = {}
8for label in ("small", "large"):
9 points = [point for point, point_label in train if point_label == label]
10 centroids[label] = tuple(sum(values) / len(values) for values in zip(*points))
11
12def predict(point):
13 return min(centroids, key=lambda label: math.dist(point, centroids[label]))
14
15correct = sum(predict(point) == label for point, label in test)
16print(f"test accuracy: {correct}/{len(test)}")test accuracy: 2/2
Keep learning
Pick the track that matches what excited you most:
- Machine Learning - logistic regression, decision trees, cross-validation, and a full capstone classifier with numpy.
- Data Science - explore, clean and visualize real data with pandas.
- Computer Vision - filters, segmentation, detection and CNNs in depth.
- Natural Language Processing - from tokenizers to word vectors and transformers.
- Large Language Models, Prompt Engineering and RAG - build with the models behind modern chatbots.
And some excellent free resources outside this site:
- Elements of AI - a friendly, no-code introduction from the University of Helsinki.
- Google’s Machine Learning Crash Course - the fundamentals in more depth, with interactive exercises.
- Harvard CS50’s Introduction to AI with Python - a project-based university course.
- 3Blue1Brown’s neural networks series - beautiful visual intuition for networks and backpropagation.
- fast.ai - a practical, code-first deep learning course.
Key takeaways
Every ML project follows the same shape: data → model → training → evaluation.
The core ideas - vectors, dot products, loss, gradient descent - reappear in every corner of AI.
Keep going with a deeper track, and build something small to make it stick.
Lesson quiz
6 questions · pass with 5 correct · up to 50 XP
Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.
Practice: write Python
Write Python in the editor and run it against sample inputs. Python runs locally in your browser using a WebAssembly runtime.
Build a complete classifier
Each input line is x y label. The last 25% of lines (rounding down the number of training lines to int(count * 0.75)) are the test set; the rest are for training.
- Train: for each label, compute the centroid (average x, average y) of its training points. Print them in alphabetical order as
LABEL: (X, Y)with 2 decimal places. - Predict each test point’s label as the label of the nearest centroid, and print
x y -> PREDICTED (actual ACTUAL)(format x and y with:g). - Print
accuracy: P%(whole percent).
When it works, look at the mistake it makes on the fruit test: weight in grams (100-170) swamps the color score (2-8) in the distance - the scaling pitfall from the k-NN lesson.
- Fruit
Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.
Questions about this lesson
Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.
Loading posts…