Train vision models: data, augmentation and transfer
Learn how vision models are trained, why they overfit, and how augmentation and pretrained models help.
- Describe a training loop and recognize overfitting
- Choose augmentations that keep the label true
- Explain transfer learning from a pretrained model
Training a vision model is a loop: show it a batch of labeled images, measure how wrong its predictions are (the loss), nudge every weight slightly in the direction that reduces the loss (gradient descent), and repeat - millions of times.
The big danger is overfitting: the model memorizes the training images - including their quirks, like a watermark that only appears on photos of one class - instead of learning what the objects look like. Its training accuracy keeps rising while its accuracy on new images stalls or drops.
Data augmentation
More varied data is the best cure for overfitting, and augmentation makes it for free: flipped, slightly rotated, cropped, brighter or darker copies of each training image. The model learns that these changes don’t change what the object is. But each augmentation must keep the label true.
Try it
Safe augmentation?
For each augmentation, decide whether the label is still correct afterwards.
“Flip a photo of a dog left-to-right”
“Flip a handwritten “6” upside down”
“Make a street photo slightly darker”
“Crop a photo so the bird is cut out of the frame”
“Flip a road sign that says “LEFT TURN ONLY” left-to-right”
“Rotate a handwritten “4” by 10 degrees”
Try it
Make augmented copies
Each operation below creates a new training example from the same letter. Which of these would be safe for a dataset of letters? (Hint: is a mirrored F still an F?)
No operations yet.
Transfer learning
You rarely train a vision model from scratch. Transfer learning starts from a model pretrained on millions of images (its early layers already detect edges, textures and parts), replaces its last layer, and fine-tunes it on your smaller dataset. A few hundred labeled images can be enough to recognize, say, your product’s defects.
Key takeaways
Training repeats: predict → measure loss → adjust weights.
Overfitting = great on training images, poor on new ones; augmentation fights it, if labels stay true.
Transfer learning fine-tunes a pretrained model, needing far less data.
Lesson quiz
6 questions · pass with 5 correct · up to 50 XP
Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.
Practice: apply computer vision with Python
Use small pixel arrays to explore vision concepts, run your code against sample images, and connect each result to the larger computer vision idea.
Count distinct flips
Read a JSON grayscale image. Make four versions: the original, a left-right flip, an upside-down flip, and both flips together. Print how many of the four are different from each other. (Symmetric images give fewer distinct copies - and less new information for training.)
- No symmetry
- Mirror symmetric
- Fully symmetric
Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.
Questions about this lesson
Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.
Loading posts…