Loading
0xD0Lesson 14 of 16

Meet modern vision: transformers, CLIP and generation

See how images become sequences of patches, and how vision now connects to language.

20 min 6-question quiz 1 code exercise
By the end of this lesson you can
  • Explain how a Vision Transformer turns an image into patch tokens
  • Describe what image-text models like CLIP make possible
  • Name what generative image models do - and their risks

For most of the 2010s, CNNs ruled vision. Then the Transformer - the architecture behind language models - arrived. The paper “An Image is Worth 16x16 Words” (Dosovitskiy et al., 2020) showed a plain Transformer can classify images very well when trained on enough data.

Images as sequences of patches

A Vision Transformer (ViT) chops the image into a grid of square patches (say 16×16 pixels), flattens each patch into a list of numbers, and turns each into a vector - a “visual word”. Add position information, and the patches go through the same self-attention layers used for text.

A 224×224 image with 16×16 patches becomes 14 × 14 = 196 patch tokens. Attention lets every patch look at every other patch from the first layer on, while a CNN only sees small neighborhoods until deep in the network.

patches.py
image_size, patch_size = 224, 16
patches_per_side = image_size // patch_size
print(patches_per_side ** 2, patch_size * patch_size * 3)
Output
196 768

Vision meets language

CLIP-style models are trained on hundreds of millions of image–caption pairs to place matching images and texts close together in one shared embedding space. That enables zero-shot classification - compare an image with the texts “a photo of a dog” and “a photo of a cat” and pick the closer one, with no task-specific training - and searching photos by description.

Generative models go the other way: diffusion models start from noise and repeatedly “denoise” it into an image that matches a text prompt. The same tools that make illustrations also make convincing fake photos, which is why provenance and watermarking matter.

Try it

Match the model to the job

Which kind of model fits each job best?

0 of 4 sortedScore 0/0
  • “Find photos in your gallery matching “sunset over the sea””

  • “Create a picture of “a fox reading a newspaper””

  • “Sort factory photos into “defect” and “no defect” with thousands of labeled examples”

  • “Label images into categories you only thought of today, with no training examples”

Key takeaways

  • A ViT splits an image into patches and treats them like words in a Transformer.

  • CLIP-style models put images and text in one space: zero-shot labels and search by description.

  • Generative models create images from text - powerful, and easy to misuse.

Lesson quiz

6 questions · pass with 5 correct · up to 50 XP

Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.

Practice: apply computer vision with Python

Use small pixel arrays to explore vision concepts, run your code against sample images, and connect each result to the larger computer vision idea.

Exercise 1

Cut an image into patches

+25 XP

Read a JSON grayscale image whose sides are multiples of the patch size, and the patch size on the next line. Print each patch, flattened row by row into one list, in reading order (left to right, then top to bottom) - one patch per line.

  • 4×4 into 2×2 patches
  • One patch
main.py
Loading editor…

Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.

Questions about this lesson

Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.

Loading posts…

Did you like the lesson? 😆👍
Consider a donation to support our work: