Meet modern vision: transformers, CLIP and generation
See how images become sequences of patches, and how vision now connects to language.
- Explain how a Vision Transformer turns an image into patch tokens
- Describe what image-text models like CLIP make possible
- Name what generative image models do - and their risks
For most of the 2010s, CNNs ruled vision. Then the Transformer - the architecture behind language models - arrived. The paper “An Image is Worth 16x16 Words” (Dosovitskiy et al., 2020) showed a plain Transformer can classify images very well when trained on enough data.
Images as sequences of patches
A Vision Transformer (ViT) chops the image into a grid of square patches (say 16×16 pixels), flattens each patch into a list of numbers, and turns each into a vector - a “visual word”. Add position information, and the patches go through the same self-attention layers used for text.
A 224×224 image with 16×16 patches becomes 14 × 14 = 196 patch tokens. Attention lets every patch look at every other patch from the first layer on, while a CNN only sees small neighborhoods until deep in the network.
image_size, patch_size = 224, 16
patches_per_side = image_size // patch_size
print(patches_per_side ** 2, patch_size * patch_size * 3)196 768
Vision meets language
CLIP-style models are trained on hundreds of millions of image–caption pairs to place matching images and texts close together in one shared embedding space. That enables zero-shot classification - compare an image with the texts “a photo of a dog” and “a photo of a cat” and pick the closer one, with no task-specific training - and searching photos by description.
Generative models go the other way: diffusion models start from noise and repeatedly “denoise” it into an image that matches a text prompt. The same tools that make illustrations also make convincing fake photos, which is why provenance and watermarking matter.
Try it
Match the model to the job
Which kind of model fits each job best?
“Find photos in your gallery matching “sunset over the sea””
“Create a picture of “a fox reading a newspaper””
“Sort factory photos into “defect” and “no defect” with thousands of labeled examples”
“Label images into categories you only thought of today, with no training examples”
Key takeaways
A ViT splits an image into patches and treats them like words in a Transformer.
CLIP-style models put images and text in one space: zero-shot labels and search by description.
Generative models create images from text - powerful, and easy to misuse.
Lesson quiz
6 questions · pass with 5 correct · up to 50 XP
Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.
Practice: apply computer vision with Python
Use small pixel arrays to explore vision concepts, run your code against sample images, and connect each result to the larger computer vision idea.
Cut an image into patches
Read a JSON grayscale image whose sides are multiples of the patch size, and the patch size on the next line. Print each patch, flattened row by row into one list, in reading order (left to right, then top to bottom) - one patch per line.
- 4×4 into 2×2 patches
- One patch
Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.
Questions about this lesson
Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.
Loading posts…