Mixture of experts and multimodal models
See how a router lets a 47B-parameter model compute like a 13B one, and how a photo turns into a few hundred “words” the model can read.
- Explain how a mixture-of-experts layer routes each token to a few experts
- Tell total parameters from active parameters, and what each one costs you
- Describe how images become tokens, and estimate how much context an image uses
Bigger models know more, but every parameter normally costs compute on every token. A mixture of experts (MoE) breaks that link. Inside each Transformer block, the single feed-forward network is replaced by several smaller ones - the experts - plus a tiny router that picks which experts each token visits.
For each token the router scores every expert, keeps the top-k (usually 1 or 2, sometimes 8 of 256), and mixes their outputs using the renormalized router weights. The other experts sit idle for that token. Mixtral, DeepSeek-V3, Qwen-MoE and (reportedly) several frontier models work this way.
1import numpy as np
2
3logits = np.array([1.2, -0.3, 2.0, 0.1, 0.4, -1.0, 0.0, 0.9]) # router scores for 8 experts
4probabilities = np.exp(logits) / np.exp(logits).sum()
5top2 = np.argsort(-probabilities)[:2]
6weights = probabilities[top2] / probabilities[top2].sum()
7for expert, weight in zip(top2, weights):
8 print(f"expert {expert}: weight {weight:.2f}")expert 2: weight 0.69 expert 0: weight 0.31
Total vs. active parameters
An MoE has two sizes. Total parameters decide how much the model can store - and how much memory you need, since every expert must be loaded. Active parameters are those a single token actually passes through - they decide compute per token and much of the speed.
With Mixtral-style numbers - about 1.58B shared parameters (attention, embeddings) and eight experts of 5.64B each, two active per token:
1shared, per_expert, experts, active = 1.58e9, 5.64e9, 8, 2
2total = shared + experts * per_expert
3active_params = shared + active * per_expert
4print(f"total {total / 1e9:.1f}B, active per token {active_params / 1e9:.1f}B")total 46.7B, active per token 12.9B
Seeing with tokens
A vision-language model reads images by turning them into something that looks like tokens:
- The image is resized and cut into a grid of small patches, say 14 × 14 pixels.
- A vision encoder (often a ViT trained CLIP-style to match images with captions) turns each patch into a vector.
- A small projector maps those vectors into the language model’s embedding space - sometimes merging neighbouring patches (2 × 2 → 1) to save tokens.
- The language model receives these image tokens in its context, mixed in with text tokens, and attends to them like any other token.
So images cost context: a 448 × 448 image in 14-pixel patches is a 32 × 32 grid = 1,024 patches, or 256 tokens after a 2 × 2 merge. A screenshot of a whole web page can cost more than the text you’d paste instead. Audio works the same way, with short slices of sound in place of patches.
Try it
Follow a photo through a vision-language model
Step through what happens when you ask “What’s in this picture?” and predict before each reveal.
Key takeaways
MoE layers route each token to its top-k experts, so parameters grow faster than compute.
Total parameters set memory; active parameters set compute per token.
Images are split into patches, encoded, projected and fed in as image tokens - which use context like text.
Lesson quiz
7 questions · pass with 5 correct · up to 50 XP
Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.
Practice: write Python
Write Python in the editor and run it against sample inputs. Python runs locally in your browser using a WebAssembly runtime.
Route tokens to experts
Each input line holds one token’s router logits, one per expert. For every token, apply softmax, pick the top 2 experts (ties go to the lower expert number), and renormalize their two probabilities to sum to 1.
Print token 1: e2 0.69, e0 0.31 (highest first, weights to 2 decimals). Then print the load - how many tokens each expert received, like load: e0=3 e1=0 e2=2 e3=1 - and idle: e1 listing experts that got nothing, comma-separated, or idle: none.
- A lazy expert
- Perfectly balanced
Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.
Image token budget
Line 1 is patch merge: the patch size in pixels and how many patches per side merge into one token (1 means no merging). Each following line is name width height.
An image becomes a grid of ceil(width / patch) × ceil(height / patch) patches. Merging turns that into ceil(cols / merge) × ceil(rows / merge) tokens. Print name: COLSxROWS patches -> T tokens per image, then total: N tokens.
- With 2 × 2 merging
- No merging
Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.
Questions about this lesson
Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.
Loading posts…