Loading
0x80Lesson 9 of 16

Mixture of experts and multimodal models

See how a router lets a 47B-parameter model compute like a 13B one, and how a photo turns into a few hundred “words” the model can read.

25 min 7-question quiz 2 code exercises
By the end of this lesson you can
  • Explain how a mixture-of-experts layer routes each token to a few experts
  • Tell total parameters from active parameters, and what each one costs you
  • Describe how images become tokens, and estimate how much context an image uses

Bigger models know more, but every parameter normally costs compute on every token. A mixture of experts (MoE) breaks that link. Inside each Transformer block, the single feed-forward network is replaced by several smaller ones - the experts - plus a tiny router that picks which experts each token visits.

For each token the router scores every expert, keeps the top-k (usually 1 or 2, sometimes 8 of 256), and mixes their outputs using the renormalized router weights. The other experts sit idle for that token. Mixtral, DeepSeek-V3, Qwen-MoE and (reportedly) several frontier models work this way.

router.py
1import numpy as np
2
3logits = np.array([1.2, -0.3, 2.0, 0.1, 0.4, -1.0, 0.0, 0.9])   # router scores for 8 experts
4probabilities = np.exp(logits) / np.exp(logits).sum()
5top2 = np.argsort(-probabilities)[:2]
6weights = probabilities[top2] / probabilities[top2].sum()
7for expert, weight in zip(top2, weights):
8    print(f"expert {expert}: weight {weight:.2f}")
Output
expert 2: weight 0.69
expert 0: weight 0.31

Total vs. active parameters

An MoE has two sizes. Total parameters decide how much the model can store - and how much memory you need, since every expert must be loaded. Active parameters are those a single token actually passes through - they decide compute per token and much of the speed.

With Mixtral-style numbers - about 1.58B shared parameters (attention, embeddings) and eight experts of 5.64B each, two active per token:

moe_size.py
1shared, per_expert, experts, active = 1.58e9, 5.64e9, 8, 2
2total = shared + experts * per_expert
3active_params = shared + active * per_expert
4print(f"total {total / 1e9:.1f}B, active per token {active_params / 1e9:.1f}B")
Output
total 46.7B, active per token 12.9B

Seeing with tokens

A vision-language model reads images by turning them into something that looks like tokens:

  1. The image is resized and cut into a grid of small patches, say 14 × 14 pixels.
  2. A vision encoder (often a ViT trained CLIP-style to match images with captions) turns each patch into a vector.
  3. A small projector maps those vectors into the language model’s embedding space - sometimes merging neighbouring patches (2 × 2 → 1) to save tokens.
  4. The language model receives these image tokens in its context, mixed in with text tokens, and attends to them like any other token.

So images cost context: a 448 × 448 image in 14-pixel patches is a 32 × 32 grid = 1,024 patches, or 256 tokens after a 2 × 2 merge. A screenshot of a whole web page can cost more than the text you’d paste instead. Audio works the same way, with short slices of sound in place of patches.

Try it

Follow a photo through a vision-language model

Step through what happens when you ask “What’s in this picture?” and predict before each reveal.

Message 1 of 4Predicted 0/0
You
Vision encoder
Projector
Language model

Key takeaways

  • MoE layers route each token to its top-k experts, so parameters grow faster than compute.

  • Total parameters set memory; active parameters set compute per token.

  • Images are split into patches, encoded, projected and fed in as image tokens - which use context like text.

Lesson quiz

7 questions · pass with 5 correct · up to 50 XP

Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.

Practice: write Python

Write Python in the editor and run it against sample inputs. Python runs locally in your browser using a WebAssembly runtime.

Exercise 1

Route tokens to experts

+25 XP

Each input line holds one token’s router logits, one per expert. For every token, apply softmax, pick the top 2 experts (ties go to the lower expert number), and renormalize their two probabilities to sum to 1.

Print token 1: e2 0.69, e0 0.31 (highest first, weights to 2 decimals). Then print the load - how many tokens each expert received, like load: e0=3 e1=0 e2=2 e3=1 - and idle: e1 listing experts that got nothing, comma-separated, or idle: none.

  • A lazy expert
  • Perfectly balanced
main.py
Loading editor…

Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.

Exercise 2

Image token budget

+25 XP

Line 1 is patch merge: the patch size in pixels and how many patches per side merge into one token (1 means no merging). Each following line is name width height.

An image becomes a grid of ceil(width / patch) × ceil(height / patch) patches. Merging turns that into ceil(cols / merge) × ceil(rows / merge) tokens. Print name: COLSxROWS patches -> T tokens per image, then total: N tokens.

  • With 2 × 2 merging
  • No merging
main.py
Loading editor…

Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.

Questions about this lesson

Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.

Loading posts…

Did you like the lesson? 😆👍
Consider a donation to support our work: