Loading
0x90Lesson 10 of 10

Capstone: build a tiny language model

Train a word-level bigram model, measure its perplexity, and generate text with temperature.

30 min 5-question quiz 2 code exercises
By the end of this lesson you can
  • Train a smoothed bigram language model from text
  • Evaluate it with perplexity on held-out text
  • Generate text greedily and with temperature

Time to build the whole loop - training, evaluation and generation - in a model small enough to read. A bigram model predicts each word from just the previous one, using counts from training text. Real LLMs replace the counts with a Transformer that looks at thousands of previous tokens, but the shape of the job is the same:

  1. Train: count which words follow which.
  2. Evaluate: perplexity on text the model hasn’t seen.
  3. Generate: pick a next word, append, repeat.

Unseen word pairs would get probability 0 (and infinite perplexity), so we use add-one smoothing: P(word | previous) = (count(previous, word) + 1) / (count(previous) + V), where V is the vocabulary size.

smoothing.py
count_pair, count_previous, vocabulary = 0, 4, 10
print((count_pair + 1) / (count_previous + vocabulary))
Output
0.07142857142857142

Key takeaways

  • Training learns next-token statistics; smoothing keeps unseen pairs possible.

  • Perplexity on held-out text measures how surprised the model is.

  • Generation is repeated next-token choice - greedy or sampled with a temperature.

Lesson quiz

5 questions · pass with 4 correct · up to 50 XP

Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.

Practice: write Python

Write Python in the editor and run it against sample inputs. Python runs locally in your browser using a WebAssembly runtime.

Exercise 1

Train and evaluate a bigram model

+25 XP

Line 1 is training text; line 2 is test text (lowercase words). V is the number of distinct words across both lines. For each consecutive pair in the test text, P = (count(prev, word) + 1) / (count(prev) + V), where count(prev) counts how often prev is followed by any word in training.

Print vocabulary: V and perplexity: X (exp of the mean −ln P, 2 decimals).

  • In-domain sentence
  • Out-of-domain sentence
main.py
Loading editor…

Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.

Exercise 2

Generate text with temperature

+25 XP

Line 1 is training text; line 2 is START WORDS TEMPERATURE. Starting from START, generate up to WORDS more words. For the current word, collect its followers and counts in training (sorted alphabetically); stop early if it has none.

  • Temperature 0: pick the most frequent follower (alphabetically first on ties).
  • Otherwise: call random.seed(0) once at the start, and pick with random.choices(followers, weights=[count ** (1 / temperature) ...])[0].

Print the generated words joined by spaces.

  • Greedy
  • Sampled at temperature 0.5
main.py
Loading editor…

Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.

Questions about this lesson

Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.

Loading posts…

Did you like the lesson? 😆👍
Consider a donation to support our work: