Language and large language models
Split text into tokens, give words meaning with embeddings, and generate text one predicted token at a time.
- Explain what tokens and embeddings are
- Compare word meanings with vector similarity
- Describe how a language model generates text and what temperature does
Natural language processing (NLP) is about getting computers to work with human language: translation, search, spam filtering, voice assistants, and the chatbots everyone is talking about.
Step one is always to cut text into tokens, the units a model works with. Tokens are often whole words, but large language models use subwords so they can handle any word, even ones they’ve never seen: “unhappiness” might become “un” + “happi” + “ness”. A rough rule of thumb for English is about 4 characters per token.
Embeddings: meaning as vectors
Next, each token becomes a vector called an embedding. These vectors are learned so that words used in similar ways end up close together: “cat” near “dog”, “Paris” near “Rome”. Directions carry meaning too - famously, .
How close are two words? Use the dot product from the math lesson, scaled by the vectors’ lengths. That’s cosine similarity: 1 means same direction, 0 unrelated, −1 opposite.
Try it
Explore a word map
Real embeddings have hundreds of dimensions; this toy map has two so you can see them.
- In Compare mode, click king and queen, then king and woman. Which pair points more the same way?
- In Solve analogies mode, guess the answer before it’s revealed. The arrow from man to woman is the “make it female” direction; adding it to king lands near queen.
Click two words on the map. The arrows point from the origin to each word; the narrower the angle between them, the higher the cosine similarity.
Nothing selected yet.
1import math
2
3embeddings = {"cat": [0.9, 0.8, 0.1], "dog": [0.8, 0.9, 0.2], "car": [0.1, 0.2, 0.9]}
4
5def cosine(a, b):
6 dot = sum(x * y for x, y in zip(a, b))
7 return dot / (math.hypot(*a) * math.hypot(*b))
8
9print(f"cat-dog: {cosine(embeddings['cat'], embeddings['dog']):.2f}")
10print(f"cat-car: {cosine(embeddings['cat'], embeddings['car']):.2f}")cat-dog: 0.99 cat-car: 0.30
Language models: predicting the next token
A language model does one thing: given the text so far, it predicts a probability for every possible next token. To write, it picks a token, appends it, and repeats. A large language model (LLM) like the ones behind modern chatbots is a huge neural network (a transformer) trained this way on trillions of tokens of text, and its billions of weights end up capturing grammar, facts and styles of reasoning.
The tiny model below learned only which words followed which in a few sentences. Build a sentence by choosing each next word.
Try it
Generate text one word at a time
Pick each next word from the options this tiny model learned. Notice that it only knows what followed what in its training text - it has no idea whether the sentence it produces is true.
the
After “the”, the corpus continues with… (choose one)
The model gives a probability to every token, but which one should it pick? Always taking the most likely makes text repetitive. Temperature controls the randomness: low temperature sharpens the probabilities toward the top choice; high temperature flattens them so unlikely tokens get picked more often.
Try it
Turn the temperature dial
These are a model’s scores for the next token. Sample 100 times at temperature 1, then at 0 (always the top token), 0.3 and 2. At which setting does “Berlin” start showing up?
The capital of France is ▁
- “ Paris”78.5%
- “ a”9.6%
- “ the”5.8%
- “ known”3.9%
- “ Lyon”1.2%
- “ not”0.6%
- “ Berlin”0.3%
Struck-through tokens were filtered out by top-k or top-p; the remaining probabilities are rescaled to add up to 100%.
Key takeaways
Text is split into tokens (often subwords) and each token becomes an embedding vector.
Similar meanings sit close together; cosine similarity measures how close.
A language model predicts the next token, appends it and repeats; temperature controls how adventurous its picks are.
Likely text can still be false: LLMs hallucinate.
Lesson quiz
6 questions · pass with 5 correct · up to 50 XP
Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.
Practice: write Python
Write Python in the editor and run it against sample inputs. Python runs locally in your browser using a WebAssembly runtime.
Build a next-word predictor
The first line is a training text. The second line is a word. Count which words follow that word in the text (lowercase, split on spaces) and print each follower with its probability as word: P (2 decimal places), most likely first and alphabetical on ties. If the word never appears with a follower, print no prediction.
- After "the"
- After "sat"
- Unknown
Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.
Find the most similar word
The first line is a query word. Every other line is word v1 v2 ..., an embedding. Print every other word with its cosine similarity to the query (2 decimal places) as word: S, most similar first, then closest: WORD.
- Pets
Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.
Questions about this lesson
Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.
Loading posts…