Compute attention step by step
Turn word-to-word scores into attention weights with softmax, and mix information by those weights.
- Turn scores into weights with softmax
- Explain queries, keys and values in plain words
- Compute attention weights from dot products
In “The animal didn’t cross the street because it was too tired”, what does it refer to? You know it’s the animal - streets don’t get tired. Attention is the mechanism that lets a model make that kind of connection: when building the new vector for it, it can draw heavily on animal.
Try it
Attention explorer
These are toy attention weights for one attention head. Click a word to see how much it attends to every word (greener = more).
- Click it: which word gets the most attention?
- Click cross: who does the crossing, and what is crossed?
- Slide the temperature: at low temperature, attention focuses on the top word; at high temperature it spreads out evenly.
Click a word to see where it looks
“it” pays the most attention to “animal” (66%). The weights add up to 100%.
- The1.2%
- animal65.8%
- didn’t1.2%
- cross1.2%
- the1.2%
- street10.9%
- because3.3%
- it5.4%
- was3.3%
- too1.2%
- tired5.4%
From scores to weights: softmax
Attention starts with a score for each pair of words: how relevant is word B to word A? Scores can be any number, so softmax turns them into weights: exponentiate each score, then divide by the total. Bigger scores get much bigger weights, every weight is positive, and they sum to 1.
1import math
2scores = [2.0, 1.0, 0.0]
3exponentials = [math.exp(score) for score in scores]
4total = sum(exponentials)
5print([round(value / total, 2) for value in exponentials])[0.67, 0.24, 0.09]
Queries, keys and values
Where do the scores come from? Each token’s vector is turned into three vectors, a bit like a library search:
- a query: what this token is looking for (it looks for a noun it could refer to),
- a key: what this token offers to others (animal offers “I’m a living noun”),
- a value: the information it passes on if chosen.
The score between two tokens is the dot product of one’s query with the other’s key (scaled by the square root of the vector size). Softmax turns the scores into weights, and the token’s new vector is the weighted sum of the values. A Transformer runs many of these attention “heads” side by side, each free to learn a different kind of link.
Key takeaways
Attention lets each token build its new vector from the tokens most relevant to it.
Softmax turns any scores into positive weights that sum to 1; temperature controls how peaked they are.
Score = query · key; output = weighted sum of values. Many heads learn different links in parallel.
Lesson quiz
6 questions · pass with 5 correct · up to 50 XP
Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.
Practice: apply NLP with Python
Try each text-processing idea in Python, run it against sample inputs, and use the results to see where the method works or falls short.
Implement softmax
Read space-separated scores on one line. Print their softmax weights, each rounded to 2 decimals, separated by one space.
- Three scores
- Equal scores
Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.
Attention weights from dot products
The first line is a query vector. The second is N, followed by N key vectors, one per line. Score each key by its dot product with the query, apply softmax, and print the weights rounded to 2 decimals, separated by one space.
- Three keys
- Two equal keys
Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.
Questions about this lesson
Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.
Loading posts…