Positions and long context
Find out how a Transformer knows word order, why RoPE took over, and why a million-token window doesn’t mean a million tokens of attention.
- Explain why attention alone ignores word order, and how positional encodings fix it
- Show that rotary embeddings (RoPE) make attention depend on relative distance
- Describe what grows with context length, and test whether a model really uses its context
“The dog bit the man” and “the man bit the dog” use the same tokens. Yet attention, as you computed it in Explore Transformers and attention, is just dot products and weighted sums - it has no idea which token came first. Shuffle the input tokens and the outputs simply shuffle along with them. A bag of words with very good taste.
So every Transformer adds position information. The idea is simple; the details decide how well a model handles long documents.
1import numpy as np
2rng = np.random.default_rng(0)
3x = rng.normal(size=(4, 8)) # 4 tokens, 8 dimensions each
4
5def attend(x):
6 scores = x @ x.T / np.sqrt(x.shape[1])
7 weights = np.exp(scores - scores.max(axis=1, keepdims=True))
8 weights /= weights.sum(axis=1, keepdims=True)
9 return weights @ x
10
11order = [2, 0, 3, 1]
12print("shuffling tokens just shuffles the outputs:", np.allclose(attend(x)[order], attend(x[order])))shuffling tokens just shuffles the outputs: True
Three ways to say “where”
- Learned absolute positions (GPT-2): a trainable vector for position 0, 1, 2… added to each token embedding. Simple - but there is simply no vector for position 2,049 if training stopped at 2,048.
- Sinusoidal positions (the original 2017 Transformer): each position gets a fixed pattern of sines and cosines at different frequencies, like the hands of a clock spinning at different speeds. Fast dimensions tell neighbors apart; slow ones track the big picture.
- Rotary position embeddings, RoPE (Su et al., 2021) - used by Llama, Mistral, Qwen and most open models today. Instead of adding a position vector, RoPE rotates each query and key by an angle proportional to its position. When a query at position meets a key at position , the rotations partly cancel, and the score depends only on the distance .
1import math
2
3def position_encoding(position, d_model):
4 values = []
5 for i in range(0, d_model, 2):
6 angle = position / 10000 ** (i / d_model)
7 values += [math.sin(angle), math.cos(angle)]
8 return values
9
10for position in [0, 1, 2, 50]:
11 print(position, [round(v, 3) for v in position_encoding(position, 4)])0 [0.0, 1.0, 0.0, 1.0] 1 [0.841, 0.54, 0.01, 1.0] 2 [0.909, -0.416, 0.02, 1.0] 50 [-0.262, 0.965, 0.479, 0.878]
Watch RoPE’s party trick. A 2-D query and key are rotated by their positions, then compared. Move both forward by the same amount and the score doesn’t budge; change the gap and it does. The model learns “the word two before me”, not “the word at slot 98” - exactly what language needs.
1import math
2
3def rotate(vector, position, theta=0.5):
4 angle = position * theta
5 x, y = vector
6 return (x * math.cos(angle) - y * math.sin(angle), x * math.sin(angle) + y * math.cos(angle))
7
8def dot(a, b):
9 return a[0] * b[0] + a[1] * b[1]
10
11query, key = (1.0, 0.5), (0.8, -0.2)
12for m, n in [(3, 1), (10, 8), (100, 98), (5, 1)]:
13 print(f"query at {m}, key at {n}: {dot(rotate(query, m), rotate(key, n)):.4f}")query at 3, key at 1: -0.1267 query at 10, key at 8: -0.1267 query at 100, key at 98: -0.1267 query at 5, key at 1: -0.8369
What grows when the context grows?
Long context isn’t free. During prefill (reading your prompt), every token attends to every earlier one: the score matrix has entries, so doubling the prompt quadruples that work. Kernels like FlashAttention avoid ever storing the full matrix, but the arithmetic is still quadratic. While generating, each new token attends to all cached tokens - linear per token - and the KV cache grows linearly too.
Some models cap the cost with sliding-window attention: each layer only looks back a fixed number of tokens (say 4,096), and information hops further back through the stacked layers.
Try it
Constant, linear or quadratic?
As the prompt length n grows, how does each cost grow? Assume a standard Transformer with a KV cache unless an item says otherwise.
“Memory for the model’s weights”
“Memory for the KV cache”
“Attention arithmetic to prefill the whole prompt”
“Attention work to generate one more token (with a KV cache)”
“The model’s parameter count”
“What an API bills you for the input tokens”
“Attention work per token with a 4,096-token sliding window, once n is past 4,096”
“The naive n × n matrix of attention scores, if you stored it”
Advertised context vs. used context
A model that accepts 200,000 tokens doesn’t necessarily use them well. Liu et al. (2023) found models answer best when the key fact sits at the start or end of a long prompt, and worst when it’s buried in the middle - lost in the middle.
The standard check is a needle in a haystack test: hide one fact (“the secret ingredient is cardamom”) at different depths of a long filler document and ask for it. Plot the hit rate by depth and context length, and you’ll see where your model goes blind. Harder variants hide several needles or need reasoning across them (RULER), because finding one quote is much easier than really reading.
Key takeaways
Attention is order-blind; positional encodings put order back in.
RoPE rotates queries and keys by position so scores depend on relative distance; interpolation tricks stretch it to longer windows.
Prefill attention grows quadratically with prompt length; the KV cache and per-token decoding grow linearly.
Long windows aren’t always used well - test with needles at different depths, and keep key facts and questions at the edges.
Lesson quiz
7 questions · pass with 5 correct · up to 50 XP
Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.
Practice: write Python
Write Python in the editor and run it against sample inputs. Python runs locally in your browser using a WebAssembly runtime.
Prove RoPE is relative
Line 1 is a query vector and line 2 a key vector, both with an even number of dimensions. Each following line is a pair of positions m n.
Apply RoPE to both: split each vector into pairs (v[0], v[1]), (v[2], v[3]), …, and rotate pair i (counting from 0 in steps of 2) by angle position × 10000 ** (-i / d), where d is the vector length. The query rotates by m, the key by n.
Print m=M n=N: S with the dot product of the rotated vectors to 4 decimals. Pairs with the same gap should agree.
- Same gap, same score
- Far along the sequence
Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.
Needle in a haystack report
Each input line is one needle-test result: depth found, where depth is how far into the document the needle was hidden (0–100, as a percent) and found is 1 or 0.
Group results into start (0–33), middle (34–66) and end (67–100), and print start: F/T found (P%) for each in that order (whole percent). Then print weakest: <group> - the lowest hit rate, ties going to the earlier group.
- Lost in the middle
- A model that fades at the end
Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.
Questions about this lesson
Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.
Loading posts…