Loading
0x70Lesson 8 of 16

Positions and long context

Find out how a Transformer knows word order, why RoPE took over, and why a million-token window doesn’t mean a million tokens of attention.

25 min 7-question quiz 2 code exercises
By the end of this lesson you can
  • Explain why attention alone ignores word order, and how positional encodings fix it
  • Show that rotary embeddings (RoPE) make attention depend on relative distance
  • Describe what grows with context length, and test whether a model really uses its context

“The dog bit the man” and “the man bit the dog” use the same tokens. Yet attention, as you computed it in Explore Transformers and attention, is just dot products and weighted sums - it has no idea which token came first. Shuffle the input tokens and the outputs simply shuffle along with them. A bag of words with very good taste.

So every Transformer adds position information. The idea is simple; the details decide how well a model handles long documents.

order_blind.py
1import numpy as np
2rng = np.random.default_rng(0)
3x = rng.normal(size=(4, 8))       # 4 tokens, 8 dimensions each
4
5def attend(x):
6    scores = x @ x.T / np.sqrt(x.shape[1])
7    weights = np.exp(scores - scores.max(axis=1, keepdims=True))
8    weights /= weights.sum(axis=1, keepdims=True)
9    return weights @ x
10
11order = [2, 0, 3, 1]
12print("shuffling tokens just shuffles the outputs:", np.allclose(attend(x)[order], attend(x[order])))
Output
shuffling tokens just shuffles the outputs: True

Three ways to say “where”

  1. Learned absolute positions (GPT-2): a trainable vector for position 0, 1, 2… added to each token embedding. Simple - but there is simply no vector for position 2,049 if training stopped at 2,048.
  2. Sinusoidal positions (the original 2017 Transformer): each position gets a fixed pattern of sines and cosines at different frequencies, like the hands of a clock spinning at different speeds. Fast dimensions tell neighbors apart; slow ones track the big picture.
  3. Rotary position embeddings, RoPE (Su et al., 2021) - used by Llama, Mistral, Qwen and most open models today. Instead of adding a position vector, RoPE rotates each query and key by an angle proportional to its position. When a query at position mm meets a key at position nn, the rotations partly cancel, and the score depends only on the distance m−nm - n.
sinusoidal.py
1import math
2
3def position_encoding(position, d_model):
4    values = []
5    for i in range(0, d_model, 2):
6        angle = position / 10000 ** (i / d_model)
7        values += [math.sin(angle), math.cos(angle)]
8    return values
9
10for position in [0, 1, 2, 50]:
11    print(position, [round(v, 3) for v in position_encoding(position, 4)])
Output
0 [0.0, 1.0, 0.0, 1.0]
1 [0.841, 0.54, 0.01, 1.0]
2 [0.909, -0.416, 0.02, 1.0]
50 [-0.262, 0.965, 0.479, 0.878]

Watch RoPE’s party trick. A 2-D query and key are rotated by their positions, then compared. Move both forward by the same amount and the score doesn’t budge; change the gap and it does. The model learns “the word two before me”, not “the word at slot 98” - exactly what language needs.

rope.py
1import math
2
3def rotate(vector, position, theta=0.5):
4    angle = position * theta
5    x, y = vector
6    return (x * math.cos(angle) - y * math.sin(angle), x * math.sin(angle) + y * math.cos(angle))
7
8def dot(a, b):
9    return a[0] * b[0] + a[1] * b[1]
10
11query, key = (1.0, 0.5), (0.8, -0.2)
12for m, n in [(3, 1), (10, 8), (100, 98), (5, 1)]:
13    print(f"query at {m}, key at {n}: {dot(rotate(query, m), rotate(key, n)):.4f}")
Output
query at 3, key at 1: -0.1267
query at 10, key at 8: -0.1267
query at 100, key at 98: -0.1267
query at 5, key at 1: -0.8369

What grows when the context grows?

Long context isn’t free. During prefill (reading your prompt), every token attends to every earlier one: the score matrix has n2n^2 entries, so doubling the prompt quadruples that work. Kernels like FlashAttention avoid ever storing the full matrix, but the arithmetic is still quadratic. While generating, each new token attends to all nn cached tokens - linear per token - and the KV cache grows linearly too.

Some models cap the cost with sliding-window attention: each layer only looks back a fixed number of tokens (say 4,096), and information hops further back through the stacked layers.

Try it

Constant, linear or quadratic?

As the prompt length n grows, how does each cost grow? Assume a standard Transformer with a KV cache unless an item says otherwise.

0 of 8 sortedScore 0/0
  • “Memory for the model’s weights”

  • “Memory for the KV cache”

  • “Attention arithmetic to prefill the whole prompt”

  • “Attention work to generate one more token (with a KV cache)”

  • “The model’s parameter count”

  • “What an API bills you for the input tokens”

  • “Attention work per token with a 4,096-token sliding window, once n is past 4,096”

  • “The naive n × n matrix of attention scores, if you stored it”

Advertised context vs. used context

A model that accepts 200,000 tokens doesn’t necessarily use them well. Liu et al. (2023) found models answer best when the key fact sits at the start or end of a long prompt, and worst when it’s buried in the middle - lost in the middle.

The standard check is a needle in a haystack test: hide one fact (“the secret ingredient is cardamom”) at different depths of a long filler document and ask for it. Plot the hit rate by depth and context length, and you’ll see where your model goes blind. Harder variants hide several needles or need reasoning across them (RULER), because finding one quote is much easier than really reading.

Key takeaways

  • Attention is order-blind; positional encodings put order back in.

  • RoPE rotates queries and keys by position so scores depend on relative distance; interpolation tricks stretch it to longer windows.

  • Prefill attention grows quadratically with prompt length; the KV cache and per-token decoding grow linearly.

  • Long windows aren’t always used well - test with needles at different depths, and keep key facts and questions at the edges.

Lesson quiz

7 questions · pass with 5 correct · up to 50 XP

Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.

Practice: write Python

Write Python in the editor and run it against sample inputs. Python runs locally in your browser using a WebAssembly runtime.

Exercise 1

Prove RoPE is relative

+25 XP

Line 1 is a query vector and line 2 a key vector, both with an even number of dimensions. Each following line is a pair of positions m n.

Apply RoPE to both: split each vector into pairs (v[0], v[1]), (v[2], v[3]), …, and rotate pair i (counting from 0 in steps of 2) by angle position × 10000 ** (-i / d), where d is the vector length. The query rotates by m, the key by n.

Print m=M n=N: S with the dot product of the rotated vectors to 4 decimals. Pairs with the same gap should agree.

  • Same gap, same score
  • Far along the sequence
main.py
Loading editor…

Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.

Exercise 2

Needle in a haystack report

+25 XP

Each input line is one needle-test result: depth found, where depth is how far into the document the needle was hidden (0–100, as a percent) and found is 1 or 0.

Group results into start (0–33), middle (34–66) and end (67–100), and print start: F/T found (P%) for each in that order (whole percent). Then print weakest: <group> - the lowest hit rate, ties going to the earlier group.

  • Lost in the middle
  • A model that fades at the end
main.py
Loading editor…

Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.

Questions about this lesson

Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.

Loading posts…

Did you like the lesson? 😆👍
Consider a donation to support our work: