Loading
0xB0Lesson 12 of 15

Represent meaning with word vectors

Place words as points in space so that similar words sit close together.

22 min 6-question quiz 1 code exercise
By the end of this lesson you can
  • Explain the distributional hypothesis behind word vectors
  • Compare words with cosine similarity
  • Solve analogies with vector arithmetic - and spot the bias it can reveal

Counts treat every word as unrelated: to a bag of words, cat is as different from kitten as from carburetor. Word vectors (also called embeddings) fix that by giving each word a list of numbers - a point in space - so that words with similar meanings land near each other.

Where do the numbers come from? From the distributional hypothesis, summed up by the linguist J. R. Firth: “You shall know a word by the company it keeps.” Words that appear in similar contexts (coffee and tea both get brewed, poured, hot) get similar vectors. Methods like word2vec and GloVe learn these vectors from billions of words of text.

Closeness as an angle

Cosine similarity compares the directions two vectors point in: 1 means the same direction, 0 means unrelated (at right angles). It ignores length, which mostly reflects how frequent a word is rather than what it means.

Try it

Word vector map

Real word vectors have hundreds of dimensions; this toy map has two, so you can see them.

  • In Compare mode, click king and queen, then king and woman. Which pair points more the same way?
  • In Solve analogies mode, guess the answer before it is revealed. The blue arrow is the “move” from the second word to the first; the green arrow repeats that move from the third word.
kingqueenprinceprincessuncleauntmanwoman

Click two words on the map. The arrows point from the origin to each word; the narrower the angle between them, the higher the cosine similarity.

Nothing selected yet.

analogy.py
vectors = {"king": (8, 9), "man": (8, 2), "woman": (2, 2), "queen": (2, 9)}
target = tuple(king - man + woman for king, man, woman in zip(vectors["king"], vectors["man"], vectors["woman"]))
print(target, target == vectors["queen"])
Output
(2, 9) True

Key takeaways

  • Word vectors place words in space; similar contexts → nearby points (distributional hypothesis).

  • Cosine similarity compares directions: near 1 = similar, near 0 = unrelated.

  • Vector arithmetic can capture relationships (king − man + woman ≈ queen) - and social biases.

Lesson quiz

6 questions · pass with 5 correct · up to 50 XP

Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.

Practice: apply NLP with Python

Try each text-processing idea in Python, run it against sample inputs, and use the results to see where the method works or falls short.

Exercise 1

Find the most similar word

+25 XP

The first line is N. Each of the next N lines is a word and its vector, like king 8 9. The last line is a query word. Print the other word with the highest cosine similarity to the query (ties: alphabetically first).

  • Royal neighbors
  • Opposite corner
main.py
Loading editor…

Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.

Questions about this lesson

Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.

Loading posts…

Did you like the lesson? 😆👍
Consider a donation to support our work: