Represent documents with word counts
Build a bag-of-words view and learn what frequency leaves out.
- Count tokens and compare frequency with document importance.
A bag-of-words representation counts tokens and ignores their order. It is a useful baseline for search and classification, but sentences with the same words in different orders look identical. Frequent words like “the” may be less informative than terms that distinguish one document from a collection.
from collections import Counter
words = "cats chase cats".split()
print(Counter(words))Counter({'cats': 2, 'chase': 1})TF-IDF combines a term’s frequency in one document with how uncommon it is across a collection. A term used often in one document but rarely in others can receive a higher weight. It is a useful baseline, not an understanding of meaning.
Key takeaways
Count tokens and compare frequency with document importance.
Simple baselines help make ideas concrete.
Interpret language tools in context and check important results.
Lesson quiz
4 questions · pass with 3 correct · up to 50 XP
Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.
Practice: apply NLP with Python
Try each text-processing idea in Python, run it against sample inputs, and use the results to see where the method works or falls short.
Count the words
Read text, lowercase it, extract alphabetic words, then print each distinct word and count in alphabetical order as word: count.
- Repeated words
- Three tokens
Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.
Questions about this lesson
Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.
Loading posts…