Loading
0x00Lesson 1 of 6

Turn text into tokens

Normalize text and split it into pieces a language system can process.

14 min 4-question quiz 1 code exercise
By the end of this lesson you can
  • Normalize a sentence and explain what tokens represent.

Natural Language Processing (NLP) studies ways to work with human language. A computer starts with text as characters; many NLP tasks split it into tokens. Tokens may be words, punctuation, characters, or subword pieces. Subword tokenizers can represent uncommon words by combining familiar pieces.

example.py
text = "  Cats, DOGS, and cats!  "
normalized = " ".join(text.lower().split())
print(normalized)
Output
cats, dogs, and cats!

Normalization makes text more consistent. Lowercasing can combine Cat and cat; trimming whitespace removes padding. The right choices depend on the task: punctuation and capitalization can carry useful information, so do not remove them automatically.

Key takeaways

  • Normalize a sentence and explain what tokens represent.

  • Simple baselines help make ideas concrete.

  • Interpret language tools in context and check important results.

Lesson quiz

4 questions · pass with 3 correct · up to 50 XP

Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.

Practice: apply NLP with Python

Try each text-processing idea in Python, run it against sample inputs, and use the results to see where the method works or falls short.

Exercise 1

Normalize a sentence

+25 XP

Read one line, lowercase it, trim the ends, and collapse repeated whitespace to one space. Print the result.

  • Mixed case and spaces
  • Already tidy
main.py
Loading editor…

Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.

Questions about this lesson

Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.

Loading posts…

Did you like the lesson? 😆👍
Consider a donation to support our work: