Loading
0x10Lesson 2 of 6

Choose a tokenization strategy

Split sentences consistently and make punctuation handling explicit.

14 min 4-question quiz 1 code exercise
By the end of this lesson you can
  • Extract word tokens while describing what a tokenizer keeps or discards.

Tokenization divides text into units. A whitespace split is quick, but leaves punctuation attached ("hello,"). A regular expression can extract word-like spans. Real tokenizers handle contractions, writing systems, and subwords with more detailed rules; there is no single best boundary for every task.

example.py
1import re
2text = "NLP is useful, isn't it?"
3tokens = re.findall(r"[a-z]+(?:'[a-z]+)?", text.lower())
4print(tokens)
Output
['nlp', 'is', 'useful', "isn't", 'it']

Whether to keep punctuation or split contractions depends on what the model should learn. The regular expression in this example keeps a simple apostrophe contraction together.

Key takeaways

  • Extract word tokens while describing what a tokenizer keeps or discards.

  • Simple baselines help make ideas concrete.

  • Interpret language tools in context and check important results.

Lesson quiz

4 questions · pass with 3 correct · up to 50 XP

Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.

Practice: apply NLP with Python

Try each text-processing idea in Python, run it against sample inputs, and use the results to see where the method works or falls short.

Exercise 1

Extract word tokens

+25 XP

Read a line, lowercase it, extract words with re.findall(r"[a-z]+", text), and print them separated by one space. Ignore punctuation.

  • Punctuation
  • Repeated spaces
main.py
Loading editor…

Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.

Questions about this lesson

Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.

Loading posts…

Did you like the lesson? 😆👍
Consider a donation to support our work: