Turn text into tokens
Normalize text and split it into pieces a language system can process.
- Normalize a sentence and explain what tokens represent.
Natural Language Processing (NLP) studies ways to work with human language. A computer starts with text as characters; many NLP tasks split it into tokens. Tokens may be words, punctuation, characters, or subword pieces. Subword tokenizers can represent uncommon words by combining familiar pieces.
text = " Cats, DOGS, and cats! "
normalized = " ".join(text.lower().split())
print(normalized)cats, dogs, and cats!
Normalization makes text more consistent. Lowercasing can combine Cat and cat; trimming whitespace removes padding. The right choices depend on the task: punctuation and capitalization can carry useful information, so do not remove them automatically.
Key takeaways
Normalize a sentence and explain what tokens represent.
Simple baselines help make ideas concrete.
Interpret language tools in context and check important results.
Lesson quiz
4 questions · pass with 3 correct · up to 50 XP
Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.
Practice: apply NLP with Python
Try each text-processing idea in Python, run it against sample inputs, and use the results to see where the method works or falls short.
Normalize a sentence
Read one line, lowercase it, trim the ends, and collapse repeated whitespace to one space. Print the result.
- Mixed case and spaces
- Already tidy
Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.
Questions about this lesson
Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.
Loading posts…