Tokenization splits text into smaller units called tokens (words, subwords, characters). It is the first step in processing language for NLP models.