THE CORE AIM OF VOCABULARY MASTERY · TOKENIZATION · TEXT → UNITS → IDS → MODEL → TEXT
Tokenization vocabulary unlocks a surprisingly common source of confusion in modern AI: a token is not necessarily a word. Words may be split into pieces; punctuation, spaces and special markers may be represented in their own ways; and a model handles numerical token IDs rather than reading letters as a person does. Understanding tokens, tokenizers, subwords, vocabulary, byte pair encoding, WordPiece, Unigram and decoding makes language-model behaviour much easier to explain.
The core aim of vocabulary mastery for tokenization is to connect ordinary language with the computational units used by models. Learners should be able to describe how text is segmented, how segments map to IDs, why token counts differ from word counts, and how special tokens help organise model tasks.
This page belongs to the eduKateSG Vocabulary Hub, linking human language study to Natural Language Processing Vocabulary and Large Language Model Vocabulary.
Central proposition: Tokenization vocabulary is mastered when you can explain the difference between the text a person sees and the units and IDs a model actually receives.
The 60-second tokenization router
- Segmentation: word token, subword, character, byte.
- Vocabulary: token vocabulary, token ID, unknown token.
- Algorithms: BPE, WordPiece, Unigram.
- Structure: BOS, EOS, padding token, separator.
- Processing: tokenization, encoding, truncation, attention mask.
- Return: decoding, detokenization, special-token handling.
Words, subwords and tokens compared
| Concept | Example or role | Important caution |
|---|---|---|
| Word | Students may perceive ‘unbelievable’ as one word. | It can take multiple model tokens. |
| Subword | A possible split is ‘un’ + ‘believ’ + ‘able’. | Actual splits vary by tokenizer; this is illustrative. |
| Token ID | An integer assigned to a vocabulary entry. | IDs are model/tokenizer-specific. |
| Special token | A marker for boundaries or special functions. | Different models use different markers. |
| Byte-level token | A unit related to encoded byte sequences. | Character count and byte count are different. |
Why subword tokenization exists
A vocabulary that assigns one entry to every complete word can become enormous and struggles with new names or rare forms. Character-only systems cover many forms but often create longer sequences. Subword methods occupy a practical middle ground: common units can remain compact while unfamiliar words are assembled from smaller pieces.
Byte pair encoding, WordPiece and Unigram
Byte pair encoding (BPE)
BPE learns a vocabulary by repeatedly combining common neighbouring units according to its algorithm and training data. A frequent sequence may receive its own token while a less common one remains split. Byte-level BPE begins from byte representations to help handle a broad range of text.
WordPiece
WordPiece is another family of subword-tokenization methods. Its training objective and merge-selection procedures differ from BPE. BERT-family examples commonly use WordPiece-style segmentation, but the exact tokenizer is determined by the model.
Unigram
Unigram methods begin with a candidate set of subwords and select a useful vocabulary probabilistically. A text sequence may admit more than one possible segmentation, with the tokenizer choosing or sampling among them according to its configuration.
A worked example: two tokenizers, one sentence
Consider the sentence The children re-read the story. One tokenizer might treat children as a single unit and split re-read into several units; another may segment them differently. Neither segmentation reveals whether the child understood the sentence. Model tokens are engineering units—not a measure of human vocabulary mastery.
Try the same experiment with a name, an emoji, a long technical term and a multilingual phrase. The exercise is especially useful for Singapore learners moving between English and languages with different writing and spacing systems.
Special tokens and input structure
Special tokens can mark beginnings, endings, separators, padding or task-specific transitions. The symbols and their meanings vary by model. A common mistake is copying a special token from one model family and assuming another model uses it identically.
Context window versus word count
A context window is typically specified in model tokens, not English words. The number of tokens per sentence depends on language, punctuation, formatting and the tokenizer. A prompt with many unusual identifiers or code fragments may use a different token budget from ordinary prose of the same apparent length.
Tokenization versus embeddings
Tokenization splits and encodes the input as token IDs. An embedding layer converts those IDs into numerical representations used by the network. Token IDs are discrete symbols; embeddings are learned vectors. The processes follow one another but are not synonyms.
How to practise tokenization vocabulary
- Find a tokenizer viewer or educational example and examine how a sentence is segmented.
- Count visible words and compare them with token count.
- Try punctuation, contractions, repeated spaces and multilingual text.
- Identify any start, end or separator markers.
- Describe what encoding and decoding each do.
- Explain why a model’s token count should not be equated with a learner’s vocabulary size.
Common mistakes and repairs
- One token equals one word: check subword segmentation before making this claim.
- Tokenization equals meaning: segmentation is a representation step, not full comprehension.
- Every tokenizer gives the same count: vocabulary and algorithms differ.
- All languages cost the same number of tokens: different scripts and vocabularies can yield different token counts.
- Token ID equals embedding: separate the index from its learned vector representation.
Frequently asked questions
What is tokenization in NLP?
It is the process of dividing text or other model input into units that can be mapped to the model’s vocabulary.
What is a subword token?
It is a token representing part of a word or text sequence, allowing compact vocabularies to represent less familiar forms.
Is a token the same as a word?
Not necessarily. A word may become several tokens, and punctuation or spacing may contribute to tokenization.
Why do token counts matter?
They affect context limits, processing work and, in some systems, usage-based costs.
What is detokenization?
It is the reconstruction of readable output from token IDs or token representations.
Reference and further reading
For technical comparisons of BPE, WordPiece, Unigram and related approaches, see the Hugging Face Tokenization Algorithms guide. Implementation details depend on the tokenizer and model.
Where this fits in eduKateSG
- Vocabulary Hub
- Natural Language Processing Vocabulary
- Large Language Model Vocabulary
- Embeddings Vocabulary
The tokenization vocabulary standard
A learner has mastered tokenization vocabulary when they can show how readable text becomes tokens and IDs, explain the main subword approaches and distinguish model token count from human words and meaning.
