VIEW THIS AS

Auto mode follows the Route Engine until you choose a viewpoint.

YOU ARE HERE

ROUTE CHECK

CONNECTED TO

WHAT NEXT

Use the canonical route for this room, or HELP if you are unsure.

The Core Aim of Vocabulary Mastery | Tokenization Vocabulary

Three learners review open books together at a classroom table, with stacks of textbooks, stationery and a whiteboard in the bright room.

THE CORE AIM OF VOCABULARY MASTERY · TOKENIZATION · TEXT → UNITS → IDS → MODEL → TEXT

Tokenization vocabulary unlocks a surprisingly common source of confusion in modern AI: a token is not necessarily a word. Words may be split into pieces; punctuation, spaces and special markers may be represented in their own ways; and a model handles numerical token IDs rather than reading letters as a person does. Understanding tokens, tokenizers, subwords, vocabulary, byte pair encoding, WordPiece, Unigram and decoding makes language-model behaviour much easier to explain.

The core aim of vocabulary mastery for tokenization is to connect ordinary language with the computational units used by models. Learners should be able to describe how text is segmented, how segments map to IDs, why token counts differ from word counts, and how special tokens help organise model tasks.

This page belongs to the eduKateSG Vocabulary Hub, linking human language study to Natural Language Processing Vocabulary and Large Language Model Vocabulary.

Central proposition: Tokenization vocabulary is mastered when you can explain the difference between the text a person sees and the units and IDs a model actually receives.

The 60-second tokenization router

  • Segmentation: word token, subword, character, byte.
  • Vocabulary: token vocabulary, token ID, unknown token.
  • Algorithms: BPE, WordPiece, Unigram.
  • Structure: BOS, EOS, padding token, separator.
  • Processing: tokenization, encoding, truncation, attention mask.
  • Return: decoding, detokenization, special-token handling.

Words, subwords and tokens compared

ConceptExample or roleImportant caution
WordStudents may perceive ‘unbelievable’ as one word.It can take multiple model tokens.
SubwordA possible split is ‘un’ + ‘believ’ + ‘able’.Actual splits vary by tokenizer; this is illustrative.
Token IDAn integer assigned to a vocabulary entry.IDs are model/tokenizer-specific.
Special tokenA marker for boundaries or special functions.Different models use different markers.
Byte-level tokenA unit related to encoded byte sequences.Character count and byte count are different.

Why subword tokenization exists

A vocabulary that assigns one entry to every complete word can become enormous and struggles with new names or rare forms. Character-only systems cover many forms but often create longer sequences. Subword methods occupy a practical middle ground: common units can remain compact while unfamiliar words are assembled from smaller pieces.

Byte pair encoding, WordPiece and Unigram

Byte pair encoding (BPE)

BPE learns a vocabulary by repeatedly combining common neighbouring units according to its algorithm and training data. A frequent sequence may receive its own token while a less common one remains split. Byte-level BPE begins from byte representations to help handle a broad range of text.

WordPiece

WordPiece is another family of subword-tokenization methods. Its training objective and merge-selection procedures differ from BPE. BERT-family examples commonly use WordPiece-style segmentation, but the exact tokenizer is determined by the model.

Unigram

Unigram methods begin with a candidate set of subwords and select a useful vocabulary probabilistically. A text sequence may admit more than one possible segmentation, with the tokenizer choosing or sampling among them according to its configuration.

A worked example: two tokenizers, one sentence

Consider the sentence The children re-read the story. One tokenizer might treat children as a single unit and split re-read into several units; another may segment them differently. Neither segmentation reveals whether the child understood the sentence. Model tokens are engineering units—not a measure of human vocabulary mastery.

Try the same experiment with a name, an emoji, a long technical term and a multilingual phrase. The exercise is especially useful for Singapore learners moving between English and languages with different writing and spacing systems.

Special tokens and input structure

Special tokens can mark beginnings, endings, separators, padding or task-specific transitions. The symbols and their meanings vary by model. A common mistake is copying a special token from one model family and assuming another model uses it identically.

Context window versus word count

A context window is typically specified in model tokens, not English words. The number of tokens per sentence depends on language, punctuation, formatting and the tokenizer. A prompt with many unusual identifiers or code fragments may use a different token budget from ordinary prose of the same apparent length.

Tokenization versus embeddings

Tokenization splits and encodes the input as token IDs. An embedding layer converts those IDs into numerical representations used by the network. Token IDs are discrete symbols; embeddings are learned vectors. The processes follow one another but are not synonyms.

How to practise tokenization vocabulary

  • Find a tokenizer viewer or educational example and examine how a sentence is segmented.
  • Count visible words and compare them with token count.
  • Try punctuation, contractions, repeated spaces and multilingual text.
  • Identify any start, end or separator markers.
  • Describe what encoding and decoding each do.
  • Explain why a model’s token count should not be equated with a learner’s vocabulary size.

Common mistakes and repairs

  • One token equals one word: check subword segmentation before making this claim.
  • Tokenization equals meaning: segmentation is a representation step, not full comprehension.
  • Every tokenizer gives the same count: vocabulary and algorithms differ.
  • All languages cost the same number of tokens: different scripts and vocabularies can yield different token counts.
  • Token ID equals embedding: separate the index from its learned vector representation.

Frequently asked questions

What is tokenization in NLP?

It is the process of dividing text or other model input into units that can be mapped to the model’s vocabulary.

What is a subword token?

It is a token representing part of a word or text sequence, allowing compact vocabularies to represent less familiar forms.

Is a token the same as a word?

Not necessarily. A word may become several tokens, and punctuation or spacing may contribute to tokenization.

Why do token counts matter?

They affect context limits, processing work and, in some systems, usage-based costs.

What is detokenization?

It is the reconstruction of readable output from token IDs or token representations.

Reference and further reading

For technical comparisons of BPE, WordPiece, Unigram and related approaches, see the Hugging Face Tokenization Algorithms guide. Implementation details depend on the tokenizer and model.

Where this fits in eduKateSG

The tokenization vocabulary standard

A learner has mastered tokenization vocabulary when they can show how readable text becomes tokens and IDs, explain the main subword approaches and distinguish model token count from human words and meaning.

Discover more from eduKate Singapore

Subscribe now to keep reading and get access to the full archive.

Continue reading