THE CORE AIM OF VOCABULARY MASTERY · LARGE LANGUAGE MODEL VOCABULARY · TOKEN → CONTEXT → ATTENTION → INFERENCE → EVALUATION
Large language model vocabulary is the language used to describe models trained to process and generate sequences of language tokens at large scale. Terms such as token, context window, attention, transformer, parameter, pretraining, fine-tuning, inference, embedding and hallucination matter because LLM behaviour is easier to understand when training, runtime context and output generation are kept separate.
The core aim of vocabulary mastery for large language model vocabulary is model-behaviour clarity. Learners should be able to explain what tokens are, what information fits in context, how attention relates information across a sequence, what training changes and what happens when the trained model performs inference.
This page is the Large Language Model Vocabulary owner inside the eduKateSG Vocabulary hub. It connects directly to Generative AI Vocabulary, Natural Language Processing Vocabulary and Machine Learning Vocabulary.
Central proposition: LLM vocabulary is mastered when the learner can distinguish what was learned during training from what information is available during a particular inference.
The 60-Second LLM Vocabulary Router
- Units: token, tokenizer, vocabulary, sequence.
- Architecture: transformer, attention, layer, parameter.
- Training: pretraining, objective, fine-tuning, alignment.
- Runtime: prompt, context window, inference, completion.
- Representation: embedding, vector, hidden state.
- Evaluation: benchmark, factuality, hallucination, robustness.
Token and Word Are Different
A token is a model-processing unit, not necessarily a whole word. Tokenizers may divide one word into several subwords, and token counts therefore do not map perfectly to word counts.
Context Window
The context window is the amount of tokenized information a model can consider in one inference, subject to the implementation. It may include instructions, conversation history, retrieved documents and the output being generated.
Attention
Attention is a mechanism that lets a transformer compute relationships among positions in a sequence. It helps the model condition each representation on relevant surrounding information; it should not be confused with human attention or awareness.
Pretraining and Fine-Tuning
Pretraining learns broad statistical patterns from large datasets. Fine-tuning further adjusts a model for particular behaviours, domains or tasks. Runtime prompting is different again: it supplies temporary context without necessarily changing model parameters.
Inference
Inference is the runtime process in which a trained model receives input and produces outputs. This is distinct from training, where parameters are adjusted.
How to Learn LLM Vocabulary
- Inspect tokenization examples.
- Separate training from inference.
- Map prompt and retrieved material into context.
- Study transformer and attention diagrams conceptually.
- Compare embeddings with generated text.
- Evaluate outputs against evidence.
Common LLM Vocabulary Mistakes
Treating token as word
Repair: inspect actual tokenizer segmentation.
Confusing context with training data
Repair: distinguish temporary runtime information from learned parameters.
Treating attention as consciousness
Repair: use the technical meaning: a computation over representations.
Assuming larger context guarantees better answers
Repair: relevance, organisation and model behaviour still matter.
Frequently Asked Questions
What is LLM vocabulary?
It is the specialised language used for tokens, transformers, attention, context, training, inference, embeddings and evaluation in large language models.
What terms should beginners learn first?
Start with token, tokenizer, context window, transformer, attention, parameter, pretraining, fine-tuning and inference.
What is a context window?
It is the token-limited information available to a model during a particular inference.
The Large Language Model Vocabulary Standard
LLM vocabulary reaches its core aim when the learner can explain what was learned during training, what enters context at runtime and how the model produces and evaluates an output.
That is the standard: model language precise enough to separate architecture, training and use.
