THE CORE AIM OF VOCABULARY MASTERY · MACHINE LEARNING VOCABULARY · FEATURES → TRAINING → MODEL → VALIDATION → GENERALISATION
Machine learning vocabulary is the language used to describe systems that learn patterns from data. Terms such as feature, label, training set, loss function, parameter, hyperparameter, overfitting, regularisation and generalisation matter because they describe distinct parts of the learning process.
The core aim of vocabulary mastery for machine learning vocabulary is model-building precision. Learners should understand where each term sits in the workflow, distinguish data concepts from model concepts, separate training performance from real-world generalisation and interpret evaluation language carefully.
This page is the Machine Learning Vocabulary owner inside the eduKateSG Vocabulary hub. For the wider analytical workflow, use Data Science Vocabulary. For broader AI language, use Artificial Intelligence Vocabulary.
Central proposition: Machine learning vocabulary is mastered when the learner can explain how data becomes a model and how the model is tested on data it did not simply memorise.
The 60-Second Machine Learning Vocabulary Router
- Data: feature, target, label, sample, class.
- Training: loss, optimisation, gradient, parameter.
- Model choice: algorithm, architecture, hyperparameter.
- Validation: validation set, test set, cross-validation.
- Generalisation: overfitting, underfitting, regularisation.
- Evaluation: accuracy, precision, recall, F1, ROC-AUC.
The Machine Learning Vocabulary Architecture
| Stage | Core terms | Core question |
|---|---|---|
| Input | feature, label, sample | What information is the model learning from? |
| Learning | loss, optimisation, gradient | How are parameters adjusted? |
| Model | parameter, hyperparameter | What is learned and what is chosen? |
| Validation | validation, cross-validation | How are choices tuned? |
| Testing | test set, benchmark | How does the final model perform? |
| Generalisation | overfitting, regularisation | Will performance hold on new data? |
Parameter and Hyperparameter Are Different
A parameter is learned from data during training. A hyperparameter is chosen or configured outside the learned parameter set, such as a learning rate or regularisation strength. The exact boundary can vary by method, but the practical distinction is important.
A Worked Example: Overfitting
A model that performs extremely well on training data but poorly on new data may be overfitting. The vocabulary should connect to the observable gap between seen-data and unseen-data performance rather than to a vague idea of “too much learning.”
A Worked Example: Regularisation
Regularisation refers to techniques that constrain or modify learning in ways intended to improve generalisation. Different methods regularise in different ways, so the term should be learned with the actual mechanism being used.
Supervised and Unsupervised Learning
Supervised learning uses examples with target information to learn a mapping. Unsupervised learning looks for structure without the same kind of labelled target. Other paradigms exist, but this contrast helps organise the field for beginners.
Machine Learning Vocabulary and Metrics
Evaluation words such as accuracy, precision, recall and F1 score answer different questions. The correct metric depends on class balance and the cost of different errors, so metric vocabulary should always be linked to the decision problem.
How to Learn Machine Learning Vocabulary
- Map terms onto one end-to-end model workflow.
- Use tiny datasets and simple models first.
- Compare confusing term pairs directly.
- Attach equations where they clarify meaning.
- Read training and evaluation output together.
- Explain the model in plain language.
- Use small projects to retrieve vocabulary from action.
Common Machine Learning Vocabulary Mistakes
Confusing parameter and hyperparameter
Repair: ask whether the value is learned from data or set externally.
Treating training accuracy as final performance
Repair: check validation and test behaviour.
Using one metric for every problem
Repair: connect metrics to error costs and task goals.
Calling every model “AI” without specifying method
Repair: name the actual model or learning approach when useful.
Frequently Asked Questions
What is machine learning vocabulary?
It is the specialised language used to describe data, training, models, validation, generalisation and evaluation in machine learning.
What terms should beginners learn first?
Start with feature, label, training set, model, parameter, loss, validation, test set and overfitting.
What is overfitting?
It is when a model fits training data too specifically and does not generalise well to unseen data.
What is the difference between a parameter and a hyperparameter?
Parameters are learned during training; hyperparameters are configuration choices that influence the learning process.
How can I learn ML vocabulary?
Use small model-building projects and attach every term to a step you can see in code or output.
Where This Article Fits in the eduKateSG Vocabulary Ecosystem
- Vocabulary Hub — the broad route.
- Data Science Vocabulary — wider analytical workflow.
- Artificial Intelligence Vocabulary — broader AI system language.
- Research Vocabulary — evidence and evaluation.
- Technical Vocabulary — specialist-language method.
The Machine Learning Vocabulary Standard
Machine learning vocabulary reaches its core aim when terms let the learner trace the path from data to model to evaluation without confusing what was learned, what was configured and what was merely measured.
That is the standard: vocabulary precise enough to build and judge models.
