THE CORE AIM OF VOCABULARY MASTERY · COMPUTER VISION VOCABULARY · IMAGE → FEATURE → DETECT → SEGMENT → INTERPRET
Computer vision vocabulary is the language used to describe how computational systems analyse images and video. Terms such as pixel, image classification, object detection, bounding box, segmentation, feature, convolution, annotation and confidence score matter because vision tasks differ according to what the system must identify and where it must identify it.
The core aim of vocabulary mastery for computer vision vocabulary is visual-task clarity. Learners should be able to distinguish classifying an entire image from locating objects, separating pixels into regions, tracking movement or estimating geometry—and explain how labels and evaluation metrics correspond to the task.
This page is the Computer Vision Vocabulary owner inside the eduKateSG Vocabulary hub. For neural networks, use Deep Learning Vocabulary. For wider AI, use Artificial Intelligence Vocabulary.
Central proposition: Computer vision vocabulary is mastered when the learner can state what visual information the model must produce and how that output is compared with labelled reality.
The 60-Second Computer Vision Vocabulary Router
- Image: pixel, channel, resolution, frame.
- Classification: class, label, confidence score.
- Detection: object, bounding box, localization.
- Segmentation: mask, semantic segmentation, instance segmentation.
- Training: annotation, augmentation, dataset, feature.
- Evaluation: precision, recall, IoU, accuracy.
The Computer Vision Vocabulary Architecture
| Task | Core terms | Core question |
|---|---|---|
| Classification | label, class, confidence | What is in the image? |
| Detection | bounding box, object | What objects are present and where? |
| Segmentation | mask, pixel class | Which pixels belong to what? |
| Tracking | frame, track, identity | Where does an object move over time? |
| Geometry | depth, pose, keypoint | Where are important spatial structures? |
| Evaluation | IoU, precision, recall | How closely does output match labels? |
Classification and Detection Are Different
Image classification assigns a label to an image or region. Object detection identifies objects and estimates their locations, commonly with bounding boxes. Detection therefore answers both “what?” and “where?”
A Worked Example: Segmentation
Segmentation assigns labels at the pixel level. Semantic segmentation labels pixels by category, while instance segmentation also distinguishes separate objects of the same category.
A Worked Example: IoU
Intersection over Union, or IoU, compares overlap between a predicted region and a reference region. It is commonly used in detection and segmentation evaluation because location quality matters, not only class identity.
Annotation and Ground Truth
Annotation is the process of creating labels for training or evaluation. Ground truth refers to the reference labels treated as the correct target for a task, although human annotation itself can contain ambiguity or error.
Data Augmentation
Data augmentation creates modified training examples through transformations such as cropping, flipping or colour changes. The goal is to expose the model to useful variation while preserving the label meaning.
How to Learn Computer Vision Vocabulary
- Compare classification, detection and segmentation on one image.
- Draw bounding boxes manually.
- Create simple masks.
- Inspect confidence scores.
- Compare predictions with annotations.
- Calculate overlap conceptually.
- Connect image transformations to augmentation.
Common Computer Vision Vocabulary Mistakes
Calling every vision task recognition
Repair: identify whether the task is classification, detection, segmentation, tracking or geometry.
Confusing bounding box and mask
Repair: a box approximates location; a mask labels pixels.
Treating confidence as probability of truth
Repair: understand it as model output that requires calibration and evaluation.
Assuming ground truth is infallible
Repair: account for annotation quality and ambiguity.
Frequently Asked Questions
What is computer vision vocabulary?
It is the specialised language used for images, classification, detection, segmentation, annotation and visual-model evaluation.
What terms should beginners learn first?
Start with pixel, image, label, classification, detection, bounding box, segmentation, annotation and confidence score.
What is object detection?
It is a task that identifies object categories and estimates where those objects appear in an image or frame.
What is segmentation?
It is a task that assigns labels to image pixels or regions.
How can I learn computer vision vocabulary?
Use real images and label what classification, detection and segmentation outputs would look like.
Where This Article Fits in the eduKateSG Vocabulary Ecosystem
- Vocabulary Hub — the broad route.
- Deep Learning Vocabulary — neural networks.
- Artificial Intelligence Vocabulary — broader AI.
- Robotics Vocabulary — perception in physical systems.
- Generative AI Vocabulary — multimodal generation.
The Computer Vision Vocabulary Standard
Computer vision vocabulary reaches its core aim when the learner can name the visual task, describe its output and explain how that output is evaluated against labelled evidence.
That is the standard: visual AI language precise enough to distinguish seeing from merely naming.
