VIEW THIS AS

Auto mode follows the Route Engine until you choose a viewpoint.

YOU ARE HERE

ROUTE CHECK

CONNECTED TO

WHAT NEXT

Use the canonical route for this room, or HELP if you are unsure.

The Core Aim of Vocabulary Mastery | Computer Vision Vocabulary

eduKate Secondary students reviewing open books for How Super Intelligence Works: Embeddings.

THE CORE AIM OF VOCABULARY MASTERY · COMPUTER VISION VOCABULARY · IMAGE → FEATURE → DETECT → SEGMENT → INTERPRET

Computer vision vocabulary is the language used to describe how computational systems analyse images and video. Terms such as pixel, image classification, object detection, bounding box, segmentation, feature, convolution, annotation and confidence score matter because vision tasks differ according to what the system must identify and where it must identify it.

The core aim of vocabulary mastery for computer vision vocabulary is visual-task clarity. Learners should be able to distinguish classifying an entire image from locating objects, separating pixels into regions, tracking movement or estimating geometry—and explain how labels and evaluation metrics correspond to the task.

This page is the Computer Vision Vocabulary owner inside the eduKateSG Vocabulary hub. For neural networks, use Deep Learning Vocabulary. For wider AI, use Artificial Intelligence Vocabulary.

Central proposition: Computer vision vocabulary is mastered when the learner can state what visual information the model must produce and how that output is compared with labelled reality.


The 60-Second Computer Vision Vocabulary Router

  • Image: pixel, channel, resolution, frame.
  • Classification: class, label, confidence score.
  • Detection: object, bounding box, localization.
  • Segmentation: mask, semantic segmentation, instance segmentation.
  • Training: annotation, augmentation, dataset, feature.
  • Evaluation: precision, recall, IoU, accuracy.

The Computer Vision Vocabulary Architecture

TaskCore termsCore question
Classificationlabel, class, confidenceWhat is in the image?
Detectionbounding box, objectWhat objects are present and where?
Segmentationmask, pixel classWhich pixels belong to what?
Trackingframe, track, identityWhere does an object move over time?
Geometrydepth, pose, keypointWhere are important spatial structures?
EvaluationIoU, precision, recallHow closely does output match labels?

Classification and Detection Are Different

Image classification assigns a label to an image or region. Object detection identifies objects and estimates their locations, commonly with bounding boxes. Detection therefore answers both “what?” and “where?”

A Worked Example: Segmentation

Segmentation assigns labels at the pixel level. Semantic segmentation labels pixels by category, while instance segmentation also distinguishes separate objects of the same category.

A Worked Example: IoU

Intersection over Union, or IoU, compares overlap between a predicted region and a reference region. It is commonly used in detection and segmentation evaluation because location quality matters, not only class identity.

Annotation and Ground Truth

Annotation is the process of creating labels for training or evaluation. Ground truth refers to the reference labels treated as the correct target for a task, although human annotation itself can contain ambiguity or error.

Data Augmentation

Data augmentation creates modified training examples through transformations such as cropping, flipping or colour changes. The goal is to expose the model to useful variation while preserving the label meaning.

How to Learn Computer Vision Vocabulary

  • Compare classification, detection and segmentation on one image.
  • Draw bounding boxes manually.
  • Create simple masks.
  • Inspect confidence scores.
  • Compare predictions with annotations.
  • Calculate overlap conceptually.
  • Connect image transformations to augmentation.

Common Computer Vision Vocabulary Mistakes

Calling every vision task recognition

Repair: identify whether the task is classification, detection, segmentation, tracking or geometry.

Confusing bounding box and mask

Repair: a box approximates location; a mask labels pixels.

Treating confidence as probability of truth

Repair: understand it as model output that requires calibration and evaluation.

Assuming ground truth is infallible

Repair: account for annotation quality and ambiguity.

Frequently Asked Questions

What is computer vision vocabulary?

It is the specialised language used for images, classification, detection, segmentation, annotation and visual-model evaluation.

What terms should beginners learn first?

Start with pixel, image, label, classification, detection, bounding box, segmentation, annotation and confidence score.

What is object detection?

It is a task that identifies object categories and estimates where those objects appear in an image or frame.

What is segmentation?

It is a task that assigns labels to image pixels or regions.

How can I learn computer vision vocabulary?

Use real images and label what classification, detection and segmentation outputs would look like.

Where This Article Fits in the eduKateSG Vocabulary Ecosystem

The Computer Vision Vocabulary Standard

Computer vision vocabulary reaches its core aim when the learner can name the visual task, describe its output and explain how that output is evaluated against labelled evidence.

That is the standard: visual AI language precise enough to distinguish seeing from merely naming.

Discover more from eduKate Singapore

Subscribe now to keep reading and get access to the full archive.

Continue reading