THE CORE AIM OF VOCABULARY MASTERY · AI INFERENCE VOCABULARY · INPUT → MODEL → COMPUTE → OUTPUT → LATENCY
AI inference vocabulary is the language used to describe what happens when a trained model receives new input and produces a prediction, classification, score or generated output. Terms such as inference, serving, endpoint, batch inference, online inference, latency, throughput, quantization, accelerator and context matter because running a model reliably is a different problem from training it.
The core aim of vocabulary mastery for AI inference vocabulary is runtime-model clarity. Learners should be able to explain what input enters the trained model, what computation occurs, what output is returned, how quickly it must arrive and what infrastructure trade-offs affect cost and performance.
This page is the AI Inference Vocabulary owner inside the eduKateSG Vocabulary hub. For operational lifecycle, use MLOps Vocabulary. For neural computation, use Neural Networks Vocabulary.
Central proposition: AI-inference vocabulary is mastered when the learner can separate model training from the runtime path that turns new inputs into outputs.
The 60-Second AI Inference Vocabulary Router
- Input: request, feature, token, context.
- Serving: endpoint, runtime, model server.
- Compute: CPU, GPU, accelerator, memory.
- Output: prediction, score, token, completion.
- Performance: latency, throughput, concurrency.
- Optimisation: batching, caching, quantization, compilation.
Training and Inference Are Different
Training changes model parameters using data and an optimisation process. Inference normally keeps those learned parameters fixed while applying the model to new input.
A Worked Example: Online vs Batch Inference
Online inference serves predictions in response to individual or small groups of requests, often under latency constraints. Batch inference processes larger collections on a schedule or job where throughput may matter more than immediate response.
Latency and Throughput
Latency is the delay for a request or unit of work. Throughput is the amount of work completed per unit time. Optimising one can affect the other, especially when batching or concurrency changes.
Quantization
Quantization represents model values using lower numerical precision or more compact formats. It can reduce memory and computation requirements, but the effect on model quality and hardware performance must be evaluated.
Model Serving and Endpoints
A model server hosts a model for inference. An endpoint is an interface through which applications submit inputs and receive outputs. Production serving also needs scaling, monitoring, version control and failure handling.
How to Learn AI Inference Vocabulary
- Run a trained model on new inputs.
- Measure latency separately from throughput.
- Compare online and batch serving.
- Observe memory use.
- Experiment with batching conceptually.
- Compare model versions.
- Connect runtime optimisation to quality and cost.
Common AI Inference Vocabulary Mistakes
Confusing inference with reasoning
Repair: inference here means executing a trained model to produce outputs.
Treating latency and throughput as synonyms
Repair: separate delay per request from total work rate.
Assuming smaller precision is always better
Repair: evaluate accuracy, hardware support and performance.
Frequently Asked Questions
What is AI inference?
It is the process of applying a trained model to new input to produce an output.
What is model serving?
It is the infrastructure and runtime process used to make model inference available to applications or users.
What is quantization?
It is representing model values with lower precision or compact formats to reduce resource requirements.
The AI Inference Vocabulary Standard
Mastery means explaining the input, model runtime, hardware, output, latency, throughput and optimisation trade-offs as one serving path.
That is the standard: runtime AI language that makes model execution measurable rather than mysterious.
