Inference is what happens when a trained model is actually used. Training changes parameters. Inference keeps those parameters fixed and applies the model to a new input: a prompt, image, audio clip, document, sensor reading or other representation.
For a large language model, inference commonly means preparing the input, tokenising it, running the transformer, producing next-token scores, choosing a token, appending it to the sequence and repeating until the response stops. In a complete SI application, inference can also be surrounded by retrieval, tools, memory, safety checks and orchestration.
This article explains how inference works from request arrival to response delivery: prefill, decoding, attention, KV cache, logits, sampling, batching, latency, throughput, memory, quantisation, long context, streaming, tool calls, multimodal inputs and failure diagnosis.
Hugging Face’s current LLM tutorial and cache explanation provide practical references for text generation and key-value caching. The engineering details differ across systems, but the core request lifecycle is stable enough to study clearly.
Previous: 030 — Distillation and Specialisation. This article begins Part 4: Inference, Reasoning and Planning.
The Hidden Transition: Training Has Stopped, but Computation Continues
A user asks, “Summarise this document.” The model does not retrain itself on the document. It runs a forward computation using its existing parameters and the document-derived context.
The model’s weights remain fixed during ordinary inference. What changes from request to request is the input, context, runtime state and possibly the tools available around the model.
This distinction prevents one of the most common misconceptions in SI: a model can adapt its response within a conversation without changing its trained parameters.
A Complete Inference Request
Imagine a user sends: “Explain photosynthesis in 120 words for a Primary 6 student.” The application records the request, adds any system instructions, selects relevant conversation context and passes the prepared message sequence to the tokenizer.
The tokenizer converts text into token IDs. Those IDs become embeddings. The transformer processes the sequence. The output layer produces logits over the vocabulary for the next token. A decoding method selects one token.
That token is appended. The model runs again for the next position, often reusing cached attention values from earlier tokens. The loop continues until an end condition is met.
Inference Versus Training
Training needs a forward pass, loss calculation, backward pass and optimiser update. Inference normally needs only the forward computation.
This makes inference cheaper than training per example, but production inference can still consume enormous resources because millions of users may generate long sequences.
Inference optimisation therefore becomes its own engineering field.
The Input Packet
A model rarely receives only the user’s latest sentence. The application may include system instructions, previous messages, retrieved passages, tool schemas and metadata.
The complete packet is serialised into the model’s expected format. Chat templates can insert special tokens that identify roles and boundaries.
A wrong input packet produces wrong inference even when the model itself is functioning perfectly.
Tokenisation at Runtime
The prepared text is converted into tokens according to the model’s tokenizer. Exact token boundaries depend on the vocabulary and algorithm.
Long inputs consume context-window capacity before the model generates anything. A 20,000-token document leaves less room for output than a 500-token prompt under the same total context limit.
Token count also affects cost and latency because the model must process the input representation.
Embeddings and Positional Information
Token IDs are mapped to learned embeddings. Positional information tells the model where tokens occur in the sequence.
Different architectures encode position differently, but the general requirement remains: the model must distinguish “dog bites man” from “man bites dog” even though the same words appear.
Prefill
In common language-model serving terminology, prefill is the phase where the model processes the existing input sequence before generating new tokens.
The model computes hidden states and attention-related key/value representations for the prompt. This stage can process many input tokens in parallel.
Long prompts increase prefill work. A system that appears slow before the first output token may be spending time on document retrieval, tokenisation, prefill or all three.
Decode
After prefill, decoding generates output tokens one step at a time for autoregressive models. Each new token depends on the existing sequence.
Unlike prefill, decode has a sequential dependency: token 101 cannot be generated until token 100 has been selected.
This makes output generation latency behave differently from input processing latency.
The KV Cache
Self-attention uses key and value representations from earlier tokens. Recomputing all earlier keys and values for every new token would waste work.
A key-value cache stores them so later decode steps can reuse previous computations. Hugging Face’s cache documentation explains this runtime optimisation directly.
The KV cache can consume substantial memory, especially for long contexts, large batches and large models.
KV Cache Is Not User Memory
The cache is a runtime efficiency structure, not a permanent personal memory. It normally exists to accelerate the current generation session.
Persistent user memory, saved conversation history and retrieved past facts are separate application mechanisms.
Logits
At each generation step, the model produces logits: raw scores for every token in the vocabulary.
Softmax can convert logits into probabilities. A decoding strategy then determines which candidate token becomes part of the output.
The logits express the model’s next-token preferences under the current context. They are not direct probabilities that the entire answer is true.
Greedy Decoding
Greedy decoding selects the highest-scoring next token at each step.
This is simple and deterministic under fixed conditions, but local best choices do not guarantee the globally best sequence.
Greedy decoding can also produce repetitive or overly predictable text in some tasks.
Sampling
Sampling chooses a token according to a probability distribution rather than always taking the maximum.
Temperature, top-k and top-p can reshape or restrict the candidate distribution. These controls affect diversity and repeatability.
They do not add factual evidence.
Beam Search
Beam search keeps several candidate sequences alive and expands them, retaining the highest-scoring beams under the search rule.
It can improve sequence quality for some tasks such as translation, but it increases computation and is not always preferred for open-ended dialogue.
Stop Conditions
Generation stops when the model emits an end token, reaches a configured length, matches a stop sequence or the application interrupts the process.
The stop condition belongs to the task. A truncated answer can be incomplete even though the model stopped normally at a token limit.
Streaming
Applications can display generated tokens incrementally instead of waiting for the full response.
Streaming improves perceived latency because the user sees progress earlier. It does not make the underlying computation faster by itself.
The interface should distinguish a partially streamed answer from a completed task if later tool calls or verification are still pending.
Time to First Token
Time to first token measures how long the user waits before output generation becomes visible. It includes some combination of queueing, preparation and prefill.
Long context, overloaded servers and slow retrieval can all increase this metric.
Time Per Output Token
After generation begins, systems can measure average decode time per token or tokens per second.
This depends on model size, hardware, batch size, KV cache behaviour and serving software.
Throughput
Throughput measures how much work the serving system handles over time: tokens per second, requests per second or another workload metric.
A system can have low latency for one user but poor throughput under load, or high aggregate throughput with slower individual responses.
Latency and throughput must be evaluated together.
Batching
Inference servers can process multiple requests together so accelerator computation is used efficiently.
Static batching waits for a group of similarly shaped requests. Continuous batching dynamically adds and removes sequences as they arrive and finish.
Batching improves throughput but can introduce queueing and scheduling trade-offs.
Variable-Length Sequences
Users send prompts of different lengths and request different output lengths. Naive batching can waste compute on padding.
Modern serving systems group or schedule requests to reduce this waste.
Memory During Inference
Inference memory includes model weights, activations, KV cache, runtime buffers and framework overhead.
Large models can be memory-bandwidth bound during decode because each token generation step reads large amounts of parameter data.
Quantisation
Quantisation stores weights or activations in lower-precision numerical formats. This can reduce memory and increase inference speed.
Lower precision can degrade output quality, especially for some layers or tasks. Quantised models therefore need evaluation, not only memory measurements.
Speculative Decoding
Speculative decoding uses a faster draft model to propose several tokens, then a stronger model verifies or accepts them under a defined algorithm.
When predictions align, several tokens can be advanced with fewer expensive target-model steps.
The technique trades extra draft computation for lower latency on the main model.
Prefix Caching
If many requests share a long prefix—such as the same system prompt or document—servers can cache processed representations and reuse them.
This reduces repeated prefill work. Correct cache identity is essential; a cached prefix from the wrong version can make the entire response stale.
Context Length and Inference Cost
Longer context does not merely use more memory. Attention and prefill computation also increase.
Applications should retrieve the relevant material instead of stuffing every available document into context by default.
More context is useful only when the model can reliably use the additional information.
Long Context and Attention Dilution
Even when information fits technically, a model may fail to attend to the relevant passage among thousands of tokens.
Placement, redundancy, conflicting versions and distractors can affect performance. Context-window capacity and effective context use are different properties.
Retrieval Before Inference
A RAG system searches or retrieves relevant information before model inference. The retrieved passages become part of the input packet.
Inference then synthesises the evidence. Retrieval failure and model inference failure should be tested separately.
Tools During Inference
A model can generate a structured tool request instead of a final answer. The surrounding application executes the tool and feeds the result into a later model invocation.
One user request can therefore involve multiple inference calls separated by external operations.
Worked Tool Example
User: “What is 18.7% of 4,250?” The model decides a calculator is appropriate and requests 0.187 × 4250.
The calculator returns 794.75. A second inference step formats the answer.
The exact arithmetic belongs to the calculator; inference handles task interpretation and explanation.
Multimodal Inference
A multimodal model can receive image, audio or video representations alongside text. Encoders or tokenisation schemes convert those modalities into representations the model can process.
The inference lifecycle still includes input preparation, forward computation and output generation, but the representation and compute paths can be more complex.
Inference on Encoders
Not every model generates tokens. An encoder can process an input once and output an embedding or class representation.
A classifier might run one forward pass and return probabilities without an autoregressive decode loop.
Inference on Diffusion Models
Image-generation diffusion models use a different iterative process: start from noise or another latent state and repeatedly denoise under conditioning.
This is still inference because parameters remain fixed while the model computes a new output.
The mechanics differ from next-token generation, which is why “inference” is broader than LLM decoding.
Determinism
Even with fixed weights and prompt, outputs may vary if sampling is enabled or if hardware and numerical kernels are nondeterministic.
Setting temperature to zero or using greedy decoding increases repeatability but does not guarantee universal bit-for-bit determinism across all systems.
Random Seeds
Some generation APIs expose random seeds. A seed can make stochastic choices more reproducible under fixed software and model conditions.
A changed model version, prompt, hardware path or serving implementation can still change the result.
Model Versioning
Inference behaviour depends on the exact checkpoint and serving configuration. Applications should record model versions when reproducibility matters.
A silent model upgrade can change wording, tool selection or task performance even if user prompts remain identical.
Inference Under Load
Production systems queue requests, allocate accelerators and balance traffic. High load can increase latency without changing model quality.
A timeout may therefore be an infrastructure problem rather than a reasoning failure.
Fallback Models
Applications can route requests to a smaller model when the primary model is overloaded or unavailable.
Fallbacks should be evaluated because capability can change. The user should not receive a lower-confidence result disguised as identical service if the difference matters.
Model Routing
Different tasks can be routed to different models. Simple classification may use a small model; difficult reasoning may use a larger one.
Routing itself becomes a prediction problem and must be evaluated for cost, latency and quality.
Inference Failure Map
Level 1: wrong input packet. Level 2: tokenisation or truncation error. Level 3: wrong model/version. Level 4: prefill overload. Level 5: decode instability. Level 6: tool-call failure. Level 7: cache contamination. Level 8: serving timeout. Level 9: completion message overstates what happened.
This eduKateSG map separates runtime failures from training failures.
Worked Diagnosis: Slow Before Any Text Appears
Inspect queueing, retrieval, tokenisation and prefill. The model’s decode speed may be fine once generation begins.
Worked Diagnosis: Fast Start, Slow Long Answer
Decode throughput may be the bottleneck. Inspect batch size, KV cache pressure, model size and hardware utilisation.
Worked Diagnosis: Wrong Answer Only on Long Documents
Inspect truncation, context assembly, retrieval and effective long-context behaviour. A generic model upgrade is not the first repair.
Worked Diagnosis: Duplicate Tool Action
A timeout occurs after an external write, then the inference loop retries blindly. The failure belongs to action-state handling and idempotency, not token generation alone.
A Practical Inference Benchmark
Measure time to first token, tokens per second, total latency, throughput, memory use and task quality across short, medium and long contexts.
Include concurrent load because a system that is fast for one request may degrade sharply at production volume.
Also test failures: unavailable tool, malformed document, long prompt, maximum output and fallback routing.
Independent Exercise 1: Training or Inference?
The model answers a user question without changing weights. Which phase is running?
Answer
Inference.
Independent Exercise 2: Prefill or Decode?
A 50,000-token document causes a long delay before the first output token. Which phase is likely expensive?
Answer
Prefill and possibly upstream document preparation.
Independent Exercise 3: KV Cache
Why does caching earlier keys and values help autoregressive generation?
Answer
It avoids recomputing attention representations for all previous tokens at every decode step.
Independent Exercise 4: Streaming
Does streaming reduce the model’s actual total compute?
Answer
Not necessarily. It mainly exposes partial output sooner, improving perceived responsiveness.
Independent Exercise 5: Tool Use
A model says “file saved” but no file tool ran. Is inference complete?
Answer
The text generation finished, but the requested task did not complete. External action requires tool execution and verification.
The Runtime Boundary: What Must Already Exist Before Inference Begins
Inference assumes several things are already available: a trained checkpoint, a tokenizer or input encoder, model configuration, runtime software, compatible hardware and any application instructions or tool definitions. If one of these is wrong, the request can fail before the model performs meaningful work.
A checkpoint without the matching tokenizer can corrupt input interpretation. A model loaded with incompatible quantisation settings can produce degraded output. A tool schema whose field names changed can turn valid model intent into failed execution.
This is why inference should be treated as a complete runtime environment rather than one mathematical forward pass isolated from the application.
Model Loading
Before requests can be served, model weights must be loaded into accelerator or system memory. Large models may be sharded across several devices.
Cold-start time can be substantial. Production services often keep models resident in memory so individual requests do not reload billions of parameters.
A serverless or elastic environment may trade idle cost against cold-start latency. That is an infrastructure decision around the inference engine.
Warm Models and Cold Models
A warm model is already loaded and ready to accept work. A cold model requires loading, compilation or cache creation before useful inference begins.
Two services using the same checkpoint can therefore have very different first-request latency. Model capability is identical while runtime architecture differs.
Compilation and Kernel Selection
Inference frameworks can compile graphs, fuse operations or choose specialised GPU kernels for attention and matrix multiplication.
These optimisations do not change the high-level model definition, but numerical precision and kernel behaviour can affect exact outputs or performance.
Serving benchmarks should therefore identify the runtime stack, not only the model name.
Prefill Parallelism Versus Decode Seriality
Prefill can process many prompt positions together because the full input is already known. Decode cannot know the next token until the current token is chosen.
This difference explains why long prompts and long outputs stress hardware differently. Prompt-heavy workloads need strong prefill throughput; generation-heavy workloads need efficient sequential decode.
Arithmetic Intuition for Prefill Cost
Suppose Request A has 500 input tokens and 100 output tokens. Request B has 20,000 input tokens and the same 100-token output. The output length is identical, but Request B requires far more prompt processing and KV-cache memory.
A pricing or latency model that counts only output tokens misses a major part of runtime cost.
Arithmetic Intuition for Decode Cost
Now compare two requests with identical 500-token prompts. One generates 50 tokens; the other generates 2,000. The longer answer performs many more sequential decode steps.
Even if time to first token is identical, total completion time can differ dramatically.
Paged KV-Cache Management
Serving systems can manage KV-cache memory in blocks or pages so completed requests release memory efficiently and variable-length sequences do not fragment accelerator memory badly.
This is an implementation optimisation around the same logical cache. It becomes important when many concurrent conversations have different lengths.
Cache Eviction
When memory pressure rises, an inference server may need to evict cached prefixes or runtime state. Eviction can force recomputation on later requests.
Cache policy therefore changes latency without changing model quality.
Batch Scheduling
Continuous batching decides which active sequences receive compute at each step. A scheduler can prioritise short requests, fairness, throughput or service-level objectives.
A poor scheduler can create long tail latency even when average throughput looks excellent.
Head-of-Line Blocking
If one giant prompt shares a batch with many tiny requests under a naive scheduler, the long request can delay the smaller ones.
Production serving therefore needs workload-aware batching and limits.
Admission Control
An overloaded service can accept every request and let queues grow, or reject/route work when capacity is exhausted.
Admission control protects latency and prevents memory exhaustion. A graceful overload response is part of inference reliability.
Concurrency Limits
Tool-using agents can generate several model calls from one user request. Without limits, one complex task can consume disproportionate server capacity.
Applications can cap concurrent subcalls, sequence length, retries and total test-time compute.
Token Budgets
A runtime token budget limits how much input and output a task can consume. The budget may be expressed in context length, maximum generated tokens or total agent-loop usage.
Budgets prevent runaway cost but can also truncate legitimate work. The task should know when to summarise, retrieve selectively or stop honestly.
Prompt Caching and Version Identity
Shared prompts—policies, long system instructions or large reference documents—can be cached after prefill.
The cache key should include content identity and model/runtime version. Reusing a cache for an outdated instruction packet can create systematic stale behaviour.
Inference-Time Retrieval
Some systems retrieve information once before the first model call. Others alternate retrieval and inference several times.
The timing matters. Early retrieval can ground the whole answer. Later retrieval can answer a missing subquestion discovered during reasoning.
Each retrieval changes the context and therefore the next inference state.
Context Compression
When available information exceeds context capacity, an application can summarise, rank, chunk or compress evidence before model inference.
Compression saves tokens but can remove qualifiers. A good system retains source links or structured evidence so important claims remain checkable.
Prompt Injection at Inference Time
Retrieved documents may contain instructions that conflict with the user’s task. The model sees both as tokens unless the application preserves trust boundaries.
Permissions, tool restrictions and instruction hierarchy must be enforced by the surrounding system. Runtime inference alone cannot guarantee that untrusted text never influences output.
Structured Output
Inference can be constrained to produce JSON, XML, SQL, function arguments or another grammar. This makes downstream integration easier.
Structural validity is not semantic validity. A perfectly valid JSON object can still contain the wrong student ID or date.
Grammar-Constrained Decoding
Some serving systems restrict token choices so only strings allowed by a grammar can be produced.
This can guarantee syntax for formats such as JSON, but it cannot guarantee that field values correspond to reality.
Function Calling as an Inference Output Mode
Instead of prose, the model may emit a tool name and structured arguments. The application validates and executes them.
A tool call is therefore one kind of generated output, not evidence that the tool has already succeeded.
Multiple Inference Passes
A system can ask one model to draft, another to critique and a third pass to revise. The user sees one answer, but runtime compute may involve several forward sequences.
Quality can improve if the passes add useful evidence or verification. Repetition without new signal simply increases cost.
Self-Consistency
For some reasoning tasks, the system can sample several candidate solutions and select the answer supported by the most consistent paths.
This uses more inference compute to reduce variance. Agreement is helpful evidence but not proof; multiple samples can share the same misconception.
Verifier Models
A separate model can score candidate outputs for correctness, policy compliance or quality.
The verifier itself has error rates and must be evaluated. A weak verifier can reward polished wrong answers.
Search During Inference
Inference can branch into explicit search over candidate actions, proofs or solutions. The model proposes alternatives and a scoring function chooses among them.
This shifts some capability from one forward pass into a larger test-time algorithm.
Inference-Time Scaling
Some systems improve hard-task performance by spending more compute at inference: longer reasoning, more samples, verification or search.
The later Test-Time Compute article will separate these techniques from ordinary single-pass generation.
Speculative Decoding in More Detail
A small draft model proposes a block of candidate tokens. The target model evaluates those candidates efficiently and accepts a prefix that matches the target distribution under the algorithm.
If draft and target agree often, target-model sequential steps are reduced. If they disagree frequently, the speed benefit shrinks.
Model Parallel Inference
A model too large for one accelerator can split tensor operations or layers across devices.
Every decode step may then require communication between devices. High-bandwidth interconnects become part of token latency.
Pipeline Parallel Inference
Layers can be divided into stages across devices. Batches move through the pipeline.
This can improve throughput but creates pipeline scheduling complexity and bubbles.
Mixture-of-Experts Inference
A mixture-of-experts model contains many expert sub-networks but routes each token through only a subset.
This can increase parameter capacity without activating every parameter for every token, but routing and communication become inference challenges.
Quantised KV Cache
KV-cache entries can also be stored at lower precision to reduce memory, especially for long-context serving.
Compression may affect attention accuracy, so long-context and generation quality should be evaluated after the change.
Offloading
When accelerator memory is insufficient, weights or cache blocks can be offloaded to CPU memory or storage.
Offloading expands capacity but adds transfer latency. A model that technically fits through offloading may become too slow for interactive use.
Local Inference
Small or compressed models can run on laptops and phones. This reduces network dependence and can improve privacy.
Local hardware imposes stricter memory, power and thermal constraints. Distillation, quantisation and specialised runtimes become important.
Cloud Inference
Cloud serving offers powerful accelerators and elastic scale but introduces network latency, account security and centralised operational dependencies.
The choice between local and cloud is a system-design trade-off rather than a measure of model intelligence.
Inference and Privacy
Inputs can contain sensitive text, images or records. A runtime system should minimise what is sent, restrict logging and respect retention policies.
Caching must also avoid leaking one user’s context into another request. Isolation is an infrastructure requirement.
Inference and Observability
Useful logs include request ID, model version, input/output token counts, latency stages, tool calls, errors and resource usage.
Operational observability does not require storing private chain-of-thought. The goal is to diagnose system behaviour from measurable events.
Tail Latency
Average response time can look good while a small percentage of requests are extremely slow. Users experience the tail, especially in interactive products.
Measure percentiles such as p95 or p99 latency across realistic workloads.
Quality–Latency Trade-Offs
A larger model may improve answer quality but increase latency. More reasoning passes may improve difficult tasks while slowing simple ones.
Routing can allocate expensive inference only where expected value justifies it.
Quality–Cost Trade-Offs
The cheapest request is not necessarily the cheapest completed task if weak output requires human rework.
Measure end-to-end usefulness, not only cost per token.
Inference Regression Testing
After changing runtime libraries, quantisation, kernels or batch settings, rerun the same model evaluation. Serving changes can alter output even when weights are unchanged.
Performance optimisation should therefore be treated as a model-affecting release from the user’s perspective.
A Complete Runtime Trace
Request arrives → authenticated context assembled → retrieval runs → prompt serialised → tokens created → request queued → prefill executes → KV cache allocated → decode begins → tool call generated → tool validated and executed → result inserted → second prefill/decode pass → final response streamed → task completion verified.
This trace is a practical debugging map. A problem can be attached to a concrete stage.
Independent Exercise 6: Cold Start
The first request after deployment takes 30 seconds, later requests take 2 seconds. What runtime issue should you inspect?
Answer
Model loading, compilation or cache warmup. The model itself may not be slower on the first request once fully loaded.
Independent Exercise 7: Tail Latency
Average latency is 2 seconds, but 5% of users wait 20 seconds. Is the average enough?
Answer
No. Inspect percentile latency, queueing and workload outliers.
Independent Exercise 8: Cache Safety
A cached policy prefix is reused after the policy changes. What happens?
Answer
Inference can be consistently grounded in stale instructions. Cache keys and invalidation must track source versions.
Independent Exercise 9: Structured Output
JSON parses correctly but contains a nonexistent room ID. Is the output valid?
Answer
Syntactically yes, semantically no. Validate field values against authoritative state.
Inference in Classification Systems
Not every inference task generates prose. A classifier can receive an input, run one forward pass and return a probability distribution over classes. There may be no decode loop, no KV cache and no streaming tokens.
Example: a support message is encoded and the model outputs scheduling 0.72, payment 0.18, curriculum 0.07 and feedback 0.03. The application may apply a threshold or choose the top class.
This is still inference because the trained parameters are being applied to new data without ordinary weight updates.
Inference in Embedding Models
An embedding model converts text, images or other inputs into vector representations. The output is not a natural-language answer but a point in a learned vector space.
Those vectors can support semantic search, clustering, retrieval or similarity comparison. Runtime inference may therefore finish after one encoder pass.
This reminds us that “inference” is a general machine-learning term, not a synonym for chatbot generation.
Inference in Ranking Models
Search and recommendation systems run models that score candidate items. A ranking model may process query-document pairs and return relevance scores.
The final user-visible order can combine those scores with business rules, freshness filters or permissions. Model inference contributes one signal inside a larger ranking system.
Inference in Computer Vision
An image classifier processes pixels or visual tokens and returns a class distribution. An object detector returns bounding boxes and labels. A segmentation model assigns categories across pixels.
The trained network performs inference once per image or sequence of images. The output may then feed downstream software instead of a language generator.
Inference in Speech Systems
Speech recognition can convert audio features into token or character probabilities, then decode them into text. Text-to-speech models perform the reverse kind of task, generating acoustic representations from text.
Latency requirements can be stricter because users expect live transcription or conversational turn-taking.
Realtime Inference
Realtime systems process continuous streams rather than one static request. Examples include speech, sensors, robotics and interactive control.
The system must meet deadlines. An answer that arrives after the required control window can be useless even if mathematically correct.
Realtime inference therefore adds scheduling, buffering and worst-case latency constraints.
Edge Cases in Context Assembly
A prompt can exceed the context window, contain invalid encoding, mix languages, include malformed tool schemas or reference attachments that failed to load.
The runtime should detect these cases before claiming the model “cannot understand” the request. Many failures are input-pipeline defects.
Truncation Strategies
When input exceeds context capacity, applications can drop oldest messages, truncate documents, summarise history or retrieve only relevant passages.
Each strategy changes the effective task. Dropping the wrong message can remove a crucial instruction; summarisation can erase qualifiers.
Truncation policy should therefore be tested like any other model-facing transformation.
Output-Length Planning
A long input plus a long requested output can exceed the total context budget. The application can reserve output capacity before sending the prompt.
If a user asks for 3,000 output tokens but only 800 remain, the system should warn, compress context or split the task rather than silently cutting off the answer.
Request Cancellation
Users can cancel a long generation. The inference server should stop unnecessary compute and release KV-cache memory.
External tool actions already executed may still need separate handling. Cancelling text generation does not undo a database write.
Timeouts
A timeout is an application boundary, not proof that the model failed internally. The request may still be running, queued or waiting for a tool.
Timeout handling should preserve whether external side effects are known, unknown or complete.
Retries
Read-only inference calls can often be retried safely. Tool-using inference that may have changed state needs idempotency or status checks before retrying.
The retry policy therefore depends on the entire task, not only the model call.
Inference and Authentication
A model may infer that a user wants account information, but the application must authenticate identity before retrieving that data.
Authentication belongs outside the model. Inference can interpret the request after the system knows who the user actually is.
Inference and Authorisation
Authorisation decides which resources or operations the authenticated user may access. A capable model cannot override those rules.
Tool definitions should expose only permitted operations, reducing the consequence of model error.
Inference and Data Minimisation
The runtime should send only information needed for the task. A scheduling question rarely requires payment notes, medical history or unrelated documents.
Smaller contexts can improve privacy, latency and focus simultaneously.
Inference and Logging
Logs can record token counts, model version, tool calls, latency and error codes without storing sensitive full prompts indefinitely.
Logging policy should be designed around debugging needs and privacy obligations.
Inference and Monitoring
Production monitoring looks for latency spikes, error rates, model-quality regressions, unusual token use, tool failures and infrastructure saturation.
A service can remain technically online while quality degrades. Operational health and model quality both need signals.
Canary Releases
A new runtime or model version can be sent to a small percentage of traffic first. Compare quality, latency and errors before wider rollout.
Canary deployment reduces the blast radius of regressions.
A/B Testing
Two inference configurations can be compared under matched traffic: different models, prompts, decoding settings or caching strategies.
A/B tests need task-relevant success metrics. Faster responses are not improvements if they require more correction.
Shadow Testing
A new system can receive copies of live requests without its outputs reaching users. This measures performance on realistic traffic safely.
Sensitive-data policy still applies because shadow systems receive real inputs.
Inference Cost Accounting
A useful cost model separates input processing, output generation, retrieval, tool calls and human review.
One short answer may be expensive if it required a long prompt and several hidden passes. One longer answer may be cheap if it came from a small model with cached context.
A Reader’s Runtime Inspection Checklist
Ask: which model ran? What context was supplied? Was retrieval used? How many model calls occurred? Which tools ran? Was the output sampled or deterministic? Did any cache or fallback model participate? What evidence confirms external actions?
These questions turn “the AI answered” into a traceable inference record.
Independent Exercise 10: Classification
A sentiment classifier returns one probability distribution and stops. Is there a decode phase?
Answer
Not necessarily. Classification inference can finish after a single forward pass.
Independent Exercise 11: Cancellation
The user cancels after an email tool already sent a message. Does cancelling generation unsend it?
Answer
No. External state must be handled separately. The system should report what already occurred.
Independent Exercise 12: Truncation
The oldest system instruction is removed to fit a long document. What risk appears?
Answer
The task boundary can change because a governing instruction disappeared. Context truncation must preserve priority and relevance.
Frequently Asked Questions About Inference
What is model inference?
Using trained model parameters to process a new input and produce an output without ordinary parameter training.
Does inference change model weights?
Not in standard deployment. Fine-tuning or online learning would add training updates.
What is prefill?
The initial processing of the existing input sequence before autoregressive output generation.
What is decoding?
The repeated process of selecting and appending output tokens based on model scores.
What is KV cache?
Stored attention key/value representations from previous tokens used to avoid redundant computation during decoding.
Why is long context expensive?
It increases prefill computation and memory, particularly KV-cache usage.
Why do outputs vary?
Sampling, model updates, context differences and numerical nondeterminism can all change generated sequences.
Can inference include multiple model calls?
Yes. Agents and tool-using workflows often alternate model inference with retrieval or external operations.
Inference Is the Runtime Engine of SI
Training creates capability. Inference spends that capability on a particular request. The runtime system must prepare the right context, execute the model efficiently, manage caches and tools, and report completion honestly.
Understanding inference turns a chatbot from a black box into a sequence of measurable stages: input preparation, prefill, decode, external operations, verification and delivery.
Continue through the How Super Intelligence Works hub. Next: 032 — Next-Token Prediction.
Inference Completion Standard
A production inference task is complete only when the requested output exists, required tools have returned known outcomes, relevant runtime errors are surfaced and the final message does not claim more than the observed state. Token generation ending is one technical event; task completion is a wider system judgment.
For a simple explanation, generation may be enough. For a saved file, completion requires the file operation and a verified resource identity. For a database update, completion requires the committed state. For a research answer, completion may require retrieved evidence and citations. The definition follows the task.
This standard is the practical bridge from inference to reasoning and agents. Runtime intelligence becomes dependable when the application can separate model output, tool execution, external state and verification instead of collapsing them into one confident sentence.
What Mastery of Inference Looks Like
A reader has mastered inference when they can separate prompt preparation from prefill, prefill from decode, token generation from tool execution, cache state from persistent memory, and model latency from application latency. They should also be able to explain why the same weights can behave differently under different contexts, decoding rules, serving stacks and external tools.
That distinction is the foundation for the next articles. Next-token prediction explains the local generation mechanism. Machine reasoning examines structured problem solving across several steps. Test-time compute asks how much additional runtime work improves difficult tasks. All three begin from the inference engine described here.
Inference therefore becomes inspectable rather than mysterious. The user-visible answer is the end of a measurable runtime chain whose stages can be timed, tested, constrained and repaired independently. That is the standard needed before adding more elaborate reasoning loops on top of the model.
