VIEW THIS AS

Auto mode follows the Route Engine until you choose a viewpoint.

YOU ARE HERE

ROUTE CHECK

CONNECTED TO

WHAT NEXT

Use the canonical route for this room, or HELP if you are unsure.

How Large Language Models Work | Tokens, Transformers, Attention, Training, Retrieval and Tools

eduKateSG · World & Knowledge · How Large Language Models Work · 29 September 2026

A large language model can produce fluent explanations, code, summaries and plans, which makes it tempting to explain the system using human metaphors: it knows, remembers, thinks, wants or understands. Those words can be convenient in ordinary conversation, but they are poor starting points for mechanism. A better explanation begins with tokens, learned parameters, attention, prediction, post-training, context, retrieval and tools.

This guide owns WEP169 as a mechanism page under the existing World & Knowledge Hub. It deepens, rather than replaces, How AI Works | From Data and Models to Useful Outputs. It also connects to the existing Logit Lens and the capability-divide/job-redesign owners without turning one product generation into the definition of the entire field.

OpenAI’s current model-development explanation describes foundation models as systems trained on large amounts of information, with text converted into tokens and model parameters adjusted so the system becomes better at predicting likely continuations. That is a useful public mechanism boundary. Modern LLMs are typically built from Transformer-family architectures descended from the attention-based design introduced in the 2017 paper Attention Is All You Need. The details vary by model, and current frontier systems may add mixture-of-experts routing, multimodal processing, retrieval, tools and extensive post-training.

Language models predict sequences

Language models predict sequences is mainly about learning statistical structure in sequences so the model can estimate what token should come next. The mechanism matters because a user’s mental model changes how they interpret output. If a person imagines a hidden database, they may trust a fabricated citation. If they imagine human memory, they may misunderstand context limits. If they imagine consciousness, they may attribute motives where a simpler explanation is token prediction shaped by training and context.

A practical example is predicting the next token in a sentence from the tokens already present. The example should be read mechanically. What information is represented? Which operation transforms it? What is learned during training and what arrives only at inference time? Which part is probabilistic? Which part comes from an external tool? Breaking the system into these questions prevents the word intelligence from doing too much explanatory work.

The common failure is imagining the model begins by looking up a complete stored answer. LLMs are especially easy to anthropomorphise because language is the medium humans use to reveal thought. A machine that produces coherent language therefore feels as though it must contain a human-like inner process. Mechanistic literacy does not require denying that the system is capable; it requires resisting claims that exceed the evidence.

The acceptance test is that the mechanism is understood as conditional prediction over tokens. A good explanation should predict behaviour. It should help explain why the model can be fluent and wrong, why wording and context matter, why retrieval helps current factual tasks, why sampling creates variability, why a tool can improve arithmetic, and why a model can perform strongly on one benchmark while failing unexpectedly on a nearby task.

This is also the boundary between mechanism and product behaviour. A specific deployed assistant may include search, memory, code execution, safety systems and routing across different models. Those features matter, but they are not the base language model itself. Keep the stack visible: model → context → post-training → retrieval/tools → product policies → user task.

Tokens are the working units

Tokens are the working units is mainly about turning text into smaller discrete pieces that the model can process. The mechanism matters because a user’s mental model changes how they interpret output. If a person imagines a hidden database, they may trust a fabricated citation. If they imagine human memory, they may misunderstand context limits. If they imagine consciousness, they may attribute motives where a simpler explanation is token prediction shaped by training and context.

A practical example is splitting an uncommon word into several subword tokens while a common word may be one token. The example should be read mechanically. What information is represented? Which operation transforms it? What is learned during training and what arrives only at inference time? Which part is probabilistic? Which part comes from an external tool? Breaking the system into these questions prevents the word intelligence from doing too much explanatory work.

The common failure is assuming one token always equals one word. LLMs are especially easy to anthropomorphise because language is the medium humans use to reveal thought. A machine that produces coherent language therefore feels as though it must contain a human-like inner process. Mechanistic literacy does not require denying that the system is capable; it requires resisting claims that exceed the evidence.

The acceptance test is that the reader understands that model context and cost are measured in tokens rather than ordinary words. A good explanation should predict behaviour. It should help explain why the model can be fluent and wrong, why wording and context matter, why retrieval helps current factual tasks, why sampling creates variability, why a tool can improve arithmetic, and why a model can perform strongly on one benchmark while failing unexpectedly on a nearby task.

This is also the boundary between mechanism and product behaviour. A specific deployed assistant may include search, memory, code execution, safety systems and routing across different models. Those features matter, but they are not the base language model itself. Keep the stack visible: model → context → post-training → retrieval/tools → product policies → user task.

Embeddings turn tokens into vectors

Embeddings turn tokens into vectors is mainly about representing token identity and context as numerical patterns that neural layers can transform. The mechanism matters because a user’s mental model changes how they interpret output. If a person imagines a hidden database, they may trust a fabricated citation. If they imagine human memory, they may misunderstand context limits. If they imagine consciousness, they may attribute motives where a simpler explanation is token prediction shaped by training and context.

A practical example is placing related language patterns in representations where useful relationships can be learned. The example should be read mechanically. What information is represented? Which operation transforms it? What is learned during training and what arrives only at inference time? Which part is probabilistic? Which part comes from an external tool? Breaking the system into these questions prevents the word intelligence from doing too much explanatory work.

The common failure is thinking the model stores dictionary definitions in one location. LLMs are especially easy to anthropomorphise because language is the medium humans use to reveal thought. A machine that produces coherent language therefore feels as though it must contain a human-like inner process. Mechanistic literacy does not require denying that the system is capable; it requires resisting claims that exceed the evidence.

The acceptance test is that meaning is distributed through learned numerical representations. A good explanation should predict behaviour. It should help explain why the model can be fluent and wrong, why wording and context matter, why retrieval helps current factual tasks, why sampling creates variability, why a tool can improve arithmetic, and why a model can perform strongly on one benchmark while failing unexpectedly on a nearby task.

This is also the boundary between mechanism and product behaviour. A specific deployed assistant may include search, memory, code execution, safety systems and routing across different models. Those features matter, but they are not the base language model itself. Keep the stack visible: model → context → post-training → retrieval/tools → product policies → user task.

Position matters

Position matters is mainly about giving the model information about sequence order so the same tokens in different arrangements are distinguishable. The mechanism matters because a user’s mental model changes how they interpret output. If a person imagines a hidden database, they may trust a fabricated citation. If they imagine human memory, they may misunderstand context limits. If they imagine consciousness, they may attribute motives where a simpler explanation is token prediction shaped by training and context.

A practical example is distinguishing ‘dog bites man’ from ‘man bites dog’ despite shared vocabulary. The example should be read mechanically. What information is represented? Which operation transforms it? What is learned during training and what arrives only at inference time? Which part is probabilistic? Which part comes from an external tool? Breaking the system into these questions prevents the word intelligence from doing too much explanatory work.

The common failure is assuming a bag of words is enough for language modelling. LLMs are especially easy to anthropomorphise because language is the medium humans use to reveal thought. A machine that produces coherent language therefore feels as though it must contain a human-like inner process. Mechanistic literacy does not require denying that the system is capable; it requires resisting claims that exceed the evidence.

The acceptance test is that the representation preserves useful order information. A good explanation should predict behaviour. It should help explain why the model can be fluent and wrong, why wording and context matter, why retrieval helps current factual tasks, why sampling creates variability, why a tool can improve arithmetic, and why a model can perform strongly on one benchmark while failing unexpectedly on a nearby task.

This is also the boundary between mechanism and product behaviour. A specific deployed assistant may include search, memory, code execution, safety systems and routing across different models. Those features matter, but they are not the base language model itself. Keep the stack visible: model → context → post-training → retrieval/tools → product policies → user task.

Attention moves information across the context

Attention moves information across the context is mainly about letting each position combine information from other positions according to learned relevance. The mechanism matters because a user’s mental model changes how they interpret output. If a person imagines a hidden database, they may trust a fabricated citation. If they imagine human memory, they may misunderstand context limits. If they imagine consciousness, they may attribute motives where a simpler explanation is token prediction shaped by training and context.

A practical example is a pronoun drawing information from an earlier noun phrase while generating the next token. The example should be read mechanically. What information is represented? Which operation transforms it? What is learned during training and what arrives only at inference time? Which part is probabilistic? Which part comes from an external tool? Breaking the system into these questions prevents the word intelligence from doing too much explanatory work.

The common failure is describing attention as conscious focus or human awareness. LLMs are especially easy to anthropomorphise because language is the medium humans use to reveal thought. A machine that produces coherent language therefore feels as though it must contain a human-like inner process. Mechanistic literacy does not require denying that the system is capable; it requires resisting claims that exceed the evidence.

The acceptance test is that attention is treated as a mathematical information-routing operation. A good explanation should predict behaviour. It should help explain why the model can be fluent and wrong, why wording and context matter, why retrieval helps current factual tasks, why sampling creates variability, why a tool can improve arithmetic, and why a model can perform strongly on one benchmark while failing unexpectedly on a nearby task.

This is also the boundary between mechanism and product behaviour. A specific deployed assistant may include search, memory, code execution, safety systems and routing across different models. Those features matter, but they are not the base language model itself. Keep the stack visible: model → context → post-training → retrieval/tools → product policies → user task.

Self-attention is repeated in layers

Self-attention is repeated in layers is mainly about building richer contextual representations through many transformations. The mechanism matters because a user’s mental model changes how they interpret output. If a person imagines a hidden database, they may trust a fabricated citation. If they imagine human memory, they may misunderstand context limits. If they imagine consciousness, they may attribute motives where a simpler explanation is token prediction shaped by training and context.

A practical example is a later layer combining syntax, reference and task context derived from earlier processing. The example should be read mechanically. What information is represented? Which operation transforms it? What is learned during training and what arrives only at inference time? Which part is probabilistic? Which part comes from an external tool? Breaking the system into these questions prevents the word intelligence from doing too much explanatory work.

The common failure is assuming one attention step contains the model’s full reasoning. LLMs are especially easy to anthropomorphise because language is the medium humans use to reveal thought. A machine that produces coherent language therefore feels as though it must contain a human-like inner process. Mechanistic literacy does not require denying that the system is capable; it requires resisting claims that exceed the evidence.

The acceptance test is that capability emerges from many learned transformations. A good explanation should predict behaviour. It should help explain why the model can be fluent and wrong, why wording and context matter, why retrieval helps current factual tasks, why sampling creates variability, why a tool can improve arithmetic, and why a model can perform strongly on one benchmark while failing unexpectedly on a nearby task.

This is also the boundary between mechanism and product behaviour. A specific deployed assistant may include search, memory, code execution, safety systems and routing across different models. Those features matter, but they are not the base language model itself. Keep the stack visible: model → context → post-training → retrieval/tools → product policies → user task.

Transformers enable scalable sequence processing

Transformers enable scalable sequence processing is mainly about using attention-based architecture to model long-range relationships efficiently enough for large-scale training. The mechanism matters because a user’s mental model changes how they interpret output. If a person imagines a hidden database, they may trust a fabricated citation. If they imagine human memory, they may misunderstand context limits. If they imagine consciousness, they may attribute motives where a simpler explanation is token prediction shaped by training and context.

A practical example is processing many token positions in parallel during training. The example should be read mechanically. What information is represented? Which operation transforms it? What is learned during training and what arrives only at inference time? Which part is probabilistic? Which part comes from an external tool? Breaking the system into these questions prevents the word intelligence from doing too much explanatory work.

The common failure is treating Transformer as a brand name rather than an architecture family. LLMs are especially easy to anthropomorphise because language is the medium humans use to reveal thought. A machine that produces coherent language therefore feels as though it must contain a human-like inner process. Mechanistic literacy does not require denying that the system is capable; it requires resisting claims that exceed the evidence.

The acceptance test is that the reader connects modern LLMs to Transformer-style computation. A good explanation should predict behaviour. It should help explain why the model can be fluent and wrong, why wording and context matter, why retrieval helps current factual tasks, why sampling creates variability, why a tool can improve arithmetic, and why a model can perform strongly on one benchmark while failing unexpectedly on a nearby task.

This is also the boundary between mechanism and product behaviour. A specific deployed assistant may include search, memory, code execution, safety systems and routing across different models. Those features matter, but they are not the base language model itself. Keep the stack visible: model → context → post-training → retrieval/tools → product policies → user task.

Parameters are learned numbers

Parameters are learned numbers is mainly about storing model behaviour in large sets of weights adjusted during training. The mechanism matters because a user’s mental model changes how they interpret output. If a person imagines a hidden database, they may trust a fabricated citation. If they imagine human memory, they may misunderstand context limits. If they imagine consciousness, they may attribute motives where a simpler explanation is token prediction shaped by training and context.

A practical example is gradient-based updates changing weights so future predictions improve. The example should be read mechanically. What information is represented? Which operation transforms it? What is learned during training and what arrives only at inference time? Which part is probabilistic? Which part comes from an external tool? Breaking the system into these questions prevents the word intelligence from doing too much explanatory work.

The common failure is imagining parameters as a database of documents. LLMs are especially easy to anthropomorphise because language is the medium humans use to reveal thought. A machine that produces coherent language therefore feels as though it must contain a human-like inner process. Mechanistic literacy does not require denying that the system is capable; it requires resisting claims that exceed the evidence.

The acceptance test is that weights encode patterns rather than a simple archive of training texts. A good explanation should predict behaviour. It should help explain why the model can be fluent and wrong, why wording and context matter, why retrieval helps current factual tasks, why sampling creates variability, why a tool can improve arithmetic, and why a model can perform strongly on one benchmark while failing unexpectedly on a nearby task.

This is also the boundary between mechanism and product behaviour. A specific deployed assistant may include search, memory, code execution, safety systems and routing across different models. Those features matter, but they are not the base language model itself. Keep the stack visible: model → context → post-training → retrieval/tools → product policies → user task.

Pre-training builds a broad language model

Pre-training builds a broad language model is mainly about learning from large corpora before task-specific adaptation. The mechanism matters because a user’s mental model changes how they interpret output. If a person imagines a hidden database, they may trust a fabricated citation. If they imagine human memory, they may misunderstand context limits. If they imagine consciousness, they may attribute motives where a simpler explanation is token prediction shaped by training and context.

A practical example is predicting continuations across many domains until broad linguistic and factual patterns are represented. The example should be read mechanically. What information is represented? Which operation transforms it? What is learned during training and what arrives only at inference time? Which part is probabilistic? Which part comes from an external tool? Breaking the system into these questions prevents the word intelligence from doing too much explanatory work.

The common failure is assuming the model is trained directly on every final user task. LLMs are especially easy to anthropomorphise because language is the medium humans use to reveal thought. A machine that produces coherent language therefore feels as though it must contain a human-like inner process. Mechanistic literacy does not require denying that the system is capable; it requires resisting claims that exceed the evidence.

The acceptance test is that pre-training is separated from later alignment or adaptation. A good explanation should predict behaviour. It should help explain why the model can be fluent and wrong, why wording and context matter, why retrieval helps current factual tasks, why sampling creates variability, why a tool can improve arithmetic, and why a model can perform strongly on one benchmark while failing unexpectedly on a nearby task.

This is also the boundary between mechanism and product behaviour. A specific deployed assistant may include search, memory, code execution, safety systems and routing across different models. Those features matter, but they are not the base language model itself. Keep the stack visible: model → context → post-training → retrieval/tools → product policies → user task.

Training data shapes capability and limitation

Training data shapes capability and limitation is mainly about learning patterns from the distribution and quality of examples seen during development. The mechanism matters because a user’s mental model changes how they interpret output. If a person imagines a hidden database, they may trust a fabricated citation. If they imagine human memory, they may misunderstand context limits. If they imagine consciousness, they may attribute motives where a simpler explanation is token prediction shaped by training and context.

A practical example is a model becoming strong in common domains and weaker in rare or poorly represented ones. The example should be read mechanically. What information is represented? Which operation transforms it? What is learned during training and what arrives only at inference time? Which part is probabilistic? Which part comes from an external tool? Breaking the system into these questions prevents the word intelligence from doing too much explanatory work.

The common failure is assuming scale removes all data bias or coverage gaps. LLMs are especially easy to anthropomorphise because language is the medium humans use to reveal thought. A machine that produces coherent language therefore feels as though it must contain a human-like inner process. Mechanistic literacy does not require denying that the system is capable; it requires resisting claims that exceed the evidence.

The acceptance test is that data composition is treated as one source of capability and error. A good explanation should predict behaviour. It should help explain why the model can be fluent and wrong, why wording and context matter, why retrieval helps current factual tasks, why sampling creates variability, why a tool can improve arithmetic, and why a model can perform strongly on one benchmark while failing unexpectedly on a nearby task.

This is also the boundary between mechanism and product behaviour. A specific deployed assistant may include search, memory, code execution, safety systems and routing across different models. Those features matter, but they are not the base language model itself. Keep the stack visible: model → context → post-training → retrieval/tools → product policies → user task.

The loss function gives a learning signal

The loss function gives a learning signal is mainly about measuring prediction error and using it to update parameters. The mechanism matters because a user’s mental model changes how they interpret output. If a person imagines a hidden database, they may trust a fabricated citation. If they imagine human memory, they may misunderstand context limits. If they imagine consciousness, they may attribute motives where a simpler explanation is token prediction shaped by training and context.

A practical example is penalising low probability assigned to the actual next token during training. The example should be read mechanically. What information is represented? Which operation transforms it? What is learned during training and what arrives only at inference time? Which part is probabilistic? Which part comes from an external tool? Breaking the system into these questions prevents the word intelligence from doing too much explanatory work.

The common failure is imagining the network is explicitly told a symbolic rule for every pattern. LLMs are especially easy to anthropomorphise because language is the medium humans use to reveal thought. A machine that produces coherent language therefore feels as though it must contain a human-like inner process. Mechanistic literacy does not require denying that the system is capable; it requires resisting claims that exceed the evidence.

The acceptance test is that learning is driven by optimisation over examples. A good explanation should predict behaviour. It should help explain why the model can be fluent and wrong, why wording and context matter, why retrieval helps current factual tasks, why sampling creates variability, why a tool can improve arithmetic, and why a model can perform strongly on one benchmark while failing unexpectedly on a nearby task.

This is also the boundary between mechanism and product behaviour. A specific deployed assistant may include search, memory, code execution, safety systems and routing across different models. Those features matter, but they are not the base language model itself. Keep the stack visible: model → context → post-training → retrieval/tools → product policies → user task.

Gradient descent changes parameters incrementally

Gradient descent changes parameters incrementally is mainly about using calculated gradients to reduce loss across training. The mechanism matters because a user’s mental model changes how they interpret output. If a person imagines a hidden database, they may trust a fabricated citation. If they imagine human memory, they may misunderstand context limits. If they imagine consciousness, they may attribute motives where a simpler explanation is token prediction shaped by training and context.

A practical example is millions or billions of updates gradually improving predictions. The example should be read mechanically. What information is represented? Which operation transforms it? What is learned during training and what arrives only at inference time? Which part is probabilistic? Which part comes from an external tool? Breaking the system into these questions prevents the word intelligence from doing too much explanatory work.

The common failure is assuming one training example rewrites a stored fact directly. LLMs are especially easy to anthropomorphise because language is the medium humans use to reveal thought. A machine that produces coherent language therefore feels as though it must contain a human-like inner process. Mechanistic literacy does not require denying that the system is capable; it requires resisting claims that exceed the evidence.

The acceptance test is that parameter change is understood as distributed optimisation. A good explanation should predict behaviour. It should help explain why the model can be fluent and wrong, why wording and context matter, why retrieval helps current factual tasks, why sampling creates variability, why a tool can improve arithmetic, and why a model can perform strongly on one benchmark while failing unexpectedly on a nearby task.

This is also the boundary between mechanism and product behaviour. A specific deployed assistant may include search, memory, code execution, safety systems and routing across different models. Those features matter, but they are not the base language model itself. Keep the stack visible: model → context → post-training → retrieval/tools → product policies → user task.

Scale changes behaviour but does not explain everything

Scale changes behaviour but does not explain everything is mainly about recognising that larger data, models and compute can produce stronger capabilities without making mechanisms trivial. The mechanism matters because a user’s mental model changes how they interpret output. If a person imagines a hidden database, they may trust a fabricated citation. If they imagine human memory, they may misunderstand context limits. If they imagine consciousness, they may attribute motives where a simpler explanation is token prediction shaped by training and context.

A practical example is larger models showing better in-context learning on some tasks. The example should be read mechanically. What information is represented? Which operation transforms it? What is learned during training and what arrives only at inference time? Which part is probabilistic? Which part comes from an external tool? Breaking the system into these questions prevents the word intelligence from doing too much explanatory work.

The common failure is assuming every improvement is caused by parameter count alone. LLMs are especially easy to anthropomorphise because language is the medium humans use to reveal thought. A machine that produces coherent language therefore feels as though it must contain a human-like inner process. Mechanistic literacy does not require denying that the system is capable; it requires resisting claims that exceed the evidence.

The acceptance test is that architecture, data, objectives, post-training and tools remain part of the explanation. A good explanation should predict behaviour. It should help explain why the model can be fluent and wrong, why wording and context matter, why retrieval helps current factual tasks, why sampling creates variability, why a tool can improve arithmetic, and why a model can perform strongly on one benchmark while failing unexpectedly on a nearby task.

This is also the boundary between mechanism and product behaviour. A specific deployed assistant may include search, memory, code execution, safety systems and routing across different models. Those features matter, but they are not the base language model itself. Keep the stack visible: model → context → post-training → retrieval/tools → product policies → user task.

Context is temporary working input

Context is temporary working input is mainly about distinguishing prompt-time information from long-term model parameters. The mechanism matters because a user’s mental model changes how they interpret output. If a person imagines a hidden database, they may trust a fabricated citation. If they imagine human memory, they may misunderstand context limits. If they imagine consciousness, they may attribute motives where a simpler explanation is token prediction shaped by training and context.

A practical example is placing a document in the context so the model can use it for one response. The example should be read mechanically. What information is represented? Which operation transforms it? What is learned during training and what arrives only at inference time? Which part is probabilistic? Which part comes from an external tool? Breaking the system into these questions prevents the word intelligence from doing too much explanatory work.

The common failure is assuming anything in a prompt is permanently learned into the model. LLMs are especially easy to anthropomorphise because language is the medium humans use to reveal thought. A machine that produces coherent language therefore feels as though it must contain a human-like inner process. Mechanistic literacy does not require denying that the system is capable; it requires resisting claims that exceed the evidence.

The acceptance test is that context and training are separated. A good explanation should predict behaviour. It should help explain why the model can be fluent and wrong, why wording and context matter, why retrieval helps current factual tasks, why sampling creates variability, why a tool can improve arithmetic, and why a model can perform strongly on one benchmark while failing unexpectedly on a nearby task.

This is also the boundary between mechanism and product behaviour. A specific deployed assistant may include search, memory, code execution, safety systems and routing across different models. Those features matter, but they are not the base language model itself. Keep the stack visible: model → context → post-training → retrieval/tools → product policies → user task.

The context window is finite

The context window is finite is mainly about recognising that only a bounded amount of input can be processed at once. The mechanism matters because a user’s mental model changes how they interpret output. If a person imagines a hidden database, they may trust a fabricated citation. If they imagine human memory, they may misunderstand context limits. If they imagine consciousness, they may attribute motives where a simpler explanation is token prediction shaped by training and context.

A practical example is summarising or retrieving relevant passages when a document set exceeds the window. The example should be read mechanically. What information is represented? Which operation transforms it? What is learned during training and what arrives only at inference time? Which part is probabilistic? Which part comes from an external tool? Breaking the system into these questions prevents the word intelligence from doing too much explanatory work.

The common failure is assuming the model has unlimited active memory. LLMs are especially easy to anthropomorphise because language is the medium humans use to reveal thought. A machine that produces coherent language therefore feels as though it must contain a human-like inner process. Mechanistic literacy does not require denying that the system is capable; it requires resisting claims that exceed the evidence.

The acceptance test is that prompt design respects context limits. A good explanation should predict behaviour. It should help explain why the model can be fluent and wrong, why wording and context matter, why retrieval helps current factual tasks, why sampling creates variability, why a tool can improve arithmetic, and why a model can perform strongly on one benchmark while failing unexpectedly on a nearby task.

This is also the boundary between mechanism and product behaviour. A specific deployed assistant may include search, memory, code execution, safety systems and routing across different models. Those features matter, but they are not the base language model itself. Keep the stack visible: model → context → post-training → retrieval/tools → product policies → user task.

In-context learning is not parameter retraining

In-context learning is not parameter retraining is mainly about using examples in the prompt to shape behaviour temporarily. The mechanism matters because a user’s mental model changes how they interpret output. If a person imagines a hidden database, they may trust a fabricated citation. If they imagine human memory, they may misunderstand context limits. If they imagine consciousness, they may attribute motives where a simpler explanation is token prediction shaped by training and context.

A practical example is showing two labelled examples before asking for a third classification. The example should be read mechanically. What information is represented? Which operation transforms it? What is learned during training and what arrives only at inference time? Which part is probabilistic? Which part comes from an external tool? Breaking the system into these questions prevents the word intelligence from doing too much explanatory work.

The common failure is calling few-shot prompting training in the same sense as updating weights. LLMs are especially easy to anthropomorphise because language is the medium humans use to reveal thought. A machine that produces coherent language therefore feels as though it must contain a human-like inner process. Mechanistic literacy does not require denying that the system is capable; it requires resisting claims that exceed the evidence.

The acceptance test is that temporary conditioning and model training remain distinct. A good explanation should predict behaviour. It should help explain why the model can be fluent and wrong, why wording and context matter, why retrieval helps current factual tasks, why sampling creates variability, why a tool can improve arithmetic, and why a model can perform strongly on one benchmark while failing unexpectedly on a nearby task.

This is also the boundary between mechanism and product behaviour. A specific deployed assistant may include search, memory, code execution, safety systems and routing across different models. Those features matter, but they are not the base language model itself. Keep the stack visible: model → context → post-training → retrieval/tools → product policies → user task.

Post-training changes behaviour

Post-training changes behaviour is mainly about using supervised examples, preference signals and other methods to make a pretrained model more useful and aligned. The mechanism matters because a user’s mental model changes how they interpret output. If a person imagines a hidden database, they may trust a fabricated citation. If they imagine human memory, they may misunderstand context limits. If they imagine consciousness, they may attribute motives where a simpler explanation is token prediction shaped by training and context.

A practical example is teaching the model to follow instructions and avoid some unsafe outputs. The example should be read mechanically. What information is represented? Which operation transforms it? What is learned during training and what arrives only at inference time? Which part is probabilistic? Which part comes from an external tool? Breaking the system into these questions prevents the word intelligence from doing too much explanatory work.

The common failure is assuming next-token pre-training alone explains the assistant interface. LLMs are especially easy to anthropomorphise because language is the medium humans use to reveal thought. A machine that produces coherent language therefore feels as though it must contain a human-like inner process. Mechanistic literacy does not require denying that the system is capable; it requires resisting claims that exceed the evidence.

The acceptance test is that base modelling and assistant behaviour are separated. A good explanation should predict behaviour. It should help explain why the model can be fluent and wrong, why wording and context matter, why retrieval helps current factual tasks, why sampling creates variability, why a tool can improve arithmetic, and why a model can perform strongly on one benchmark while failing unexpectedly on a nearby task.

This is also the boundary between mechanism and product behaviour. A specific deployed assistant may include search, memory, code execution, safety systems and routing across different models. Those features matter, but they are not the base language model itself. Keep the stack visible: model → context → post-training → retrieval/tools → product policies → user task.

Reasoning behaviour can be scaffolded

Reasoning behaviour can be scaffolded is mainly about using training, prompting and tools to improve multi-step problem solving. The mechanism matters because a user’s mental model changes how they interpret output. If a person imagines a hidden database, they may trust a fabricated citation. If they imagine human memory, they may misunderstand context limits. If they imagine consciousness, they may attribute motives where a simpler explanation is token prediction shaped by training and context.

A practical example is decomposing a task and checking intermediate results. The example should be read mechanically. What information is represented? Which operation transforms it? What is learned during training and what arrives only at inference time? Which part is probabilistic? Which part comes from an external tool? Breaking the system into these questions prevents the word intelligence from doing too much explanatory work.

The common failure is assuming fluent explanation proves every internal step is faithful. LLMs are especially easy to anthropomorphise because language is the medium humans use to reveal thought. A machine that produces coherent language therefore feels as though it must contain a human-like inner process. Mechanistic literacy does not require denying that the system is capable; it requires resisting claims that exceed the evidence.

The acceptance test is that observable reasoning performance is separated from claims about hidden mechanism. A good explanation should predict behaviour. It should help explain why the model can be fluent and wrong, why wording and context matter, why retrieval helps current factual tasks, why sampling creates variability, why a tool can improve arithmetic, and why a model can perform strongly on one benchmark while failing unexpectedly on a nearby task.

This is also the boundary between mechanism and product behaviour. A specific deployed assistant may include search, memory, code execution, safety systems and routing across different models. Those features matter, but they are not the base language model itself. Keep the stack visible: model → context → post-training → retrieval/tools → product policies → user task.

Sampling turns probabilities into output

Sampling turns probabilities into output is mainly about selecting tokens from a predicted distribution rather than always choosing one deterministic next token. The mechanism matters because a user’s mental model changes how they interpret output. If a person imagines a hidden database, they may trust a fabricated citation. If they imagine human memory, they may misunderstand context limits. If they imagine consciousness, they may attribute motives where a simpler explanation is token prediction shaped by training and context.

A practical example is producing different valid phrasings from the same prompt. The example should be read mechanically. What information is represented? Which operation transforms it? What is learned during training and what arrives only at inference time? Which part is probabilistic? Which part comes from an external tool? Breaking the system into these questions prevents the word intelligence from doing too much explanatory work.

The common failure is treating variability as evidence that the model has moods or intentions. LLMs are especially easy to anthropomorphise because language is the medium humans use to reveal thought. A machine that produces coherent language therefore feels as though it must contain a human-like inner process. Mechanistic literacy does not require denying that the system is capable; it requires resisting claims that exceed the evidence.

The acceptance test is that randomness is understood as part of decoding. A good explanation should predict behaviour. It should help explain why the model can be fluent and wrong, why wording and context matter, why retrieval helps current factual tasks, why sampling creates variability, why a tool can improve arithmetic, and why a model can perform strongly on one benchmark while failing unexpectedly on a nearby task.

This is also the boundary between mechanism and product behaviour. A specific deployed assistant may include search, memory, code execution, safety systems and routing across different models. Those features matter, but they are not the base language model itself. Keep the stack visible: model → context → post-training → retrieval/tools → product policies → user task.

Temperature changes sampling behaviour

Temperature changes sampling behaviour is mainly about adjusting how strongly output favours high-probability tokens. The mechanism matters because a user’s mental model changes how they interpret output. If a person imagines a hidden database, they may trust a fabricated citation. If they imagine human memory, they may misunderstand context limits. If they imagine consciousness, they may attribute motives where a simpler explanation is token prediction shaped by training and context.

A practical example is using lower-temperature decoding for more stable outputs in some settings. The example should be read mechanically. What information is represented? Which operation transforms it? What is learned during training and what arrives only at inference time? Which part is probabilistic? Which part comes from an external tool? Breaking the system into these questions prevents the word intelligence from doing too much explanatory work.

The common failure is assuming temperature changes the underlying knowledge in the model. LLMs are especially easy to anthropomorphise because language is the medium humans use to reveal thought. A machine that produces coherent language therefore feels as though it must contain a human-like inner process. Mechanistic literacy does not require denying that the system is capable; it requires resisting claims that exceed the evidence.

The acceptance test is that decoding settings are separated from learned parameters. A good explanation should predict behaviour. It should help explain why the model can be fluent and wrong, why wording and context matter, why retrieval helps current factual tasks, why sampling creates variability, why a tool can improve arithmetic, and why a model can perform strongly on one benchmark while failing unexpectedly on a nearby task.

This is also the boundary between mechanism and product behaviour. A specific deployed assistant may include search, memory, code execution, safety systems and routing across different models. Those features matter, but they are not the base language model itself. Keep the stack visible: model → context → post-training → retrieval/tools → product policies → user task.

Hallucination is a generation failure, not a mystical event

Hallucination is a generation failure, not a mystical event is mainly about producing plausible text that is unsupported, false or mismatched to the requested evidence. The mechanism matters because a user’s mental model changes how they interpret output. If a person imagines a hidden database, they may trust a fabricated citation. If they imagine human memory, they may misunderstand context limits. If they imagine consciousness, they may attribute motives where a simpler explanation is token prediction shaped by training and context.

A practical example is inventing a citation because the continuation pattern looks plausible. The example should be read mechanically. What information is represented? Which operation transforms it? What is learned during training and what arrives only at inference time? Which part is probabilistic? Which part comes from an external tool? Breaking the system into these questions prevents the word intelligence from doing too much explanatory work.

The common failure is assuming fluent language guarantees factual grounding. LLMs are especially easy to anthropomorphise because language is the medium humans use to reveal thought. A machine that produces coherent language therefore feels as though it must contain a human-like inner process. Mechanistic literacy does not require denying that the system is capable; it requires resisting claims that exceed the evidence.

The acceptance test is that truth-sensitive tasks include retrieval or verification. A good explanation should predict behaviour. It should help explain why the model can be fluent and wrong, why wording and context matter, why retrieval helps current factual tasks, why sampling creates variability, why a tool can improve arithmetic, and why a model can perform strongly on one benchmark while failing unexpectedly on a nearby task.

This is also the boundary between mechanism and product behaviour. A specific deployed assistant may include search, memory, code execution, safety systems and routing across different models. Those features matter, but they are not the base language model itself. Keep the stack visible: model → context → post-training → retrieval/tools → product policies → user task.

Retrieval-Augmented Generation adds external evidence

Retrieval-Augmented Generation adds external evidence is mainly about retrieving documents and placing relevant passages into context before generation. The mechanism matters because a user’s mental model changes how they interpret output. If a person imagines a hidden database, they may trust a fabricated citation. If they imagine human memory, they may misunderstand context limits. If they imagine consciousness, they may attribute motives where a simpler explanation is token prediction shaped by training and context.

A practical example is answering a policy question from current official material rather than model memory alone. The example should be read mechanically. What information is represented? Which operation transforms it? What is learned during training and what arrives only at inference time? Which part is probabilistic? Which part comes from an external tool? Breaking the system into these questions prevents the word intelligence from doing too much explanatory work.

The common failure is assuming retrieval automatically makes every answer correct. LLMs are especially easy to anthropomorphise because language is the medium humans use to reveal thought. A machine that produces coherent language therefore feels as though it must contain a human-like inner process. Mechanistic literacy does not require denying that the system is capable; it requires resisting claims that exceed the evidence.

The acceptance test is that retrieval quality, context relevance and answer faithfulness are checked separately. A good explanation should predict behaviour. It should help explain why the model can be fluent and wrong, why wording and context matter, why retrieval helps current factual tasks, why sampling creates variability, why a tool can improve arithmetic, and why a model can perform strongly on one benchmark while failing unexpectedly on a nearby task.

This is also the boundary between mechanism and product behaviour. A specific deployed assistant may include search, memory, code execution, safety systems and routing across different models. Those features matter, but they are not the base language model itself. Keep the stack visible: model → context → post-training → retrieval/tools → product policies → user task.

Tool use extends the model beyond text generation

Tool use extends the model beyond text generation is mainly about allowing a model to call calculators, search, code or external systems under controlled interfaces. The mechanism matters because a user’s mental model changes how they interpret output. If a person imagines a hidden database, they may trust a fabricated citation. If they imagine human memory, they may misunderstand context limits. If they imagine consciousness, they may attribute motives where a simpler explanation is token prediction shaped by training and context.

A practical example is using a calculator for exact arithmetic instead of relying on token prediction. The example should be read mechanically. What information is represented? Which operation transforms it? What is learned during training and what arrives only at inference time? Which part is probabilistic? Which part comes from an external tool? Breaking the system into these questions prevents the word intelligence from doing too much explanatory work.

The common failure is assuming the model internally becomes the tool it calls. LLMs are especially easy to anthropomorphise because language is the medium humans use to reveal thought. A machine that produces coherent language therefore feels as though it must contain a human-like inner process. Mechanistic literacy does not require denying that the system is capable; it requires resisting claims that exceed the evidence.

The acceptance test is that language-model planning and external execution are separated. A good explanation should predict behaviour. It should help explain why the model can be fluent and wrong, why wording and context matter, why retrieval helps current factual tasks, why sampling creates variability, why a tool can improve arithmetic, and why a model can perform strongly on one benchmark while failing unexpectedly on a nearby task.

This is also the boundary between mechanism and product behaviour. A specific deployed assistant may include search, memory, code execution, safety systems and routing across different models. Those features matter, but they are not the base language model itself. Keep the stack visible: model → context → post-training → retrieval/tools → product policies → user task.

Agents combine models, tools and loops

Agents combine models, tools and loops is mainly about letting a model choose actions, observe results and continue toward a goal. The mechanism matters because a user’s mental model changes how they interpret output. If a person imagines a hidden database, they may trust a fabricated citation. If they imagine human memory, they may misunderstand context limits. If they imagine consciousness, they may attribute motives where a simpler explanation is token prediction shaped by training and context.

A practical example is searching, reading, updating a document and checking completion across multiple steps. The example should be read mechanically. What information is represented? Which operation transforms it? What is learned during training and what arrives only at inference time? Which part is probabilistic? Which part comes from an external tool? Breaking the system into these questions prevents the word intelligence from doing too much explanatory work.

The common failure is calling any chatbot an autonomous agent. LLMs are especially easy to anthropomorphise because language is the medium humans use to reveal thought. A machine that produces coherent language therefore feels as though it must contain a human-like inner process. Mechanistic literacy does not require denying that the system is capable; it requires resisting claims that exceed the evidence.

The acceptance test is that agency is defined by an action loop and permissions. A good explanation should predict behaviour. It should help explain why the model can be fluent and wrong, why wording and context matter, why retrieval helps current factual tasks, why sampling creates variability, why a tool can improve arithmetic, and why a model can perform strongly on one benchmark while failing unexpectedly on a nearby task.

This is also the boundary between mechanism and product behaviour. A specific deployed assistant may include search, memory, code execution, safety systems and routing across different models. Those features matter, but they are not the base language model itself. Keep the stack visible: model → context → post-training → retrieval/tools → product policies → user task.

Memory systems are usually external to base model weights

Memory systems are usually external to base model weights is mainly about storing selected information in databases, files or retrieval stores for later use. The mechanism matters because a user’s mental model changes how they interpret output. If a person imagines a hidden database, they may trust a fabricated citation. If they imagine human memory, they may misunderstand context limits. If they imagine consciousness, they may attribute motives where a simpler explanation is token prediction shaped by training and context.

A practical example is saving a user preference in an external memory layer. The example should be read mechanically. What information is represented? Which operation transforms it? What is learned during training and what arrives only at inference time? Which part is probabilistic? Which part comes from an external tool? Breaking the system into these questions prevents the word intelligence from doing too much explanatory work.

The common failure is assuming the base model remembers every previous conversation by default. LLMs are especially easy to anthropomorphise because language is the medium humans use to reveal thought. A machine that produces coherent language therefore feels as though it must contain a human-like inner process. Mechanistic literacy does not require denying that the system is capable; it requires resisting claims that exceed the evidence.

The acceptance test is that product memory architecture and model parameters are not conflated. A good explanation should predict behaviour. It should help explain why the model can be fluent and wrong, why wording and context matter, why retrieval helps current factual tasks, why sampling creates variability, why a tool can improve arithmetic, and why a model can perform strongly on one benchmark while failing unexpectedly on a nearby task.

This is also the boundary between mechanism and product behaviour. A specific deployed assistant may include search, memory, code execution, safety systems and routing across different models. Those features matter, but they are not the base language model itself. Keep the stack visible: model → context → post-training → retrieval/tools → product policies → user task.

Fine-tuning changes weights for a narrower objective

Fine-tuning changes weights for a narrower objective is mainly about training further on curated examples to change behaviour or domain performance. The mechanism matters because a user’s mental model changes how they interpret output. If a person imagines a hidden database, they may trust a fabricated citation. If they imagine human memory, they may misunderstand context limits. If they imagine consciousness, they may attribute motives where a simpler explanation is token prediction shaped by training and context.

A practical example is adapting a model to a specialised response format. The example should be read mechanically. What information is represented? Which operation transforms it? What is learned during training and what arrives only at inference time? Which part is probabilistic? Which part comes from an external tool? Breaking the system into these questions prevents the word intelligence from doing too much explanatory work.

The common failure is using fine-tuning to upload a constantly changing knowledge base. LLMs are especially easy to anthropomorphise because language is the medium humans use to reveal thought. A machine that produces coherent language therefore feels as though it must contain a human-like inner process. Mechanistic literacy does not require denying that the system is capable; it requires resisting claims that exceed the evidence.

The acceptance test is that weight adaptation is used for behaviour or capability rather than every current fact. A good explanation should predict behaviour. It should help explain why the model can be fluent and wrong, why wording and context matter, why retrieval helps current factual tasks, why sampling creates variability, why a tool can improve arithmetic, and why a model can perform strongly on one benchmark while failing unexpectedly on a nearby task.

This is also the boundary between mechanism and product behaviour. A specific deployed assistant may include search, memory, code execution, safety systems and routing across different models. Those features matter, but they are not the base language model itself. Keep the stack visible: model → context → post-training → retrieval/tools → product policies → user task.

Embeddings and semantic search support retrieval

Embeddings and semantic search support retrieval is mainly about representing passages and queries numerically so related material can be found. The mechanism matters because a user’s mental model changes how they interpret output. If a person imagines a hidden database, they may trust a fabricated citation. If they imagine human memory, they may misunderstand context limits. If they imagine consciousness, they may attribute motives where a simpler explanation is token prediction shaped by training and context.

A practical example is retrieving a policy passage whose wording differs from the user’s question. The example should be read mechanically. What information is represented? Which operation transforms it? What is learned during training and what arrives only at inference time? Which part is probabilistic? Which part comes from an external tool? Breaking the system into these questions prevents the word intelligence from doing too much explanatory work.

The common failure is assuming embedding similarity proves factual relevance. LLMs are especially easy to anthropomorphise because language is the medium humans use to reveal thought. A machine that produces coherent language therefore feels as though it must contain a human-like inner process. Mechanistic literacy does not require denying that the system is capable; it requires resisting claims that exceed the evidence.

The acceptance test is that retrieval results are still inspected. A good explanation should predict behaviour. It should help explain why the model can be fluent and wrong, why wording and context matter, why retrieval helps current factual tasks, why sampling creates variability, why a tool can improve arithmetic, and why a model can perform strongly on one benchmark while failing unexpectedly on a nearby task.

This is also the boundary between mechanism and product behaviour. A specific deployed assistant may include search, memory, code execution, safety systems and routing across different models. Those features matter, but they are not the base language model itself. Keep the stack visible: model → context → post-training → retrieval/tools → product policies → user task.

Longer context does not guarantee better use

Longer context does not guarantee better use is mainly about recognising that models can struggle to use every relevant detail in a very long prompt. The mechanism matters because a user’s mental model changes how they interpret output. If a person imagines a hidden database, they may trust a fabricated citation. If they imagine human memory, they may misunderstand context limits. If they imagine consciousness, they may attribute motives where a simpler explanation is token prediction shaped by training and context.

A practical example is placing critical instructions and evidence clearly instead of dumping an entire archive. The example should be read mechanically. What information is represented? Which operation transforms it? What is learned during training and what arrives only at inference time? Which part is probabilistic? Which part comes from an external tool? Breaking the system into these questions prevents the word intelligence from doing too much explanatory work.

The common failure is assuming more tokens always improve accuracy. LLMs are especially easy to anthropomorphise because language is the medium humans use to reveal thought. A machine that produces coherent language therefore feels as though it must contain a human-like inner process. Mechanistic literacy does not require denying that the system is capable; it requires resisting claims that exceed the evidence.

The acceptance test is that context quality and structure matter as well as size. A good explanation should predict behaviour. It should help explain why the model can be fluent and wrong, why wording and context matter, why retrieval helps current factual tasks, why sampling creates variability, why a tool can improve arithmetic, and why a model can perform strongly on one benchmark while failing unexpectedly on a nearby task.

This is also the boundary between mechanism and product behaviour. A specific deployed assistant may include search, memory, code execution, safety systems and routing across different models. Those features matter, but they are not the base language model itself. Keep the stack visible: model → context → post-training → retrieval/tools → product policies → user task.

Mixture-of-experts changes compute allocation

Mixture-of-experts changes compute allocation is mainly about activating subsets of model components for a token in some architectures. The mechanism matters because a user’s mental model changes how they interpret output. If a person imagines a hidden database, they may trust a fabricated citation. If they imagine human memory, they may misunderstand context limits. If they imagine consciousness, they may attribute motives where a simpler explanation is token prediction shaped by training and context.

A practical example is an MoE model routing a token through selected experts rather than all parameters. The example should be read mechanically. What information is represented? Which operation transforms it? What is learned during training and what arrives only at inference time? Which part is probabilistic? Which part comes from an external tool? Breaking the system into these questions prevents the word intelligence from doing too much explanatory work.

The common failure is assuming all models use the same architecture internally. LLMs are especially easy to anthropomorphise because language is the medium humans use to reveal thought. A machine that produces coherent language therefore feels as though it must contain a human-like inner process. Mechanistic literacy does not require denying that the system is capable; it requires resisting claims that exceed the evidence.

The acceptance test is that architecture-specific claims are kept bounded. A good explanation should predict behaviour. It should help explain why the model can be fluent and wrong, why wording and context matter, why retrieval helps current factual tasks, why sampling creates variability, why a tool can improve arithmetic, and why a model can perform strongly on one benchmark while failing unexpectedly on a nearby task.

This is also the boundary between mechanism and product behaviour. A specific deployed assistant may include search, memory, code execution, safety systems and routing across different models. Those features matter, but they are not the base language model itself. Keep the stack visible: model → context → post-training → retrieval/tools → product policies → user task.

Multimodal models extend the same learning idea

Multimodal models extend the same learning idea is mainly about learning relationships across text, images, audio or other representations. The mechanism matters because a user’s mental model changes how they interpret output. If a person imagines a hidden database, they may trust a fabricated citation. If they imagine human memory, they may misunderstand context limits. If they imagine consciousness, they may attribute motives where a simpler explanation is token prediction shaped by training and context.

A practical example is using image patches and text tokens inside one model family. The example should be read mechanically. What information is represented? Which operation transforms it? What is learned during training and what arrives only at inference time? Which part is probabilistic? Which part comes from an external tool? Breaking the system into these questions prevents the word intelligence from doing too much explanatory work.

The common failure is assuming multimodal understanding works by translating every input to words first. LLMs are especially easy to anthropomorphise because language is the medium humans use to reveal thought. A machine that produces coherent language therefore feels as though it must contain a human-like inner process. Mechanistic literacy does not require denying that the system is capable; it requires resisting claims that exceed the evidence.

The acceptance test is that the representation can span multiple modalities. A good explanation should predict behaviour. It should help explain why the model can be fluent and wrong, why wording and context matter, why retrieval helps current factual tasks, why sampling creates variability, why a tool can improve arithmetic, and why a model can perform strongly on one benchmark while failing unexpectedly on a nearby task.

This is also the boundary between mechanism and product behaviour. A specific deployed assistant may include search, memory, code execution, safety systems and routing across different models. Those features matter, but they are not the base language model itself. Keep the stack visible: model → context → post-training → retrieval/tools → product policies → user task.

Capabilities can emerge from combination

Capabilities can emerge from combination is mainly about seeing useful behaviour as the interaction of learned representations, scale, context, post-training and tools. The mechanism matters because a user’s mental model changes how they interpret output. If a person imagines a hidden database, they may trust a fabricated citation. If they imagine human memory, they may misunderstand context limits. If they imagine consciousness, they may attribute motives where a simpler explanation is token prediction shaped by training and context.

A practical example is a model solving a coding task through pretrained patterns, instructions, retrieved docs and execution feedback. The example should be read mechanically. What information is represented? Which operation transforms it? What is learned during training and what arrives only at inference time? Which part is probabilistic? Which part comes from an external tool? Breaking the system into these questions prevents the word intelligence from doing too much explanatory work.

The common failure is searching for one magic component that explains intelligence. LLMs are especially easy to anthropomorphise because language is the medium humans use to reveal thought. A machine that produces coherent language therefore feels as though it must contain a human-like inner process. Mechanistic literacy does not require denying that the system is capable; it requires resisting claims that exceed the evidence.

The acceptance test is that mechanism is treated as a system. A good explanation should predict behaviour. It should help explain why the model can be fluent and wrong, why wording and context matter, why retrieval helps current factual tasks, why sampling creates variability, why a tool can improve arithmetic, and why a model can perform strongly on one benchmark while failing unexpectedly on a nearby task.

This is also the boundary between mechanism and product behaviour. A specific deployed assistant may include search, memory, code execution, safety systems and routing across different models. Those features matter, but they are not the base language model itself. Keep the stack visible: model → context → post-training → retrieval/tools → product policies → user task.

Benchmarks are samples, not reality

Benchmarks are samples, not reality is mainly about using tests as bounded evidence of performance. The mechanism matters because a user’s mental model changes how they interpret output. If a person imagines a hidden database, they may trust a fabricated citation. If they imagine human memory, they may misunderstand context limits. If they imagine consciousness, they may attribute motives where a simpler explanation is token prediction shaped by training and context.

A practical example is comparing models on a defined benchmark while checking real task behaviour separately. The example should be read mechanically. What information is represented? Which operation transforms it? What is learned during training and what arrives only at inference time? Which part is probabilistic? Which part comes from an external tool? Breaking the system into these questions prevents the word intelligence from doing too much explanatory work.

The common failure is treating one score as universal intelligence. LLMs are especially easy to anthropomorphise because language is the medium humans use to reveal thought. A machine that produces coherent language therefore feels as though it must contain a human-like inner process. Mechanistic literacy does not require denying that the system is capable; it requires resisting claims that exceed the evidence.

The acceptance test is that measurement claims stay tied to tasks and conditions. A good explanation should predict behaviour. It should help explain why the model can be fluent and wrong, why wording and context matter, why retrieval helps current factual tasks, why sampling creates variability, why a tool can improve arithmetic, and why a model can perform strongly on one benchmark while failing unexpectedly on a nearby task.

This is also the boundary between mechanism and product behaviour. A specific deployed assistant may include search, memory, code execution, safety systems and routing across different models. Those features matter, but they are not the base language model itself. Keep the stack visible: model → context → post-training → retrieval/tools → product policies → user task.

Evaluation should include failure cases

Evaluation should include failure cases is mainly about testing where the model is likely to be confidently wrong. The mechanism matters because a user’s mental model changes how they interpret output. If a person imagines a hidden database, they may trust a fabricated citation. If they imagine human memory, they may misunderstand context limits. If they imagine consciousness, they may attribute motives where a simpler explanation is token prediction shaped by training and context.

A practical example is using adversarial examples, rare cases and current facts in evaluation. The example should be read mechanically. What information is represented? Which operation transforms it? What is learned during training and what arrives only at inference time? Which part is probabilistic? Which part comes from an external tool? Breaking the system into these questions prevents the word intelligence from doing too much explanatory work.

The common failure is showing only successful demonstrations. LLMs are especially easy to anthropomorphise because language is the medium humans use to reveal thought. A machine that produces coherent language therefore feels as though it must contain a human-like inner process. Mechanistic literacy does not require denying that the system is capable; it requires resisting claims that exceed the evidence.

The acceptance test is that reliability includes knowing the boundary. A good explanation should predict behaviour. It should help explain why the model can be fluent and wrong, why wording and context matter, why retrieval helps current factual tasks, why sampling creates variability, why a tool can improve arithmetic, and why a model can perform strongly on one benchmark while failing unexpectedly on a nearby task.

This is also the boundary between mechanism and product behaviour. A specific deployed assistant may include search, memory, code execution, safety systems and routing across different models. Those features matter, but they are not the base language model itself. Keep the stack visible: model → context → post-training → retrieval/tools → product policies → user task.

LLMs do not automatically know current events

LLMs do not automatically know current events is mainly about recognising that training and live information access are different. The mechanism matters because a user’s mental model changes how they interpret output. If a person imagines a hidden database, they may trust a fabricated citation. If they imagine human memory, they may misunderstand context limits. If they imagine consciousness, they may attribute motives where a simpler explanation is token prediction shaped by training and context.

A practical example is using search or retrieval for a policy changed this week. The example should be read mechanically. What information is represented? Which operation transforms it? What is learned during training and what arrives only at inference time? Which part is probabilistic? Which part comes from an external tool? Breaking the system into these questions prevents the word intelligence from doing too much explanatory work.

The common failure is assuming a model’s general fluency means current knowledge. LLMs are especially easy to anthropomorphise because language is the medium humans use to reveal thought. A machine that produces coherent language therefore feels as though it must contain a human-like inner process. Mechanistic literacy does not require denying that the system is capable; it requires resisting claims that exceed the evidence.

The acceptance test is that freshness-sensitive tasks use live sources. A good explanation should predict behaviour. It should help explain why the model can be fluent and wrong, why wording and context matter, why retrieval helps current factual tasks, why sampling creates variability, why a tool can improve arithmetic, and why a model can perform strongly on one benchmark while failing unexpectedly on a nearby task.

This is also the boundary between mechanism and product behaviour. A specific deployed assistant may include search, memory, code execution, safety systems and routing across different models. Those features matter, but they are not the base language model itself. Keep the stack visible: model → context → post-training → retrieval/tools → product policies → user task.

LLMs do not have human experience

LLMs do not have human experience is mainly about avoiding anthropomorphic claims about feelings, intentions or lived understanding. The mechanism matters because a user’s mental model changes how they interpret output. If a person imagines a hidden database, they may trust a fabricated citation. If they imagine human memory, they may misunderstand context limits. If they imagine consciousness, they may attribute motives where a simpler explanation is token prediction shaped by training and context.

A practical example is describing an output pattern without asserting inner subjective experience. The example should be read mechanically. What information is represented? Which operation transforms it? What is learned during training and what arrives only at inference time? Which part is probabilistic? Which part comes from an external tool? Breaking the system into these questions prevents the word intelligence from doing too much explanatory work.

The common failure is inferring consciousness from natural conversation. LLMs are especially easy to anthropomorphise because language is the medium humans use to reveal thought. A machine that produces coherent language therefore feels as though it must contain a human-like inner process. Mechanistic literacy does not require denying that the system is capable; it requires resisting claims that exceed the evidence.

The acceptance test is that mechanistic explanation stays separate from philosophical speculation. A good explanation should predict behaviour. It should help explain why the model can be fluent and wrong, why wording and context matter, why retrieval helps current factual tasks, why sampling creates variability, why a tool can improve arithmetic, and why a model can perform strongly on one benchmark while failing unexpectedly on a nearby task.

This is also the boundary between mechanism and product behaviour. A specific deployed assistant may include search, memory, code execution, safety systems and routing across different models. Those features matter, but they are not the base language model itself. Keep the stack visible: model → context → post-training → retrieval/tools → product policies → user task.

Interpretability is incomplete

Interpretability is incomplete is mainly about acknowledging that researchers can inspect activations, features and circuits without possessing a full human-readable map of model cognition. The mechanism matters because a user’s mental model changes how they interpret output. If a person imagines a hidden database, they may trust a fabricated citation. If they imagine human memory, they may misunderstand context limits. If they imagine consciousness, they may attribute motives where a simpler explanation is token prediction shaped by training and context.

A practical example is using tools such as logit-lens or sparse-autoencoder analyses as partial probes. The example should be read mechanically. What information is represented? Which operation transforms it? What is learned during training and what arrives only at inference time? Which part is probabilistic? Which part comes from an external tool? Breaking the system into these questions prevents the word intelligence from doing too much explanatory work.

The common failure is claiming every learned concept can be located cleanly. LLMs are especially easy to anthropomorphise because language is the medium humans use to reveal thought. A machine that produces coherent language therefore feels as though it must contain a human-like inner process. Mechanistic literacy does not require denying that the system is capable; it requires resisting claims that exceed the evidence.

The acceptance test is that the page distinguishes evidence from open questions. A good explanation should predict behaviour. It should help explain why the model can be fluent and wrong, why wording and context matter, why retrieval helps current factual tasks, why sampling creates variability, why a tool can improve arithmetic, and why a model can perform strongly on one benchmark while failing unexpectedly on a nearby task.

This is also the boundary between mechanism and product behaviour. A specific deployed assistant may include search, memory, code execution, safety systems and routing across different models. Those features matter, but they are not the base language model itself. Keep the stack visible: model → context → post-training → retrieval/tools → product policies → user task.

Safety is partly a systems problem

Safety is partly a systems problem is mainly about combining model training, policy, tools, monitoring and permission boundaries. The mechanism matters because a user’s mental model changes how they interpret output. If a person imagines a hidden database, they may trust a fabricated citation. If they imagine human memory, they may misunderstand context limits. If they imagine consciousness, they may attribute motives where a simpler explanation is token prediction shaped by training and context.

A practical example is restricting what an agent can execute even if the language model proposes an unsafe action. The example should be read mechanically. What information is represented? Which operation transforms it? What is learned during training and what arrives only at inference time? Which part is probabilistic? Which part comes from an external tool? Breaking the system into these questions prevents the word intelligence from doing too much explanatory work.

The common failure is expecting one prompt or one classifier to solve all risk. LLMs are especially easy to anthropomorphise because language is the medium humans use to reveal thought. A machine that produces coherent language therefore feels as though it must contain a human-like inner process. Mechanistic literacy does not require denying that the system is capable; it requires resisting claims that exceed the evidence.

The acceptance test is that controls exist at multiple layers. A good explanation should predict behaviour. It should help explain why the model can be fluent and wrong, why wording and context matter, why retrieval helps current factual tasks, why sampling creates variability, why a tool can improve arithmetic, and why a model can perform strongly on one benchmark while failing unexpectedly on a nearby task.

This is also the boundary between mechanism and product behaviour. A specific deployed assistant may include search, memory, code execution, safety systems and routing across different models. Those features matter, but they are not the base language model itself. Keep the stack visible: model → context → post-training → retrieval/tools → product policies → user task.

Useful LLM literacy is operational

Useful LLM literacy is operational is mainly about helping users decide when to trust, verify, retrieve, calculate or escalate. The mechanism matters because a user’s mental model changes how they interpret output. If a person imagines a hidden database, they may trust a fabricated citation. If they imagine human memory, they may misunderstand context limits. If they imagine consciousness, they may attribute motives where a simpler explanation is token prediction shaped by training and context.

A practical example is checking a generated answer against a primary source when the stakes are high. The example should be read mechanically. What information is represented? Which operation transforms it? What is learned during training and what arrives only at inference time? Which part is probabilistic? Which part comes from an external tool? Breaking the system into these questions prevents the word intelligence from doing too much explanatory work.

The common failure is memorising architecture terms without changing how the tool is used. LLMs are especially easy to anthropomorphise because language is the medium humans use to reveal thought. A machine that produces coherent language therefore feels as though it must contain a human-like inner process. Mechanistic literacy does not require denying that the system is capable; it requires resisting claims that exceed the evidence.

The acceptance test is that mechanism knowledge improves decisions. A good explanation should predict behaviour. It should help explain why the model can be fluent and wrong, why wording and context matter, why retrieval helps current factual tasks, why sampling creates variability, why a tool can improve arithmetic, and why a model can perform strongly on one benchmark while failing unexpectedly on a nearby task.

This is also the boundary between mechanism and product behaviour. A specific deployed assistant may include search, memory, code execution, safety systems and routing across different models. Those features matter, but they are not the base language model itself. Keep the stack visible: model → context → post-training → retrieval/tools → product policies → user task.

The shortest useful mechanism

  • Text is converted into tokens.
  • Tokens are represented as vectors and processed through many Transformer-style layers.
  • Attention operations let positions exchange information across the context.
  • Training adjusts large sets of parameters to reduce prediction error across enormous numbers of examples.
  • At generation time, the model predicts a probability distribution for the next token.
  • A decoding rule selects a token, appends it to the context and repeats the process.
  • Post-training shapes instruction following, style and safety behaviour.
  • Retrieval and tools can provide current evidence or capabilities the base model does not reliably supply on its own.

This simplified account omits many engineering details, but it is powerful because it explains the central paradox of LLMs: very simple local training and generation objectives can support complex global behaviour after enough representation learning, scale, post-training and system design. The behaviour is real; the explanation still needs to remain precise.

What this mechanism does not prove

  • It does not prove that every generated statement is true.
  • It does not prove that model reasoning is identical to human reasoning.
  • It does not prove that a model is conscious or has subjective experience.
  • It does not prove that larger models are always better for every task.
  • It does not prove that retrieval eliminates hallucination.
  • It does not prove that benchmark scores transfer perfectly to real-world use.
  • It does not prove that a model has current knowledge unless a live information path is present.
  • It does not prove that natural language explanations reveal the model’s exact internal causal process.

A practical LLM reliability workflow

  • Classify the task: creative, explanatory, computational, factual-current, high-stakes or action-taking.
  • Identify what the base model can plausibly do from learned patterns and what requires external evidence or tools.
  • For current facts, retrieve primary sources.
  • For exact arithmetic or code execution, use appropriate tools and inspect results.
  • For consequential decisions, define verification and human accountability before generation.
  • Ask for assumptions and uncertainty when the task is underspecified.
  • Test changed conditions instead of accepting one polished answer.
  • Keep product-specific claims dated because interfaces, models and tool access change quickly.

How this connects to the wider eduKateSG AI map

Use How AI Works for the broader artificial-intelligence system, How the Logit Lens Works for one interpretability probe, and Job Redesign and AI for the work-design consequences. The mechanism page stays narrow enough that these owners can remain distinct.

Frequently asked questions

Does an LLM search the internet every time it answers?

No. A base model generates from learned parameters and the current context. A product may separately add web search or retrieval. When live search is present, the external evidence path should be distinguished from model memory.

Does an LLM store copies of its training documents?

The useful default model is learned parameters rather than a document database. OpenAI’s current explanation says models adjust weights to learn patterns rather than retaining copies of training items as a simple store. Memorisation can still be a separate research and privacy concern, so do not turn this simplification into an absolute claim.

What is a token?

A token is a unit the model processes. It may correspond to a word, part of a word, punctuation or another text fragment depending on the tokenizer.

What does attention do?

Attention is a mathematical mechanism that lets token representations combine information from other positions in the context. It is not the same thing as human conscious attention.

Why do LLMs hallucinate?

The generation objective produces plausible continuations, not a built-in guarantee of factual truth. When the model lacks reliable evidence or the task demands current facts, it can generate fluent but unsupported content. Retrieval, tools and verification reduce risk but do not make it zero.

What is the difference between pre-training and post-training?

Pre-training builds broad language and representation capability from large-scale prediction objectives. Post-training further shapes behaviour for instructions, preferences, safety or specific tasks.

What is Retrieval-Augmented Generation?

RAG retrieves external documents and places relevant information into the model’s context before generation. Quality depends on retrieval, context selection and whether the final answer stays faithful to the evidence.

Are all LLMs Transformers?

Most modern large language models use Transformer-family ideas, but architectures differ. Some use mixture-of-experts routing, sparse attention or multimodal components. Avoid assuming one implementation describes every model.

Do LLMs reason?

They can perform tasks that require multi-step problem solving and can produce useful reasoning-like behaviour. What internal processes deserve the word reasoning is an active scientific question. It is safer to describe observed capability and mechanism separately.

Where can I read a current public explanation?

See OpenAI’s How ChatGPT and our foundation models are developed and the original Transformer paper Attention Is All You Need.

The quiet conclusion: fluency is the surface, mechanism is the discipline

Large language models are remarkable because next-token prediction, learned at scale inside deep neural networks, becomes capable of supporting behaviours that look far richer than the training objective sounds. The result can write, explain, translate, code, plan and interact with tools. That capability deserves serious study precisely because it is easy to misunderstand.

The most useful literacy is neither awe nor dismissal. It is the ability to locate the mechanism and then choose the right control. Tokens explain context. Prediction explains variability. Training explains broad pattern knowledge. Retrieval explains access to fresh evidence. Tools explain exact external actions. Verification explains why fluent output still needs checking. Once the stack is visible, the user can stop asking whether the machine is simply smart and start asking a better question: what part of this result came from which mechanism, and what evidence is strong enough to trust it?