VIEW THIS AS

Auto mode follows the Route Engine until you choose a viewpoint.

YOU ARE HERE

ROUTE CHECK

CONNECTED TO

WHAT NEXT

Use the canonical route for this room, or HELP if you are unsure.

How Super Intelligence Works | Pretraining — How Models Learn Broad Capability Before the User Arrives

eduKate Secondary students reviewing open books for How Super Intelligence Works: Attention.

Pretraining is the stage where a model develops broad capability before a particular user arrives with a particular task. A large language model is not normally taught every future question one by one. Instead, it is trained on large collections of data using an objective that rewards better prediction. The resulting parameters become a reusable foundation for many downstream tasks.

This is why the same model can later summarise an article, classify a message, draft code, compare policies or answer a question without being retrained from scratch for every interaction. The broad patterns were learned earlier. Runtime context tells the model what to do now.

This article explains how pretraining works: data preparation, token sequences, prediction objectives, loss, batches, parameter updates, epochs, checkpoints, validation, scaling, transfer and the difference between pretraining, fine-tuning, post-training and in-context learning. It also explains what pretraining does not guarantee.

The Stanford report On the Opportunities and Risks of Foundation Models describes broadly trained models that can be adapted to many downstream tasks. The word foundation is useful because pretraining creates a general base; it does not finish every application.

Previous: 020 — Context Windows. This article begins Part 3 of the series: Training and Machine Learning.


The Hidden Transition: From Raw Examples to Reusable Capability

Imagine teaching a child every possible future sentence by memorising answers. That would be impossible. A model instead encounters enormous numbers of sequences and learns statistical regularities that help it predict what comes next, reconstruct missing information or solve another training objective.

During pretraining, the model does not receive a list saying “this neuron means grammar” or “this parameter stores Python”. Training adjusts many parameters together so that the model’s predictions improve across examples.

The useful capability is distributed. A later prompt can activate patterns learned from many earlier examples without locating one exact training sentence.

Pretraining Versus Inference

Pretraining changes the model’s parameters. Inference uses those parameters to process a current input. This is one of the most important distinctions in SI.

If a user tells the model, “Use British spelling in this answer,” the next output can change because the instruction is part of the current context. That does not prove the model weights were updated.

The GPT-3 paper Language Models are Few-Shot Learners demonstrated tasks performed from instructions and examples in context without gradient updates during those evaluations. In-context adaptation and parameter training are different mechanisms.

A Simple Autoregressive Pretraining Objective

Consider the sentence: “The cat sat on the mat.” An autoregressive model can be trained to predict each next token from the tokens before it. After “The”, predict “cat”. After “The cat”, predict “sat”. The training objective measures how much probability the model assigned to the correct continuation.

Real training uses tokenized sequences rather than whole English words, and the corpus is vastly larger. The principle remains: the model repeatedly predicts targets, receives a loss signal and updates parameters so future predictions improve.

This objective creates enormous amounts of training signal from ordinary text. The data does not need a human to label every sentence with a task category.

Loss Turns Prediction Quality Into a Training Signal

A training objective needs a numerical measure of error. For language modeling, cross-entropy loss is commonly used to penalise low probability on the correct target token.

Suppose the correct next token is “mat”. Model A assigns it probability 0.8 while Model B assigns 0.1. The loss for Model A is lower because it placed more probability on the observed target.

The optimiser uses gradients of this loss to determine how parameters should change. Training is therefore a repeated loop: predict, calculate loss, compute gradients, update parameters, continue.

One Toy Training Step

Take a tiny fictional vocabulary: cat, dog, sat, ran, mat. The training example is “cat sat”. Assume the model initially predicts the next-token distribution after “cat” as sat 0.30, ran 0.25, dog 0.20, mat 0.15, cat 0.10.

The observed target is “sat”. The loss penalises the model because only 0.30 probability was assigned to the correct target. Backpropagation computes how each relevant parameter contributed to that loss.

The optimiser then changes parameters a small amount. On a later pass, the model might assign “sat” probability 0.36 in similar contexts. One step teaches almost nothing by itself; learning emerges from enormous numbers of updates across diverse sequences.

Training Examples Are Usually Packed Into Batches

Modern accelerators work efficiently on groups of examples. Training therefore processes batches of token sequences, computes an aggregate loss and updates parameters from that batch.

Batch size affects memory use, gradient noise and optimisation behaviour. Large distributed training may divide batches across many accelerators and then combine gradient information.

The exact engineering is complex, but the conceptual unit remains stable: a batch supplies examples, the model predicts, loss is computed, gradients are accumulated and parameters are updated.

An Epoch Is a Pass Through a Defined Dataset

In smaller supervised datasets, an epoch means one full pass through the training set. Large-scale language pretraining can use enormous corpora, mixtures and sampling schemes, so the simple classroom picture of “ten clean epochs over one static file” may not describe the whole run.

Still, the term is useful for understanding repeated exposure. Seeing examples repeatedly can improve learning but can also increase memorisation or overfitting if the data and objective are poorly managed.

Pretraining Data Is a Mixture, Not One Textbook

Large models can be trained on mixtures of web pages, books, code, reference material and other permitted datasets, depending on the model and developer. Different mixtures produce different strengths and biases.

A code-heavy mixture can improve programming capability. Multilingual data can broaden language coverage. High-quality mathematical material can help mathematical representation. Data composition is therefore part of model design.

The next article, Training Data, will examine collection, filtering, duplication, provenance and contamination in depth.

Self-Supervised Learning Creates Labels From the Data Itself

Language-model pretraining is often described as self-supervised because the learning targets are derived from the data rather than manually annotated one by one. In next-token prediction, the next token already present in the sequence becomes the target.

Masked-language approaches use another pattern. BERT’s pretraining objective masks parts of text and trains the model to recover them from surrounding context. The BERT paper describes pretraining deep bidirectional representations from unlabeled text.

Different objectives create different learning pressures, but both show how raw text can generate large amounts of training supervision.

Autoregressive Versus Masked Pretraining

Autoregressive models predict future tokens from previous context. This aligns naturally with generation because inference can repeatedly predict the next token.

Masked models recover hidden parts using context on both sides and historically became powerful encoders for classification, question answering and representation tasks.

The categories are not a ranking. They represent different objectives and architectures. Modern systems use many variations, mixtures and additional objectives.

Pretraining Learns Representations, Not a List of Answers

If training only stored exact answers, the model would fail whenever wording changed. Instead, useful models learn representations that capture recurring relationships across many examples.

A model can therefore translate a sentence it never saw exactly, complete unfamiliar code or classify a new message. Generalisation is the point of training.

Memorisation can still occur, especially for repeated or distinctive sequences. Pretraining contains both learned generalities and the possibility of remembered fragments. Evaluation must distinguish useful transfer from rote reproduction.

Generalisation: Performance Beyond the Training Examples

A model is valuable when it performs well on appropriate unseen examples. Training loss alone cannot establish this because a model could simply fit the training data.

Validation and test sets estimate performance on held-out data. The evaluation set should reflect the task distribution and remain isolated enough that training contamination does not invalidate the result.

The GPT-3 paper explicitly studied data contamination because benchmark material can appear on the public web. This illustrates why evaluation provenance matters for large-scale pretraining.

Training, Validation and Test Splits

A training set updates parameters. A validation set helps choose hyperparameters, checkpoints or training decisions. A test set is reserved for final evaluation under the defined protocol.

If the model repeatedly sees the test set during tuning, the test result stops being an independent estimate. This is conceptually similar to a student practising with the exact exam paper and then being evaluated on the same questions.

Checkpoints Preserve Training Progress

Large training runs periodically save model checkpoints. A checkpoint captures parameter state at a particular stage so training can resume, models can be evaluated and earlier states can be compared.

The final checkpoint is not automatically the best for every downstream task. Developers may choose among checkpoints using validation behaviour, stability and safety evaluations.

The Optimiser Controls Parameter Updates

Gradient descent is the broad idea of moving parameters in a direction that reduces loss. Practical training commonly uses more sophisticated optimisers that adapt update behaviour.

The optimiser does not decide what the model should value. The loss function and training data define the learning signal. Optimisation simply searches parameter space under those signals.

Learning Rate Controls Update Size

A learning rate determines how strongly gradients change parameters. Google’s Machine Learning Crash Course explains that a learning rate that is too small can make convergence extremely slow, while one that is too high can cause training to bounce around or diverge.

Large-scale pretraining often uses learning-rate schedules that change over the run. The principle remains: update size is a training hyperparameter, not something learned in the same way as ordinary model weights.

Parameters Versus Hyperparameters

Parameters are learned values such as weights and biases. Hyperparameters are configuration choices set by the training process: learning rate, batch size, optimiser settings, model depth, context length and many other choices.

Changing a hyperparameter can alter what parameters the model ultimately learns. Training is therefore the interaction between data, objective, architecture, optimiser and hyperparameter choices.

Scaling Pretraining

Pretraining capability can increase when models, datasets and compute scale appropriately. Larger models can represent more complex patterns; more data can expose broader variation; more compute permits more optimisation work.

But scaling is not a universal guarantee. Poor data, unsuitable objectives, weak evaluation or system bottlenecks can limit value. The correct question is what performance changed under which controlled conditions.

Distributed Training

Very large models and datasets do not fit comfortably on one accelerator. Distributed training divides computation, parameters, data or both across many devices.

This introduces systems challenges: communication overhead, synchronisation, fault recovery, checkpointing and efficient memory use. The intelligence that users see rests on a substantial infrastructure layer.

Article 083 later in this series will treat distributed training as its own topic.

Pretraining Is Expensive; Inference Reuses the Result

One economic reason foundation models matter is that broad capability is learned during a costly training process and then reused across many downstream tasks.

A user asking for a summary does not launch a new global pretraining run. The application uses the already-trained model, adds current context and performs inference.

This separation enables one foundation model to support many products and workflows.

Pretraining Versus Fine-Tuning

Pretraining builds broad capability from large general datasets. Fine-tuning continues training on a narrower dataset to adapt behaviour toward a particular task or domain.

A pretrained model might know general legal language. Fine-tuning could adapt it for document classification. The downstream training changes parameters, unlike ordinary prompting.

Fine-tuning can improve specialised performance but can also overfit, narrow behaviour or introduce regressions. It needs its own evaluation.

Pretraining Versus Post-Training

Post-training is a broader term for additional training stages after base pretraining. It can include supervised instruction tuning, preference optimisation, safety training and task adaptation.

The base model learns broad predictive structure. Post-training shapes how that capability is exposed to users: following instructions, refusing certain requests, formatting answers and using tools.

Article 026 will examine post-training in detail.

Pretraining Versus In-Context Learning

In-context learning happens at inference time. The user supplies instructions or examples inside the context window, and the model changes its output without a parameter update.

Example: show three sentiment examples labelled POSITIVE or NEGATIVE, then present a fourth sentence. The model can continue the pattern.

The capability to do this was created by training, but the new task specification lives in context. Training creates adaptability; context directs it.

Worked Example: Pretraining to Few-Shot Classification

Imagine pretraining has exposed a model to broad English text, reviews, labels and many patterns. At runtime, the prompt says: “Classify sentiment. ‘Loved the lesson’ → POSITIVE. ‘Too confusing’ → NEGATIVE. ‘Very clear explanation’ → POSITIVE. ‘I could not follow the examples’ → ?”

The model can infer the task and output NEGATIVE without updating weights. The prompt examples are not enough to teach English from scratch. They exploit representations learned during pretraining.

What Pretraining Can Learn From Code

Code pretraining provides sequences where syntax, variable relationships, library patterns, documentation and tests interact. Predicting code tokens rewards understanding of structural regularities.

This helps later code completion, explanation and generation. Execution is still required to prove a program works.

What Pretraining Can Learn From Multilingual Data

Multilingual corpora expose relationships across languages, scripts and translations. Shared representations can support transfer between languages, especially when data coverage is strong.

Performance can vary sharply by language and domain because training data is not evenly distributed. Broad multilingual capability should therefore be measured rather than assumed.

Pretraining and Bias

A model learns from patterns in its training data, including undesirable stereotypes, imbalances and historical biases. Scaling data can broaden coverage without automatically removing those patterns.

Filtering, data balancing, post-training, evaluation and application safeguards can reduce some harms. Bias is not a single defect with one universal repair.

Pretraining and Privacy

Large datasets can accidentally contain personal or sensitive information depending on collection and filtering. Models can sometimes memorise rare sequences.

Data governance therefore matters before training, and privacy evaluation matters after training. The fact that a model stores distributed representations does not make training data provenance irrelevant.

Pretraining and Copyright

Training-data rights and licensing are legal and policy questions that depend on jurisdiction, source and use. A technical description of pretraining should not pretend those questions disappear because the objective is self-supervised.

Responsible model development needs data governance alongside optimisation.

Pretraining Does Not Make Knowledge Permanently Current

A model can only learn information available through its training and update process. Events after the training cutoff require retrieval, tools or later training to become available.

This is why a current policy, price or schedule should often come from a live source even when the model understands the domain perfectly.

Pretraining Does Not Guarantee Truth

The objective rewards prediction of training sequences, not direct philosophical truth. Training data can contain errors, fiction, persuasion and disagreement.

A model can learn to produce plausible falsehoods because false statements also occur in text. Grounding and verification remain separate system responsibilities.

Pretraining Does Not Guarantee Instruction Following

A base model trained only for next-token prediction may continue text rather than behave like a polished assistant. Instruction-following behaviour usually depends on additional post-training and application design.

This is why base model and chat model are not interchangeable terms.

Pretraining Does Not Guarantee Tool Use

A model can describe a calculator without knowing a particular application’s tool-call schema. Tool use usually requires demonstrations, post-training, prompting and an execution environment.

The tool itself remains external to model weights.

A Pretraining Failure Map

Level 1: unsuitable data mixture. Level 2: poor filtering or duplication. Level 3: weak tokenisation or representation. Level 4: objective mismatch. Level 5: optimisation instability. Level 6: overfitting or memorisation. Level 7: weak generalisation. Level 8: contaminated evaluation. Level 9: downstream system misattributes the failure to pretraining.

This diagnostic ladder is an eduKateSG teaching device. It encourages precise investigation rather than blaming “the data” or “the model” generically.

Worked Diagnosis: Great Training Loss, Weak New Tasks

Suppose training loss continues falling, but held-out evaluation stops improving. The model may be overfitting or the validation distribution may differ from training.

The repair is not automatically “train longer”. Inspect data diversity, regularisation, optimisation and evaluation design.

Worked Diagnosis: Strong Benchmark, Weak Real Users

Suppose a model scores highly on a benchmark but struggles with current internal documents. The benchmark may not represent the deployment domain, or the application may need retrieval.

Pretraining created broad capability; the downstream system still needs task-specific evidence and evaluation.

Worked Diagnosis: Apparent Benchmark Success From Contamination

If benchmark examples or close copies appear in training data, the model may reproduce them rather than generalise. This can inflate evaluation results.

Deduplication, contamination analysis and protected test sets help preserve meaningful measurement.

A Practical Pretraining Audit

Ask: what is the model family and objective? What kinds of data were used? How were sources filtered and deduplicated? Which languages and domains are strong or weak? What validation sets were used? How was contamination assessed?

Then ask: which capabilities come from base pretraining, which from post-training, and which from tools or retrieval? This prevents the entire application from being credited to one training stage.

Independent Exercise 1: Training or Context?

A user gives three examples of a new classification task and the model performs the fourth without a weight update. Did pretraining happen during the conversation?

Answer

No. The model uses in-context learning. Pretraining created parameters that can adapt to patterns in context, but the current examples do not by themselves update those parameters.

Independent Exercise 2: Base Model or Chat Behaviour?

A pretrained model predicts continuations well but ignores direct user instructions. What training stage might be missing?

Answer

Instruction-oriented post-training may be needed. Broad predictive capability and assistant-style instruction following are different behaviours.

Independent Exercise 3: Freshness

A model was pretrained before a policy changed yesterday. Should you retrain the entire model just to answer the current policy question?

Answer

Usually not. Retrieve the current authoritative policy at inference time. Retraining may be appropriate for broader persistent changes, but live retrieval is more direct for current facts.

Independent Exercise 4: Evaluation

Training loss is excellent, but test performance is poor. What does this suggest?

Answer

The model may not be generalising to the test distribution, or the evaluation pipeline may differ. Inspect overfitting, data mismatch and test integrity rather than celebrating training loss alone.

Independent Exercise 5: Data Coverage

A multilingual model performs well in English but poorly in a low-resource language. Is model size alone enough to diagnose the problem?

Answer

No. Inspect training-data coverage, tokenisation, evaluation quality and language-specific representation. Scale is only one factor.

A Complete Miniature Pretraining Run

To make the process concrete, build a toy corpus with four short sequences: “cats chase mice”, “dogs chase balls”, “cats sleep quietly”, and “dogs sleep indoors”. A real model would use tokenisation, embeddings, many layers and a much larger corpus, but this tiny dataset exposes the training loop.

First, convert each sequence into token IDs. Next, create input-target pairs for next-token prediction. From “cats chase mice”, one pair is input “cats” with target “chase”; another is input “cats chase” with target “mice”. Repeat across the corpus.

The model performs a forward pass and produces a probability distribution over the vocabulary for each target position. Suppose after “cats” it predicts chase 0.25, sleep 0.25, dogs 0.20, mice 0.15 and other tokens 0.15. The observed target is chase, so the loss is relatively high.

Backpropagation computes gradients showing how parameters contributed to the loss. The optimiser updates those parameters. Over many passes, the model increases the probability of patterns that match the training examples while trying to preserve performance across the entire dataset.

This example is deliberately tiny. A large model does not simply count bigrams. Deep layers can represent complex contextual interactions. The miniature run is useful because it shows the same control loop: data → prediction → loss → gradient → update → new prediction.

Sequence Packing and Why Training Data Is Not Always One Document at a Time

Large-scale pretraining often packs many token sequences into long training examples to use accelerator capacity efficiently. Document boundaries, special tokens and attention masks can preserve which tokens belong together.

Packing matters because the model should not accidentally treat the end of one unrelated document as the natural beginning of another unless the training format intentionally permits that. Data engineering shapes the examples the model sees.

The training pipeline therefore includes more than “download text and start learning”. It must decide how documents are cleaned, tokenised, shuffled, sampled, packed and presented.

Curriculum and Sampling Mixtures

When training data comes from several domains, the developer can sample them at different rates. A small high-quality code dataset might be sampled more often than its raw size would imply. A dominant web source might be down-weighted so it does not overwhelm smaller domains.

This creates an implicit curriculum. The model’s exposure is not always proportional to the raw number of bytes in each source. Sampling choices influence which patterns receive more gradient updates.

A responsible model report should therefore distinguish raw data composition from effective training mixture when possible.

Data Order and Randomness

Stochastic optimisation relies on batches that provide varied examples. If training examples are ordered badly—thousands of near-identical documents together, for instance—gradients can become less representative of the broader objective.

Shuffling and sampling reduce these correlations. Distributed systems also need deterministic-enough seeding and checkpoint state so training can be reproduced or resumed when required.

Randomness does not mean lack of control. It is an engineered part of the optimisation process.

Validation During Pretraining

Developers do not need to wait until the end to discover whether training is working. Periodic validation measures loss and other metrics on held-out data while checkpoints are saved.

If training loss falls but validation loss rises, overfitting or distribution mismatch may be emerging. If both losses stop improving, optimisation may have plateaued. If loss suddenly spikes, numerical instability or bad data may be involved.

Validation curves turn an enormous training run into observable evidence rather than a blind wait for the final checkpoint.

Learning Curves and What They Can Tell You

Plot training and validation loss against training steps. A healthy early run often shows both falling. The gap between them provides clues about generalisation.

Suppose training loss continues falling sharply while validation loss flattens. Adding more passes over the same data may mainly increase memorisation. More diverse data or stronger regularisation may help more than simply training longer.

Suppose both losses remain high. The model may be too small, the objective may be difficult, the learning rate may be poor or the representation pipeline may be broken. Diagnosis needs more evidence than one curve.

Perplexity and Language Modeling

Perplexity is a common transformation of average cross-entropy loss used in language modeling. Lower perplexity means the model assigns more probability to observed sequences under the evaluated distribution.

Perplexity is useful for comparing related language models under consistent tokenisation and datasets. It is not a universal intelligence score. Different tokenizers and domains can make raw perplexity values difficult to compare directly.

A model with lower perplexity on general text can still be worse at a specialised downstream task if the relevant capability or evaluation differs.

Why Training Data Deduplication Matters

Large web corpora contain copies, mirrors and quoted fragments. If one sequence appears thousands of times, it can receive disproportionate training weight and increase memorisation risk.

Deduplication removes exact or near-duplicate material according to chosen rules. This can improve data efficiency and make evaluation contamination easier to control.

Deduplication is not perfect. Two pages can paraphrase the same information. The goal is to reduce obvious repetition, not prove that every semantic idea occurs only once.

Contamination: When Evaluation Leaks Into Training

A benchmark is intended to test generalisation on unseen examples. If benchmark questions and answers are present in training data, the model may reproduce them from memory.

The GPT-3 paper explicitly reports contamination analysis for public benchmarks because web-scale training can accidentally include benchmark material. This remains a general challenge for large models trained on broad public data.

Contamination does not automatically invalidate every result, but it changes what the score can be interpreted to mean. Protected private evaluation sets are valuable because they reduce the chance of prior exposure.

Pretraining and Long-Tail Knowledge

A large corpus contains both common and rare topics. Frequent patterns receive many updates; rare facts or niche terminology may be seen only a few times.

This helps explain uneven capability. A model can be fluent about broad topics but unreliable in a specialised local domain. Retrieval or domain-specific adaptation can compensate without rebuilding the entire base model.

Pretraining and Rare Tokens

Tokenisation interacts with data frequency. Common words or character sequences can receive compact token representations, while rare terms may split into many subwords.

A specialised scientific identifier that becomes ten tokens uses more context and may be harder to model than a common one-token word. Tokenisation and data coverage therefore interact before the transformer even begins contextual processing.

Pretraining Across Modalities

The pretraining idea extends beyond text. Vision models can learn from images, audio models from sound, and multimodal models from aligned text-image or text-audio data.

The learning objective changes: predict masked patches, align image and text representations, reconstruct audio features or model multimodal sequences. The general principle remains broad representation learning before a specific downstream task.

Transfer Learning: Reusing the Foundation

A pretrained model can be adapted instead of training a specialised model from random initialisation. Fine-tuning, adapters, prompt tuning and other methods reuse learned representations.

This lowers the amount of task-specific data and compute required in many cases. The foundation has already learned general structure; downstream training focuses on the new task.

Catastrophic Forgetting

Further training on a narrow dataset can improve one domain while degrading previously learned capabilities. This is called catastrophic forgetting in some settings.

A post-training or fine-tuning process should therefore evaluate both target improvement and regression on important existing behaviours.

The goal is not simply lower loss on the new dataset. It is the right capability profile for the deployed system.

Checkpoint Selection Is a Decision, Not an Automatic Ending

The latest checkpoint may not have the best validation behaviour. Developers can compare checkpoints across loss, downstream tasks, safety evaluations and stability.

A later checkpoint could improve coding while regressing factual calibration or multilingual performance. Selection criteria should match the intended use.

Training Infrastructure Failures

Large runs can encounter hardware failures, communication errors, corrupt batches, checkpoint failures and numerical instability. Infrastructure therefore becomes part of model quality.

If gradients contain NaNs or a data shard is malformed, the run can silently degrade. Monitoring must detect both optimisation and systems problems.

Precision Formats and Numerical Stability

Large training often uses reduced-precision arithmetic to save memory and increase throughput. Lower precision must be managed carefully because gradients can underflow or overflow.

Techniques such as mixed precision, gradient scaling and clipping help make optimisation stable. These are engineering choices around the same mathematical training loop.

Gradient Accumulation

When a desired batch is too large to fit in device memory, training can accumulate gradients across several smaller micro-batches before applying an optimiser step.

This changes how computation is scheduled without changing the conceptual objective. It is one example of infrastructure adapting to the scale of the model.

Warmup and Learning-Rate Schedules

Large models often begin with a learning-rate warmup so early updates do not destabilise randomly initialised parameters. The learning rate can later decay as training progresses.

The schedule affects convergence. A fixed high rate can overshoot; a rate that falls too quickly can stop useful learning. Hyperparameter choices are therefore part of the experiment.

Regularisation During Pretraining

Regularisation discourages the model from overfitting specific training examples. Techniques can include weight decay, dropout, data augmentation and architectural choices.

The right methods depend on the model and training regime. More regularisation is not automatically better; it can also reduce useful fitting.

Why Training Can Produce Surprising Capabilities

A model optimised across diverse data can develop capabilities that were not specified as separate supervised tasks. Code generation, translation or few-shot classification can emerge from shared representations and scale.

The safest scientific language is to measure these capabilities rather than assume them. A task that appears suddenly on one benchmark may show smoother improvement under another metric.

Pretraining and Homogenisation Risk

Foundation models create leverage because one base model can power many applications. The same leverage can spread defects widely. Biases, factual weaknesses or security vulnerabilities in the base model may be inherited downstream.

The Stanford foundation-model report highlights this homogenisation concern. System builders should therefore understand which risks originate in the foundation and which can be mitigated at the application layer.

What a Good Pretraining Report Would Tell a Builder

A useful report describes architecture family, tokenizer, objective, parameter scale, broad data mixture, training compute or steps, validation approach, contamination controls and known capability limitations.

It should also separate facts that are public from proprietary details that are not disclosed. Absence of a detail should remain absence, not be filled by speculation.

Worked Exercise: Choose the Repair

Scenario A: the model often produces stale company policy even though retrieval is available. Is pretraining the first place to intervene?

Answer

Probably not. Ensure the application actually retrieves and supplies the current policy. Pretraining cannot remain current on every changing private rule.

Scenario B: the model fails basic syntax across many code tasks even when current documentation is supplied. A model with stronger code pretraining or targeted adaptation may be appropriate because the deficit is broad, not merely missing current facts.

Scenario C: the model handles one benchmark perfectly but fails paraphrased versions. Investigate contamination or narrow memorisation. Add private paraphrase tests and transfer cases.

Independent Exercise 6: Loss Curves

Training loss falls steadily while validation loss rises. What is the first concern?

Answer

Overfitting or distribution mismatch. More training is not automatically the repair.

Independent Exercise 7: Checkpoints

Checkpoint 10 performs better on general language; checkpoint 12 is better on code but worse on safety tests. Which is “best”?

Answer

There is no universal answer. Choose based on the intended capability and safety requirements, or continue development to improve the trade-off.

Independent Exercise 8: Fresh Data

A private handbook changes every month. Should the company put every update into global pretraining?

Answer

Usually no. Maintain the handbook as an authoritative retrieval source. Pretraining should provide broad capability; retrieval supplies changing private facts.

A Final Pretraining Checklist Before Deployment

Before attributing a deployed capability to pretraining, identify the complete route. Ask which base checkpoint is being used, whether instruction post-training changed behaviour, whether retrieval supplies current knowledge, whether tools perform calculations or actions, and whether application memory adds user-specific context. This prevents the phrase “the model learned it” from swallowing several different mechanisms.

Then separate capability from reliability. A pretrained model may demonstrate that it can write code, answer science questions or translate text. Deployment still needs tests for the exact domain, data format, latency, cost, security and failure conditions that matter to the user.

Finally, preserve a regression set across training changes. A new checkpoint that lowers generic pretraining loss can still regress on a critical downstream capability. Broad training metrics and task-level evaluations answer different questions.

Why Pretraining Is Not the End of Learning System Design

A foundation model is valuable because one expensive training process creates reusable capability. That reuse is also why the surrounding system matters so much. The same base model can power a tutoring assistant, coding tool, research agent or customer-service system depending on context, tools and controls.

Pretraining determines the space of capabilities the application can draw from. It does not determine the legitimate purpose, current source of truth or permission to act. Those remain system responsibilities.

This is the practical boundary for the rest of the series: training explains how parameters acquire useful structure; inference explains how that structure is used; retrieval explains how current evidence enters; agents explain how model decisions can become actions.

Pretraining Transfer Test: Can the Base Model Use What It Learned?

A strong pretraining run should produce capability that survives changes in surface form. To test this, build transfer tasks that were not copied from the training corpus: paraphrased questions, new combinations of familiar concepts, unseen document layouts and domain examples whose answers require applying structure rather than recalling a phrase.

If a model succeeds only when prompts resemble common training patterns, the representation may be brittle. If it handles new wording and recombines learned relationships, that is stronger evidence of reusable capability.

Transfer tests should include easy, medium and adversarial variations. The purpose is not to trick the model for entertainment; it is to locate the boundary between learned generality and memorised familiarity.

Pretraining and the Model–System Boundary

A pretraining team can improve the base model without solving every downstream problem. Current facts may still require retrieval. Exact calculations may still need tools. Private records still belong in databases. Permission still belongs to the application and institution.

This separation protects architecture. If a deployment fails because the model never received the current policy, retraining the foundation model is an expensive answer to an information-routing problem. If the model cannot understand the policy even when supplied directly, then model capability becomes a more plausible bottleneck.

The best pretraining programme therefore ends with a handoff: a documented base capability profile, known limitations, representative evals and clear guidance for downstream builders about what must still be supplied at runtime.

Pretraining Reality Check: What a Strong Base Model Still Needs

A strong base model still needs an application that supplies the right task and the right evidence. It may understand the concept of a timetable but not know today’s timetable. It may understand arithmetic but still benefit from a calculator. It may understand documents but not have permission to read a private file.

This is why model quality and system quality should be evaluated separately. Pretraining expands the space of possible capabilities; it does not complete every workflow.

The strongest evidence that pretraining worked is not that the model sounds knowledgeable. It is that held-out and transfer evaluations show reusable capability across representative tasks, while downstream systems can reliably supply current context and tools.

Pretraining Reader Exercise: Trace the Mechanism

Take one useful behaviour such as code completion. Ask which parts plausibly come from pretraining: syntax patterns, API associations, variable relationships and documentation style. Then identify what must come from runtime: the user’s repository, current library version, tests, file permissions and deployment state.

Repeat the exercise for tutoring. Pretraining may supply language and subject representations. The current worksheet, learner level, school topic and assessment date belong to context or retrieval. The student’s official grade belongs to an external record, not the model’s weights.

If you can separate those sources cleanly, you are ready for the later articles on post-training, retrieval, memory and agents.

Frequently Asked Questions About Pretraining

Does pretraining happen every time I ask a question?

No. Normal use performs inference with already-trained parameters. A provider may separately collect permitted data for future training, but that is a different process.

What is a foundation model?

A model trained broadly enough to be adapted across many downstream tasks. The Stanford foundation-model report popularised the term to emphasise both broad reuse and the model’s incomplete nature as one layer of a larger system.

Is pretraining supervised learning?

Many language-model objectives are self-supervised because targets are derived from the data itself. Other training stages can use human-labelled or preference-labelled examples.

Why does next-token prediction create broad capability?

Predicting complex sequences rewards learning reusable structure: syntax, concepts, relationships, code patterns and long-range dependencies. The objective is simple; the representations required to perform it well can be rich.

Can pretraining memorise data?

Yes, especially repeated or distinctive material. Useful training aims for generalisation, but memorisation remains an important research and privacy concern.

Is more pretraining always better?

No. Benefits depend on data, objective, model capacity, optimisation and evaluation. More compute on poor data can waste resources or reinforce defects.

What changes during pretraining?

Model parameters change through gradient-based optimisation. Hyperparameters and training infrastructure guide that process.

Does pretraining make a model current forever?

No. New events require retrieval, tools or later training.

Pretraining Builds the Base; the Rest of SI Builds the Working System

Pretraining explains how a model acquires broad reusable capability before the user’s task begins. It turns large collections of examples into parameters that can generalise across many later contexts.

But pretraining is only the beginning. Training data shapes what can be learned. Self-supervised objectives create the learning signal. Backpropagation and optimisation perform the parameter updates. Post-training shapes assistant behaviour. Retrieval and tools connect the model to current evidence and action.

Continue through the How Super Intelligence Works hub. Next: 022 — Training Data.


How Super Intelligence Works Series Navigation

Previous: 020 — Context Windows · Series Hub · Next: 022 — Training Data

Discover more from eduKate Singapore

Subscribe now to keep reading and get access to the full archive.

Continue reading