Scaling is the study of how model performance changes as we increase parameters, training data and compute. It sounds simple: make the model bigger, feed it more data and train it longer. In practice, scaling is an allocation problem. A fixed compute budget can be spent on a larger model, more training tokens, more optimisation steps or a different balance between them.
This article explains how scaling works in Super Intelligence: model size, data size, training compute, scaling laws, compute-optimal training, diminishing returns, undertraining, overtraining, inference cost, quality, efficiency and evaluation. It also explains why “bigger is better” is an incomplete rule.
The paper Scaling Laws for Neural Language Models reported empirical power-law relationships between language-model loss and model size, dataset size and compute over wide ranges. The later Training Compute-Optimal Large Language Models paper, associated with Chinchilla, argued that model size and training tokens should be scaled together more aggressively than many earlier models had done.
These are empirical relationships, not physical laws guaranteeing that every architecture, dataset or capability follows one curve forever. Scaling must be measured under the conditions being studied.
Previous: 024 — Backpropagation and Gradient Descent. Here we zoom out from one parameter update to the economics of an entire training run.
The Hidden Transition: More Compute Is Not One Thing
Suppose you have twice the training budget. You could double the number of model parameters and keep the token count fixed. You could keep the model the same size and train on twice as much data. You could increase both by smaller amounts. These choices create different outcomes even if the total compute is similar.
Scaling therefore asks not only how much resource is available, but how to allocate it. A model can be too large for its data budget, too small for its data, trained too briefly, or trained long after marginal gains become weak.
Three Core Scaling Axes
Model size
Model size is commonly described by parameter count. More parameters increase representational capacity, but also increase memory use, training cost and inference cost. Parameter count alone does not describe architecture, tokenizer, data mixture or quality.
Training data
Data scale can be measured in tokens, examples, hours of audio, image-text pairs or other units depending on the model. More data broadens exposure, but repeated low-quality data can add less value than diverse high-quality examples.
Training compute
Compute reflects the amount of numerical work performed during training. For transformers, rough compute estimates often scale with model size multiplied by the number of training tokens, although exact accounting depends on architecture and implementation.
Scaling Laws Are Empirical Curves
Kaplan and colleagues reported that language-model cross-entropy loss followed approximate power laws with parameter count, dataset size and compute over the regimes they studied. A power law means improvements continue as scale rises but with diminishing returns.
If doubling compute improves loss, the next doubling may improve it again by a smaller absolute amount. This creates a predictable-looking curve without implying infinite improvement.
The value of scaling laws is planning. Developers can run smaller experiments, fit a curve and estimate what larger runs might achieve before spending the full budget.
A Toy Power-Law Example
Imagine validation loss follows L(C) = 3.0 × C^-0.05, where C is compute measured in arbitrary units. At C = 1, loss is 3.0. At C = 32, the factor 32^-0.05 is about 0.84, so loss falls only to about 2.52.
The exact numbers are fictional. The point is diminishing returns: multiplying compute many times does not divide loss by the same factor.
Power-law exponents therefore matter. A small exponent means enormous resource increases are required for modest loss improvements.
Parameters Without Enough Data Can Be Undertrained
A very large model trained on too few tokens may not use its capacity efficiently. It has many parameters but insufficient optimisation signal to shape them well.
The Chinchilla paper challenged earlier allocation practices by showing that, under its experimental assumptions, many large models were undertrained relative to their size. The authors trained over 400 models to estimate compute-optimal model and data scaling.
Their result popularised the idea that a smaller model trained on substantially more data can outperform a larger undertrained model at the same training compute.
The Chinchilla Principle
The Chinchilla study reported that, under its setup, compute-optimal training required model size and the number of training tokens to scale roughly together: doubling one should be accompanied by roughly doubling the other.
This is not a universal constant for all later architectures, datasets and objectives. It is an empirical result that reshaped how practitioners think about allocating compute between parameters and tokens.
Worked Compute-Budget Example
Suppose training compute is approximately proportional to parameters × tokens. Budget A permits 10 billion parameter-units times 100 billion token-units, giving an arbitrary product of 1,000. Another design uses 20 billion parameters and 50 billion tokens, also giving 1,000.
The compute proxy is the same, but the data-to-parameter ratio differs by a factor of four. Which design performs better cannot be decided from arithmetic alone. Scaling experiments estimate the optimum under the specific architecture and corpus.
This is why parameter count is not a complete statement of model quality.
Undertraining Versus Overtraining
Undertraining means the model has not received enough optimisation or data to approach the useful region available to its capacity. Overtraining can mean continuing on the same distribution after marginal value becomes weak, or fitting idiosyncrasies that harm generalisation.
In large foundation-model practice, “overtraining” can also be used operationally to describe deliberately training a smaller model on more tokens than a strict compute-optimal rule would suggest because the smaller model will be cheaper to serve later.
The right trade-off depends on whether the objective is minimum training loss per training FLOP or minimum total lifecycle cost including inference.
Training-Optimal Is Not Deployment-Optimal
Imagine two models achieve similar quality. Model A is larger and required less training time; Model B is smaller but trained on more data. If the service will answer billions of requests, Model B may be cheaper to deploy even if it was not the strict compute-optimal training choice.
This creates a lifecycle optimisation problem: training compute, inference compute, memory, latency and hardware availability all matter.
Inference Cost Grows With Scale
Larger models require more parameter memory and more computation per generated token. Serving can dominate lifetime cost for a heavily used product.
A research model run once has different economics from a consumer assistant serving millions of daily requests. Scaling decisions should be tied to expected usage.
Data Quality Can Change the Scaling Curve
A billion new high-quality tokens can be more valuable than a billion near-duplicates. Scaling laws fitted on one data mixture may shift when data quality or domain coverage changes.
This is why article 022 treated training data as a first-class component. Scale multiplies whatever the dataset contains—both useful variation and repeated defects.
Mixture Scaling
A model can scale total tokens while also changing the proportion of code, mathematics, dialogue, multilingual text or scientific material. Capability changes then reflect both quantity and composition.
A clean experiment holds the mixture stable when estimating pure scale effects, then studies mixture changes separately.
Token Budget Is Not Document Count
One long book can contain more training tokens than thousands of short messages. Model training budgets are therefore often discussed in tokens rather than documents.
Tokenisation matters: the same text can produce different token counts under different tokenizers, especially across languages and code.
Parameter Count Is Not Architecture
Two 7-billion-parameter models can behave differently because they use different layer counts, widths, attention structures, tokenizers, data, context lengths and post-training.
Scaling studies try to control architecture families so parameter count is a meaningful axis.
Depth Versus Width
A parameter budget can be allocated to more layers, wider hidden states, more attention heads or other architectural dimensions. Earlier scaling work found relatively smooth performance across a range of width/depth choices within its studied regime.
Later architecture innovations can alter these trade-offs. The broad lesson is to measure rather than assume one shape is always optimal.
Context Length Is Another Scaling Axis
Increasing context length lets the model process more tokens in one inference, but attention and memory costs can grow substantially depending on architecture.
Long-context training also requires appropriate data and objectives. Simply setting a larger maximum length does not guarantee the model uses distant information reliably.
Vocabulary and Tokenizer Scale
A larger vocabulary can shorten sequences for common patterns but increases embedding and output projection costs. A smaller vocabulary shares pieces more broadly but can lengthen rare words and code identifiers.
Tokenizer design affects compute and multilingual efficiency, so scaling is not only parameter count.
Batch Size Scaling
Larger training runs often use larger effective batch sizes to keep many accelerators busy. Batch size interacts with learning rate, gradient noise and optimisation.
There are limits to useful batch growth. After a point, more examples per update may add little optimisation benefit even if hardware can process them.
Hardware Scaling
Training on more accelerators can reduce wall-clock time, but communication overhead grows. Distributed training must move activations, gradients or parameters between devices.
Speedup is therefore sublinear. Doubling hardware does not automatically halve training time.
Memory Bandwidth and Interconnect
Large models are constrained not only by arithmetic throughput but by moving data between memory and compute units and across devices.
A model that looks efficient in FLOPs can be slow if memory traffic dominates.
Scaling the Data Pipeline
A thousand accelerators waiting for data is wasted compute. Storage, preprocessing, shuffling, tokenisation and networking must feed training fast enough.
The model training stack therefore scales as a system: data input, checkpointing, monitoring and recovery all need capacity.
Checkpoint Scaling
Larger models create larger checkpoints. Saving them frequently protects against failures but consumes storage and network bandwidth.
Checkpoint strategy balances recovery time against overhead. Large runs may use sharded checkpoints and asynchronous saving.
Fault Probability Grows With Cluster Size
As training uses more machines for longer periods, hardware failures become normal rather than exceptional. Distributed systems need restart and recovery mechanisms.
Scaling compute without scaling reliability engineering can reduce effective throughput.
Power-Law Forecasting Needs Honest Uncertainty
Fitting a scaling curve to small runs and extrapolating ten or hundred times beyond observed compute introduces uncertainty. Architectural changes, data shifts and optimisation regimes can break the trend.
Forecasts should therefore include the range of tested scales and avoid presenting extrapolation as guaranteed outcome.
IsoFLOP Experiments
The Chinchilla work used experiments that compare many model sizes under fixed compute budgets. These “IsoFLOP” profiles help identify which size-data balance minimises loss for a given compute level.
This is more informative than training one giant model and guessing whether it was efficient.
Diminishing Returns
Scaling can continue improving performance while marginal gains shrink. An organisation should ask whether the next unit of compute produces enough task value to justify cost.
The answer differs for frontier research, consumer chat, embedded devices and scientific discovery.
Capability Thresholds Can Look Sudden
A benchmark can show near-zero success until the model crosses a capability threshold, then jump quickly. This has been described as emergent behaviour.
But thresholded metrics can make smooth underlying improvements look discontinuous. A model score moving from 0.49 to 0.51 appears as a zero-to-one jump under a 0.5 pass threshold.
The safest approach is to inspect both continuous and discrete metrics before declaring a capability fundamentally appeared from nowhere.
Emergence Is a Measurement Question Too
Some capabilities genuinely become useful only after enough component skills coexist. Others appear sudden because the benchmark or prompting changed.
Scaling research should therefore report metric definition, prompting method, model family and sample size when discussing emergence.
Scaling and Memorisation
Larger models can have greater capacity to memorise rare sequences while also generalising better. More data can reduce reliance on any one example but also increase the chance of encountering sensitive or benchmark content.
Privacy and contamination analysis need to scale alongside the model.
Scaling and Bias
A larger model trained on more diverse data can reduce some errors through broader coverage while learning more subtle stereotypes from the same world.
Bias does not monotonically disappear with parameter count. Subgroup evaluation is necessary.
Scaling and Calibration
Model probabilities can become sharper as models improve, but calibration is not guaranteed. A larger model can be more accurate and still overconfident on certain distributions.
Evaluate probability quality separately from raw benchmark accuracy.
Scaling and Hallucination
Broad scaling can improve factual recall and reasoning, yet unsupported generation remains possible because the language objective rewards plausible continuation.
Retrieval and verification remain useful even for larger models.
Scaling and Tool Use
A larger base model may understand tool descriptions and plan better, but reliable action also depends on post-training, schemas, permissions and execution feedback.
Tool capability is therefore not a pure scaling result.
Scaling and Long-Horizon Agents
Agent performance compounds many decisions. A small per-step improvement can create a large end-to-end gain over long tasks, while a small error rate can also compound into failure.
Scaling the model can help, but agent architecture and verification become increasingly important as horizon length grows.
Scaling Evaluation Suites
As capability broadens, evaluation must broaden too. A model that improves coding while regressing multilingual safety cannot be summarised adequately by one score.
Maintain suites across knowledge, reasoning, code, instruction following, safety, calibration, latency and domain-specific tasks.
Scaling the Human Evaluation Process
Some outputs require human judgment. Evaluating millions of generations manually is impossible, so teams sample, use expert panels, model-based graders and automated checks.
Each evaluation method has its own biases. Scaling evaluation requires measurement engineering as much as model engineering.
Model-Based Evaluation at Scale
A stronger model can grade candidate responses, compare outputs or detect obvious defects. This reduces cost but risks grader bias and correlated mistakes.
Human spot checks and benchmark anchors remain important.
Scaling and Data Scarcity
High-quality human-created data is finite. As models consume more tokens, developers increasingly rely on filtering, synthetic data, domain partnerships and repeated epochs.
Data scarcity can shift the optimum away from simply “train on more unique internet text”.
Repeated Data
Training for multiple epochs over the same dataset increases effective tokens without increasing unique information. This can be useful up to a point, especially with high-quality data, but increases overfitting and memorisation risk.
The optimal repeat count depends on dataset size, model size and objective.
Synthetic Data as a Scaling Resource
Models can generate additional training examples, especially for code, mathematics and instruction data. Synthetic scaling is attractive because targeted examples can be produced cheaply.
The quality ceiling depends on verification and diversity. Unchecked synthetic data can recycle model errors.
Distillation
A large teacher model can generate soft targets or examples used to train a smaller student model. Distillation transfers some capability while reducing serving cost.
The student does not automatically inherit every teacher strength. Evaluate what survives compression.
Mixture-of-Experts Scaling
Mixture-of-experts architectures activate only a subset of model parameters for each token. This can increase total parameter capacity without using every parameter on every inference step.
Routing, load balancing and communication become new system challenges. Total parameters and active parameters should be distinguished.
Sparse Versus Dense Scale
A dense model uses most of its parameters for each token. A sparse model may contain many more total parameters but activate fewer.
Comparing models by raw parameter count can therefore mislead. Active compute and memory also matter.
Scaling on Edge Devices
Phones, robots and embedded systems cannot host frontier-scale models easily. Scaling down through quantisation, pruning, distillation and specialised architectures is a separate optimisation problem.
The best model for the cloud is not necessarily the best model for a classroom device with limited memory and battery.
Scaling Up Versus Scaling Out
Scaling up makes one model larger. Scaling out can use several models, retrieval systems or agents working together.
A collection of specialised smaller models may outperform one giant model on cost or reliability for a bounded application.
Routing by Difficulty
A system can send easy tasks to a small cheap model and difficult tasks to a larger model. This creates conditional compute.
The router must predict difficulty reliably. If routing itself is weak, savings can come at the cost of quality.
Test-Time Scaling
Capability can also be increased by spending more compute during inference: generating multiple candidates, searching solution paths, verifying answers or using longer reasoning.
This is different from training scale. Article 034 will examine test-time compute directly.
Training Scale Versus Test-Time Scale
A smaller model with more test-time search can sometimes outperform a larger model using one quick pass on particular tasks. The economic optimum depends on query frequency and latency.
Training invests compute once; test-time compute is paid per request.
Scaling and Energy
More training compute requires more electrical energy. Data-centre efficiency, hardware generation, utilisation and energy source affect environmental impact.
A responsible comparison states the measurement boundary. Model size alone cannot determine energy use.
Scaling and Water
Some data centres use water for cooling, directly or indirectly through electricity generation. Water impact varies by location, cooling design and time.
Article 087 will treat environmental load in detail. Here the scaling lesson is that physical resources grow alongside digital capability.
Scaling the Organisation
Frontier training requires data engineering, distributed systems, evaluation, safety, security and operations teams. The organisational system scales with the model.
A bigger training cluster without stronger incident response and evaluation can create fragile progress.
A Scaling Failure Map
Level 1: wrong target metric. Level 2: weak small-scale experiments. Level 3: poor curve fit. Level 4: model-data imbalance. Level 5: data quality collapse. Level 6: optimisation instability. Level 7: distributed-system inefficiency. Level 8: inference cost ignored. Level 9: benchmark gains fail to transfer. Level 10: organisational controls do not scale.
This eduKateSG map turns “we need a bigger model” into a diagnostic problem.
Worked Diagnosis: Bigger Model, Same Data
A team doubles parameters but keeps tokens fixed and sees little improvement. Possible diagnosis: the model is undertrained relative to its capacity.
Run compute-allocation experiments or increase high-quality data before assuming the architecture is saturated.
Worked Diagnosis: More Tokens, No Gain
The team doubles tokens by adding near-duplicate low-quality pages. Loss barely improves.
The bottleneck may be information diversity rather than token count. Inspect deduplication and source quality.
Worked Diagnosis: Better Benchmark, Worse Product
A larger model improves academic benchmarks but doubles latency and cost while users notice little difference.
The deployment optimum may favour the smaller model, routing difficult tasks selectively to the larger one.
Worked Diagnosis: Scaling Curve Breaks
Small models suggested a smooth improvement, but the largest run underperforms. Check optimiser settings, learning-rate schedule, data pipeline, hardware faults and whether extrapolation crossed into a new regime.
A Practical Scaling Experiment
Train several model sizes under several token budgets. Keep architecture family, tokenizer and data quality as stable as practical. Measure validation loss and downstream tasks. Fit curves only within the observed regime.
Then choose a compute budget and identify the model-token pair with the best target performance. Finally include serving cost in the decision if the model will be heavily deployed.
Independent Exercise 1: Bigger or More Data?
A 20B-parameter model has trained on only half the tokens of a 10B model. Can parameter count alone tell which is better?
Answer
No. Training data, compute allocation, architecture and evaluation all matter. The larger model may be undertrained.
Independent Exercise 2: Diminishing Returns
If every 10× increase in compute reduces loss by a smaller absolute amount, what economic question follows?
Answer
Whether the next scale increase produces enough task value to justify its training, inference and infrastructure cost.
Independent Exercise 3: Training Versus Serving
A smaller model requires more training tokens but is half the inference cost. Which is preferable?
Answer
It depends on expected lifetime usage. Heavy serving can make the smaller model economically superior even if training it was less compute-optimal.
Independent Exercise 4: Emergence
A benchmark jumps from 0% to 80% success when a continuous score crosses a threshold. Does that prove a capability appeared discontinuously?
Answer
No. Inspect the underlying continuous metric and evaluation design. Thresholding can create apparent jumps.
Independent Exercise 5: Data Scale
Adding 100B tokens of near-duplicate text produces little gain. What should you inspect?
Answer
Data diversity, deduplication and mixture quality before assuming the model needs even more raw tokens.
Scaling Curves Need the Right Dependent Variable
A scaling law is only as meaningful as the metric on its vertical axis. Cross-entropy loss is smooth and sensitive, which makes it useful for fitting power laws. A pass/fail benchmark can hide small but real improvements until the score crosses a threshold.
For reasoning tasks, teams may track exact accuracy, partial credit, calibration and cost at the same time. One metric can improve smoothly while another remains flat. This is not a contradiction; the metrics measure different aspects of behaviour.
Scaling Across Domains
A model can scale strongly in general language while improving more slowly in mathematics, code or a low-resource language. The effective data mixture and evaluation difficulty differ by domain.
Therefore, a global validation-loss curve should be supplemented with domain-specific curves. A small average gain can conceal a large specialised gain or regression.
Scaling Across Languages
Increasing parameters does not automatically equalise language performance. A language with little training data may remain weak even in a larger model.
A better intervention may be improved tokenizer coverage, more high-quality language data or targeted adaptation. Scale magnifies exposure; it does not create missing exposure from nothing.
Scaling Across Context Lengths
A model trained primarily on short sequences may fail to use a newly enlarged context window effectively. Long-context performance depends on positional representation, training length distribution and evaluation.
The correct scaling question is not only “how many tokens fit?” but “how well does performance hold as relevant information moves farther away?”
Needle-in-a-Haystack Versus Real Long-Context Tasks
A synthetic test can hide one exact fact in a long document and ask the model to retrieve it. This measures one capability. Real long-context work may require comparing dozens of passages, resolving contradictions and tracking structure.
Scaling context length should therefore be evaluated on both retrieval-like tests and realistic synthesis tasks.
Scaling Retrieval With Model Size
A larger model can interpret retrieved evidence better, but poor retrieval still caps performance. If the right passage never enters context, more parameters may only generate a more persuasive wrong answer.
End-to-end scaling studies should separate retriever recall from generator quality so the bottleneck remains visible.
Scaling Tool Use
Tool use adds structured decisions: which tool, which arguments, when to stop and how to interpret the result. A stronger model may improve each decision probability, but long workflows multiply error opportunities.
If one step succeeds with probability 0.98 and a task requires 50 independent critical steps, the probability of all 50 succeeding under a simplistic independence model is about 0.98^50, roughly 0.36. Real dependencies are more complex, but the example shows why long-horizon reliability matters.
Compounding Error and Agent Scale
Improving a per-step error rate from 2% to 1% can matter more for long tasks than a small benchmark gain suggests. Conversely, an agent that seems excellent on one-step tests can fail often over 100 actions.
Scaling agent capability requires verification, recovery and checkpointing in addition to stronger models.
Compute Budgets and Opportunity Cost
Training one large model consumes resources that could instead fund several experiments, better data curation, evaluation or smaller specialised models.
The optimal research portfolio is therefore broader than the compute-optimal model. Exploration and measurement have value because they reduce the risk of spending the entire budget on the wrong scaling direction.
Frontier Runs Versus Ablation Runs
A frontier run tests the largest scale a team can afford. Ablation runs deliberately remove or vary one factor—data quality, context length, optimiser, architecture—to identify causality.
Without ablations, a successful large model can tell us that the whole recipe worked but not which ingredient produced the gain.
Scaling Research Is Experimental Science
Good scaling work specifies hypotheses, controlled variables, measurement uncertainty and extrapolation limits. “We made it bigger and it got better” is an observation, not a complete scaling study.
Small pilot runs can estimate slopes, reveal instability and compare mixtures before committing to expensive training.
Scaling Law Residuals
After fitting a power law, inspect residuals—the difference between predicted and observed performance. Systematic residual patterns may reveal architecture transitions, data-quality changes or optimisation issues.
A smooth fit can hide subgroups where the model behaves differently. Residual analysis keeps the empirical law honest.
Confidence Intervals on Scaling Fits
A fitted exponent is an estimate. Limited experiments and noisy measurements create uncertainty. Reporting a range is more informative than pretending the exponent is exact.
Extrapolation uncertainty grows as the target scale moves farther beyond observed experiments.
Scaling and Benchmark Saturation
A benchmark can stop being useful when strong models approach its ceiling. Small score changes become dominated by annotation noise, and the task no longer distinguishes frontier systems.
Evaluation must scale too: harder examples, fresh private tests and measures of robustness rather than only more repetitions of saturated questions.
Scaling and Benchmark Gaming
When a benchmark becomes a target for optimisation, developers can tune prompts, training data and system settings specifically for it. The benchmark then measures optimisation toward the test as much as general capability.
Private or rotating evaluations help preserve diagnostic value.
Scaling and Human Baselines
Comparing model scores with human performance can be useful, but human populations vary by expertise, time budget and access to tools. “Human-level” should specify which humans under which conditions.
A model allowed multiple attempts and search should not be compared casually with a human given one unaided attempt.
Scaling and Sample Efficiency
One benefit of larger models observed in some studies is better sample efficiency: they can achieve lower loss with fewer data examples than smaller models.
However, the total compute used to train the larger model may still be greater. Sample efficiency and compute efficiency are different metrics.
Scaling and Parameter Efficiency
Parameter-efficient adaptation methods such as LoRA change a small subset or low-rank component of the model instead of all parameters. This lets one large foundation support many specialised variants at lower storage cost.
The base model still carries the full inference cost unless additional compression is used.
Scaling and Quantisation
Quantisation stores weights or activations with fewer bits. It can reduce memory and accelerate inference while introducing approximation error.
Quantisation changes deployment economics rather than the number of learned parameters. A 70B model quantised to fewer bits is still a 70B-parameter model, but the hardware footprint is smaller.
Scaling and Pruning
Pruning removes weights, heads, channels or structures judged less important. The goal is to preserve useful behaviour while reducing compute or memory.
Pruned models need re-evaluation because removed components can affect rare capabilities even when average performance stays stable.
Scaling and Distillation Economics
A frontier model can act as a teacher for many smaller students. If the students serve high-volume tasks, the expensive teacher run can amortise across deployments.
This creates a two-stage scaling strategy: scale up to discover or generate capability, then scale down for efficient use.
Scaling and Specialisation
A smaller specialised model trained on relevant data can outperform a general giant model on a narrow task, especially when latency and cost matter.
The decision should compare end-to-end quality, not prestige. A specialised classifier does not need to write poetry.
Scaling and Safety Evaluations
New scale can produce stronger capabilities and new failure modes. Safety evaluation should run before wider deployment, not only after users discover the boundary.
A regression suite should include misuse resistance, privacy, tool permissions and high-impact domain behaviour appropriate to the model.
Scaling and Security
Larger models can understand more complex tool environments, which can improve productivity and expand attack surface. Prompt injection, data exfiltration and excessive permissions are application risks that can increase with capability.
Security controls must scale at least as fast as autonomous reach.
Scaling and Organisational Latency
A faster model does not automatically make the organisation faster. Human approval queues, slow databases and manual handoffs can dominate end-to-end task time.
Measure the whole workflow. Model latency may be only a small part of the user’s wait.
Scaling and Reliability Budgets
An organisation can define a reliability budget for a workflow: maximum error rate, timeout rate, cost, human-review load and allowed external impact.
Scaling decisions then target the binding constraint. If reliability is already sufficient but cost is high, a smaller model may be the next improvement.
A Lifecycle Cost Equation
A simplified lifecycle cost can be written as training cost + deployment cost per request × number of requests + evaluation/maintenance cost. This is not a financial accounting standard, but it makes the trade-off explicit.
For a model used only a thousand times, training dominates. For a model used a trillion times, serving efficiency can dominate.
Worked Lifecycle Example
Model A costs 100 training units and 2 units per million requests. Model B costs 150 training units but 1 unit per million requests. After 50 million requests, A totals 200 units while B totals 200. Beyond that simplified break-even point, B becomes cheaper.
Quality must still be comparable. Cost arithmetic cannot justify a weaker model for a task whose failures are expensive.
Scaling and Latency Percentiles
Average latency can hide bad tail behaviour. Users notice p95 or p99 delays when queues build or long prompts arrive.
Serving larger models should be evaluated under realistic concurrent load, not only one benchmark request on an idle accelerator.
Scaling and Throughput
Throughput measures how much work a service handles per unit time. Batching can increase throughput while increasing latency for individual requests.
Production systems tune this trade-off according to user expectations and hardware utilisation.
A Compute-Optimal Thought Experiment
Suppose three candidate runs use the same compute. Run A: 2B parameters × 200B tokens. Run B: 4B × 100B. Run C: 8B × 50B. The raw product is equal in this toy approximation.
Small pilot experiments could reveal B has the lowest validation loss under the chosen mixture. That result is more useful than assuming C wins because it has the most parameters.
A Deployment-Optimal Thought Experiment
Now assume C is most accurate but too slow for an on-device tutor. A compressed version of A may be the only model that fits memory and responds within 100 milliseconds.
The “best” scale is therefore conditional on the deployment envelope.
Independent Exercise 6: Curve Fit
A power-law fit uses only three tiny models and predicts a 1000× larger run. What is the main concern?
Answer
Extrapolation uncertainty. The fitted regime is too narrow to justify confident prediction far beyond observed scale.
Independent Exercise 7: Lifecycle Cost
A model is 20% more expensive to train but 50% cheaper per query. What additional fact is needed before choosing it?
Answer
Expected deployment volume, along with quality and latency. Serving savings matter only if the model will be used enough.
Independent Exercise 8: Long-Horizon Agent
A model improves one-step tool accuracy from 95% to 98%. Why can this matter greatly for 50-step tasks?
Answer
Errors compound across steps. Even a few percentage points of per-step improvement can produce a much larger end-to-end reliability gain.
A Final Scaling Audit Before Spending More Compute
Before approving a larger run, write down the bottleneck in one sentence. Is the present model too inaccurate on a defined task, too slow, too expensive, too weak in one domain, or too unreliable over long workflows? “We want a bigger model” is not yet an engineering objective.
Next, ask whether scale is actually the likely repair. If the model lacks current private policy, retrieval may help more than additional parameters. If arithmetic is failing, a calculator may help more. If the dataset is duplicated, curation may dominate. If tool permissions are unsafe, architecture rather than scale is the priority.
Then define the measurement that would justify the new run. Specify the validation loss, downstream task score, long-horizon success rate, latency, inference cost and safety evaluations that matter. This creates a falsifiable reason for scaling instead of a prestige target.
Finally, define the alternative use of the same resources. A larger run competes with better data, more ablations, larger evaluation sets, specialised models and deployment engineering. The economically rational choice is the one that improves the complete SI system, not necessarily the one with the most parameters.
What Mastery of Scaling Looks Like
A reader has understood scaling when they can explain why model size, data and compute must be considered together; why compute-optimal training can differ from deployment-optimal design; why scaling laws are empirical; and why benchmark jumps need careful measurement before being called emergence.
They should also be able to diagnose a weak result without reflexively demanding a larger model. Sometimes scale is the repair. Sometimes the first unstable point is data quality, retrieval, optimisation, context, tools or evaluation. Scaling is powerful precisely when it is used as a measured systems decision.
A Practical Scaling Decision Before Spending More Compute
Before scaling a model, write down the bottleneck. Is performance limited by parameter capacity, data quality, data quantity, optimisation, context length, retrieval, tool access or evaluation design? Bigger training runs only address some of these constraints.
Then define the scaling hypothesis in testable form: “Doubling effective training compute while preserving data quality should improve held-out code generation by X without regressing safety and multilingual performance.” A hypothesis forces the team to name the expected benefit and the regression checks before spending the compute.
Finally, measure cost per useful improvement rather than model size alone. A smaller architecture with better data, routing or retrieval can outperform a larger poorly integrated system on the actual task. Scaling is therefore one lever inside system design, not a universal substitute for diagnosis.
The Clementi-style test is simple: identify the first unstable point, scale only if scale plausibly repairs that point, and keep the earlier success cases in the regression set. If the same failure remains after a larger run, the evidence should move attention elsewhere.
Frequently Asked Questions About Scaling
Does a larger model always perform better?
No. Data, compute allocation, architecture, optimisation and task fit matter. A smaller model trained more effectively can outperform a larger undertrained one.
What is a scaling law?
An empirical relationship describing how a performance measure changes with model size, data or compute over an observed regime.
What did Chinchilla change?
It showed under its experimental setup that compute-optimal language-model training used more data relative to model size than many earlier large models, popularising balanced model-and-token scaling.
What does compute-optimal mean?
The model/data allocation that gives the best measured performance for a fixed training compute budget under the chosen assumptions.
Why not train the biggest model possible?
A model can be too large for its data budget, expensive to serve and slower than a smaller well-trained model.
Does more data always help?
No. Quality, duplication, rights, relevance and distribution matter. More tokens are not automatically more information.
What is test-time scaling?
Spending more computation during inference through search, multiple candidates, verification or longer reasoning rather than increasing only training scale.
Can sparse models have huge parameter counts cheaply?
They can activate only part of the model per token, reducing active compute relative to total parameters, but routing and communication introduce new complexity.
Scaling Is Resource Allocation, Not a Magic Spell
Scaling works because larger models, more data and more compute can reduce predictive error and broaden capability. But the gains depend on balance. Parameters without enough data are wasteful; data without enough capacity can be underused; compute without evaluation can optimise the wrong goal.
The mature question is not “How big can we make it?” It is “What allocation of model, data, training compute and inference compute delivers the required capability at acceptable cost and risk?”
Continue through the How Super Intelligence Works hub. Previous: 024 — Backpropagation and Gradient Descent. Next: 026 — Post-Training.
Scaling Decision Record
For any major scaling decision, record the current model, dataset, training-token count, compute estimate, evaluation suite, serving profile and the specific limitation the larger run is intended to repair. After training, compare the predicted gain with the measured gain. This creates an institutional learning loop: future teams can see where scaling-law extrapolations held, where they broke and which bottlenecks moved elsewhere in the system.
A decision record also prevents hindsight bias. If a larger model improves one benchmark but fails the original latency or cost requirement, the project should not quietly redefine success. The training objective, deployment envelope and acceptance criteria belong in the same record.
Scaling Reality Check
A scaling claim should always name the regime. A result obtained for autoregressive transformer language models on one data mixture does not automatically apply unchanged to multimodal models, retrieval systems, robotics or sparse expert architectures. The correct habit is to preserve the empirical scope: architecture family, token budget, compute range, benchmark and training recipe.
The same caution applies to business conclusions. A model that is more compute-efficient during training may be more expensive to serve; a larger model that wins a benchmark may fail the latency envelope; a smaller model may become superior once retrieval and tools are added. Scaling is therefore one optimisation layer inside the complete Super Intelligence stack.
For readers, the practical test is simple: when somebody says “this capability appeared because the model got bigger,” ask what else changed. Did the dataset grow? Did post-training improve? Did context length increase? Were tools added? Was the benchmark reformulated? Causal attribution requires more than observing two model sizes.
