Neural networks are learnable functions built from layers of numerical operations. They take input vectors, multiply and combine them with trainable weights and biases, apply nonlinear transformations, and produce outputs that can be compared with a training objective. Learning changes the parameters so future outputs become more useful for the task.
This is the computational engine beneath many modern Super Intelligence systems. Embedding layers are neural-network parameters. Transformers are neural-network architectures. Image models, speech models, classifiers and many generative systems are constructed from networks whose parameters are learned through optimisation.
This article explains how neural networks work from first principles. We will build a neuron by hand, connect neurons into layers, compute a forward pass, define a loss, follow gradients backward with the chain rule, update parameters with gradient descent and examine generalisation, overfitting, regularisation and training failure.
In the eduKateSG series, Super Intelligence, or SI, is our umbrella term for AI technologies. Standard terms such as weight, bias, activation, loss, gradient, backpropagation, optimisation, batch, epoch and parameter remain visible so readers can connect the explanation to current machine-learning practice.
Previous: 014 — Vector Space. Return to the How Super Intelligence Works hub for the full sequence.
Neural Networks at a Glance
- A neuron computes a weighted combination of inputs, adds a bias and usually applies a nonlinear activation.
- A layer applies many such transformations in parallel to a vector or tensor.
- Weights and biases are trainable parameters; training changes them to reduce a loss function.
- Backpropagation uses the chain rule to compute how the loss changes with respect to parameters.
- An optimiser uses those gradients to update parameters.
- Training performance and real-world generalisation are different; validation and test data measure behaviour beyond the training examples.
- Deep learning means composing many learned layers, not giving a machine a human biological brain.
The Simplest Neuron: Weighted Sum Plus Bias
Suppose a model receives two inputs x1 and x2. A simple neuron computes z = w1x1 + w2x2 + b. The weights w1 and w2 control how strongly each input contributes; the bias b shifts the result.
Let x=[2,3], w=[0.5,-1] and b=0.25. The weighted sum is 0.5×2 + (-1)×3 + 0.25 = 1 – 3 + 0.25 = -1.75.
This value is called a pre-activation in many explanations. A nonlinear activation function can then transform it into the neuron’s output.
Why Weights Matter
Weights are learned numerical parameters. If w1 becomes larger and positive, increases in x1 push the neuron’s pre-activation upward more strongly. A negative weight makes an input pull the result in the opposite direction.
One weight does not necessarily have a simple human interpretation in a large network. Useful behaviour emerges from many parameters interacting across many layers.
Training is the process that changes these weights based on a loss signal rather than a programmer manually setting every value.
Why Bias Matters
Without a bias, the simple linear transformation w·x must pass through the origin. A bias lets the model shift the function.
In a binary linear classifier, w defines the orientation of a decision boundary and b shifts its position. This extra degree of freedom makes even simple models more flexible.
Activation Functions Add Nonlinearity
If every layer only performed linear transformations, stacking many layers would still collapse into one linear transformation. The network could not represent functions that require nonlinear boundaries.
Activation functions break that collapse. They transform pre-activation values before the next layer receives them.
Common activation families include ReLU, sigmoid, tanh and smoother functions such as GELU. Different architectures choose activations according to optimisation and modelling needs.
ReLU
The Rectified Linear Unit is ReLU(z)=max(0,z). Positive inputs pass through; negative inputs become zero.
For z=-1.75 from our toy neuron, ReLU returns 0. For z=2.3, it returns 2.3.
ReLU became popular partly because it is simple and avoids some saturation problems of sigmoid-like activations in deep networks, although modern architectures use many variants.
Sigmoid
The sigmoid function maps any real number to a value between zero and one. It is useful when an output needs a probability-like binary score.
Large positive values approach one; large negative values approach zero. Around zero, the function changes more strongly.
Sigmoid can saturate at extreme values, producing small gradients that slow learning in deep hidden layers. It remains useful in particular output layers and gating mechanisms.
Tanh
Hyperbolic tangent maps values roughly between -1 and 1. It is zero-centered compared with sigmoid and has appeared extensively in recurrent networks.
Like sigmoid, tanh can saturate at large magnitudes. Architecture and optimisation choices determine where it remains useful.
GELU and Smooth Activations
Transformer architectures commonly use smooth activations such as GELU or related variants in feed-forward networks. These functions scale values smoothly rather than applying ReLU’s hard zero boundary.
The activation choice affects optimisation and representational behaviour, but no single activation is universally best for every architecture.
A Layer Is a Vector Transformation
Instead of one neuron, a layer computes many neurons at once. In matrix form, y = activation(Wx + b). If x has four dimensions and the layer has six output units, W has shape 6×4 and b has six entries.
The output y is a six-dimensional vector. It becomes input to the next layer.
This matrix view connects Article 014’s vector space directly to neural networks: each layer learns a transformation between representation spaces.
A Worked Two-Neuron Layer
Let x=[1,2]. Neuron A has weights [1,0.5] and bias 0. Neuron B has weights [-1,1] and bias 0.5.
Neuron A pre-activation = 1×1 + 0.5×2 = 2. ReLU gives 2. Neuron B pre-activation = -1×1 + 1×2 + 0.5 = 1.5. ReLU gives 1.5. The layer output is [2,1.5].
That output can now be transformed by another layer. Deep networks repeat this process many times.
Hidden Layers
Layers between raw input and final output are called hidden layers because their activations are internal representations rather than directly observed labels.
A hidden layer can transform the input into a space where the desired output is easier to compute. Earlier layers may detect local patterns; later layers can combine them into more task-specific representations.
The exact interpretation depends on architecture and domain. Avoid assuming each hidden neuron has one simple concept.
Why Depth Helps
Composing nonlinear layers allows a network to represent complex functions efficiently. One layer can create intermediate features that the next layer recombines.
In image networks, low-level patterns can be combined into shapes and objects. In language transformers, token representations are repeatedly transformed through attention and feed-forward blocks.
Depth is not a guarantee of quality. Training data, optimisation, architecture and evaluation all matter.
The XOR Example
XOR is a classic demonstration that one linear decision boundary is insufficient. Inputs [0,0] and [1,1] belong to one class; [0,1] and [1,0] belong to the other.
No single straight line separates the classes in the original 2D plane. A network with a nonlinear hidden layer can transform the points into a representation where the final decision is separable.
The lesson is not that XOR is difficult. It shows why nonlinearity and hidden representation learning matter.
Forward Propagation
A forward pass applies the current parameters to input data layer by layer until the network produces an output.
PyTorch’s current tutorial describes neural networks as nested functions defined by parameters, with forward propagation producing a prediction before backward propagation computes gradients. See A Gentle Introduction to torch.autograd.
During inference, the system performs forward computation without using the result to update parameters unless a separate training process is running.
The Loss Function
Training needs a numerical objective measuring how undesirable the current output is. That objective is the loss.
For regression, mean squared error is a common example. For classification, cross-entropy-style losses are common. Language models often train with token prediction losses based on the probability assigned to the correct next token.
The loss function defines what optimisation rewards. If the objective does not represent the real task, training can optimise the wrong behaviour very effectively.
Worked Mean Squared Error
Suppose a regression model predicts 8 when the target is 10. Squared error is (8−10)² = 4. If another example predicts 4 when the target is 5, its squared error is 1. Mean squared error across the two is (4+1)/2 = 2.5.
Training adjusts parameters to reduce this aggregate objective over batches of examples.
Cross-Entropy Intuition
In a three-class classification problem, the model outputs probabilities for A, B and C. If the correct class is B, cross-entropy strongly penalises assigning B a very low probability and rewards assigning it a high probability.
The loss is differentiable with respect to logits, allowing gradients to flow into earlier layers.
Gradients
A gradient tells us how a scalar loss changes with small changes in parameters. For one parameter w, the derivative dL/dw measures local sensitivity.
For millions or billions of parameters, the gradient is a vector containing one partial derivative per trainable value.
The sign indicates which local direction increases the loss; optimisation typically moves in the opposite direction.
Backpropagation
Backpropagation efficiently computes gradients through a network by applying the chain rule from the loss backward through the operations that produced it.
PyTorch’s current automatic differentiation tutorial describes backpropagation as adjusting parameters according to gradients of the loss, while autograd computes those gradients through the recorded computational graph.
Backpropagation is not a biological claim. It is an algorithm for differentiating composed functions.
The Chain Rule
If y=f(x) and loss L=g(y), then dL/dx = dL/dy × dy/dx. A deep network is a long composition of functions, so the chain rule passes sensitivity backward from the loss to each parameter.
Automatic differentiation frameworks record the operations needed to compute these derivatives efficiently rather than requiring programmers to derive every gradient by hand.
A Tiny Gradient Example
Let prediction y = wx, target t, and loss L=(y−t)². Choose x=2, w=3, t=10. Prediction y=6. Loss=(6−10)²=16.
dL/dy=2(y−t)=-8. dy/dw=x=2. By the chain rule, dL/dw=-16.
A small gradient-descent update with learning rate 0.1 gives w_new = 3 − 0.1×(-16) = 4.6. The update moves w upward because the model predicted too low.
Gradient Descent
Gradient descent updates parameters in the direction that locally reduces loss: θ_new = θ_old − learning_rate × gradient.
The learning rate controls step size. Too small can make training slow; too large can overshoot useful regions or destabilise training.
Real training uses variants such as stochastic gradient descent, Adam and other optimisers that manage momentum, adaptive scaling or additional state.
Stochastic Gradient Descent
Computing the gradient over an entire huge dataset for every update can be expensive. Stochastic or mini-batch methods estimate gradients using subsets of data.
This introduces noise into updates, which can be computationally efficient and can sometimes help optimisation.
Batch
A batch is a set of examples processed together for one forward/backward update cycle. Matrix operations make batch processing efficient on accelerators.
Batch size affects memory usage, gradient noise and training dynamics. It is an optimisation parameter, not merely a throughput setting.
Epoch
An epoch is one pass through the training dataset under a conventional definition. Large-scale training sometimes tracks steps or tokens rather than epochs because datasets are enormous or dynamically sampled.
More epochs do not guarantee better generalisation. Repeated exposure can eventually lead to overfitting.
Learning Rate Schedules
Training often changes the learning rate over time. Warmup can begin with smaller updates; decay can reduce step size later.
The schedule interacts with optimiser, batch size and model architecture. Training instability can come from a poor schedule even when the model design is sound.
Parameter Initialisation
Training starts from initial parameter values. If all neurons begin identically in a symmetric layer, they can learn identical features. Random initialisation breaks that symmetry.
Initialisation scale also affects activation and gradient magnitude. Techniques such as Xavier/Glorot or He initialisation are designed for particular activation and architecture assumptions.
Vanishing Gradients
When many small derivatives multiply through deep networks, gradients reaching early layers can become extremely small. Those layers then learn slowly.
Saturating activations and deep chains historically made this a major problem. ReLU-like activations, residual connections, normalisation and improved initialisation help modern deep networks train.
Exploding Gradients
The opposite problem occurs when derivatives multiply into very large values. Parameter updates become unstable and numerical values can overflow.
Gradient clipping, careful initialisation, normalisation and architecture design can help contain exploding gradients.
Normalization Layers
Normalization techniques transform activation statistics to make optimisation more stable. Batch Normalization and Layer Normalization operate differently and suit different architectures.
Transformers commonly use Layer Normalization variants because sequence models need stable representation scales across deep blocks.
Residual Connections
A residual connection adds a layer’s input to its transformed output. Instead of forcing the network to learn an entire replacement representation, a block can learn a modification.
Residual paths help gradients flow through deep networks and are central to modern transformer architectures.
The operation is possible because the representations being added have compatible shapes in the residual stream.
Dropout
Dropout randomly removes selected activations during training according to a configured probability. The network cannot rely on one exact pathway every time.
This can act as regularisation and reduce overfitting in some settings. During standard inference, dropout is usually disabled and scaling is handled according to the framework implementation.
Weight Decay
Weight decay penalises or shrinks large parameter values depending on optimiser formulation. It is another regularisation technique intended to improve generalisation.
The best regularisation strategy depends on model scale, architecture, dataset and objective.
Overfitting
A model overfits when it becomes very good at the training data but performs poorly on new relevant examples. Memorisation of quirks can reduce generalisation.
A large network can have enough capacity to fit training examples extremely well. This is why evaluation on held-out data matters.
Underfitting
Underfitting occurs when the model or training process cannot capture the patterns needed even on the training set. The network may be too small, insufficiently trained or mismatched to the task.
Underfitting and overfitting require different repairs. Increasing capacity can help one and worsen the other.
Training, Validation and Test Sets
A training set drives parameter updates. A validation set helps choose hyperparameters and training decisions. A test set estimates final behaviour under conditions not used to tune the model.
Repeatedly tuning on the test set leaks information and weakens its role as an independent estimate.
Generalisation
Generalisation is the ability to perform usefully on relevant examples beyond the training data. It is the real reason training matters.
Generalisation is not one scalar property independent of population. A model can generalise well within one distribution and fail after domain shift.
Data Quality
Neural networks learn from the data and objective they receive. Incorrect labels, duplicated examples, missing populations and spurious correlations can shape learned behaviour.
Better architecture cannot reliably repair every data problem. Dataset documentation and evaluation design are part of the complete training system.
Spurious Correlations
A network may learn a shortcut correlated with the label rather than the intended concept. An image classifier could rely on background patterns; a text classifier could rely on formatting or source-specific wording.
Performance can collapse when the shortcut disappears. Counterexamples and targeted evaluation help reveal the dependence.
Class Imbalance
If one class dominates training data, a model can achieve high accuracy by favouring the majority class while performing poorly on rare cases.
Metrics such as precision, recall, confusion matrices and per-class performance can reveal what aggregate accuracy hides.
Regression Networks
A regression network outputs continuous quantities, such as an estimated demand or temperature. The output layer and loss are designed accordingly.
Prediction intervals or uncertainty estimation require additional modelling; one number is not automatically a calibrated confidence statement.
Classification Networks
A classification network produces scores or probabilities over categories. The final class may be chosen by argmax or a threshold depending on the task.
Article 005’s probability lesson applies: model output and decision policy remain separate.
Embedding Layers Are Neural-Network Layers
Article 013 treated embeddings as vectors. In a language model, the embedding table is itself a trainable parameter matrix.
Token IDs select rows, then the resulting vectors enter deeper neural computation. Gradients can update those embedding values during training.
Neural Networks Versus Traditional Software
Traditional software expresses explicit rules; a neural network learns numerical parameters from data. The two approaches are complementary.
A network can classify messy text while explicit code validates the output schema. A model can predict an action while software enforces permission boundaries.
Neural Networks Versus Databases
A neural network transforms input into learned predictions or representations. A database stores authoritative state.
Do not use model weights as a substitute for a transaction ledger, and do not ask a database to infer fuzzy language meaning without an appropriate model.
Training Versus Inference
Training computes forward outputs, loss, gradients and parameter updates. Inference uses a trained parameter state to compute outputs without those training updates.
A user correcting a chatbot response does not by itself prove that backpropagation changed the deployed model. The correction may simply enter context or application memory.
Parameter Count
Parameter count measures how many learned values a model contains. More parameters increase capacity and resource requirements but do not alone determine usefulness.
Architecture, data, training compute, objective, context and system tools all influence final capability.
Deep Versus Wide
A wide layer has many units; a deep network has many sequential layers. Different architectures trade depth, width, compute and inductive bias.
There is no universal rule that deeper always means better. Optimisation and task structure matter.
Convolutional Neural Networks
CNNs use convolutional operations with shared weights and locality biases, historically central to image processing and still important in many systems.
They demonstrate that “neural network” is a broad family, not one architecture.
Recurrent Neural Networks
RNNs process sequences with recurrent state. LSTM and GRU variants introduced gates to manage information and gradients across time.
Transformers displaced RNNs for many large language tasks because attention enables more parallel training and different long-range information flow, but recurrent architectures remain useful in some domains.
Transformers Are Neural Networks
A transformer is not an alternative to neural networks. It is a particular neural-network architecture built from embeddings, attention, feed-forward networks, normalisation and residual connections.
Article 016 will assemble those parts into the full transformer block.
Computation Graphs
A neural network can be viewed as a directed graph of tensor operations. Inputs flow through nodes to produce outputs and loss.
Automatic differentiation systems record enough of this computation to propagate gradients backward. PyTorch’s current Autograd mechanics documentation describes the dynamic graph and reverse automatic differentiation process.
Accelerators
Training large neural networks relies heavily on GPUs and other accelerators because tensor operations can be parallelised.
Hardware affects feasible model size, batch size and training speed. It does not change the mathematical definition of a gradient or layer.
Numerical Precision
Networks can train and run using different numerical formats such as float32, float16, bfloat16 or lower-precision schemes with appropriate techniques.
Reduced precision can improve speed and memory use but requires attention to numerical stability.
A Complete Worked Training Loop
Task: learn y≈2x from examples (1,2), (2,4), (3,6). Use one parameter w with prediction y_hat=wx and start w=0.
For x=1, prediction is 0, target 2, squared loss 4. Gradient dL/dw=2(wx−y)x = 2(0−2)×1 = -4. With learning rate 0.1, w becomes 0.4.
For x=2 using w=0.4, prediction 0.8, target 4, gradient = 2(0.8−4)×2 = -12.8. The next update raises w again. Repeated examples move w toward 2 under this simplified sequential setup.
Real training uses batches, vectorised operations, more parameters and sophisticated optimisers, but the loop is the same: predict, measure loss, compute gradients, update, repeat.
A Complete Classification Example
Suppose a network classifies short feedback into “concept gap” or “execution error”. Text first becomes tokens and embeddings. Hidden layers transform those representations. An output layer produces two logits.
Softmax converts logits to class probabilities. Cross-entropy compares the correct label with those probabilities. Backpropagation computes gradients through output, hidden layers and embeddings. Optimisation updates all trainable parameters.
After training, inference produces logits and probabilities without updating weights. A later workflow can use thresholds or human review according to consequence.
Reality Check: A Neural Network Learns a Function, Not Understanding by Definition
- Low training loss does not guarantee good generalisation.
- A large parameter count does not establish reasoning, truthfulness or safe autonomy.
- A model can learn shortcuts and spurious correlations.
- The loss function rewards what it measures, not every quality humans care about.
- Backpropagation computes gradients; it does not explain the human meaning of learned representations.
- A neural network can be one component inside a system whose retrieval, tools or permissions still fail.
Neural-Network Failure Class 1: Wrong Objective
The model optimises a measurable proxy that differs from the real goal. Training can succeed numerically while the application performs poorly.
Repair begins by revisiting labels, loss, reward and evaluation criteria.
Failure Class 2: Data Leakage
Information from validation or test examples leaks into training, producing artificially optimistic evaluation.
Separate datasets and preprocessing carefully. Group related examples when random splitting would leak near-duplicates.
Failure Class 3: Overfitting
Training loss keeps improving while validation performance worsens. Consider regularisation, more data, earlier stopping, architecture changes or a better task formulation.
Failure Class 4: Optimisation Instability
Loss becomes NaN, spikes violently or fails to decrease. Inspect learning rate, gradient magnitude, data scaling, numerical precision and implementation.
Failure Class 5: Distribution Shift
The deployed population differs from training. Monitor representative outcomes and retrain or adapt only after identifying the changed conditions.
Repair Pathway
When training fails, separate model capacity from optimisation and data. Can the network overfit a tiny clean sample? If not, architecture or implementation may be broken. If it can, but full-data validation remains poor, investigate data, regularisation and generalisation.
This small-sample overfit test is a practical debugging technique because a model that cannot memorise a few examples often has a more basic problem.
Stabilisation Pathway
Track training and validation loss, gradient norms, activation statistics and task metrics. Preserve seed/configuration/version information so regressions can be reproduced.
Add evaluation slices for important subgroups and known failure modes rather than relying on one aggregate metric.
Extension Pathway
After a network is stable, increase scale, add modalities, distil to smaller models or connect the network to retrieval and tools. Re-run the original evaluation set after each extension.
Independent Exercise 1: Neuron Output
Input x=[2,1], weights w=[0.5,2], bias=-1. What is z=w·x+b?
Answer
0.5×2 + 2×1 −1 = 2.
Independent Exercise 2: ReLU
What are ReLU outputs for [-2,0,3]?
Answer
[0,0,3].
Independent Exercise 3: Gradient Direction
If dL/dw is positive and the learning rate is positive, does gradient descent increase or decrease w?
Answer
It decreases w because the update subtracts learning_rate×gradient.
Independent Exercise 4: Training Versus Inference
A deployed model answers one user request. Does ordinary inference perform backpropagation and update all model weights?
Answer
No. Standard inference performs forward computation. Weight updates require a training process.
Independent Exercise 5: Overfitting
Training accuracy is 99% and validation accuracy is 72%. What concern does this raise?
Answer
Overfitting or distribution mismatch. Inspect the validation split, data leakage, regularisation and whether validation represents deployment.
Independent Exercise 6: Objective Mismatch
A writing model is trained only to predict next tokens well. Does low token-prediction loss guarantee every factual answer is verified?
Answer
No. The training objective does not itself perform source verification. System-level retrieval and checking can add that responsibility.
A Practical Neural-Network Training Checklist
- Define the task and target before choosing architecture.
- Verify labels, splits and preprocessing on a small sample.
- Check tensor shapes through the forward pass.
- Confirm the model can overfit a tiny clean batch as a debugging test.
- Track training and validation metrics separately.
- Inspect learning rate, gradient norms and numerical stability.
- Preserve model, optimiser, tokenizer and dataset versions.
- Evaluate important subgroups and known edge cases.
- Test inference through the complete application, not only the model.
- Keep external permissions and authoritative state outside learned parameters.
A checklist cannot guarantee good machine learning, but it prevents many avoidable category errors. The recurring principle is to separate representation, model capacity, optimisation, data quality and system behaviour so each can be tested directly.
The Jacobian: Many Outputs, Many Sensitivities
When a function maps several inputs to several outputs, the derivatives can be arranged in a Jacobian matrix. Each entry describes how one output changes locally with respect to one input.
Deep networks implicitly compose many Jacobians during backpropagation. The chain rule multiplies these local sensitivities so gradients can reach earlier layers.
This provides another view of vanishing and exploding gradients: repeated multiplication through transformations can shrink or amplify sensitivity dramatically.
The Hessian and Curvature
The gradient tells us the local slope of a scalar loss. The Hessian contains second derivatives and describes local curvature with respect to parameters.
Full Hessians are too large to construct explicitly for modern large networks, but curvature ideas influence optimisation research, learning-rate choices and approximations.
A region with steep curvature can make large updates unstable. A flatter region can tolerate larger local movement, although “flatness” itself depends on parameterisation and should not be oversimplified.
Loss Landscapes
Imagine the loss as terrain over parameter space. Each coordinate corresponds to a model parameter, so a real network’s landscape has millions or billions of dimensions. Training moves through this space seeking lower loss.
The landscape can contain valleys, saddle points and directions with very different curvature. Optimisers do not literally search every point; they use local gradient information and accumulated state.
Two parameter settings can achieve similar training loss yet generalise differently. This is another reason the lowest training loss alone is not the complete goal.
Momentum
Plain gradient descent uses the current gradient. Momentum methods accumulate a moving direction from recent gradients so updates can continue through noisy or shallow regions.
A ball rolling downhill is a common analogy, but the actual algorithm is a numerical recurrence over gradient history. Momentum can accelerate progress along consistent directions and reduce oscillation across steep dimensions.
Adam
Adam maintains moving estimates related to first and second moments of gradients and adapts update scale per parameter. It is widely used in deep learning and transformer training.
The optimiser introduces its own state in addition to model parameters. Saving a training checkpoint for exact continuation can therefore require weights, optimiser state, scheduler state and other run metadata.
Adam is not universally superior to every optimiser. Training recipes depend on architecture, data scale and objective.
Optimizer State Is Not Model Knowledge
At inference time, a deployed model usually needs its learned parameters but not the full optimiser state used during training. Momentum buffers and adaptive-moment estimates are training machinery.
This distinction matters for storage: an active training checkpoint can occupy substantially more memory than model weights alone.
Learning-Rate Warmup
Large networks are often trained with a warmup phase in which the learning rate starts small and rises over initial steps. Early activations and optimiser statistics can be unstable before useful scale develops.
Warmup is a training heuristic with empirical value in many large-model recipes. It is not part of the mathematical definition of a neural network.
Gradient Clipping
Gradient clipping limits gradient magnitude before the optimiser applies an update. Norm clipping rescales the gradient vector when its norm exceeds a threshold.
This can contain occasional exploding gradients, especially in recurrent or large sequence models. Excessive clipping can also hide deeper optimisation problems, so monitor how often clipping occurs.
Gradient Accumulation
If hardware cannot hold a desired large batch in memory, training can process several smaller micro-batches, accumulate their gradients and perform one optimiser step after the effective batch is reached.
The method changes memory requirements without necessarily changing the mathematical effective batch, though details such as normalisation and optimiser scheduling must be handled consistently.
Mixed-Precision Training
Modern accelerators can compute lower-precision formats faster and with less memory. Mixed-precision training uses reduced precision for many operations while retaining higher precision where needed for stability.
Loss scaling and hardware-aware kernels can help prevent tiny gradients from underflowing. The exact techniques depend on numerical format and framework.
Precision is therefore an engineering dimension of training, not a statement about the conceptual intelligence of the network.
Data Parallel Training
Data parallelism replicates the model across multiple devices, gives each device a different mini-batch and combines gradients before the update.
This can increase training throughput but introduces communication cost. The model is conceptually one parameter state even though computation is distributed.
Model Parallelism
When one model is too large for one accelerator, its parameters or operations can be partitioned across devices. Tensor parallelism, pipeline parallelism and other strategies split work in different ways.
Distributed training belongs to infrastructure around the neural network. The forward and backward mathematics remain connected, but communication schedules become crucial to performance.
Checkpointing
Training runs can fail after hours or weeks. Checkpoints save model state and often optimiser, scheduler and training progress so the run can resume.
Activation checkpointing is a different concept: it saves memory by recomputing selected forward activations during backward propagation rather than storing them all.
Clear terminology avoids confusing durable training checkpoints with temporary activation-memory techniques.
Early Stopping
If validation performance stops improving while training loss continues falling, a training process can stop and preserve an earlier checkpoint.
Early stopping is one way to reduce overfitting, though large foundation-model training often follows predetermined compute budgets and uses broader evaluation suites rather than one validation scalar.
Hyperparameters Versus Parameters
Parameters are learned through optimisation: weights, biases and embeddings. Hyperparameters are chosen by the training process or experimenter: learning rate, batch size, layer count, hidden dimension, dropout rate and others.
Some hyperparameters can themselves be tuned by automated search, but they are not updated through the ordinary backpropagation step in the same role as network weights.
Ablation Studies
To understand what a component contributes, remove or change it while keeping other conditions as stable as practical. This is an ablation.
For example, compare a network with and without residual connections under the same training setup. Differences can provide evidence about the component’s contribution, though interactions mean no ablation explains every causal effect alone.
Seed Variance
Training includes randomness from initialisation, data order, dropout and distributed operations. Two runs with identical high-level settings can produce different final metrics.
For small experimental differences, repeat runs or report uncertainty rather than treating one seed as definitive evidence.
Reproducibility
A reproducible training record includes code version, model configuration, dataset version, preprocessing, random seeds, optimiser settings, learning-rate schedule, precision and hardware details where relevant.
Exact bitwise reproducibility can be difficult across hardware and software versions, but disciplined records make behavioural changes far easier to diagnose.
Calibration After Training
A classifier can achieve high accuracy while producing overconfident probabilities. Calibration evaluates whether predicted probabilities correspond to observed frequencies.
Post-hoc calibration methods can adjust output probabilities without relearning the entire representation. Article 005 explains why probability quality and decision thresholds are separate from raw class accuracy.
Transfer Learning
A network trained on a broad source task can provide parameters useful for a narrower target task. Fine-tuning continues training on target data; feature extraction can freeze much of the network and train only a smaller head.
Foundation models extend this principle dramatically: broad pretrained capability is adapted through prompting, fine-tuning, retrieval or tool-connected systems.
Fine-Tuning
Fine-tuning updates some or all pretrained parameters using new data and objectives. It can specialise behaviour but can also degrade capabilities if data, learning rate or objective are poorly chosen.
Evaluation should therefore include both the new target task and important capabilities that must remain intact.
Parameter-Efficient Fine-Tuning
Methods such as adapters and low-rank adaptation introduce or train a smaller set of parameters rather than updating every base weight.
These techniques reduce training memory and make multiple specialisations easier to store. They still require evaluation; a small trainable parameter count does not guarantee a small behavioural change.
Distillation
Knowledge distillation trains a smaller student model to reproduce useful behaviour or distributions from a larger teacher, often together with ordinary labels.
The student can become cheaper to serve, but some capabilities may be lost. Distillation quality should be evaluated on the deployment task and safety boundaries.
Pruning
Pruning removes or zeros selected parameters, channels or structures to reduce model size or computation. The challenge is preserving useful behaviour while gaining efficiency.
Unstructured sparsity may not translate directly into hardware speed unless the serving stack can exploit it.
Quantisation
Quantisation represents weights or activations with fewer bits. It can reduce memory and accelerate inference, but approximation error can affect output.
Quantisation-aware training or careful post-training calibration can preserve quality better than naive rounding.
Neural Networks Can Memorise
Large neural networks can memorise training examples as well as learn generalisable patterns. Memorisation is not automatically a defect; exact rare facts or examples can be useful. It becomes problematic when the model relies on memorisation instead of transferable structure or exposes sensitive training data.
Privacy and generalisation analysis should therefore look beyond aggregate test loss.
Neural Networks Can Be Brittle Off Distribution
A network trained within one range can behave unpredictably far outside it. A regression model trained on temperatures from 10°C to 30°C may extrapolate poorly at 100°C.
Applications should detect unsupported operating regions where possible and avoid interpreting any numerical output as equally trustworthy.
Adversarial Examples
Small input changes can sometimes cause large output changes in neural networks, including changes that appear insignificant to humans. Adversarial robustness is its own evaluation problem.
This illustrates that learned decision boundaries can differ from human perceptual boundaries.
Interpretability and Probing
Researchers can inspect activations, train probes, trace attention patterns or perform interventions to study learned representations. These tools provide evidence, but one visualization rarely yields a complete mechanistic explanation.
A probe’s ability to decode information from an activation does not by itself prove the model uses that information causally for the output.
A Neural-Network Failure Matrix
- Training loss will not fall: inspect implementation, learning rate, data scaling, initialization and model capacity.
- Training good, validation poor: inspect overfitting, leakage, distribution mismatch and regularisation.
- Validation good, deployment poor: inspect distribution shift, input pipeline and system integration.
- Gradients explode: inspect learning rate, architecture, normalization and clipping.
- Model predicts confidently but incorrectly: inspect calibration, hard negatives and out-of-distribution cases.
- Fine-tune fixes target but breaks old tasks: run multi-capability regression tests and revise adaptation strategy.
Independent Exercise 7: Momentum
Why can momentum help when gradients repeatedly point in a similar direction but contain noise?
Answer
The accumulated update direction reinforces consistent gradient components while reducing sensitivity to one noisy step, allowing faster movement through the shared direction.
Independent Exercise 8: Hyperparameter
Is the learning rate a model parameter learned by ordinary backpropagation in a standard training run?
Answer
No. It is a hyperparameter controlling optimiser updates unless a separate meta-learning scheme explicitly learns it.
Independent Exercise 9: Transfer Learning
A pretrained image network is reused for a small new classification task. Why might this require fewer examples than training from random initialization?
Answer
Earlier layers may already contain useful visual representations learned from a larger dataset, so the target task can adapt existing features rather than learn every pattern from scratch.
Independent Exercise 10: Mixed Precision
Why can reduced-precision training require special stability techniques?
Answer
Lower-precision formats have smaller representable ranges or fewer significant bits, so tiny or huge values can underflow, overflow or accumulate rounding error. Mixed precision retains higher precision where necessary.
Frequently Asked Questions About Neural Networks
What is a neural network?
A parameterised function composed of layers of numerical operations, typically including learned weights, biases and nonlinearities.
What is a parameter?
A learned numerical value such as a weight, bias or embedding entry that training can update.
What is a neuron?
A simplified computational unit that forms a weighted combination of inputs, adds a bias and often applies an activation function.
Why are activation functions necessary?
Without nonlinearities, stacked linear layers collapse into one linear transformation and cannot represent many complex functions.
What is backpropagation?
An efficient application of the chain rule that computes gradients of the loss with respect to parameters through a computational graph.
What is gradient descent?
A family of optimisation methods that update parameters using gradients to reduce loss.
Does training happen every time I use a model?
Not necessarily. Standard inference uses fixed trained parameters. Context and memory can change outputs without retraining.
What is overfitting?
When a model fits training data much better than it performs on relevant unseen data.
Are transformers neural networks?
Yes. Transformers are neural-network architectures built from attention, feed-forward layers, embeddings, normalization and residual connections.
Does a bigger neural network always perform better?
No. Scale can expand capacity, but data, objective, optimisation, architecture, compute and evaluation all matter.
Selected Technical References
- PyTorch Learn the Basics — current neural-network workflow tutorials.
- Automatic Differentiation with torch.autograd — gradients and backpropagation.
- PyTorch Autograd Mechanics — current reverse automatic differentiation details.
- Google Machine Learning Crash Course: Neural Networks — neural-network fundamentals.
Neural Networks Learn Transformations Through Error and Gradients
The complete learning loop is now visible. Inputs become vectors. Layers compute weighted transformations and nonlinear activations. The network produces an output. A loss compares the output with the training objective. Backpropagation computes gradients. An optimiser updates parameters.
Repeated across data and compute, that loop can learn extraordinarily complex functions. It still needs appropriate objectives, data and evaluation, and it remains only one layer of a complete SI system.
Next: 016 — Transformers, where embeddings, vector projections, attention, feed-forward networks, residual connections and normalisation are assembled into the architecture behind many modern language models.
How Super Intelligence Works Series Navigation
Previous: 014 — Vector Space · Series Hub · Next: 016 — Transformers.
