VIEW THIS AS

Auto mode follows the Route Engine until you choose a viewpoint.

YOU ARE HERE

ROUTE CHECK

CONNECTED TO

WHAT NEXT

Use the canonical route for this room, or HELP if you are unsure.

How Inverse Problems Work | From Noisy Measurements Back to Hidden Causes, Regularisation and Uncertainty

An inverse problem starts with observations and asks what hidden state, object or parameter could have produced them through a specified process. A reliable answer must do more than reproduce the observations: it must account for missing information, noise, modelling assumptions and the possibility of other explanations.

This is a Mathematics long read about reconstruction, not a promise that every hidden cause can be recovered. It begins with arithmetic that a school student can follow, develops the linear algebra behind unstable answers, and reaches regularisation, probabilistic uncertainty, imaging and experimental design. The main calculations use deliberately small examples so that the reader can inspect the mechanism rather than take a software output on trust.

Aisha, Ben, Mira, Ryan, Clara, Ethan, Adrian and Jo appear in clearly fictional teaching scenes. All numerical datasets, classroom exchanges and measurement devices constructed for this article are illustrative, not reports of experiments or student outcomes. The medical-imaging passages explain mathematical ideas; they do not provide clinical interpretation or equipment-operating instructions.

Choose a reading route. For the central idea, read Chapters 1–5 and 25. For the calculations, follow Chapters 6–13 and the reproducible laboratory in Chapter 24. For uncertainty and trustworthy evidence, read Chapters 14–16, 20–22 and 25. For imaging and physical systems, follow Chapters 17–19 after Chapter 8. For teaching and independent practice, start with Chapters 1–3, then Chapters 23–25.

The existing How Mathematics Works library provides the wider route. The preceding Uncertainty Quantification article asks how uncertainty travels through a model. Here the direction changes: observations have arrived, and we want to know what they genuinely reveal about something we cannot directly see.

1. The reading that did not reveal what was inside

Aisha has two covered boxes on a table. There are counters inside them, but she is not allowed to open either lid. Ben is allowed one measurement: he can find the combined number of counters. The display reads twelve. He writes six beneath each box.

His arithmetic is impeccable. Six plus six is twelve. His inference is not yet justified. The boxes could contain five and seven, or two and ten, or zero and twelve. Even if both boxes must contain at least one counter, there are still several possibilities. The display has measured a total. Ben has supplied a division of that total. Something important happened between the observation and the answer, and it was not another measurement.

That small gap is the entrance to inverse problems. In the forward direction, we choose the contents and calculate the reading. Given five counters in one box and seven in the other, the display must show twelve under our idealised counting rule. In the inverse direction, we start with twelve and try to recover the contents. The forward operation has combined information. Reversing the arithmetic does not automatically separate what was combined.

Let the hidden counts be x₁ and x₂. The observation is y = x₁ + x₂. Writing y = 12 gives one equation. It does not give x₁ = 6 and x₂ = 6. That pair becomes justified only if another condition enters, such as reliable knowledge that the boxes contain equal numbers. The equality condition might be true, but it is additional information. It did not come from the total alone.

Now imagine that Ben is building a system to return an answer every time. He programs it to split each total equally. The system is fast, consistent and easy to understand. It will always reproduce the measured total. None of those properties establishes that its reconstructed contents match the actual contents. The system has selected one answer from a set of compatible answers. A useful selection rule has been mistaken for a discovery.

Aisha changes the question. Instead of asking how many counters are in each box, she asks how many counters there are altogether. Now the same observation is sufficient. The experiment has not improved. The question has become one the experiment can answer. This is an important possibility throughout the article: sometimes the responsible response is to recover a coarser quantity accurately rather than manufacture a detailed reconstruction.

Suppose the display reports the total only to the nearest two counters. A reading of twelve may then represent more than one actual total, depending on the rounding convention. We now have two uncertainties: which total was present, and how that total was divided. They are not the same problem. A better display could reduce uncertainty about the total while leaving the division completely unresolved. Buying more precision would not necessarily buy the information Ben needs.

The obvious repair is a different measurement. We could measure the count in one box separately, measure the difference between the boxes, or use a device whose response to the first box differs from its response to the second. The best extra observation is not simply another copy of the first. It is one that distinguishes alternatives that the first observation treats as identical.

This explains why inverse problems are as much about observation design as calculation. Before asking which algorithm will recover the answer, ask what the measuring process has preserved. A process can erase detail, blend it, average it, sample it incompletely or attenuate it until noise dominates. The algorithm inherits those limitations. It does not stand outside them.

The same broad forward–inverse distinction appears in university treatments of inverse problems and imaging: a mathematical model predicts measurements from an object, while reconstruction tries to infer the object from measurements. That formal framework is developed in Tristan van Leeuwen’s introductory lectures. The boxes here are an original teaching construction, chosen because the missing information is visible without advanced notation. [1]

There is a humane lesson in Ben’s mistake. He did not fail because he could not add. He answered a question that the available evidence had not settled. A learner can make that mistake with a word problem; a research team can make it with a sophisticated model. What changes is the size of the machinery around the inference, not the logical obligation.

By the end of this article, six and six will return in several disguises: the minimum-length vector, a smooth image, a probable reconstruction, an early-stopped iteration. Each can be useful. Each must be described honestly. The central question will remain the same: which part of this answer was measured, which part was inferred, and which part was supplied by a preference about what the hidden world should look like?

2. Start with the journey from the object to the observation

The cleanest way to understand an inverse problem is to write the forward process first. Let x represent the hidden object, F the measurement model, and y the ideal observation. Then y = F(x). An actual observation is often written y_obs = F(x) + ε, where ε represents a specified observation error. This equation is a starting model, not a declaration that every real error is additive or independent.

The object x could be one number, such as a spring constant. It could be a vector containing several concentrations. It could be an image, a temperature field, a boundary shape or a function varying over time. The observation y may have a very different form. A picture can produce a list of intensity measurements; an underground structure can produce arrival times; a changing system can produce a short sequence of sensor readings.

F tells us how a candidate hidden object would become observable. That description may include physical evolution, a sensor response, sampling, averaging and units conversion. In a useful decomposition, x first becomes a physical signal, the signal passes through an instrument, and the instrument produces recorded data. Combining these stages into one F is convenient, but the individual stages remain important when something goes wrong.

Consider an idealised concentration measurement y = gx + b. Here x is the concentration, g is the instrument gain, and b is an offset. When g and b are known, recovering x from y is straightforward if g is nonzero: x = (y − b)/g. When g is unknown too, a single observation cannot generally separate concentration from gain. A reading of ten could arise from x = 5 and g = 2, or x = 2 and g = 5, when the offset is zero.

The algebra has not suddenly become difficult. The information structure has changed. An inverse problem with a known sensor is not the same problem as one that must infer both the hidden object and the sensor. Calling both tasks “find x” hides an essential difference. Later, this distinction will reappear as calibration, operator uncertainty and blind deconvolution.

There is another distinction between an inverse problem and an inverse function. A function has an ordinary inverse on an appropriate domain only when it maps different inputs to different outputs and the requested outputs belong to its range. Many measurement maps do not satisfy those conditions. Adding two numbers cannot be inverted into a unique ordered pair. Squaring a real number cannot reveal its sign. Sampling a signal may not reveal its behaviour between samples.

Nevertheless, inverse problems can still be useful and solvable in a qualified sense. We may recover a set of possible objects, a stable coarse feature, a regularised estimate, or a probability distribution conditional on assumptions. The term “inverse” names the direction of the question. It does not guarantee that there is a single algebraic undo button.

For a candidate reconstruction x_hat, the natural first check is forward prediction: compute F(x_hat) and compare it with the observation. This is often called a data-consistency or residual check. If the predicted observation is incompatible with what was measured, the reconstruction needs explanation. But passing the check is only necessary in the chosen model, not sufficient for truth. Ben’s six-and-six answer passes it perfectly.

A useful habit is to keep three objects on the page: the observation y_obs, the reconstruction x_hat, and the predicted observation F(x_hat). Students often collapse the last two. They say that an answer “matches the data” when they mean that the answer looks plausible. The actual check requires sending the reconstructed object through the same process that produced the observations.

This forward check can also reveal a misleading story. Suppose a reconstructed image is visually sharp but its simulated measurements disagree with the recorded measurements. Its appearance does not resolve that disagreement. Conversely, an image may reproduce the measurements while containing arbitrary detail in unobserved directions. The two checks address different questions: compatibility with observations and justification of unobserved structure.

For the rest of the article, “hidden causes” should therefore be read carefully. Sometimes we are recovering a physical parameter inside an already justified causal model. Sometimes we are reconstructing a state without establishing why that state arose. Sometimes we are estimating a mathematical quantity with no causal interpretation at all. Reconstruction and causal discovery overlap in some applications, but they are not synonyms.

The practical sequence is now clear. Name the hidden object. Describe how that object produces observations. Name the quantities assumed known. Specify what can vary and how measurement error is represented. Only then decide what it would mean to recover an answer. This sequence keeps a mathematical solution attached to the question it is supposed to answer, rather than allowing an attractive formula to quietly choose the question for us.

3. What exactly is unknown, and what exactly is being asked?

Before solving anything, write a small contract for the problem. It should name the hidden quantity, the observations, the measurement process, the allowed values and the desired result. This is not administrative decoration. Changing any one of these can change whether recovery is possible, which method is appropriate and what confidence the answer deserves.

Take the two boxes again. If x₁ and x₂ are nonnegative integers, the solutions to x₁ + x₂ = 12 are a finite set. If they are nonnegative real masses, the solutions form a continuous line segment. If negative values are allowed because the quantities represent signed corrections, the set extends without bound. The observation is unchanged, but the admissible world is different.

This is why domain restrictions belong inside the mathematics. Positivity may be a genuine physical condition for a mass. It may be inappropriate for an electrical potential measured relative to a reference. An image intensity might be nonnegative after a particular calibration, while a background-subtracted signal can be negative without being erroneous. A constraint should be justified by the quantity’s meaning, not selected because it makes the answer look tidier.

Units deserve the same care. Suppose a vector combines a temperature and a length. Penalising x₁² + x₂² without a scale convention gives a numerical preference that changes when metres become centimetres. The physical situation is identical, but the penalty can now favour a different answer. A model that claims to reveal the object should not silently change its preference because the report changed units.

One repair is nondimensionalisation. Divide each quantity by a meaningful reference scale before combining it in a norm or penalty. Another is a covariance model that encodes the expected scale and dependence of each component. Neither removes judgement. They make the judgement inspectable. Later chapters will show that a regularisation parameter only has meaning alongside the scaling of the data, the operator and the penalty.

It is also essential to distinguish the object from the quantity of interest. An image may contain thousands of unknown values, while the actual decision concerns their average over one region. A system may have many unknown parameters, while the useful output is a predicted total. We might be unable to identify the entire object but still identify the quantity of interest.

A simple original example makes this precise. Suppose x is a vector of four compartment amounts and the only reliable observation is their total. Recovering all four amounts is impossible without additional information. Recovering their mean is easy: divide the total by four. Recovering the largest compartment is different again. Nonnegativity supplies bounds, but not an exact largest value. The measurement supports different answers to different questions.

When Mira asks whether the reconstruction is accurate, Adrian asks her to finish the sentence: accurate about what? Correct total, correct location, correct peak height, correct sign, or correct classification? These can diverge. A smoothed reconstruction may preserve total mass while reducing the peak. A shifted reconstruction may preserve shape while giving the wrong location. A low average error can coexist with a serious error in one small region.

The error measure must therefore match the promised answer. Squared error across every component treats the task differently from maximum error, relative error, edge-location error or error in a derived total. A method may be strong under one measure and weak under another. Ranking methods without stating the measure is like comparing journeys without saying whether distance, time or cost matters.

The measurement model also needs an error convention. Does “accurate to 0.1” mean a guaranteed bound, a standard deviation, instrument resolution, or a rounding interval? Those are different statements. A standard deviation is not an absolute maximum. A resolution specification does not automatically include calibration bias. A noise distribution must not be invented merely because the software expects one.

Finally, decide what kind of answer is acceptable. A unique exact object is one possibility. A family of compatible objects is another. A regularised estimate with an explicit preference, a posterior distribution under a stated model, or a bound on a useful feature may be more defensible. There is no virtue in demanding a detailed point estimate when the observations support only a range.

A good problem contract therefore protects both ambition and restraint. It allows us to attempt difficult reconstruction while preventing the answer from expanding beyond the evidence. The purpose of the contract is not to make the problem easy. It is to make clear what a successful solution would actually establish, and what would remain unresolved even after the computation has finished.

4. Three different ways a reconstruction can fail

The traditional well-posedness questions are existence, uniqueness and stability. Does an admissible solution exist? Is it the only one? Does it change continuously when the data change slightly, in the mathematical spaces being used? These questions frame much of the theory of inverse problems. They are not three names for the same defect. [2]

For existence, imagine two instruments supposedly measuring the same exact constant with no noise. One reports eight and the other nine. Under the declared exact model, there is no value that satisfies both readings. The problem is inconsistent. We should not accuse the solver of weakness for failing to find the nonexistent value. We should revisit the model, error assumptions or data.

In practice, an inconsistent equation often becomes a fitting problem. We choose a quantity that makes the predictions as compatible as possible with the observations according to a defined error measure. That is what least squares will do. But the change matters: “minimises disagreement” is not the same claim as “satisfies every equation exactly.” The mathematical wording must follow the actual task.

For uniqueness, return to x₁ + x₂ = 12. A solution exists, but there are many. Extra numerical precision in the total does not solve this. Even knowing the total as exactly twelve leaves the division unknown. An additional condition can select an answer, but we must record where that condition came from.

Stability is subtler because a unique solution can still be extremely sensitive. Consider the two equations x₁ + x₂ = 8 and x₁ + 1.001x₂ = 8.002. Subtracting gives 0.001x₂ = 0.002, so x₂ = 2 and x₁ = 6. Now change only the last recorded number from 8.002 to 8.003. The inferred values become x₂ = 3 and x₁ = 5. A change of one thousandth in one reading has moved both inferred quantities by one.

The computation is correct in both cases. The problem is that the measurements are almost redundant. Their tiny difference carries the information needed to separate the unknowns, so measurement error in that difference is greatly amplified. An answer may have a unique algebraic formula and still be too fragile for a particular measurement precision.

There is a useful technical qualification. A fixed invertible matrix in finite dimensions has a continuous inverse. We usually describe a matrix like the one above as ill-conditioned, not literally discontinuous. In infinite-dimensional problems, such as recovering an entire past temperature field, an inverse can genuinely fail to be continuous in the chosen norms. Numerical practice sometimes uses “ill-posed” broadly, but the distinction is worth preserving.

Conditioning asks how the mathematical answer responds to changes in the input. Stability of an algorithm asks what extra damage the computational procedure introduces. An excellent algorithm cannot manufacture information that the measurements barely contain. A poor algorithm can also damage an otherwise manageable problem. Both checks are needed, and solving one does not settle the other.

A scalar example helps separate absolute and relative sensitivity. If y = 0.001x, then an absolute observation error of 0.01 becomes an absolute reconstruction error of ten. Yet a scalar linear map has relative condition number one away from zero when errors are measured relatively in the corresponding input and output. Calling the problem simply “badly conditioned” without specifying scales can obscure this. What matters operationally is the error level relative to the measured signal and the requested accuracy.

We should also avoid assuming that every inverse problem is difficult. Recovering x from y = 3x + 2, with known coefficients and modest noise, is straightforward. Some inverse maps are stable, some have finite ambiguity that additional information resolves, and some are severely unstable. “Inverse” identifies the direction of inference; the measurement operator determines the difficulty.

When a solver gives an unexpected answer, these distinctions provide a diagnostic sequence. First check whether the observations are compatible with the exact model. Then ask whether more than one admissible object could fit. Then test how the answer changes under realistic perturbations. Finally inspect whether the algorithm solves the formulated problem accurately. Skipping directly to a new optimiser can waste effort on the wrong failure.

A strong reconstruction report should say which of these questions has actually been answered. “The optimiser converged” may answer none of them completely. “The regularised objective has a unique minimiser” answers a question about a modified problem, not necessarily about the original hidden object. Being precise here is not pessimism. It is how the reader learns what has been established and which next step could genuinely improve the inference.

5. The invisible directions inside a measurement

The two-box problem becomes more revealing when we stop listing possible answers and describe the difference between them. Starting with six counters in each box, move one counter from the first box to the second. The hidden arrangement changes, but the measured total does not. Move two counters instead. Again, the instrument remains silent. There is a whole direction of change that the measurement cannot see.

In linear algebra, the collection of changes that produce no change in the observations is called the nullspace. If the forward model is y = Ax, a vector z belongs to the nullspace when Az = 0. Consequently, A(x + z) = Ax. Objects separated by such a change are observationally indistinguishable under that instrument. The language is compact, but the meaning is concrete: the apparatus has a blind direction. [2]

For A = [1, 1], the change z = (1, −1) is invisible because 1 − 1 = 0. Every multiple of that change is also invisible. We can move along x + tz without altering the total, provided the resulting quantities remain in the allowed domain. Nonnegativity limits the journey, but it does not usually reduce the entire compatible line segment to one point.

This gives a precise test for whether a particular quantity can be recovered even when the full object cannot. Suppose we want q = cᵀx, where c is a vector of known weights. If cᵀz = 0 for every invisible change z, then all exact solutions compatible with the observations give the same q. The quantity is unaffected by the ambiguity. We do not need every component of x to determine it.

For the two boxes, the total corresponds to c = (1, 1). Its dot product with the invisible change is zero. The contrast corresponds to c = (1, −1), whose dot product with that same change is two. The total is recoverable; the contrast is not. These are different questions asked of the same data, not different opinions about the instrument.

There is a useful computational equivalent. In a finite-dimensional linear problem, the recoverable linear quantities are those whose coefficient vectors lie in the row space of A. If cᵀ = wᵀA for some known weights w, then q = cᵀx = wᵀAx = wᵀy. We have expressed the desired quantity directly in terms of the observations. This is a short proof of recoverability, not merely a successful numerical experiment.

Now imagine adding a second instrument that reports twice the total. Its matrix row is [2, 2]. We have twice as many readings, but the second row supplies no new direction of information. It is a multiple of the first. It can help check consistency or reduce some random uncertainty when measured independently, but it cannot reveal the split. Measurement count and information count are not the same.

By contrast, an instrument that reports the first box minus the second adds the row [1, −1]. Together, total and contrast determine both contents: the first is half their sum, and the second is half their difference. The useful new instrument does not merely give another number. It responds differently to arrangements that the original instrument could not distinguish.

This is why matrix rank matters. Rank measures the number of independent linear measurement directions, not simply the number of rows written on the page. Four measurements can contain only three independent relationships. A hundred repetitions of one relationship can still leave the same exact blind direction. Later, our miniature tomography example will make this visible without requiring a large scanner or a large matrix.

A common response to nonuniqueness is to choose the solution with smallest Euclidean norm. For a fixed sum of twelve, that choice gives six and six. It is mathematically well defined and can be useful. But its interpretation must remain honest: among compatible answers, it prefers the one closest to zero in the chosen coordinates. The data did not independently report that the boxes were balanced.

Changing coordinates can change the meaning of that preference. If one unknown is recorded in metres and another in millimetres, an unscaled norm can favour changes in one component simply because its numerical units are different. The preference is therefore part of the model. It deserves an explanation just as much as the measurement equations do.

This chapter provides a practical way to ask a better question when reconstruction seems impossible. Identify what the instrument cannot see, then ask whether the decision actually depends on those invisible directions. Sometimes a new measurement is needed. Sometimes a justified constraint will help. Sometimes the right outcome is to report a well-determined total rather than an unsupported detailed picture. Recovering less can produce a more informative answer when the smaller answer is the one the evidence genuinely supports.

6. Noise is not one universal kind of disturbance

When Aisha repeats a measurement, the displayed number changes slightly. Ben proposes averaging ten readings. That is sensible for some kinds of noise. It is almost useless for others. Before averaging, we need to know what changes between readings and what stays wrong in the same way each time.

Consider a scale whose zero point is shifted upward by two units. Repeating a measurement a thousand times can estimate its biased average very precisely. The common offset remains. By contrast, independent zero-mean fluctuations around a correctly calibrated reading can average down. The difference is not visible in the word “noise”; it is visible in the mechanism that generated the observations.

A simple model separates these effects as yᵢ = x + b + εᵢ. Here x is the wanted quantity, b is a common offset, and εᵢ varies between repetitions. With only repeated readings from that device, x and b are confounded: increasing x while decreasing b by the same amount leaves the predicted readings unchanged. Independent calibration information is needed to separate them.

Under the additional assumption that the εᵢ are independent with mean zero and common variance σ², their average has variance σ²/n. The standard deviation of the average is therefore σ/√n. This familiar improvement is a conditional mathematical result. It does not apply unchanged when the errors share a drifting offset, are strongly correlated, or are selected in a biased way.

Noise can also depend on signal strength. A counting instrument may behave differently when it receives ten events and when it receives ten thousand. An imaging system may clip values above a maximum. A communication device may drop observations entirely. An error model that allows every real value with equal variance is not automatically suitable for counts, censored measurements or missing data.

For Gaussian additive errors with known positive-definite covariance matrix Σ, the discrepancy between prediction and observation is naturally measured by rᵀΣ⁻¹r, where r = Ax − y. Residuals in high-variance directions count less than equally sized residuals in low-variance directions. Correlations matter because the joint pattern of errors, not just their individual sizes, affects how surprising a residual is. This is one reason statistical formulations of inverse problems explicitly describe the observation process. [3]

The same idea can be written using a whitening matrix W satisfying WᵀW = Σ⁻¹. Then the discrepancy becomes ‖W(Ax − y)‖². Whitening changes the geometry in which fit is assessed. A residual that looks large in raw units may be ordinary after measurement uncertainty is considered; another that looks small may be serious for a highly precise instrument.

For an elementary example, suppose two independent devices report the same quantity as 10 and 12. Their standard deviations are known to be 1 and 2. Weighted least squares gives the first reading weight 1 and the second weight 1/4. The estimate is (10 + 12/4)/(1 + 1/4) = 10.4, not the unweighted average 11. This result follows from the stated uncertainty model, not from a general rule that one particular device is always trustworthy.

A bound is a different object. Saying that an error lies between −0.1 and 0.1 does not imply a Gaussian distribution with standard deviation 0.1. It does not even imply that central values are more likely than edge values. Bounds support worst-case compatibility questions. Probability models support likelihood and probability questions. Converting one into the other requires additional assumptions that should be visible.

Outliers create another choice. Squared residuals give very large discrepancies disproportionate influence. An absolute-value loss or a robust loss can reduce that influence, but robustness is not permission to erase inconvenient data. An extreme reading may be a recording error, a genuine event, or evidence that the forward model is missing something important. The analyst should investigate before deciding what mathematical treatment is appropriate.

Preprocessing can also change the noise. Averaging neighbouring pixels creates dependence between neighbouring outputs. Taking logarithms can turn additive measurement error into a complicated signal-dependent error. Normalising every observation by the same uncertain reference couples their errors. A likelihood written for the raw instrument should not be reused automatically after the observations have been transformed.

For a reader, the practical question is not “Did the article mention noise?” It is “Did the analysis connect its error assumptions to the measurement process?” A strong reconstruction records units, calibration, missing observations, dependence and preprocessing. It explains which assumptions are simplifications and tests the consequential ones. More sophisticated reconstruction cannot compensate for pretending that the instrument produced a cleaner, simpler kind of evidence than it actually did.

7. Least squares finds a compromise, not a hidden guarantee

Suppose Clara measures a straight-line relationship at three input settings. The settings are 0, 1 and 2. The observed outputs are 1, 2 and 4. She proposes the model y = a + bt, with an unknown intercept a and slope b. No single straight line passes through all three observations. The inverse problem is now to infer the two parameters from imperfect data.

Ordinary least squares chooses a and b to minimise the sum of squared residuals. With these observations, the objective is (a − 1)² + (a + b − 2)² + (a + 2b − 4)². Every candidate line receives a score. The method selects the line with the smallest score according to this particular definition of disagreement.

We can solve this small problem exactly. The mean input is one and the mean output is 7/3. The slope numerator is (−1)(−4/3) + 0(−1/3) + 1(5/3) = 3. The denominator is 1 + 0 + 1 = 2. Thus b = 3/2, and a = 7/3 − 3/2 = 5/6. The fitted line is ŷ = 5/6 + 3t/2.

At the three measurement settings, this line predicts 5/6, 7/3 and 23/6. Using observed minus predicted residuals, the discrepancies are 1/6, −1/3 and 1/6. Their squared sum is 1/36 + 1/9 + 1/36 = 1/6. We can check the arithmetic without needing a statistical claim about the world.

The residuals have two important orthogonality properties. Their sum is zero. Their input-weighted sum is also zero: 0(1/6) + 1(−1/3) + 2(1/6) = 0. These are the conditions saying that a small change in intercept or slope cannot reduce the objective to first order. Geometrically, the fitted observation vector is the orthogonal projection of the measured vector onto the space of observation vectors that the line model can produce.

In matrix notation, A contains rows [1, 0], [1, 1] and [1, 2], and x contains a and b. Least squares minimises ‖Ax − y‖². At a minimiser, Aᵀ(Ax − y) = 0. These normal equations explain the projection, but they are not always the best computational recipe. Forming AᵀA can square the two-norm condition number for a full-column-rank matrix, making an already sensitive problem harder to handle numerically.

For actual computation, a least-squares solver based on suitable matrix decompositions is usually preferable to explicitly calculating (AᵀA)⁻¹Aᵀy. Our accompanying laboratory uses numpy.linalg.lstsq with an explicit rcond choice rather than implementing an inverse of the normal matrix. The documentation describes the solution and rank information returned by that routine; it does not validate the measurement model supplied to it. [11]

Least squares also needs a uniqueness qualification. If the columns of A are linearly independent, the quadratic objective has a unique minimiser. If they are dependent, different parameter vectors can produce the same fitted observations. A numerical routine may return a minimum-norm solution, but that is a particular selection from the compatible minimisers. The number returned does not make the original parameters identifiable.

The method’s statistical interpretation requires further assumptions. Independent Gaussian errors with a common known variance lead to ordinary least squares as a maximum-likelihood parameter estimate. Unequal variances lead naturally to weighting. Correlated errors call for their joint covariance. Without those assumptions, least squares can still be a useful geometric fitting criterion, but its uncertainty interpretation must not be imported automatically.

Clara should now ask whether a straight line is a reasonable model. Three observations are not enough to establish that nature is linear. A curved relationship could also explain them. Least squares has answered the question posed—find the best straight-line fit under squared discrepancy—not the larger question of whether a straight line is the correct mechanism.

Prediction and parameter recovery may also have different reliability. Several combinations of intercept and slope can predict almost the same outputs over a narrow observed range while diverging far outside it. A good fit near the data does not license unrestricted extrapolation. The uncertainty of a future prediction depends on where it is requested and which parts of the model are being trusted.

The proper conclusion is therefore modest but useful: under this line model and this discrepancy measure, the selected parameters minimise the observed disagreement. We can then inspect residuals, compare alternative models, collect observations at new settings and quantify uncertainty. Least squares is a powerful stage in reconstruction precisely when we remember what it does not finish. It supplies a controlled compromise with the observations, not an automatic certificate that the hidden process has been discovered.

8. Singular values show which details the instrument weakens

Return to two hidden quantities, but replace the total-only instrument with a pair of mixed readings. The first device records 51% of the first quantity plus 49% of the second. The second records 49% of the first plus 51% of the second. The instruments are different, but only slightly. They should be able to distinguish the quantities in exact arithmetic. The question is how well they can do it with imperfect readings.

Write the matrix as A = [[0.51, 0.49], [0.49, 0.51]]. Instead of looking at the original coordinates, describe the hidden object by its sum s = x₁ + x₂ and contrast d = x₁ − x₂. Adding the two measurement equations gives y₁ + y₂ = s. Subtracting them gives y₁ − y₂ = 0.02d. The total passes through unchanged; the contrast is reduced by a factor of fifty.

This is the whole difficulty in a form a student can understand. The measurement pair contains two kinds of information with very different strengths. The total is prominent. The difference between the quantities is squeezed into a small difference between almost equal readings. Recovering the contrast requires multiplying that small observed difference by fifty, which also multiplies error in the difference.

Normalised vectors make the geometry exact. Let u₊ = (1, 1)/√2 and u₋ = (1, −1)/√2. They are perpendicular unit vectors. The matrix maps u₊ to itself and maps u₋ to 0.02u₋. For this symmetric positive matrix, these are both eigen-directions and singular-vector directions. Its singular values are 1 and 0.02, so its two-norm condition number is fifty.

The singular value decomposition generalises this picture. A matrix can be written A = UΣVᵀ, where the columns of V describe orthogonal input directions, the columns of U describe corresponding output directions, and Σ records the gains applied to them. A small singular value identifies an input pattern that produces only a weak measurement signal. A zero singular value identifies an invisible pattern. [2]

In exact inversion along a nonzero singular direction, the observed coefficient is divided by the singular value. If that coefficient contains measurement error, the error is divided too. The problem is not that division is mathematically illegal. It is that small singular values make the division operationally expensive in accuracy. Weakly measured detail is where noise can become a large reconstructed feature.

Suppose the true hidden quantities are six and two. The ideal measurements are 4.04 and 3.96. Now deliberately perturb them to 4.06 and 3.94. The measured total is still eight, but the measured difference is 0.12 rather than 0.08. Exact inversion gives contrast six, and therefore hidden quantities seven and one. The measurements changed by only 0.02 in each component; the reconstructed components each changed by one.

The reconstruction fits the perturbed observations perfectly. That is not reassuring by itself. A zero residual means that the returned object reproduces the supplied data under A. Because the supplied data contain the assigned perturbation, perfect fitting has also reproduced the perturbation’s amplified effect. In this example, more exact optimisation would not improve the answer. We have already solved the unregularised problem exactly.

A common misunderstanding is to call small singular values “small objects.” They are not. They describe weak response to particular directions of change. A hidden pattern can have a large amplitude and still produce a small observational effect. In imaging, such directions may resemble fine alternating structure. In parameter estimation, they may represent compensating changes in several parameters. Their appearance depends on the operator.

The singular spectrum also depends on scaling and discretisation. Changing the units of parameters changes the numerical matrix. Adding much finer pixels changes the number and type of unknown patterns. A condition number should therefore be reported with its mathematical setting, not used as a context-free rating of an instrument. The physically important question remains which plausible hidden changes are distinguishable at the actual error level.

This perspective suggests two broad repairs. We can improve the measurement system so that relevant directions are less suppressed, or we can limit how aggressively the reconstruction tries to recover weak directions. The first is experimental design. The second is regularisation. They are complementary, but they do different jobs: a better measurement can add evidence, while a reconstruction preference changes how existing evidence is interpreted.

Before moving on, notice how much has been gained without a complicated algorithm. A change of coordinates separated a reliable total from a fragile contrast. That tells us what to report, where to test sensitivity, what kind of additional sensor would help and what a penalty should be designed to control. The purpose of linear algebra here is not symbolic display. It is to reveal where the evidence is strong and where the reconstruction is being asked to do more than the instrument comfortably supports.

9. Regularisation makes the extra assumption visible

Regularisation begins with an admission: fitting the observations as closely as possible may not recover the object as well as possible. When data are noisy, incomplete or weakly informative, an unrestricted fit can chase features that the measurements do not support reliably. We therefore introduce additional structure into the reconstruction problem and make that structure explicit.

One common formulation balances a discrepancy term against a penalty. The discrepancy rewards agreement with the observations. The penalty expresses a preference about the reconstructed object. For a linear model, a useful general form is to minimise ‖W(Ax − y)‖² + λ‖L(x − x₀)‖², subject to any declared admissibility constraints. W accounts for measurement scaling or uncertainty, L defines what departures are penalised, x₀ is a reference object, and λ controls the balance.

The notation contains several independent choices. Setting L to the identity penalises overall departure from x₀. Setting L to a difference operator penalises changes between neighbouring components. Choosing a reference of zero is not the same as choosing a reference supplied by a previous calibrated measurement. A penalty should be explained in terms of the object, not described vaguely as “making the result better.”

The parameter convention also matters. Some books write λ² in front of the penalty; others write λ. This article uses λ directly. A numerical value copied between the two conventions changes the strength of regularisation. A reproducible result therefore records the objective itself, not only a parameter name and value.

In the unconstrained quadratic case, differentiating gives (AᵀWᵀWA + λLᵀL)x = AᵀWᵀWy + λLᵀLx₀. This is a mathematical characterisation of a minimiser. For computation, we can instead stack the weighted observation equations above the penalty equations and solve an augmented least-squares problem. The augmented matrix is [WA; √λL], and the augmented right-hand side is [Wy; √λLx₀].

A useful uniqueness condition can be derived directly. A direction z receives zero curvature from the objective precisely when WAz = 0 and Lz = 0. For λ greater than zero, the minimiser is unique if the only such direction is z = 0. In compact notation, ker(WA) ∩ ker(L) = {0}. Thus adding a penalty does not automatically guarantee uniqueness if the measurements and penalty share an uncontrolled direction.

For example, a neighbouring-difference penalty does not penalise constant vectors. That is often desirable: a smoothness preference should not necessarily suppress an overall level. But the measurements must then determine that overall level, or another constraint must do so. Otherwise an entire family of constant shifts can remain unresolved despite the presence of a regularisation term.

Regularisation changes the question. Instead of asking only “Which object fits the measurements best?”, it asks “Which admissible object achieves the chosen balance between fit and the declared preference?” That can be exactly the useful question. It becomes misleading only when the preference disappears from the explanation and the selected object is presented as if it came from the data alone.

The simplest ridge form, with W and L equal to identity and x₀ = 0, makes the filtering effect visible through singular values. In singular direction i, the reconstructed coefficient is σᵢ/(σᵢ² + λ) times the observed coefficient. Compared with unregularised division by σᵢ, the extra factor σᵢ²/(σᵢ² + λ) damps weak directions. It approaches one for strong directions relative to λ and approaches zero when σᵢ² is very small.

This is a trade, not a free repair. Suppressing a weak direction reduces amplified noise, but it also suppresses genuine signal in that direction. A smoother or smaller reconstruction may be more stable and less accurate for some true objects. The appropriateness of the preference depends on what objects are plausible and what features the task needs to preserve.

Regularisation theory studies more than finite parameter tuning. It asks how a family of approximate inverses behaves as noise decreases and regularisation is adjusted. Under suitable assumptions, a carefully chosen parameter rule can balance improving data fit against increasing instability. Those guarantees concern stated classes of operators, objects and noise; they do not promise universal reconstruction of arbitrary detail. The variational and regularisation lecture routes provide background for these distinctions. [2] [4]

For practical reading, three questions now become indispensable. What extra information or preference was introduced? Which measurement directions does it affect? What evidence supports its strength for this task? A result can be valuable even when those answers include judgement. What it cannot legitimately do is hide the judgement inside a smooth image, a default setting or a confident numerical output.

10. A complete experiment with two hidden quantities

We can now run the paired-sensor problem from beginning to end. Everything in this experiment is deliberately constructed. The hidden vector is x = (6, 2), the forward matrix is A = [[0.51, 0.49], [0.49, 0.51]], and the ideal readings are y = (4.04, 3.96). We assign a perturbation e = (0.02, −0.02), producing observed readings y_obs = (4.06, 3.94). These are not measurements from a real device or evidence about an actual student.

The unregularised answer is (7, 1), as the previous chapter showed. Its predicted readings exactly match y_obs. Its Euclidean distance from the known hidden truth is √[(7 − 6)² + (1 − 2)²] = √2. Because this is a synthetic experiment, we can calculate that reconstruction error. In a real inverse problem, the hidden truth would usually be unavailable.

Now choose a penalty that acts on contrast but leaves the total alone. Let P = ½[[1, −1], [−1, 1]]. This matrix projects a vector onto the contrast direction. We minimise ‖Ax − y_obs‖² + λ‖Px‖². The stated preference is therefore balance between the two hidden components, not small total quantity. This matters because the total is already well measured.

Express the candidate object through s = x₁ + x₂ and d = x₁ − x₂. Write the measured sum as s_y = 8 and measured contrast as d_y = 0.12. Since A preserves the sum and scales contrast by α = 0.02, the squared residual is ½(s − 8)² + ½(αd − 0.12)². The penalty is λd²/2. We have reduced the problem to two separate one-variable quadratics.

The optimal sum is ŝ = 8 for every nonnegative λ. For contrast, differentiating ½(αd − d_y)² + λd²/2 gives α(αd − d_y) + λd = 0. Hence d̂ = αd_y/(α² + λ). The components follow as x̂₁ = (8 + d̂)/2 and x̂₂ = (8 − d̂)/2. This derivation makes every parameter’s effect visible.

The resulting values form a useful comparison table. The residual norm below measures disagreement with the observed readings, while reconstruction error measures distance from the deliberately known truth. They are different quantities and move differently as the penalty changes.

Penalty λEstimated contrastReconstructed pairResidual normReconstruction error
06.0(7.0, 1.0)01.414214
0.00014.8(6.4, 1.6)0.0169710.565685
0.00024.0(6.0, 2.0)0.0282840
0.00043.0(5.5, 2.5)0.0424260.707107
0.00161.2(4.6, 3.4)0.0678821.979899

The second column shrinks steadily. The residual increases steadily in this example. Reconstruction error first decreases and then increases. That is the central lesson: improved data fit and improved recovery are not the same objective when the observations contain error. The best fit here is not the best reconstruction.

The apparently perfect result at λ = 0.0002 requires special care. We selected a simple hidden truth and a particular assigned perturbation. For this one case, that penalty exactly offsets the noise-induced contrast error. It would be wrong to advertise 0.0002 as a universal optimum or to select it in a real experiment by secretly consulting the unknown truth. This is a retrospective demonstration of the trade, not a deployable parameter-selection rule.

Change the perturbation and the story changes. With the opposite perturbation, (−0.02, 0.02), the unregularised contrast is two instead of four. Any positive contrast-shrinkage penalty moves that already underestimated contrast even nearer zero. The same regulariser can therefore make this particular realisation worse. Its performance must be evaluated over a stated family of objects and noise conditions, not judged from one favourable picture.

There is also a distinction between the assigned perturbation and a stochastic error model. The pair (0.02, −0.02) has exactly zero sum because we chose it that way. It does not demonstrate that real sensor errors are independent, Gaussian or always opposite. A later simulation can investigate such assumptions, but its draws must be labelled as simulated. A fixed pedagogical perturbation should not quietly become an empirical noise distribution.

What should a reader conclude from the experiment? The measurements reliably determine the total in the exact assigned-noise construction. They determine contrast much less reliably. Regularisation can reduce a weak direction’s noise amplification, but does so by imposing a preference. A single successful reconstruction does not validate that preference generally. All these conclusions can be checked by direct substitution rather than accepted on authority.

This small experiment is deliberately reusable. Change the matrix, truth, noise or penalty and predict the consequences before running the code. Ask which features remain stable and which move. Compare residual and reconstruction error only when synthetic truth is available. That sequence turns an attractive demonstration into a laboratory for understanding, and it prepares the reader to recognise the same information problem inside a much larger image, sensor network or scientific model.

11. Choosing the penalty without looking at the answer

The synthetic experiment allowed us to see the true reconstruction error. Real work usually does not. That is why choosing regularisation is not merely drawing a graph and picking the point nearest the hidden truth. We need a rule that uses information legitimately available before the reconstruction is judged.

The first question is what the noise information means. A deterministic upper bound on ‖ε‖ is not the same as a standard deviation for each observation. If the error is bounded by δ in the same norm used for discrepancy, one strategy is to seek a fit whose residual is consistent with that bound rather than forcing it to zero. This is the idea behind discrepancy-based selection. It requires a credible error estimate and an adequate forward model.

For our assigned perturbation, the Euclidean norm is √(0.02² + 0.02²) ≈ 0.028284. The regularised result at λ = 0.0002 has exactly that residual norm. This is consistent with the deliberately designed example, but it should not be read backwards as proof that the discrepancy principle always reveals truth. Another signal, perturbation or model error can change the relationship between residual and recovery.

With independent Gaussian errors of variance σ², the squared norm of the raw noise has expected value mσ² for m observations. That expectation can help reason about a typical residual scale, but it is not a guaranteed bound for one realisation. Fitted residuals also have their own distribution because parameter estimation consumes directions in the data. A credible implementation must use the appropriate statistical construction, not casually equate every residual target with √mσ.

Cross-validation offers a different route. Fit using part of the available observations and evaluate predictions on withheld observations. In favourable settings, this discourages fitting noise that does not transfer. But the validation measurements must be designed and separated meaningfully. Holding out adjacent pixels whose errors are strongly dependent can give an overly optimistic picture of generalisation.

The distinction between predicting observations and recovering objects remains. A penalty can predict withheld measurements well while suppressing a feature the sensor scarcely observes. If the task is to detect that feature, prediction performance alone is insufficient. Validation must be aligned with the desired quantity, and the experimental design must provide information about it.

An L-curve plots a measure of fit against a measure of regularity, often on logarithmic axes, as the parameter varies. A pronounced bend can suggest a transition between chasing data and imposing excessive regularity. The curve is a diagnostic, not an oracle. It may have no clear corner, several ambiguous bends or a corner that does not optimise the actual task.

Generalised cross-validation is another parameter-selection approach, motivated by predictive assessment for suitable linear estimators. Its behaviour depends on the model, noise and estimator structure. The lesson for a non-specialist is not to memorise a preferred acronym. It is to ask what each selection rule is estimating and under which assumptions that estimate is relevant. The regularisation lecture offers entry points to these established approaches. [2]

A defensible workflow can compare several plausible parameter choices rather than pretending there is one sacred value. Inspect how the main conclusion changes. Does a broad feature remain visible throughout a reasonable range? Does a tiny feature appear only under one finely tuned setting? Does the residual retain systematic structure? Such checks distinguish a robust inference from a setting-sensitive picture.

Parameters also have scale. Multiplying all observations and predictions by a constant multiplies the squared discrepancy by that constant’s square. Unless the penalty is transformed consistently, the same numerical λ no longer expresses the same balance. Changing voxel size or parameter units creates related issues. A reproducible report states scaling, discretisation and objective conventions alongside the selected value.

There are times when no parameter can repair the situation. Suppose every reconstruction either fits badly or produces an implausible object. The problem may be a wrong operator, missing physics, inappropriate constraints or corrupted observations. Searching a finer parameter grid will not necessarily help. Regularisation strength is only one possible failure point in the larger inference chain.

The final answer should record the selection rule, evidence used, candidate range and sensitivity of the main conclusion. Where uncertainty remains, show it. “We chose the smoothest image that looked plausible” is a judgement and should be identified as such. “We chose the parameter using independently withheld measurements under this error model” is a more specific claim. Neither should be inflated into proof that every recovered detail is real.

12. Different priors protect different kinds of structure

There is no single natural-looking reconstruction independent of the object. A temperature field may reasonably be smooth over some scales. A manufactured object may have sharp boundaries. A spectrum may contain a few isolated peaks. A distribution of material density may be nonnegative. Each case invites a different preference, and each preference can preserve some structures while damaging others.

A penalty on ‖x‖² favours small amplitudes in the chosen coordinates. A penalty on neighbouring differences favours slowly varying components. A penalty on second differences favours changes that themselves vary slowly. These are different statements. Calling all of them “smoothing” can conceal what shapes the method actually prefers.

For a one-dimensional vector, the first-difference operator produces x₂ − x₁, x₃ − x₂ and so on. Squaring and summing these differences makes a large jump expensive. A broad gradual change may be cheaper than a sharp jump of the same total size. That can stabilise a noisy reconstruction, but it can also blur a genuine boundary. The trade follows from the penalty’s arithmetic.

Total variation instead adds the absolute magnitudes of local differences. In a simple one-dimensional setting, a single jump of size four and four same-direction jumps of size one have the same total variation. This helps explain why total-variation methods can preserve edges differently from squared-difference methods. It does not mean they preserve every edge correctly or introduce no artefacts. [4]

Sparsity is another kind of structure. An ℓ₁ penalty, the sum of absolute coefficient magnitudes, can encourage many coefficients to become exactly zero in suitable optimisation problems. But sparsity must be specified in a representation. A signal may be sparse in a frequency or wavelet basis while being dense in ordinary sample coordinates. “The object is sparse” is incomplete without saying where.

A tiny example makes the representation issue concrete. The vector (1, 1, 1, 1) has four nonzero entries in its ordinary coordinates. In a basis whose first vector is proportional to (1, 1, 1, 1), it needs only one nonzero coefficient. The physical object did not change. The description changed. A sparsity preference therefore carries an assumption about which patterns should be simple.

Nonnegativity is not merely an aesthetic preference when the unknown is, for example, a physical amount that cannot be negative. It can rule out algebraically compatible but inadmissible answers. Yet nonnegativity does not generally guarantee uniqueness. In the two-box sum problem, every nonnegative split of twelve remains compatible. In our tomography example, positivity will narrow a family without identifying one member.

A known total can also be enforced separately from smoothness. This is what our contrast penalty effectively respects: the measured total is left to the observations, while balance is penalised. An indiscriminate ridge penalty would shrink both total and contrast. Neither formulation is universally superior. The appropriate choice depends on which parts of the object are already well measured and what additional knowledge is justified.

Reference-based regularisation deserves particular attention. Penalising departure from a previous image or earlier parameter estimate can be useful when change is expected to be limited. It can also make a new, genuinely unusual event harder to see. The old reference is information, not a guarantee that the present must resemble the past. Testing should include plausible changes that violate the usual pattern.

A prior can be wrong in a way that still produces a convincing answer. Suppose a smoothing rule removes a narrow peak. The resulting curve may look cleaner and fit noisy measurements adequately, especially if the sensor is weakly sensitive to narrow peaks. Visual plausibility can therefore reward the very loss that matters to the task. Validation should ask whether important features survive, not only whether the result appears orderly.

The responsible response is not to avoid priors. An underdetermined problem often needs additional structure before any useful answer is possible. The response is to expose the structure, test its consequences and compare alternatives where the decision depends on them. A useful reconstruction can say, “This feature is stable under several plausible priors,” or, “This detail is strongly prior-dependent.” Those statements are more informative than an unqualified sharp picture.

For teaching, ask learners to reconstruct the same data under two different explicit rules. Let one rule favour equality and another favour a known reference. Then compare the outputs and forward predictions. The exercise makes visible a principle that is easy to lose inside advanced software: two reconstructions can agree with the measurements equally well while disagreeing because their extra assumptions differ. Understanding the disagreement is part of the mathematics, not a reason to hide it.

13. Why an algorithm can improve before it starts to overfit

For a small matrix, a decomposition can solve the reconstruction problem directly. For a large image or a model involving a differential equation, repeatedly improving a candidate may be more practical. The basic loop is simple: predict the observations from the current candidate, compare them with the measurements, and use the discrepancy to update the candidate.

For the objective ½‖Ax − y‖², the gradient is Aᵀ(Ax − y). A gradient step therefore has the form xₖ₊₁ = xₖ + ηAᵀ(y − Axₖ). In this linear least-squares setting the iteration is commonly called Landweber iteration. The step size η matters: for a nonzero finite matrix A, a sufficient standard range for convergence is 0 < η < 2/‖A‖₂². The formula is not a licence to use an arbitrarily large update. [6]

The transpose Aᵀ has a particular role. It sends discrepancies in observation space back into directions in parameter space. It is not generally the inverse of A. Calling it “the inverse” obscures both the algorithm and the geometry. Backprojection can be a useful step toward reconstruction while remaining very different from complete inversion.

The singular directions reveal how the iteration behaves. Start from zero and consider one direction with singular value σ. After k steps, its reconstructed coefficient is [1 − (1 − ησ²)ᵏ] times the unregularised inverse coefficient, when σ is nonzero. This follows by solving the scalar recurrence. The factor approaches one as iteration proceeds, but it does so at different speeds for different singular values.

Strongly measured directions, with larger σ, are recovered quickly for a suitable step size. Weakly measured directions may require many iterations. Early in the run, the algorithm can therefore recover large-scale reliable information while leaving weak, noise-sensitive detail suppressed. Continuing indefinitely eventually fits more of that weak information, including its noise.

In our paired-sensor example, choose η = 1. The sum direction has σ = 1, so it is recovered in the first step. The contrast direction has σ = 0.02, so its filter factor is 1 − 0.9996ᵏ. After ten iterations, that factor is only about 0.004. After a thousand, it is about 0.33. The total becomes reliable long before the contrast approaches the unregularised value.

This is the mechanism behind early stopping as a form of regularisation. The stopping point controls which measurement directions have been substantially inverted. It is not merely a computational convenience. The iteration count becomes a parameter with the same kind of evidence obligation as an explicit penalty strength.

For our deliberately perturbed data, the limiting contrast is six while the hidden truth has contrast four. At an intermediate iteration, the contrast passes near four and the reconstruction error can be small. But using the known truth to select that iteration would again be an oracle choice confined to a synthetic experiment. In practical work, a stopping rule should use available noise information, validation data or another justified criterion.

There are two different notions of progress. The discrepancy with the observed data can keep improving while distance from the hidden truth first improves and then worsens. A solver log showing a steadily decreasing objective proves only progress on that objective. It does not show that every later image is more truthful. This behaviour is often called semiconvergence in iterative treatment of ill-posed problems.

Adding an explicit regularisation term changes the gradient and the limiting problem. For ½‖Ax − y‖² + λ‖L(x − x₀)‖²/2, the gradient includes λLᵀL(x − x₀). With suitable assumptions and algorithm settings, the iteration can converge to a stable minimiser of that regularised objective. We should still distinguish convergence to the selected mathematical solution from empirical recovery of the actual object.

Nonlinear forward models introduce additional difficulties. The sensitivity of F can change with x, different starting points can approach different local minima, and a local quadratic approximation may be poor far from the current candidate. Line searches, trust regions and other controls manage numerical steps, but they do not remove nonidentifiability or guarantee that the selected minimum corresponds to the real mechanism.

The computational record should therefore include the objective, initialisation, step rules, stopping criterion and achieved residual, alongside checks for feasibility and numerical sensitivity. A statement such as “the algorithm ran for one hundred iterations” is not meaningful without knowing why one hundred was appropriate. Reliable computation is not the longest run or the smallest displayed loss. It is a controlled route whose numerical result and inferential meaning have both been checked.

14. Bayesian reconstruction gives a distribution, conditional on a model

A point estimate returns one candidate. A Bayesian formulation instead describes uncertainty about the hidden object using a prior distribution, combines it with a likelihood for the observations, and obtains a posterior distribution. The posterior can represent several plausible answers and how their relative support changes after data arrive. It is conditional on the entire probabilistic model, including the prior and observation assumptions. [3]

The basic identity is p(x | y) proportional to p(y | x)p(x). Here p(x) is the prior, p(y | x) is the likelihood, and p(x | y) is the posterior. The likelihood describes how observations would be generated from a candidate object. It is not, by itself, a probability distribution over objects after the data are observed.

A fully worked scalar example keeps the assumptions visible. Suppose x has a Gaussian prior with mean zero and variance four. The observation model is y = 0.2x + ε, with independent Gaussian error ε of mean zero and variance 0.01. We observe y = 0.6. The small gain 0.2 and the stated noise scale determine how informative that observation is relative to the prior.

The likelihood contributes a quadratic term (0.6 − 0.2x)²/(2 × 0.01) to the negative log density. The prior contributes x²/(2 × 4). Adding them gives a quadratic in x. The coefficient of x² determines posterior precision, while the linear coefficient determines the posterior mean. We can obtain both by completing the square rather than invoking a black box.

The posterior precision is 0.2²/0.01 + 1/4 = 4.25. Thus its variance is 1/4.25 = 4/17, approximately 0.235294. The posterior mean is (0.2 × 0.6/0.01)/4.25 = 48/17, approximately 2.823529. Its standard deviation is approximately 0.485071. A central interval using 1.96 standard deviations is approximately [1.8728, 3.7743].

Without the prior, exact inversion of the central observation gives x = 3. The posterior mean is slightly nearer zero because the prior favours values near zero. This is not a defect hidden in the method. It is the intended effect of combining the two information sources. The correct explanation tells the reader how much of the answer came from each.

The maximum a posteriori estimate, or MAP estimate, maximises posterior density. In this Gaussian example it equals the posterior mean. Multiplying the negative log posterior by a positive constant shows that the MAP estimate minimises (0.6 − 0.2x)² + 0.0025x². The regularisation coefficient is the noise variance divided by the prior variance: 0.01/4. This establishes a specific connection between a quadratic prior and a quadratic penalty.

That connection is not a statement that every regulariser automatically provides a calibrated posterior. A penalty can be used purely as an optimisation preference. To interpret it probabilistically, one must specify a likelihood, prior, measure and scaling consistently. A method that returns only a MAP image has not automatically quantified uncertainty merely because someone can write down a related probability model.

The distinction becomes important in weak directions. In a multivariate linear Gaussian model with positive-definite prior covariance C₀ and noise covariance Σ, posterior precision is AᵀΣ⁻¹A + C₀⁻¹. The likelihood adds information along directions the measurements see. Along an exact null direction of A, it adds no direct quadratic information. A narrow posterior there can reflect a strong prior or prior dependence, not direct observational determination.

Prior dependence deserves care. If components are correlated under the prior, data about one component can change beliefs about another even when the second is not directly measured. This is coherent within the joint model. It should not be described as the instrument directly observing the second component. The inferential route passes through the assumed relationship.

A credible interval such as the one above has a posterior-probability interpretation under the stipulated model. Its validity does not survive arbitrary changes to the prior, likelihood or data-generating process. Frequentist confidence procedures answer a different repeated-sampling question. A numerical interval without its construction is therefore incomplete information, regardless of how familiar the percentage looks.

Bayesian reconstruction is valuable because it can make uncertainty and prior dependence explicit. It also creates new checks: whether the likelihood represents the instrument, whether the prior permits important unusual objects, whether computation explores the posterior adequately, and whether predictive behaviour agrees with independent evidence. A distribution is a richer answer than a single image only when the reader can understand what that distribution is conditional on.

15. Two plausible explanations cannot always be averaged into one

Consider a different observation model: y = x² + ε. Suppose the observed value is close to four and the noise is small. Positive two and negative two both predict nearly the same observation. Unless the domain or other information resolves the sign, the posterior may have two separated regions of high probability. The ambiguity is structural, not a numerical accident.

If the prior is symmetric around zero and the noise model treats the two signs equally, the posterior is also symmetric. Its mean can be zero. Yet x = 0 predicts y = 0, which is a poor explanation of an observation near four. Averaging two plausible objects has created a point that may be implausible as an explanation. This original example shows why a summary statistic should be checked against the forward model.

A single Gaussian approximation centred at one mode can miss the other mode entirely. An optimisation method started near positive two may return that solution repeatedly, creating the appearance of uniqueness. A sampling method that never crosses the low-probability region between the modes may report misleadingly narrow uncertainty. Algorithmic consistency does not settle whether the uncertainty representation is complete.

The right response depends on the question. If only x² matters for a decision, the sign ambiguity may be irrelevant. If the sign matters, the analysis needs another informative measurement or an explicit prior-based statement. If physics genuinely restricts x to nonnegative values, that domain restriction resolves the ambiguity legitimately. The condition belongs in the problem contract, not in a hidden postprocessing step.

Uncertainty also has dependence between components. Return to two boxes with a nearly exact total of twelve. If the first box contains more, the second must contain less. Their separate ranges can both be broad, but their sum is tightly constrained. Reporting only independent error bars would discard this relationship and could suggest combinations that are incompatible with the data.

A joint uncertainty set or distribution preserves such dependence. In the exact-total case, compatible pairs lie along a line rather than filling a square. With small measurement error in the total, they occupy a narrow band. A marginal interval tells us about one coordinate; it does not describe which combinations of coordinates can occur together.

This matters when calculating a new quantity from the reconstructed object. The uncertainty of x₁ + x₂ depends on covariance as well as individual variances. Specifically, Var(x₁ + x₂) = Var(x₁) + Var(x₂) + 2Cov(x₁, x₂). Strong negative covariance can make the sum precise even when both components are uncertain. Ignoring dependence can badly misrepresent the decision-relevant uncertainty.

The opposite can happen with shared errors. If two measurements contain the same unknown calibration offset, combining them may preserve that uncertainty rather than average it away. A posterior or confidence calculation that assumes independence can become too narrow. The issue is not the mathematical sophistication of the final formula; it is whether its joint model matches the way information was acquired.

Uncertainty conditional on a fixed operator is also not the same as uncertainty about the operator. Suppose the paired sensors’ contrast gain is α, but α is only approximately known. The contrast estimate divides the measured difference by α. Uncertainty in that denominator can be substantial when α is small. A reconstruction interval computed while fixing α at its nominal value omits a relevant source.

Sometimes a model family is more informative than one posterior. Fit several defensible forward models or error structures and compare the conclusions they support. Agreement across models can strengthen a particular feature’s credibility, while disagreement reveals sensitivity. It does not establish that the models are collectively exhaustive or that their spread is a complete probability distribution over truth.

A report should therefore distinguish at least three things: uncertainty within the selected model, sensitivity to alternative assumptions, and unresolved limitations outside those models. These are not interchangeable error bars. Their separation prevents a sophisticated uncertainty calculation from becoming a false assurance that every important source of doubt has been measured.

The most useful answer may be a set of competing explanations accompanied by a proposal for an observation that distinguishes them. In the square example, a measurement responsive to the sign would help. In the two-box example, a contrast measurement would help. The purpose of uncertainty analysis is not to decorate an estimate with a percentage. It is to show where the evidence leaves room for different worlds and how a better question could narrow that room.

16. Test the measurement model, not only the reconstruction code

Our paired-sensor laboratory knows the matrix exactly because we invented it. A real apparatus must earn that knowledge. Its gains, offsets, geometry and response can change. A reconstruction may solve the supplied equations accurately while those equations describe the device only approximately. This is model error, and it can resemble noise or be absorbed into misleading parameter estimates.

Suppose the true contrast gain is 0.03, but the reconstruction uses 0.02. With true contrast four, the ideal observed difference is 0.12. A noiseless reconstruction using the wrong gain returns contrast six. The same incorrect answer appears as in our earlier noisy example, but the cause is different. More repeated observations do not repair a fixed gain error.

This gives an important diagnostic contrast. Random independent measurement fluctuations may shrink through repetition. A shared calibration defect may persist. An incorrect forward relationship may produce systematic residual patterns. A nullspace may remain even with perfect calibration. These problems can coexist, so “collect more data” should be followed by the question: more data of what kind?

Verification asks whether the software implements the intended mathematics. We can check matrix products by hand, compare direct formulas with a least-squares solver, test gradients with finite differences and check known limiting cases. Such tests are essential. They should not be described as validation of the physical model, because all of them can succeed while the model itself is inappropriate.

Validation asks whether the model and inference procedure perform adequately for a specified use when confronted with relevant independent evidence. That might involve a calibrated test object, observations from another instrument or a new acquisition configuration. The test should challenge the parts of the system that matter to the intended conclusion.

Synthetic tests deserve explicit labelling. If data are generated by exactly the same discretised operator used in reconstruction, the test removes an important source of real-world difficulty. It can be excellent for checking algebra and implementation, but it can overstate practical performance when presented as evidence about an actual apparatus. This issue is commonly discussed under the name “inverse crime”; the useful lesson is to distinguish software verification from realistic validation. [12]

Our small experiment deliberately uses the same operator in generation and inversion. That is acceptable because its stated job is to explain noise amplification and regularisation. It is not a benchmark claiming physical robustness. A stronger simulated validation could generate data on a finer grid, alter calibration within defensible ranges, change the noise model and evaluate objects different from those used in tuning.

Independence also applies to parameter selection. A test set used repeatedly to choose the best regularisation strength becomes part of development. Reporting its performance afterward as untouched validation exaggerates independence. A transparent workflow records which data informed the model, which informed tuning, and which were reserved for final evaluation.

Model discrepancy should not automatically be hidden by inflating noise. If a straight-line model systematically misses curvature, a larger variance can make the observations appear less surprising without explaining the missing mechanism. Sometimes a broader error model is justified; sometimes the forward model needs revision. Residual structure and independent physical knowledge help distinguish these possibilities.

There can also be nuisance parameters such as sensor offsets or unknown timing. Estimating them jointly with the object may be appropriate, but it enlarges the inverse problem. New ambiguities can appear. For example, an unknown multiplicative gain and an unknown overall object amplitude can compensate for one another. A joint optimiser cannot resolve that symmetry without additional information or a stated constraint.

A practical validation plan names perturbations before admiring the result. Change noise level, acquisition geometry, calibration, discretisation and object class separately where possible. Record which changes break the inference and whether they are plausible in use. A method does not need to survive every imaginable world to be useful. It does need an honest operating domain.

Finally, preserve a reproducible record of observations, preprocessing, operator version, parameters and output. When a result changes, the record should reveal whether the cause was new evidence, a repaired bug or a different prior. Without that record, a visually improved reconstruction can conceal a material change in assumptions. Trust grows when changes can be traced, not when a workflow merely produces the same confident picture twice.

17. A blurred image is a record of mixing

A blurred photograph invites a tempting story: the sharp image is still inside the blurry one, and the correct software merely needs to uncover it. Sometimes substantial detail can be recovered. Sometimes several sharp images are compatible with the same recorded data. To understand the difference, we must describe what the imaging process did to the light before asking an algorithm to reverse it.

A simple one-dimensional blur replaces each sample by a weighted average of its neighbours. Suppose the weights are one quarter, one half and one quarter. Then the observed value at position j is yⱼ = xⱼ₋₁/4 + xⱼ/2 + xⱼ₊₁/4, before noise is added. Each measurement mixes three nearby unknown values. The inverse problem is to infer the original sequence from these mixtures.

Consider a constant sequence of value five. The blurred sequence also has value five, because the weights sum to one. Now add an alternating pattern of amplitude one: six, four, six, four, and so on. With periodic boundaries on an even-length sequence, every blurred value is again five. The alternating component cancels exactly. Both original sequences are nonnegative and produce the same observations.

This is an explicit nullspace example inside a blur model. More precise recording of the blurred value five does not identify whether the alternating detail was present. An algorithm that produces the constant sequence is making a selection, possibly a reasonable one, but the observations alone did not exclude the alternating alternative. A sharp-looking reconstruction cannot overturn that algebra.

The frequency response of this particular blur can be derived by applying it to a complex sinusoid exp(iωj). The resulting gain is exp(−iω)/4 + 1/2 + exp(iω)/4 = (1 + cosω)/2 = cos²(ω/2). Low-frequency patterns pass with gain near one. The highest alternating frequency, ω = π, has gain zero. Frequencies near that point are strongly attenuated.

Other blur kernels have different responses. Some suppress certain frequencies without making them exactly zero; some create additional zeros. The corresponding inverse may therefore suffer weak recovery or exact ambiguity depending on the operator and domain. “Deblurring is difficult” is less informative than identifying which patterns the particular blur suppresses at the actual noise level.

Boundary conditions are part of the model too. Periodic wrapping, zero padding and reflective boundaries produce different edge behaviour. Our cancellation example explicitly used periodic boundaries to keep the calculation exact and simple. A real image does not necessarily wrap around. Reusing that assumption in an application without justification can create edge artefacts or misrepresent the data.

Blur is not the only loss. Sampling records values at discrete locations. Suppose samples are taken at tⱼ = j/10 for integer j. A sinusoid with frequency two and another with frequency twelve satisfy sin(2π × 12 × j/10) = sin(2π × 2 × j/10 + 2πj), so their sampled values are identical. The extra cycles occur between samples and disappear from this measurement record.

This is a direct example of aliasing. It does not require noisy observations. An additional assumption restricting the allowable frequencies can remove that particular ambiguity, and appropriate sampling theory explains when reconstruction is possible for a defined signal class. But the assumption must be stated. The same ten-samples-per-unit record does not determine arbitrary continuous signals.

Pixel count can consequently be misleading. Reconstructing a million-pixel image from a much smaller information content creates more unknowns, not automatically more observed detail. A finer grid can improve the numerical representation of a model, but it can also introduce many weakly constrained directions. Display resolution and evidence-supported resolution are different quantities.

A task-specific notion of resolution is more useful. Can the acquisition distinguish one broad object from two nearby objects? Can it identify the sign of a contrast? Can it detect a small feature of a specified size and amplitude under plausible noise? The answer depends on operator, noise, prior and task. It should not be reduced to the number of pixels in the exported file.

Regularisation determines what happens to poorly observed frequencies. Quadratic smoothing suppresses them. Sparsity may prefer a small number of sharp components. Learned reconstruction may favour patterns represented in its training distribution. These choices can be effective, but each brings information beyond the particular observation. A trustworthy account separates observed constraints from assumed image structure.

A useful validation experiment deliberately includes features that are unusual but admissible. Reconstruct a sharp edge, a small isolated peak, a gradual background and two closely spaced objects. Vary the noise and blur calibration. Look for features that disappear, move or split. A method that produces attractive average images can still fail on the rare feature that matters most to the user.

For students, the alternating-sequence example is especially valuable. They can calculate both forward images by hand and establish exact observational equality. No amount of admiration for a reconstruction can erase the compatible alternative. The example does not prove that image recovery is hopeless. It identifies the missing information and shows why additional assumptions, different measurements or a narrower question are needed.

The broader lesson is practical: before asking software to sharpen an image, ask what the acquisition measured and what it mixed away. The answer tells us which improvements could be evidence-based, which rely strongly on prior structure, and which cannot be settled by the present data. That distinction is the difference between recovering a feature and merely rendering a plausible one.

18. Tomography: recovering an interior from measurements across it

Computed tomography offers a familiar application of inverse thinking. An X-ray source and detectors collect measurements from different directions, and reconstruction methods use those observations to estimate an internal image. The measurements are not ordinary photographs of every internal point. They encode how radiation has interacted with material along paths through the object. This description is about mathematical imaging, not clinical interpretation or instructions for operating equipment. [8]

An idealised monoenergetic, straight-ray attenuation model writes the received intensity as I = I₀ exp(−∫μ ds), where μ is an attenuation coefficient along the ray. Taking the negative logarithm of I/I₀ gives the line integral of μ. That transformation connects a physical observation to a mathematical projection model. Real acquisition requires additional treatment of effects such as spectrum, scatter, geometry and noise; the ideal formula is a starting model rather than a complete scanner specification. [5]

A four-cell puzzle shows the informational issue without pretending to reproduce a clinical scanner. Arrange four nonnegative unknowns as a square: a and b in the top row, c and d in the bottom row. Our fictional device records the two row sums and two column sums. The observations are a + b = 10, c + d = 6, a + c = 9 and b + d = 7.

There are four unknowns and four written equations. It is tempting to announce that the problem must therefore be determined. But the equations are not independent. Adding the two row equations gives a + b + c + d = 16. Adding the two column equations gives the same total. One relation duplicates information already available in the others.

Solve the system symbolically. Set a = t. Then b = 10 − t, c = 9 − t and d = t − 3. The final column equation is automatically satisfied. Nonnegativity requires t at least three and at most nine. Every t in the interval [3, 9] therefore gives a nonnegative object with exactly the observed row and column sums.

For t = 6, the object is [[6, 4], [3, 3]]. For t = 4, it is [[4, 6], [5, 1]]. Both give the same four observations. Their interiors differ substantially, yet our instrument cannot distinguish them. This is not a failure of arithmetic; it is a rank-three measurement system for four unknown quantities.

The invisible change is again easy to write: add δ to a and d while subtracting δ from b and c. In vector order (a, b, c, d), that is δ(1, −1, −1, 1). Every row and column contains one addition and one subtraction, so all four sums stay fixed. The measurement operator has a null direction with a clear spatial pattern.

Now add a fifth fictional measurement, a + d = 9. Substituting the family gives 2t − 3 = 9, so t = 6. We recover [[6, 4], [3, 3]]. The new measurement works because it responds to the previously invisible pattern: the change in a + d is 2δ. It supplies information in exactly the direction the earlier observations missed.

This diagonal-sum measurement is an educational abstraction. Real tomography involves path lengths, geometry, attenuation and many measurements, not simply selecting arbitrary cells and adding them. The puzzle’s purpose is narrower: it proves that number of measurements alone does not establish independence, and that an informative new view must distinguish objects already compatible with the existing views.

Noise changes exact equalities into approximate consistency. If the row sums add to sixteen but the column sums add to 16.1, no four-cell object fits all readings exactly. We need a discrepancy model, possibly weighted by measurement uncertainty. The residual can then reveal the inconsistency, but a fitted object should not be described as satisfying all observations exactly.

Regularisation can select one object before the fifth measurement is added. Minimum norm, smoothness, nonnegativity or a prior reference each imposes additional structure. Yet positivity alone already allowed the whole interval from three to nine. A reconstruction that returns one interior should state which preference resolved the remaining ambiguity.

The distinction between limited views and sparse views also matters conceptually. Having fewer measurement directions can leave certain structures poorly constrained. Having a restricted angular range can make some directions of structure especially difficult to recover. The precise consequences depend on the acquisition geometry and object class. Increasing display resolution cannot substitute for measurements that were never collected.

Validation of a tomographic method should therefore examine identifiable features, acquisition conditions and uncertainty, not only average visual similarity. A reconstruction may preserve total attenuation well while misplacing a small component. A task focused on location needs a different evaluation from one focused on total amount. Mathematical performance measures should follow the actual question.

For a learner, the four-cell model creates a bridge from simultaneous equations to scientific reconstruction. They can find the compatible family, prove its limits, propose an additional measurement and check it by substitution. The same habits scale: inspect independence, identify invisible directions, separate prior selection from observation, and ask whether a new measurement changes the set of possible interiors.

An internal image is therefore an inference with a measurement history. Its usefulness can be extraordinary, but the inference remains conditioned by the physical model, acquisition geometry, reconstruction method and validation evidence. Understanding those conditions does not diminish tomography. It explains why so much mathematics is needed to turn indirect observations into a dependable view of what lies inside.

19. Why running heat backwards amplifies tiny errors

Imagine a temperature pattern along a circular ring. At an initial moment, one part is warmer, another cooler, and small ripples are superimposed on the broad pattern. Heat flow gradually smooths the differences. If we know the initial state and physical model, predicting the later state is a forward problem. Recovering the earlier state from a later measurement is an inverse problem.

To see the mathematics clearly, use a periodic coordinate ξ from zero to 2π and the normalised heat equation ∂u/∂t = κ∂²u/∂ξ². A sinusoidal spatial mode sin(kξ) evolves by multiplication with exp(−κk²t). Differentiating the proposed solution verifies this directly: the second spatial derivative introduces −k², matching the time derivative’s decay factor. The analysis of inverse heat problems provides a classical example of instability. [6]

The factor contains k². Fine spatial ripples, corresponding to large k, decay much more strongly than broad ripples. The forward process does not treat every detail equally. It preserves the constant mode, weakens broad differences and rapidly suppresses fine variation. Later measurements may therefore contain reliable information about the average while retaining very little about fine initial structure.

Choose κt = 0.02 for a numerical illustration. The mode with k = 10 is multiplied by exp(−2), about 0.1353. The mode with k = 30 is multiplied by exp(−18), about 1.523 × 10⁻⁸. The mode with k = 50 is multiplied by exp(−50), about 1.929 × 10⁻²². An initial pattern of amplitude one can leave a later signature far below a practical observation’s accuracy.

Exact reversal divides by those decay factors. The corresponding amplification factors are exp(2), about 7.389; exp(18), about 65.66 million; and exp(50), about 5.185 × 10²¹. A tiny error in a fine-frequency coefficient can therefore become an enormous error in the inferred earlier temperature field.

This is more than a large condition number for one small matrix. Consider the sequence of initial functions vₖ(ξ) = sin(kξ). Their L² norms on [0, 2π] all equal √π. At the later time, their images are exp(−0.02k²)sin(kξ), whose norms tend to zero as k increases. Thus arbitrarily small changes in the later data correspond to initial changes with a fixed nonzero size.

That sequence demonstrates failure of continuity of the inverse in these L² norms. It makes precise what “unstable” means in this infinite-dimensional setting. The issue is not that an exact solution formula is missing. We have one. The issue is that the formula is not continuously dependable on imperfect data in the chosen spaces.

A finite computational grid truncates the possible frequencies. On a fixed finite set of nonzero modes, inversion is a finite-dimensional map and is continuous, but its worst amplification can be enormous. Refining the grid introduces higher frequencies and can make the unregularised inverse increasingly sensitive. A more detailed discretisation does not automatically create a more trustworthy recovered past.

Regularisation can limit the recovered frequency range or damp the most unstable modes. A truncated reconstruction may recover broad initial structure while declining to infer fine ripples below the evidence level. This is not simply throwing away inconvenient information. The discarded directions are precisely those in which the observations may be dominated by error after reversal.

We can derive a rough resolution rule for this teaching model. Suppose an initial mode of amplitude M produces a later amplitude M exp(−0.02k²), and observations cannot reliably distinguish amplitudes below δ. Requiring the later signature to exceed δ gives k² less than approximately log(M/δ)/0.02. This is a conditional signal-versus-error argument, not a universal imaging resolution theorem. It shows how allowed initial amplitude and measurement precision jointly limit recoverable frequency.

For M = 1 and δ = 0.001, the right-hand side is about 345.4, giving k below about 18.6. Improving δ by a factor of one thousand changes the logarithm from about 6.9 to about 13.8, so the corresponding frequency threshold grows only by a factor of √2. Severe smoothing can make large improvements in measurement accuracy buy surprisingly modest increases in recoverable detail.

Prior restrictions change what can be established. If physics or independent knowledge excludes rapid initial oscillations, the relevant object class becomes smaller. Stability may improve within that class. But the restriction is not a consequence of the later temperature measurement alone. It should be justified and tested against the intended application.

The quantity of interest matters again. The average temperature is associated with k = 0 and is unchanged in this isolated periodic model. Recovering that average is far easier than recovering every fine initial fluctuation. A report that says only “the past temperature is uncertain” misses this useful distinction between stable and unstable features.

The model itself also has limits. Real thermal systems can include uncertain boundary conditions, sources, spatially varying material properties and measurement effects. Our ring example isolates one mathematical mechanism. It does not justify reconstructing an actual historical temperature field without a validated physical model and appropriate uncertainty analysis.

This chapter connects the elementary two-sensor example to a genuine functional problem. In both, the forward process weakens some patterns, and inversion magnifies the associated errors. The heat equation makes that weakening increasingly severe at finer scales. Once the mechanism is visible, regularisation is no longer an arbitrary act of smoothing. It becomes a controlled decision about how much of the lost detail the available evidence can responsibly support.

20. Better measurements can be more valuable than a better optimiser

Suppose the paired-sensor reconstruction remains too uncertain for the decision. One response is to try more algorithms. Another is to change the sensors. The second can be more powerful because it changes the information in the observations rather than only changing how existing information is processed.

Replace A = [[0.51, 0.49], [0.49, 0.51]] with B = [[0.8, 0.2], [0.2, 0.8]]. The new pair still preserves the total, but its contrast gain is 0.6 rather than 0.02. The instruments now respond much more differently to the two hidden quantities. A change in their split produces a larger measurable difference.

For the same hidden truth (6, 2), the ideal new readings are (5.2, 2.8). Add the same assigned perturbation (0.02, −0.02), giving (5.22, 2.78). The measured contrast is 2.44. Dividing by 0.6 gives an inferred hidden contrast of 4.0666667. The reconstructed quantities are approximately (6.0333333, 1.9666667), much nearer the truth than (7, 1).

The contrast error has fallen from two to about 0.0666667, a factor of thirty. That improvement follows directly from the change in gain: 0.6/0.02 = 30. It does not depend on a cleverer inversion algorithm. We used the same elementary sum-and-difference calculation, but supplied it with more discriminating observations.

Under an additional statistical model with independent equal-variance Gaussian errors in the two readings, the measured contrast has variance 2σ². Dividing it by α gives an unregularised contrast variance 2σ²/α². Increasing α from 0.02 to 0.6 reduces that variance by a factor of 900. This is a model-based precision calculation; the earlier single assigned perturbation did not establish the Gaussian assumptions.

The same expression connects to Fisher information. In this simple contrast model, information about d is proportional to α²/(2σ²). Stronger response or lower noise increases information. If the measurement gives no response to d at all, repeated measurement of the same quantity cannot create direct information about it. This connects experimental design to the geometry of the inverse map.

Repetition can still help when the direction is weak rather than exactly invisible. Under independent errors with known stable variance, averaging n readings reduces variance by n. Matching a thirtyfold improvement in standard deviation would require nine hundred independent repetitions in this simplified comparison. That arithmetic invites a practical question: is it cheaper and more reliable to repeat the old acquisition or redesign it?

The comparison changes when errors share a calibration defect. If every repetition contains the same uncertain offset or gain, averaging does not remove that component. Likewise, repeated scans can be costly, disruptive or physically constrained. The mathematical design objective should include acquisition cost, admissibility and the actual uncertainty mechanism rather than optimising an abstract precision measure in isolation.

A new measurement should be chosen to separate currently plausible explanations. In the tomography puzzle, the extra diagonal sum worked because it responded to the null vector. In the square observation y = x², another squared measurement does not reveal the sign; a sign-sensitive observation could. The general design question is: where do the competing candidate objects make different observable predictions?

Sometimes the wanted quantity is simpler than the full object. If the decision depends only on total amount, there may be little value in adding expensive contrast measurements. Conversely, a task that depends on a tiny difference may require acquisition designed specifically for that difference. Experimental design should begin with the decision-relevant quantity, not with a general desire for “more data.”

Local design criteria have limits in nonlinear problems. A derivative-based sensitivity matrix describes response near one parameter setting. If plausible objects lie far apart or produce different regimes, a locally optimal sensor configuration may fail elsewhere. Comparing several plausible states or designing for robust discrimination across them can be more appropriate than optimising only at one nominal point.

Unknown nuisance parameters can also undermine a seemingly ideal design. A sensor may be highly sensitive to the target but equally sensitive to an uncertain offset. The two effects can remain confounded. A useful design includes calibration or reference measurements that separate them. Sensitivity to the target alone is not sufficient; distinguishability from other possible changes matters.

A practical planning sheet can therefore name the target, the current compatible alternatives, the observations on which they disagree, the expected measurement uncertainty and the acquisition cost. The proposed measurement should then be evaluated by how much it could change the actual conclusion. This is an original planning aid, not a claim that one numerical criterion governs every scientific experiment.

For learners, designing the next observation is often a stronger test of understanding than repeating a solved inverse calculation. It forces them to recognise what remains unknown and which information would resolve it. A student who can explain why another total-only reading fails, but a contrast reading succeeds, has moved beyond algebraic execution into mathematical design.

The final principle is constructive. Inverse problems do not end with a warning that the data are insufficient. They can tell us exactly how the data are insufficient. That diagnosis can guide calibration, sensor placement, acquisition angles, timing or a narrower question. Sometimes the most effective reconstruction algorithm begins before the data are collected, with a measurement designed to make the important hidden difference visible.

21. Recovering parameters is not automatically discovering causation

The phrase “hidden causes” is useful when describing inverse problems, but it can also tempt us to overclaim. Inferring an unobserved parameter in a specified physical model is not automatically the same task as discovering which variable causes another. Both involve moving from observations toward explanations, yet they have different assumptions and different proof obligations.

Consider a deliberately simple logical construction. Let U be a random quantity taking values zero and one with equal probability. In Model A, set X = U and Y = X. In Model B, set Y = U and X = Y. Both models produce exactly the same observed pairs: (0, 0) and (1, 1), each half the time. No amount of observational precision in these pairs distinguishes the direction of the relationship.

Now define an intervention that sets X to one while leaving the other structural assignment unchanged. In Model A, Y = X, so Y becomes one. In Model B, Y = U, so setting X does not change Y’s distribution. The two models agree on the observed data but disagree on the intervention. This is a self-contained counterexample to the claim that observational fit alone determines causal direction.

The example does not say that causal inference is impossible. It says that the observations and assumptions must contain information adequate for the causal question. Randomised interventions, justified structural restrictions, temporal information and other design features can matter. But an optimiser that merely fits the same observational distribution has not resolved the difference by calculation alone.

Contrast that with a calibrated physical inverse problem. If a measurement operator is independently established and the relevant physical assumptions hold, estimating its hidden input can be a meaningful causal reconstruction within that model. The key phrase is “within that model.” The inference inherits the evidence supporting the operator and its application to the observed case.

A similar distinction arises in system identification. Suppose input and output signals are observed over time and we fit parameters in a dynamical equation. The parameters may be identifiable under sufficiently informative inputs and a suitable model class. If the input barely varies, several dynamical models may produce nearly identical output histories. Time stamps alone do not guarantee that the excitation was informative enough.

One can see this without advanced control theory. If a proposed static relation is y = ax + b but every observed input equals two, the observations constrain only 2a + b. Even perfectly repeated measurements at that input cannot determine a and b separately. Varying the input creates different equations and can separate the parameters. The design of the input is part of the information structure.

Predictive success also needs careful interpretation. A model can predict ordinary observations well because several mechanisms produce similar behaviour in the observed regime. When a policy, control action or operating condition changes, those mechanisms may diverge. The ability to interpolate within one regime does not by itself establish reliability under intervention or major extrapolation.

For inverse problems involving human behaviour, this boundary becomes especially important. A parameter that summarises observed performance is not necessarily an intrinsic personal trait. A learner’s answer can reflect knowledge, question interpretation, fatigue, prompting and task familiarity. A numerical model may help organise those possibilities, but a single observed outcome rarely licenses a complete causal account of the learner.

The same principle appears in our fictional classroom. Ben’s correct answer of six and six in the hidden-box task does not establish that he understood uniqueness. It may be a guess, an equality preference or a justified conclusion under an extra condition. To distinguish those explanations, Adrian asks Ben to name an alternative compatible split and explain what new observation would rule it out. The follow-up is informative because competing explanations predict different responses.

This is an educational analogy, not a validated psychological measurement model. Its value lies in a logical design habit: do not infer a hidden state from an observation without considering alternative states that could produce the same observation. The analogy should not be stretched into claims that a classroom can be fully described by a two-variable linear operator.

A reconstruction report should therefore separate parameter estimation, prediction, mechanism explanation and intervention claims. It can be strong on one and weak on another. “These parameters fit the observed signal under this model” is an important result. “Changing this variable will cause the predicted outcome in a new environment” is a stronger claim requiring additional support.

When alternatives remain, name them and design a discriminating test where feasible. If no ethical, safe or practical test is available, report the uncertainty rather than converting a preferred story into a discovered cause. Mathematical precision should make the boundary sharper, not make an unsupported causal statement sound more scientific.

The lesson is not to retreat from explanation. It is to match the explanation to the evidence. Inverse reasoning becomes more powerful when it asks which parts of the hidden account are identified, which are assumed and which would change under another plausible mechanism. That discipline supports better experiments, better teaching questions and more defensible decisions.

22. Learned reconstruction still answers to the measurement

A learned reconstruction method can use examples of objects and observations to develop a powerful mapping from data to an estimated object. The mapping may be fast and may exploit structures difficult to specify by hand. But the basic information questions do not disappear: what did the instrument observe, what did the training examples contribute, and what remains ambiguous for this particular observation?

Our exact blur counterexample provides a useful test. The constant sequence five and the alternating sequence six, four, six, four produce identical blurred data under the stated periodic averaging operator. Any deterministic algorithm receiving only those data and the same supplied metadata must return the same output for both cases. It cannot be correct on both original sequences simultaneously if correctness requires exact recovery of the full sequence.

That conclusion does not depend on whether the algorithm is linear, iterative or a neural network. It follows from identical inputs to a deterministic rule. A learned prior may make one object more probable than the other within a training distribution. That can improve average performance for that distribution, but it does not make the two objects observationally distinct.

A stochastic generative method can return different plausible objects from the same data. This can be valuable when the outputs represent genuine conditional alternatives. It also requires care: diversity of generated images is not automatically calibrated uncertainty. A collection may omit an important compatible mode, concentrate on familiar patterns or include features not adequately constrained by the observation.

Data consistency is an important check. Apply the forward operator to the reconstructed object and compare the predicted observations with the actual measurements using an appropriate discrepancy measure. A gross mismatch is evidence of a problem with the reconstruction, operator or noise account. But a small mismatch is not sufficient proof of the object’s truth when a nullspace or weakly measured directions remain.

Published research illustrates specific risks. Antun and colleagues reported image-reconstruction instabilities in tested deep-learning systems, including sensitivity to small perturbations and failures to preserve small structural changes. Their work motivates explicit stability testing. It should not be flattened into a claim that every learned reconstruction method fails, or that every conventional method is safe. The source linked here is the authors’ abstract record, not a claim that this article independently replicated their experiments. [9]

There is also constructive research connecting learning with mathematical guarantees. Mukherjee and colleagues studied learned convex regularisers within a variational reconstruction framework, combining data-adaptive structure with analysis of the resulting optimisation and regularisation procedure. This is a useful counterpoint: learning and rigorous constraints are not opposites. The guarantees belong to the stated formulation and assumptions, not to all neural reconstruction. [10]

For a practical evaluation, average reconstruction error is only one view. Test whether important features survive, whether small input perturbations create large output changes, whether calibration changes affect reliability, and whether performance transfers to different object classes. A method can improve an average score while failing in an uncommon but consequential case.

The evaluation split should also match how the model will be used. Nearly duplicate objects across training and testing can exaggerate generalisation. Simulated observations generated by one ideal operator may not test performance under a real instrument’s calibration error. A reconstruction trained on one acquisition geometry may need new evidence before use under another geometry. These are design issues, not merely demands for a larger test set.

The term hallucination should be used precisely rather than as a general insult. A reconstructed feature may be plausible under the prior yet unsupported by the particular measurement. Another feature may be an artefact of numerical instability or operator mismatch. Different causes require different repairs. The important question is what evidence supports the feature, not whether the image looks natural.

Human review helps only when the reviewer has relevant evidence and an appropriate task. A person may prefer a visually smooth image even when a real small feature has been removed. Expert inspection can identify some implausibilities, but it cannot recover information that the acquisition failed to distinguish without further knowledge. Review is one layer of validation, not a substitute for a measurement model.

Uncertainty should be assessed at the level the decision uses. Pixelwise variance may not describe uncertainty in the position of a boundary or the presence of a small object. Several samples can shift the same structure together, creating dependence that marginal summaries hide. Reporting an ensemble without analysing its decision-relevant features can give the appearance of caution without useful information.

A defensible learned reconstruction account therefore identifies the acquisition model, training source, admissible object class, loss or prior, data-consistency checks, stress tests and uncertainty limits. It distinguishes empirical performance from mathematical guarantees and does not treat either as universal. Strong performance deserves recognition when its scope is clear.

The underlying opportunity is substantial: learning can provide useful representations, efficient algorithms and informative priors. The governing constraint is equally clear. A method should not receive more inferential authority simply because its output is impressive. The measurement, model and validation evidence still determine what can responsibly be claimed about the hidden object.

23. Four workshops that teach the reasoning, not just the formula

An inverse-problems lesson can become a sequence of instructions: write a matrix, call a solver, display the answer. That sequence may teach execution without teaching why the answer deserves trust. The workshops below instead ask learners to identify what is measured, expose ambiguity, choose a repair and test whether the repair transfers. They are original teaching designs, not reports of measured classroom outcomes.

The fictional learners are not fixed ability labels. Aisha, Ben, Mira, Ryan, Clara and Ethan can each succeed at one part and struggle at another. Adrian and Jo represent adults guiding the investigation. The purpose of the cast is to make decisions concrete without pretending that one observed answer reveals everything about a real student.

Workshop A: distinguish a guess from an identified answer

Adrian places two covered containers on a table and states that together they hold twelve counters. The task is to determine how many are in each. Ben writes six and six. Rather than marking the response immediately, Adrian asks whether the observations require that split or merely allow it.

The first learning objective is not equation solving. It is understanding the strength of a conclusion. Aisha proposes five and seven. Mira proposes three and nine. Each is checked against the same total. The presence of several compatible answers proves that the original data do not uniquely determine the split.

The next task is to describe all nonnegative integer solutions without listing them randomly. Let the first count be t. The second is 12 − t, with t an integer from zero to twelve. There are thirteen ordered pairs. If the containers are distinguishable, five and seven is different from seven and five. If they are deliberately indistinguishable, the counting question changes and must be stated differently.

Jo then introduces a new statement: the first container holds two more counters than the second. The equations become x₁ + x₂ = 12 and x₁ − x₂ = 2. Adding them gives 2x₁ = 14, so the counts are seven and five. The repair was additional information, not merely a better manipulation of the original sum.

For transfer, replace counters with two continuous nonnegative quantities totalling twelve. The compatible family becomes a continuous interval rather than thirteen discrete possibilities. Ask whether a minimum-norm preference produces the same kind of evidence as the second measurement. A successful explanation says no: the preference selects, while the informative measurement constrains through an additional observation.

The diagnostic question is whether the learner can state why the original answer was underdetermined. A learner who simply memorises “two unknowns need two equations” has not finished, because dependent equations will break that rule. That leads naturally to the next workshop.

Workshop B: more equations do not always mean more information

Give the four-cell tomography puzzle from Chapter 18. Ryan counts four equations and four unknowns, then expects one answer. Ethan substitutes a = t and discovers that the remaining equations determine b, c and d in terms of t but do not determine t itself.

The crucial teaching move is to compare the sums of the equations. Both row measurements and both column measurements produce total sixteen. One equation is redundant. The class should identify the redundant relationship explicitly, rather than treating a solver’s rank output as a mysterious verdict.

Next, require two admissible interiors with identical observations. The pairs t = 4 and t = 6 are convenient, but learners can choose others in [3, 9]. Each candidate must be forward-checked. This makes nonuniqueness constructive: the learner demonstrates it by building competing objects rather than repeating a definition.

Ask which proposed fifth measurement would help. A measurement of the grand total again will not distinguish the family. A measurement of a + d will. Learners can prove the difference by applying each proposal to the null direction (1, −1, −1, 1). The first response is zero; the second is two.

For transfer, use a different target such as a + b. That quantity is already identified as ten even though the interior is not. Ask whether another measurement is necessary to answer it. The exercise separates complete reconstruction from recovery of a useful quantity and prevents “not unique” from becoming a blanket declaration that nothing is known.

A repair is complete only when the learner can choose an informative new observation for a changed problem. Repeating the same diagonal trick is not enough. The new measurement must respond to whatever variation the original observations failed to see.

Workshop C: a perfect fit can be the wrong recovery

Use the paired-sensor matrix and the observed vector (4.06, 3.94). Ask learners first to solve the equations exactly. They should obtain (7, 1). Then reveal the synthetic truth (6, 2) and the assigned perturbation. The apparent contradiction becomes the lesson: the exact solution fits the noisy observations, not the hidden truth.

Before showing regularisation, ask learners to work in sum and contrast. The sum is eight and the observed contrast is 0.12. The measurement gain on hidden contrast is 0.02. This decomposition explains the failure more directly than a long sequence of elimination steps.

Introduce the contrast penalty and calculate two candidates, perhaps λ = 0 and λ = 0.0004. The first gives (7, 1), while the second gives (5.5, 2.5). Compare both their measurement residuals and their distances from the known truth. The learner must name which metric is being used before deciding which result is “better.”

Then reverse the perturbation and repeat the reasoning. The unregularised result becomes (5, 3), and positive contrast shrinkage moves the estimate toward (4, 4). This prevents the learner from forming the false rule that smoothing always improves an answer. The benefit depends on the signal, error and preference.

For transfer, change the sensor contrast gain to 0.6 while preserving the assigned error. Ask learners to predict the error reduction before calculating it. They should connect stronger measurement response with less amplification, not merely notice that a new matrix happens to produce a nicer number.

The important assessment is explanatory. Can the learner distinguish noise amplification from an algebra mistake, regularisation from new evidence, and residual from reconstruction error? These distinctions indicate whether the underlying reasoning can survive a different numerical example.

Workshop D: design a claim that the evidence can support

Give the learner a reconstruction with three pieces of information: the total is stable across plausible settings, a broad division is moderately stable, and a small sharp feature appears only under one strong prior. Ask for a report to a non-specialist reader. No additional calculation is needed; the challenge is to align language with evidence.

A weak report says that the image reveals the true interior. A better report separates the stable total, the more uncertain division and the prior-sensitive feature. It states the model assumptions and proposes a measurement that would help test the feature. The goal is not timid writing but differentiated confidence.

Next, ask the learner to identify what kind of interval is being reported. Is it a deterministic compatibility bound, a sampling confidence interval or a Bayesian credible interval under a stated prior? The same numerical endpoints can have different meanings. A correct label must be accompanied by a correct explanation of what was assumed.

For transfer, replace the image with estimated parameters of a simple growth model. Which conclusions concern fitted data, which concern future prediction, and which concern intervention? A learner who carries over the evidence boundaries has understood something broader than image reconstruction.

These workshops can be used at different depths. Younger learners can work with counters and alternative answers. More advanced learners can derive nullspaces, regularisation formulas and posterior distributions. The common thread is not difficulty for its own sake. It is responsibility for the inferential claim.

The adult’s role is to provide the next useful question, not to praise every correct-looking answer equally. “What else could produce this observation?” tests ambiguity. “What happens when the data change slightly?” tests sensitivity. “Which part did the penalty decide?” tests prior awareness. “What new observation would help?” tests design and transfer.

A final written response should connect the measured object, mathematical operation and conclusion in ordinary language. The learner should be able to explain not only how an answer was obtained but why the available information supports that kind of answer. That is a more durable outcome than remembering the name of a regularisation method.

24. Reproduce the experiment, then try to break its conclusion

The accompanying laboratory is a complete Python script using NumPy. It generates only synthetic data, writes machine-readable results and runs explicit mathematical checks. It is designed to make this article’s calculations inspectable, not to certify a real sensor, imaging system or clinical application. The numerical release accompanying this manuscript ran on Python 3.13.5 with NumPy 2.3.5 and passed 79 recorded assertions.

The simplest way to begin is to run python inverse_problems_lab.py --out lab_results in a directory where the script is available and NumPy is installed. The output directory contains a JSON record and four CSV tables. The script does not contact an external service, require private student data or modify an existing project. Its output is an educational calculation record.

Before running anything, predict three results. The unregularised paired-sensor estimate should be (7, 1). Every contrast-penalised estimate should preserve total eight in the assigned example. The fifth tomography measurement should select (6, 4, 3, 3) from the previously compatible family. Prediction makes the run a test of understanding rather than a ceremony in which software supplies the answer.

The function paired_sensor constructs the matrix from its contrast gain. This avoids maintaining several nearly identical matrices by hand. The function checks that the gain is finite and lies in the declared interval. The restriction is part of this particular mixing-device model; it is not a theorem that every mathematical measurement operator has gains in that range.

The reconstruction function forms the augmented least-squares problem rather than explicitly inverting a normal matrix. It validates dimensions, finite inputs, nonnegative regularisation and the reference vector. The solver can return a minimum-norm minimiser when the augmented matrix is rank deficient. That behaviour is documented because a returned vector does not imply that the hidden object was uniquely identified. The NumPy reference explains the underlying routine’s contract. [11]

Several checks compare independent mathematical routes. The augmented solver’s answer is compared with the sum-and-contrast formula at five penalty strengths. Stationarity is checked separately. Singular values are checked against one and 0.02. The tomography family is tested at several admissible values of t, and the additional measurement is checked for full column rank. These are verification checks of the constructed mathematics.

The script also checks the straight-line fit, the Gaussian posterior, the stronger sensor, the exact periodic-blur ambiguity and the sampled-sinusoid alias. Invalid inputs are deliberately submitted to confirm that they are rejected. A useful laboratory should fail loudly when its mathematical contract is broken rather than silently printing a result with undefined meaning.

The iterative section compares Landweber updates with the independently derived closed form. At 1,000 iterations, the reconstructed contrast is approximately 1.9784. At 2,746 iterations, it is approximately 4.000015, close to the synthetic truth’s contrast four. At 10,000 iterations, it is approximately 5.8902 and moving toward the noisy inverse contrast six. The residual keeps shrinking while reconstruction error eventually worsens.

Again, 2,746 is not a recommended stopping count. It is a revealing point in a deliberately known-truth experiment. In deployment, choosing that count by comparing with the hidden answer would be invalid. The laboratory records it to show the mechanism of semiconvergence and to distinguish a verification example from a practical stopping rule.

A separate simulation changes the noise model. Instead of reusing the assigned opposite-sign perturbation, it draws 100,000 independent Gaussian error pairs with component standard deviation 0.02. The fixed random seed is 20260915. The same simulated draws are reused across penalty choices so that differences between choices are not needlessly obscured by different random samples.

That simulation has a different question: for the fixed true object (6, 2) and this stipulated Gaussian observation model, what is the average squared Euclidean reconstruction error? It is not estimating performance across all possible objects, and it is not inferring a noise distribution from real measurements. The truth and the distribution are inputs to the experiment.

The mean squared errors returned by this run were approximately 1.0022 at λ = 0, 0.9619 at 0.0001, 1.3350 at 0.0002, 2.2513 at 0.0004 and 5.1608 at 0.0016. Thus the parameter that exactly recovered the single assigned example, 0.0002, was not best on this tested simulation grid. The smallest grid value of mean error occurred at 0.0001. No global optimality claim follows from checking five candidates.

We can verify the simulation against a formula. Let α = 0.02, true contrast d = 4 and per-sensor noise variance σ² = 0.0004. The contrast shrinkage factor is h = α²/(α² + λ). Its bias is (h − 1)d. Its variance is 2σ²α²/(α² + λ)². The sum remains unbiased and has variance 2σ². Since component error norm squared equals half the sum of squared sum-error and contrast-error, the total expected squared error is σ² + ½[(h − 1)²d² + 2σ²α²/(α² + λ)²].

The exact expectations from that formula are 1.0004, 0.9604, 1.333733, 2.2504 and 5.1604 for the five tested parameters. The simulation estimates agree with those expectations within their recorded Monte Carlo uncertainty checks. This is stronger verification than observing that one plot looks plausible, because a separate derivation predicts the quantity being estimated.

Notice the difference between object uncertainty and Monte Carlo error. The simulation’s changing reconstructions represent the consequences of stipulated measurement noise. The Monte Carlo standard error describes uncertainty in the estimated average caused by using finitely many simulated draws. Increasing the number of draws reduces the latter; it does not improve the original sensor or make its reconstructed contrast more informative.

The first productive modification is to change the hidden truth while leaving the noise model fixed. A balanced truth has no genuine contrast to lose, so contrast shrinkage behaves differently. A strongly unbalanced truth pays more bias cost when contrast is suppressed. The best penalty depends on what object class and task the evaluation represents.

The second modification is to introduce a calibration mismatch: generate data with one contrast gain and reconstruct with another. The earlier analytical example predicts systematic error even without measurement noise. This experiment teaches why a benchmark using exactly the same operator for generation and inversion is useful for verification but insufficient for physical validation.

The third modification is to keep the instrument fixed but change the requested quantity. Evaluate error in the total rather than in the whole vector. The contrast penalty then has a very different practical significance. A method comparison can reverse when the task changes, which is why a single reconstruction score should never be treated as an all-purpose measure of usefulness.

A reproducible laboratory does not end when its assertions pass. The assertions say that specified calculations behaved as expected. They do not show that the examples exhaust the topic, that the real world follows the assumed model or that another reviewer has independently certified the article. The next intellectual step is to vary the assumptions deliberately and explain why the result changes. That is where computation becomes a tool for mathematical understanding rather than a source of borrowed confidence.

25. A final casebook: decide what the observation really tells you

The following problems bring the main ideas together. Each begins with a concrete request and ends with a worked interpretation. They are designed to test selection of the right reasoning, not just calculation. A complete answer should name the domain, show the relevant mathematics and state the strength of the conclusion. All numbers are constructed for teaching.

Case 1: a total with an additional bound

Two nonnegative real quantities sum to twenty. Independent information says the first is at least six and at most nine. What can be concluded about the second, and is either quantity uniquely determined?

Let the first quantity be t. The second is 20 − t, so it lies between eleven and fourteen. The full compatible set is (t, 20 − t) for 6 ≤ t ≤ 9. Neither component is unique, but the total is exactly identified and both component intervals are bounded. A sensible report gives the compatible relationship, not two independent intervals that permit impossible pairs.

For instance, first equal to six and second equal to eleven are individually inside their reported intervals but do not sum to twenty. The dependence matters. The correct uncertainty description should preserve the pairing. This is the elementary version of why joint posterior structure can matter more than separate error bars.

Case 2: a redundant measurement disguised as extra evidence

Suppose the two observations are x₁ + 2x₂ = 10 and 3x₁ + 6x₂ = 30. A learner argues that two equations determine two unknowns. Explain the flaw and propose one additional measurement that would resolve the ambiguity in exact arithmetic.

The second equation is three times the first, so the measurement matrix has rank one. Its invisible direction can be written (−2, 1): changing x₁ by −2δ and x₂ by δ leaves both observations unchanged. A measurement of x₂ alone would resolve the ambiguity because it changes by δ along that direction.

If the new observation is x₂ = 3, substitution gives x₁ = 4. A measurement of 2x₁ + 4x₂ would not help because it is another multiple of the original relationship. The important distinction is not that the useful measurement has a different number written on it. It measures a different direction in the unknowns.

Case 3: an exact inverse with large absolute amplification

A calibrated sensor follows y = 0.005x. Its observed reading is 0.25, and a deterministic error bound is |ε| ≤ 0.01. Find the range of compatible x values and distinguish absolute from relative sensitivity.

The true ideal reading lies between 0.24 and 0.26. Dividing by 0.005 gives 48 ≤ x ≤ 52. The central inverse estimate is fifty, with an absolute uncertainty bound of two under the stated model. The gain produces an absolute amplification factor of two hundred from observation error to object error.

However, relative uncertainty in this example is 0.01/0.25 = 4% in the reading and 2/50 = 4% in the reconstructed quantity. The scalar relative condition number is one away from zero. Both statements are true. A report that says merely “the inverse amplifies error two hundred times” without stating units and scales can be misleading about practical relative accuracy.

Case 4: a penalty that leaves the ambiguity untouched

The only measurement is x₁ − x₂ = 2. An analyst adds a penalty λ(x₁ − x₂)² and expects a unique reconstruction for λ > 0. Is that expectation justified?

No. Both the discrepancy and the penalty depend only on the contrast. A common shift (δ, δ) changes neither. The shared null direction remains, so the objective cannot determine the total. Writing s = x₁ + x₂ and d = x₁ − x₂ makes the issue immediate: the objective contains d but not s.

The optimal contrast for the objective (d − 2)² + λd² is d̂ = 2/(1 + λ). Every pair with that contrast is a minimiser unless other constraints restrict the total. A numerical minimum-norm solver may return the centred pair (d̂/2, −d̂/2), but that is a selection rule. It does not make the total observationally identified.

Case 5: a useful penalty that changes the answer

Now measure x₁ + x₂ = 10 and choose the minimum-norm compatible solution. Then repeat using a reference x₀ = (8, 2), choosing the compatible point closest to that reference. What changed?

The zero-reference minimum-norm solution is (5, 5), because for fixed sum ten the squared norm is minimised by equal components. The reference (8, 2) already satisfies the measurement, so the nearest compatible point to that reference is (8, 2). Both fit the same exact observation.

The difference comes entirely from the reference preference. Neither output is illegitimate as a clearly stated selection. What would be illegitimate is presenting both as independently discovered truths from the same sum alone. The example is a compact test of whether a reader understands the distinction between data information and regularisation information.

Case 6: noisy readings with unequal reliability

Two independent Gaussian measurements of the same unknown have observed values eight and ten. Their known standard deviations are one and three. Find the weighted least-squares estimate and its variance under the declared model.

The inverse-variance weights are one and one ninth. The estimate is (8 + 10/9)/(1 + 1/9) = 8.2. Its variance is 1/(1 + 1/9) = 0.9, so its standard deviation is approximately 0.948683. This combines independent information while giving less weight to the noisier device.

If both readings share an unknown offset, this variance formula no longer accounts for all relevant uncertainty. If the stated standard deviations were merely guesses, the precision claim inherits that weakness. The arithmetic is exact for the stipulated statistical model, but applying it to a real instrument requires evidence for that model.

Case 7: when averaging plausible objects fails

An observation close to nine is generated by y = x² + ε with small symmetric noise, and the prior is symmetric around zero. Explain why reporting only the posterior mean can be a poor summary of the hidden object.

The two sign-related regions near positive three and negative three receive symmetric support. The posterior mean is zero when the posterior has a finite mean, yet zero predicts an observation near zero rather than nine. The average of plausible explanations need not itself be a plausible explanation under a nonlinear forward map.

A better answer can describe the two modes, report the posterior for the quantity x² if that is what matters, or obtain a sign-sensitive observation. A nonnegative domain restriction would also resolve the sign if independently justified. The solution should not silently impose that restriction merely to obtain a simpler answer.

Case 8: the residual reaches zero

A reconstruction exactly fits a noisy observation vector. A second reconstruction has a small nonzero residual because it suppresses a weakly measured direction. Which one is more accurate?

The information given is insufficient to decide. The first may have fitted amplified noise, while the second may have removed genuine signal. Accuracy concerns distance to the hidden truth or performance on the actual task. Residual concerns agreement with the supplied observations under the forward model. Their relationship depends on conditioning, noise and the reconstruction preference.

In a synthetic benchmark, known truth allows direct comparison. In real work, independent validation, noise-consistent discrepancy, sensitivity analysis and task-specific evidence may help. A correct response refuses the false shortcut “zero residual means correct object” without replacing it with the equally false shortcut “smoother always means better.”

Case 9: choosing the next sensor

An instrument measures only x₁ + x₂ + x₃. The wanted quantity is x₁ − x₂. The proposed new sensors measure either twice the total or x₁ − x₂. Which proposal directly resolves the wanted quantity, and does it recover the whole object?

The contrast sensor directly measures the wanted quantity. The twice-total sensor adds no new independent direction. Yet total plus contrast gives only two independent equations for three unknowns, so the full three-component object remains underdetermined. Recovering the quantity of interest does not require recovering every component.

This is an important design economy. An additional sensor should be evaluated against the question it is intended to answer. Demanding a complete reconstruction when a simple contrast suffices can create unnecessary cost and uncertainty. Conversely, reporting the contrast as if it determined the whole object would overstate the result.

Case 10: a finer grid and a sharper picture

A reconstruction is recomputed on a grid with four times as many unknown pixels, using the same observations. The displayed image looks more detailed. Has the amount of measurement information increased?

No new observations have been collected. The representation has become finer, which may improve numerical modelling of geometry or boundaries, but it also introduces additional unknown degrees of freedom. Any claim about newly visible detail must be assessed against operator sensitivity, noise, priors and validation. Pixel count alone cannot establish evidence-supported resolution.

The same principle explains why a generated high-resolution image may look more convincing without containing more measured information about a particular object. Some added detail may be justified by structural knowledge; some may be prior-driven. The report should make the distinction rather than letting visual sharpness carry the inferential claim.

Case 11: the best parameter on one example

A penalty strength gives exactly the true answer in a single synthetic test. Can it be chosen as the default for a new class of real measurements?

Not from that result alone. The parameter may have offset one particular error realisation, as λ = 0.0002 did in our assigned paired-sensor experiment. Its average behaviour under another declared noise model was different. The appropriate selection rule must use legitimate development evidence and be evaluated on cases representative of the intended task.

A stronger study varies truth, error, calibration and acquisition conditions, separates tuning from evaluation and reports failures. It can establish bounded evidence of usefulness. It cannot become a universal theorem by accumulating attractive examples. The lesson is to generalise through a justified class of conditions, not through confidence in one demonstration.

Case 12: a reconstruction report that can be audited

Write a conclusion for a hypothetical study in which the forward model has been checked against calibration measurements, the main total is stable, component detail depends strongly on regularisation and no independent data test the smallest features.

A defensible conclusion would state that the observations support a stable estimate of the total under the calibrated model. It would describe the component reconstruction as a regularised estimate, identify the imposed preference, and show the sensitivity of the detailed split. It would explicitly withhold a strong claim about the smallest features pending relevant validation or additional measurements.

The next step would be chosen by informational value. A new observation sensitive to the unresolved component could help. Repeating the same total measurement more precisely might not. The report should therefore connect the remaining uncertainty to a specific measurement or modelling question rather than ending with a generic request for “more research.”

What the cases should leave behind

The cases share a disciplined sequence, but they do not share one universal answer. Sometimes the correct result is a unique number. Sometimes it is a compatible family, a stable total, a prior-conditioned estimate or a set of alternative explanations. Sometimes the next useful action is a new measurement rather than another calculation.

The essential habit is to keep five objects distinct: the hidden state, the measurement process, the recorded observations, the reconstruction rule and the final claim. Most serious misunderstandings arise when one is allowed to impersonate another. A recorded number is not the hidden state. A selected image is not a direct observation. A converged optimiser is not a validation experiment. A prior is not an extra sensor.

This is why inverse problems belong in a connected Mathematics library rather than only in a specialist catalogue. They bring together equations, linear algebra, probability, optimisation, approximation and proof around a question that matters outside mathematics: what can we responsibly infer from the evidence we can actually obtain?

At the beginning, Ben wrote six and six because the boxes together held twelve. By the end, the important improvement is not that he always produces a different number. It is that he can explain when six and six is determined, when it is merely one possibility, what assumption would select it, and what new observation would distinguish it from another split.

That is the lasting capability. Mathematics does not only generate answers from data. It can reveal how answers depend on data, where information has been lost, which assumptions are doing the remaining work, and how to design a better test. An inverse problem is solved responsibly when the reconstruction and its limits become understandable together.

Sources and further study

The explanations, numerical examples, workshop designs and software in this article are original educational constructions. The following sources provide established mathematical background and precisely bounded factual context. Linked research abstracts support only the research summaries identified above; they are not presented as independently replicated experiments or exhaustive literature reviews. Access was checked for this edition on 15 September 2026.

1. Tristan van Leeuwen — What is an inverse problem? An authored introduction in 10 Lectures on Inverse Problems and Imaging. Read the lecture.

2. Tristan van Leeuwen — Discrete Inverse Problems and Regularization. A technical route into matrix formulations, singular values and regularisation. The examples in this article are separately constructed. Read the lecture.

3. Tristan van Leeuwen — A Statistical Perspective on Inverse Problems. Background on likelihood, priors and posterior formulations. Read the lecture.

4. Tristan van Leeuwen — Variational Formulations for Inverse Problems. Further study of data discrepancy and regularity preferences. Read the lecture.

5. Tristan van Leeuwen — Computed Tomography. Mathematical background for projection-based reconstruction. Read the lecture.

6. Manabu Machida — Lecture Note on Inverse Problems and Reconstruction Methods. Version 1, 2024; a more technical route into ill-posedness and reconstruction. Read the HTML edition.

7. MIT — Waves and Imaging, Laurent Demanet. Course materials and a route from mathematical foundations into wave-based applications. This article does not claim to reproduce or audit every course document. Open the course page.

8. National Institute of Biomedical Imaging and Bioengineering — Computed Tomography. Primary institutional explanation of the imaging technology. Used for brief contextual description, not clinical advice. Read the overview.

9. Vegard Antun and colleagues — On Instabilities of Deep Learning in Image Reconstruction: Does AI Come at a Cost? Author abstract record, submitted 2019; journal reference PNAS, 2020. Read the abstract and publication record.

10. Subhadip Mukherjee and colleagues — Learned Convex Regularizers for Inverse Problems. Author abstract record, version 2, 2021. Read the abstract and version history.

11. NumPy — numpy.linalg.lstsq. Official documentation for the least-squares routine used in the laboratory. The release record separately identifies the installed version used for execution. Read the documentation.

12. Armand Wirgin — The Inverse Crime. Author abstract record, 2004, for the terminology concerning reuse of the same modelling ingredients to generate and invert synthetic data. Read the abstract.

Continue the Mathematics route

For the broad subject map, return to How Mathematics Works. For the next question—how measurement, parameter, model and numerical uncertainty reach the final decision—continue to How Uncertainty Quantification Works.

To strengthen the mathematical tools used here, follow Linear Algebra or Numerical Analysis. Each route develops its own subject rather than replacing the distinction between indirect evidence and the hidden object we are trying to understand.