VIEW THIS AS

Auto mode follows the Route Engine until you choose a viewpoint.

YOU ARE HERE

ROUTE CHECK

CONNECTED TO

WHAT NEXT

Use the canonical route for this room, or HELP if you are unsure.

What Is Neural Population Geometry? | How the Shape of a Population Code Controls Readout and Generalisation

Neural population geometry is becoming one of the most useful ways to understand neural population codes, representational geometry, neural manifolds, dimensionality and linear readout in modern computational neuroscience. Two neural populations can contain the same number of neurons, preserve the same distance between their average responses and even share the same noise eigenvalues—yet support very different classification, generalisation and behaviour because the signal, noise and readout are arranged differently.

In systems neuroscience and cognitive neuroscience, this problem appears whenever researchers compare neural population activity across stimuli, tasks, contexts or learning. A population code can preserve a variable while changing its representational geometry; a neural manifold can remain low-dimensional while rotating relative to a linear readout; noise correlations can either spare or damage discriminability depending on their direction; and latent structure can become easier to generalise without simply making every response farther apart.

This longform guide develops neural population geometry from first principles through neural manifolds, representational similarity, covariance, task axes, margins, linear separability, population decoding, abstraction and cross-condition generalisation. It asks not only whether information is decodable, but whether the geometry is stable across context, robust to noise, accessible to a biologically plausible readout, and supported by current evidence from visual cortex, prefrontal cortex, amygdala, olfactory memory networks and computational models.

That is the entrance to neural population geometry. Counting neurons tells us the size of a recording. Counting effective dimensions tells us how broadly activity varies. Geometry asks the next question: how are the meaningful differences, irrelevant changes, noise and receiving circuits arranged relative to one another?

The answer can change what is separable, what is transferable and what a particular downstream circuit can use. In this article, “controls” means constrains a specified computation under stated assumptions. It does not mean that the shape of an attractive plot, by itself, proves a causal mechanism in the brain.

The answer before the mathematics

Neural population geometry is the arrangement of neural response patterns and response distributions in a population activity space, considered in relation to the distinctions a task requires and the operations a receiver can perform. Its ingredients include distances between conditions, variation within conditions, the directions of task and context effects, the shapes of response manifolds, and the margins available to a proposed readout.

A “population” here means a set of neurons analysed together. A “response pattern” is a vector of measurements from those neurons. A “geometry” is not necessarily a picture drawn in two or three dimensions. It may describe relationships in a space with hundreds of measured coordinates. The population-geometric approach provides a framework for connecting such relationships to computation in biological and artificial networks. [1]

Three distinctions keep the subject clear. First, neural population geometry is not the physical arrangement of neurons in tissue. Second, it is not the same as neural dimensionality: two arrangements can use the same number of dimensions and support different classifications. Third, it is not a substitute for the readout. A representation may contain a distinction that a particular receiver cannot extract.

This guide develops those distinctions through explicit examples. The numerical populations below are constructed mathematical models, not recordings from people or animals. Negative coordinates denote centred activity or latent coordinates, not negative spike counts. Research findings are identified separately, with their experimental or theoretical scope preserved.

For the conceptual foundations, start with Population Code and Neural Dimensionality. Here we move beyond asking how much structure exists to asking what the structure permits.

1. Build the space before interpreting its shape

Suppose we record N neurons during repeated presentations of a stimulus. For one trial and one declared time window, collect the measurements into a vector r = (r₁, r₂, …, rN). Each coordinate might be a spike count, a firing-rate estimate or a measurement related to activity. The choice matters: different signals and different windows can produce different geometries.

One trial is now one point in an N-dimensional measurement space. Present the same condition repeatedly and those points form a cloud. Change the stimulus, context, movement or task rule and the cloud may move, rotate, widen or change shape. A geometric analysis asks which of those changes are relevant to the scientific question.

For condition s, write the mean response as μs and its trial-to-trial covariance as Σs. The mean tells us where the cloud is centred. The covariance describes its spread and orientation in a second-order approximation. Neither completely specifies an arbitrary distribution, but together they already reveal why a mean-response plot can conceal a serious coding problem.

It is helpful to distinguish three spaces. The measurement space contains the recorded neural coordinates. A latent space contains coordinates inferred by an analysis. A task space contains variables such as stimulus identity, context or required action. A successful model connects these spaces without pretending that they are interchangeable. One latent axis labelled “context” is an interpretation supported by its relationship to the task, not an anatomical label attached to a piece of brain.

Even centring has consequences for interpretation. Subtracting the same baseline vector from every response does not change pairwise Euclidean distances. It does change the numerical origin. Subtracting separate condition means, however, deliberately removes the between-condition differences. The first operation changes coordinates; the second changes the question the remaining data can answer.

Before asking what a neural shape means, therefore, write down what one point represents. Is it one trial, an average across trials, a time bin, an entire time course, a neuron or a condition? A surprising number of apparent disagreements disappear when the objects being compared are made explicit.

2. A population code is a collection of clouds, not only a collection of centres

Consider two classes with mean responses μA and μB. Their separation vector is Δ = μB − μA. A first attempt at analysis measures its Euclidean length, ||Δ||. That is a legitimate description of mean separation in the chosen coordinates. It is not yet a complete measure of discriminability.

Imagine two small, tight clouds whose centres are one unit apart. A boundary between them may work well. Now retain those centres but expand both clouds until they overlap heavily. The centre distance is unchanged while individual trials become harder to assign correctly. Alternatively, stretch the clouds mostly perpendicular to their separation. The total spread may become large without producing equally large overlap along the decision direction.

This gives us three different questions. How far apart are the average responses? How variable are responses to the same condition? How is that variability oriented relative to the distinction we want to read? A geometry that omits any one of these questions can be incomplete for a classification task.

The choice of readout adds a fourth question. Suppose a receiver forms z = wᵀr, a weighted sum of population activity. The difference between the projected class means is wᵀΔ. With a common within-class covariance Σ, the projected noise variance is wᵀΣw. A useful signal-to-noise quantity for that specific receiver is therefore:

d′(w) = |wᵀΔ| / √(wᵀΣw).

The numerator rewards alignment with the signal. The denominator penalises noise in the direction the receiver reads. A large neural fluctuation can be harmless to this receiver when it lies in a direction that contributes little to z. A small fluctuation can be damaging when it lies precisely along the decisive direction.

This example illustrates, rather than replaces, the broader research finding that the information consequences of neural correlations depend on their structure, not merely on their average magnitude. Information-limiting correlations provide one formal account of why particular covariance directions can constrain population information. [2]

The practical lesson is not “remove correlations.” It is to specify the represented distinction, the noise geometry and the receiver before deciding which correlations matter.

3. Worked example: the same noise budget can produce 97.7% or 69.1% accuracy

We can make the opening example exact. Assume two equally likely classes with Gaussian response distributions and the same covariance within each class. Their means are μA = (−1, 0) and μB = (1, 0). The Euclidean distance between the means is two in both of the populations we will compare.

In population A, the covariance is diag(0.25, 4). The horizontal standard deviation is 0.5, and the vertical standard deviation is 2. Most noise extends vertically, away from the horizontal distinction between classes. In population B, the covariance is diag(4, 0.25). Now the horizontal standard deviation is 2, and most uncertainty extends along the task distinction.

Both covariance matrices have the same eigenvalues, the same total variance and the same participation ratio. Only their orientation relative to the signal differs. The qualifier matters: these statements concern the within-class noise covariance. The covariance of all trials pooled across both class means need not have the same spectrum.

For the equal-covariance Gaussian model, the optimal linear separation is described by squared Mahalanobis distance:

d² = ΔᵀΣ⁻¹Δ, with Δ = (2, 0).

Population A: d² = 4 / 0.25 = 16, so d = 4.
Population B: d² = 4 / 4 = 1, so d = 1.

The optimal boundary in this symmetric example is x = 0. With equal class priors, accuracy is Φ(d/2), where Φ is the standard normal cumulative distribution function. That gives approximately 97.7% for population A and 69.1% for population B. These are analytical predictions for the constructed distributions, not reported experimental accuracies.

Why is the difference so large? In population A, the horizontal means sit four standard deviations apart. In population B, they sit only one horizontal standard deviation apart. The vertical uncertainty, however large, does not make one horizontal class look like the other when the receiver can use the appropriate boundary.

Now calculate the noise participation ratio. Both populations have D = (4 + 0.25)²/(4² + 0.25²), approximately 1.125. Equal effective noise dimensionality has not produced equal classification performance. A dimensionality statistic describes how variance is distributed across modes. It does not tell us whether those modes point along the task signal.

The example establishes a precise boundary: neuron count, mean distance, total noise and noise dimensionality together still do not specify discriminability unless their relative arrangement is known. That is a genuinely geometric fact, not a claim that one particular dimensionality level is intrinsically intelligent.

4. Distance has units, assumptions and a question attached

Suppose an analyst multiplies the first neural coordinate by ten. The distance between the means in our example increases from two to twenty. Has the population suddenly become ten times better at representing the stimulus? No. The horizontal noise was multiplied by ten as well. Changing the measuring scale did not add information.

Write this coordinate transformation as r′ = Ar, with A = diag(10, 1). Then Δ′ = AΔ and Σ′ = AΣAᵀ. For any invertible A, using the consistently transformed covariance gives:

Δ′ᵀΣ′⁻¹Δ′ = ΔᵀΣ⁻¹Δ.

The noise-normalised separation is unchanged. This is a useful safeguard against treating larger numerical distances as better codes without checking units and preprocessing. It also explains why noise whitening can be helpful: it expresses distances relative to estimated variability rather than allowing an arbitrary high-variance measurement coordinate to dominate.

Mahalanobis distance is not a universal answer either. The covariance must be appropriate to the question and estimable from available data. A common positive-definite covariance defines a symmetric metric. Distances calculated using a different reference covariance for each class answer a different question and need not be symmetric. Mixing these conventions while calling them all “distance” can hide a material change of model.

There is also a practical difference between rescaling recorded numbers and changing biological gain. In the first case, the analyst changes coordinates after measurement. In the second, the system itself may change firing, noise, thresholds or saturation. A biological gain manipulation is not generally equivalent to a reversible relabelling of the axes.

Finally, invariance under an invertible transform assumes the receiver can adapt appropriately. A fixed receiver operating in physical neural coordinates does not automatically acquire the transformed weights. Nor does a learning algorithm with a particular regularisation penalty behave identically after every rescaling. Information preservation, biological accessibility and finite-sample learnability are three different claims.

Whenever a result says that representations became “farther apart,” ask: under which metric, after which preprocessing, relative to what variability, and for which receiver? Those questions turn a decorative distance into a meaningful measurement.

5. Margins turn separation into a robustness question

A classifier can be correct and still fragile. Suppose its boundary passes extremely close to one class. A small perturbation may then change the prediction even though the original training examples were perfectly separated. A margin describes how much room exists between relevant responses and a decision boundary.

For a linear boundary wᵀr + b = 0, the signed geometric margin of a labelled response (r, y), with y equal to −1 or +1, is y(wᵀr + b)/||w||. Dividing by ||w|| matters. Without that normalisation, merely multiplying all classifier weights by a large number would appear to improve the margin while leaving the boundary unchanged.

Consider a simple bounded example. Put two class centres at (−2, 0) and (2, 0), and let each class occupy a disk of radius R. The boundary x = 0 has minimum margin γ = 2 − R. If R = 0.5, the margin is 1.5. If R = 1.8, it is 0.2. At R = 2, the closed disks meet at the boundary and there is no positive margin. Above that radius, the supports overlap.

This gives a literal perturbation guarantee. If the original response has margin γ greater than zero, any perturbation e with ||e|| less than γ cannot cross that linear boundary. The reason is the Cauchy–Schwarz inequality: |wᵀe| is no greater than ||w|| ||e||. The movement available to the score is smaller than the distance needed to reverse its sign.

Notice the assumptions. These are bounded disks, not Gaussian clouds with unbounded tails. A Gaussian model can provide an error probability, but it cannot provide the same finite-radius guarantee for every possible response. Statements about robustness must identify whether they concern worst-case bounded disturbances, average performance or a specified probability of error.

Real response manifolds need not be circular. Their orientation and anisotropy can matter as much as their apparent size. A long manifold running parallel to a boundary can remain easy to classify, while a shorter one pointing across the boundary may be troublesome. General manifold-classification theory studies how effective radius, dimension, correlations and margin jointly constrain linear classification capacity. [3]

The disk calculation is only a special case, but it teaches a durable habit: do not ask merely whether two classes can be separated. Ask how much disturbance the separation can tolerate, in which directions, and under which model of variation.

6. Angles explain why a good representation can defeat an old receiver

Now keep the class distributions simple but change the relationship between the representation and the receiver. Let the two class means be +au and −au, where u is a unit signal direction and a is positive. Assume equally likely classes and Gaussian within-class noise with covariance σ²I, where σ is positive. A receiver classifies the sign of vᵀr, where v is another unit direction; positive scores predict the class with mean +au.

If the angle between u and v is θ, the projected class means are ±a cos θ and the projected noise standard deviation is σ. The accuracy of this fixed, zero-threshold receiver is therefore Φ(a cos θ/σ). Alignment has become an explicit computational variable.

Set a/σ = 2. Perfect alignment gives approximately 97.7% accuracy. At sixty degrees, the same receiver achieves approximately 84.1%. At ninety degrees, it falls to 50% because the receiver is now orthogonal to the signal. At 180 degrees, accuracy is approximately 2.3%: the readout is systematically reversed.

The last result is especially revealing. Almost completely wrong output need not mean almost no information. The population still carries the distinction. A receiver allowed to reverse its weights could recover the original performance. What failed was the mapping between the code and the fixed decision rule.

This makes a useful diagnostic triangle. A behavioural failure might reflect a degraded representation. It might reflect a misaligned readout. Or it might reflect a changed task rule that makes the previous readout inappropriate. Those explanations can produce similar accuracy losses but require different evidence.

The mathematical example does not establish that a biological circuit literally rotates one clean task axis through these angles. It demonstrates why angle is an important quantity to test when a population representation and a receiving operation are both specified. The companion article on Neural Readout develops the biological distinction between information present in a population and information actually used downstream.

One further caution: principal components that look similar across sessions are not automatically the readout axes. The direction with the most variance may concern movement, state or another variable. A task axis must be identified through the task and tested without using the held-out data that will later be used to claim success.

7. The same four points can support abstraction—or an exclusive-or problem

We can now construct an even sharper example. Let s denote a task variable and c denote context. Each can take the value −1 or +1, giving four equally likely conditions. Compare two population representations:

Aligned representation: r(s,c) = (s,c).
Twisted representation: r(s,c) = (s×c,c).

Both representations occupy the same four corners of a square: (−1,−1), (−1,+1), (+1,−1) and (+1,+1). Both have zero global mean and covariance I when the four conditions are equally weighted. Both therefore have covariance participation ratio two. Remove the condition labels and their point clouds are identical.

Yet their task geometry differs. In the aligned representation, the sign of the first coordinate always gives s. A decoder trained in context c = +1 can use that same rule in context c = −1. The context moves points vertically while the task distinction remains horizontal.

In the twisted representation, the sign of the first coordinate gives s only when c = +1. When context changes to −1, the relationship reverses. A decoder trained only in the positive context and then kept fixed will classify every noiseless negative-context case incorrectly. Within either context, the task is easy. Across contexts, the old linear rule fails.

The full twisted representation places the two s = +1 conditions on one diagonal of the square and the two s = −1 conditions on the other. No single affine line separates the two diagonals. Each class has the origin in its convex hull, so their convex hulls intersect. This is the familiar exclusive-or structure expressed as a population-code problem.

Now add one nonlinear feature: q = r₁r₂. In the twisted representation, q = (sc)c = s because c² = 1. A receiver that reads q can recover the task variable with a linear threshold on that new feature. Nonlinear feature construction has made the desired distinction linearly accessible.

This is not evidence that a particular biological population computes exactly this product. It is a worked demonstration of why task labels and nonlinear interactions matter. Equal neuron count, equal covariance dimensionality and identical unlabelled geometry do not determine whether one linear rule will generalise. The assignment of task conditions to the points is part of the scientific geometry.

It also shows why a clean clustering picture can be misleading. The relevant issue may not be whether there are clusters at all, but whether the same task distinction runs in a compatible direction across the contexts in which the rule must be used.

8. Abstraction is a relationship across conditions, not just good decoding

Return to the aligned square. For each context, subtract the mean response for s = −1 from the mean response for s = +1. In both contexts, the difference vector points in the same direction. A common readout can therefore extract the task variable without needing a separate rule for each context.

This gives one operational geometric meaning of abstraction: a relevant distinction remains readable when other aspects of the situation change. It does not require all context information to disappear. In the aligned square, context is still perfectly represented by the second coordinate. What matters is that context and task are arranged so that one can be read without being confused with the other.

Bernardi and colleagues studied abstraction through neural population geometry in hippocampus and prefrontal cortex. Their work used cross-condition generalisation performance and related geometric measures to examine whether a decoder trained on some condition combinations could generalise to others. Importantly, conditions withheld from a decoder were not necessarily novel experiences for the animals: the animals had encountered the task conditions during training. [4]

That distinction protects the interpretation. Generalisation by an analysis model across known experimental conditions is evidence about representational organisation. It is not automatically a demonstration that the animal spontaneously solved a never-before-seen task. To establish the latter, the experimental design must include genuinely novel behavioural situations.

There are at least three tests worth separating. A within-condition test holds out trials while retaining examples of every condition during decoder training. A cross-condition test holds out combinations of variables from decoder training. A behavioural transfer test asks the organism to act in conditions it has not previously learned in the same form. Passing the first does not imply passing the second or third.

Abstraction also need not require globally low-dimensional activity. A population may preserve a stable task direction while using other dimensions to distinguish detailed combinations. The important question is whether task-relevant invariance and useful specificity coexist in a form that the receiving computation can exploit.

This is why advanced analysis should resist a single score. High decoding accuracy, high dimensionality, parallel coding directions and good transfer each describe something useful. They are not interchangeable certificates of “understanding.”

9. A code can keep all its distances while changing what a fixed circuit reads

Suppose we rotate every population response by the same orthogonal matrix Q. Orthogonality means QᵀQ = I. For any two responses r and t, the squared Euclidean distance after rotation is ||Q(r − t)||² = ||r − t||². All pairwise distances are preserved.

The rotated representation may therefore have exactly the same representational dissimilarity matrix as the original. An observer allowed to rotate its decoder accordingly can recover the same linear decisions. If the original output was wᵀr, the transformed weight vector Qw gives (Qw)ᵀQr = wᵀr.

But a biological receiver does not automatically rotate its synapses just because an analyst has identified a rotated code. If its weights remain w, it now produces wᵀQr, which can differ substantially from wᵀr. The code’s internal distance relationships may be unchanged while its accessibility to that fixed receiver changes.

This produces an important limitation on distance-based comparisons. Preserved pairwise geometry can establish a kind of representational equivalence without establishing identical implementation or identical consequences for existing downstream circuitry. Conversely, changing single-neuron tuning does not necessarily destroy all population-level relationships.

The distinction between single-neuron tuning and population geometry is a central theme in Kriegeskorte and Wei’s treatment of neural representation. Their perspective explains why these are complementary descriptions rather than competing claims about which level is “real.” [5]

For a cross-session analysis, this means that an old decoder can fail for at least two reasons. The task information may have degraded. Or the information may remain in a transformed representation that the old weights no longer align with. Retraining within the new session helps distinguish these possibilities, although it introduces its own sampling and comparability requirements.

The lesson is not that every observed representational change is merely a rotation. It is that rotation is a concrete counterexample to the inference “the previous decoder failed, therefore the information disappeared.” A valid interpretation has to eliminate such alternatives rather than ignoring them.

10. What a representational dissimilarity matrix preserves—and what it leaves out

A representational dissimilarity matrix, or RDM, records the estimated dissimilarity between every pair of conditions. It is useful because different recordings may contain different neurons or measurement channels. One can compare the relationships among conditions without requiring neuron one in one recording to correspond to neuron one in another.

For a simple Euclidean construction, put the K condition-mean response vectors into the rows of a matrix R. The Gram matrix G = RRᵀ contains their inner products. The squared distance between conditions i and j is Gii + Gjj − 2Gij. If the rows are centred across conditions, those inner-product relationships can be reconstructed from the squared-distance matrix D²:

H = I − 11ᵀ/K
G = −½ H D² H.

This algebra explains why distances and second-order representational structure are closely related. It also explains their limitations. Distances do not fix an absolute origin. They do not fix the orientation relative to particular biological neurons. A common orthogonal rotation leaves them unchanged.

Condition labels must also remain attached to the matrix. The aligned and twisted square occupy the same unlabelled set of points, but matching rows to the same task conditions reveals a different organisation. Comparing matrices after allowing arbitrary condition permutations would answer a weaker question than preserving the task identities.

Furthermore, a matrix built only from condition means does not fully describe trial distributions. Two populations can have identical mean RDMs but different within-condition covariances, as our opening example shows. A chosen distance estimator may incorporate some noise normalisation; that does not turn one summary into a complete generative description.

Encoding models, pattern-component models and representational-similarity analysis can be related through a common framework for representational structure, but their estimation procedures and inferential targets must still be specified. Diedrichsen and Kriegeskorte provide a detailed treatment of those connections. [6]

The useful stance is neither to worship the RDM nor to dismiss it. Treat it as a carefully chosen compression. State which relationships it preserves, which transformations it ignores, and which properties of the original responses require separate analysis.

11. Intrinsic dimension does not determine the difficulty of a linear readout

Imagine a response trajectory that traces a circle. Locally, one coordinate—the angle around the circle—is enough to describe position on it. Its intrinsic dimension is one. Yet the full circle is not a straight line, and a global linear projection can merge distinct points.

For example, reading only the horizontal coordinate maps the top and bottom of the circle to the same value. That readout may be sufficient for a left-versus-right task but inadequate for an upper-versus-lower distinction near the vertical axis. The geometric object is unchanged; the adequacy of the projection depends on the label and receiver.

Now place class labels on alternating arcs. Even though the underlying manifold remains one-dimensional, a single linear boundary may not implement the required classification. A nonlinear readout, a different feature transformation or a temporal history that resolves position may be necessary.

This is why “the brain uses a low-dimensional manifold” is not the end of an explanation. One must still ask about curvature, arrangement of classes, sampling coverage and the operations available to the readout. The Manifold article develops the distinction between intrinsic structure and ambient measurement space; population geometry adds the task-dependent consequences.

Visualisation can conceal this distinction. Project a curved high-dimensional structure onto two axes and separate parts may appear to cross. Such an apparent intersection might be caused by the projection, not by an actual collision of full response states. A different nonlinear embedding can also produce attractive islands whose separations depend on method settings.

The correct response is to return to the original data and the declared scientific measure. Does a classifier evaluated without leakage actually confuse the conditions? Do their estimated distances support the apparent relation? Is the projection faithful to the neighbourhoods or global distances being discussed? A picture can suggest a hypothesis; it cannot supply every premise of the argument.

High-dimensional visual responses also need not be random or uselessly complicated. Stringer and colleagues found structured high-dimensional responses to natural images in visual cortex and related their eigenspectrum to smoothness constraints. The result is a useful counterweight to the assumption that meaningful representation must always compress into a tiny number of dimensions. [7]

12. Why the geometry best suited to learning can depend on the number of examples

There is a difference between a code from which a task can theoretically be decoded and a code from which a particular learner can acquire the task with limited examples. A representation may contain many useful distinctions while demanding more data than the learner receives. Another may sacrifice some detail but make the relevant rule much easier to estimate.

A 2026 theoretical study by Wakhloo, Slatton and Chung examined tasks sharing latent structure. It analysed linear-threshold readouts learned through a specified supervised Hebbian rule and derived a generalisation theory under Gaussian assumptions. The analysis related performance to sample count, neural–latent alignment, factorisation properties and covariance dimensionality, with tests on biological and artificial representations. Its optimal geometry depended on the training regime rather than following one universal instruction to maximise or minimise dimensionality. [8]

We can understand the broader issue without importing the entire model. Suppose a representation contains one strong task-relevant direction and ninety-nine weak directions that are unrelated to the current task. With a small training sample, an unconstrained learner may mistake accidental variation in the weak directions for useful evidence. Additional coordinates have increased opportunities for fitting the sample without improving the underlying task signal.

Now suppose some of those weak directions actually carry information needed for another task. Removing them permanently could improve the first task’s small-sample performance while damaging the wider task family. Whether compression helps therefore depends on what is being learned, what must remain reusable, how many examples are available and how the readout is regularised.

The same point applies to nuisance variation. A variable is not intrinsically a nuisance. Colour is irrelevant when the task concerns shape, but essential when the task is to identify a red signal. The geometry should be evaluated relative to an actual task or task distribution, not an analyst’s unspoken preference for which features deserve to survive.

Likewise, claims of optimality require an objective. Optimal average accuracy across a specified task family is different from optimal worst-case robustness, minimal energy, fastest adaptation or preservation of information for unknown future tasks. A theoretical optimum is informative precisely when those commitments are explicit.

The advanced lesson is conditional rather than vague: ask what the learner knows, what it observes, what operation it can fit and which errors count. Geometry becomes a theory of learning only after those quantities are attached.

13. A September 2026 example: learning changes olfactory manifold geometry

A recent experiment gives the geometric perspective a concrete biological setting. Hu and colleagues trained zebrafish on odour discrimination and studied population responses in the telencephalic region pDp. Their work, published on 10 September 2026, combined behaviour with volumetric calcium imaging in ex vivo preparations after training. Learning was associated with changes in odour-response manifold geometry that improved task-relevant separation and estimated classification capacity. The results connected geometric changes to behavioural discrimination without reducing the effect to a simple increase in average response amplitude. [9]

The paper’s conclusion is more specific than “learning makes neural clouds farther apart.” Several aspects of manifold organisation contribute to classification capacity, and not every change in every geometric descriptor has the same effect. Nor does an absence of particular attractor signatures in the studied population rule out every attractor-based account elsewhere. These are observations in a particular organism, circuit, preparation and task, not a direct measurement of a child learning vocabulary. [9]

What does this example invite us to ask? Not whether learning simply increases activity, but whether training changes the shape of the response distributions in a way a relevant readout can exploit. A change might reduce variation that crosses a task boundary, preserve variation parallel to it, or change the relationships among several condition manifolds.

Our earlier disk model provides a deliberately simpler version of the reasoning. Moving two centres apart can increase the margin. Keeping the centres fixed while reducing the radius can also increase the margin. Keeping both centre distance and overall spread similar while redirecting variability away from the boundary can improve discrimination. These are different geometric routes to a superficially similar behavioural outcome.

A serious experiment should therefore measure more than one summary and resist explaining the full outcome through whichever number changed most neatly. The relevant computation may depend on a combination of centre arrangement, within-condition variation and readout orientation.

For a learning library, the value of the research is conceptual and methodological: it supplies a worked example of how to investigate a representation after experience. It does not authorise a classroom prescription such as “make every neural manifold smaller.” That sentence omits the task, the retained information, the receiver and the evidence needed to connect the laboratory model to instruction.

14. Geometry changes when time becomes part of the representation

A population response can be represented as one vector per trial, one vector per time bin, or a longer vector that concatenates several time bins. Those choices do not merely change the plotting style. They determine which temporal distinctions remain available to the analysis.

Imagine two conditions with the same total spike count in every neuron. In one condition, neuron A responds early and neuron B late. In the other, their order reverses. A whole-trial count representation can collapse the conditions. A time-resolved representation can distinguish them. The apparent geometry depends on which features of the response were retained.

The companion articles on Rate Code and Temporal Code examine this information-preservation question. Population geometry asks what the retained temporal coordinates do to distances, directions and readout compatibility.

A useful analysis is temporal generalisation. Train a decoder at one time and test it at another, maintaining the appropriate separation of training and test trials. If the decoder works across a broad interval, a compatible task direction may be preserved. If it works only near the training time, the representation may be changing—or the signal-to-noise ratio, task content or measurement quality may differ.

Our rotation example shows why the interpretation cannot stop at failure. A task axis can rotate while preserving the information available to a newly trained decoder. An old readout then fails even though within-time decoding remains good. Alternatively, both old and newly trained decoders may fail because the relevant information is weak at the later time. Those patterns motivate different hypotheses.

A temporal model must also respect the behavioural deadline. An offline classifier that sees responses after a choice may predict the choice extremely well without identifying the information that caused it. Concatenating late activity into a high-dimensional vector can create impressive accuracy while invalidating a proposed early-decision explanation.

Thus every geometric result should carry a time statement: what interval was measured, what events aligned it, what information could have arrived before the output, and whether future activity entered the analysis. In a dynamic system, “where are the points?” is incomplete without “when were the points available?”

15. The direction with the most variance need not be the direction that matters

Principal component analysis orders directions by variance. That is useful when the aim is to describe dominant fluctuations or build a compact approximation. It does not guarantee that the first component contains the variable required by the task.

Consider a synthetic response r = (100c, s), where s and c independently take the equally likely values −1 and +1. Let s be the task label and c be context. The coordinate variances are 10,000 and 1, so a one-component variance-preserving reduction selects the first coordinate and retains more than 99% of the total variance. Yet that coordinate contains no information about the independent task label. A receiver interested in s should read the second coordinate. Preserving nearly all variance can therefore lose the answer.

This is not an argument against PCA. It is an argument for aligning the method with the question. A variance summary, a task-predictive projection and a source-to-target communication model optimise different objectives. They can legitimately identify different directions in the same recording.

Demixed principal component analysis is one method designed to relate population structure to experimental task variables rather than merely pooling all variance. Its use does not make the resulting components causal mechanisms: the analysis still depends on the task design, the model and the validation strategy. [10]

The issue becomes sharper when two brain areas are considered. A source population may contain a prominent local fluctuation that contributes little to the measured target. Another, smaller source direction may predict the target more strongly. The question has shifted from “what dominates inside the source?” to “what relates source activity to this receiver?” The Communication Subspace article takes up that separate problem.

One can express the distinction with a matrix W mapping source activity to a proposed output. A source change v is immediately null for that mapping when Wv = 0. It can be large in source space while leaving the chosen instantaneous output unchanged. However, in a recurrent dynamical system, a state change that is immediately null may influence future potent activity. “Null now” does not imply “irrelevant forever.”

This last qualification is important. Static geometry describes a particular relationship at a particular level. A full mechanistic account may need the evolution law as well as the instantaneous projection.

16. Estimated geometry is not automatically true geometry

The condition means in a recording are estimated from finite trials. Even if two true means are identical, their sample estimates will usually differ. Squaring that accidental difference produces a positive distance. This creates a simple but important bias: ordinary squared distances between noisy mean estimates can overstate true separation.

To see why, let the estimated difference be δ̂ = δ + e, where δ is the true mean difference and e is a zero-mean estimation error. For a fixed positive-definite precision matrix P, the expected squared estimate is:

E[δ̂ᵀPδ̂] = δᵀPδ + E[eᵀPe].

The final term is non-negative. The expected distance contains both the true separation and a contribution from estimation noise. A small positive estimate is therefore not automatically evidence of a real representational distinction.

Now estimate the difference independently in two data splits, δ̂A = δ + eA and δ̂B = δ + eB. If the errors are independent, zero-mean and the precision matrix is validly fixed or estimated independently for this argument, then:

E[δ̂AᵀPδ̂B] = δᵀPδ.

The independent noise terms no longer contribute the same positive self-product. This is the central intuition behind crossvalidated distance estimation, including the crossvalidated Mahalanobis measure often called crossnobis. Work by Walther and colleagues examined the reliability of representational dissimilarity measures and the benefits of crossvalidation and noise normalisation. [11]

A finite crossvalidated estimate can be negative. That does not mean the underlying squared distance is physically negative. It means an estimator designed to fluctuate around a true value has sampled below zero. Clipping every negative estimate to zero changes its statistical behaviour and can reintroduce bias.

The assumptions must remain visible. Shared trial noise, dependent splits, unstable covariance estimates or preprocessing that leaks across the split can invalidate a simple unbiasedness argument. Estimating an inverse covariance from too few observations also requires care; regularisation changes the estimation problem and must be part of the reported method.

The general lesson reaches beyond one estimator. A geometric picture is inferred from measurements. Every distance, angle and dimension should come with an account of uncertainty, sampling and the transformations used to estimate it.

17. Choose the test split before choosing the attractive result

A geometric analysis can fail long before its final classifier is evaluated. Suppose the analyst inspects every trial, selects the neurons that best separate the labels, finds the most informative time window and chooses the projection with the clearest clusters. Only then are trials split into training and test sets. The nominal test set has already influenced the representation.

This is a form of circular analysis. The test data helped choose the features or hypothesis whose performance they are then asked to validate. Kriegeskorte and colleagues’ discussion of double dipping in systems neuroscience remains a useful warning about how selection and evaluation can become entangled. [12]

A better workflow begins with the intended generalisation claim. To predict new trials from the same conditions, use an appropriate trial-level split while respecting temporal dependence. To predict across sessions, hold out sessions. To claim generalisation across animals, the experimental unit must reflect animals rather than treating every neuron as an independent organism. To test abstraction, withhold the relevant condition combinations from decoder training.

Preprocessing learned from data belongs inside the training procedure when it would otherwise use information from the test observations. This includes feature selection, supervised axis fitting, hyperparameter choices and data-dependent scaling. Even an unsupervised transform fitted on all data can give the method access to the held-out distribution; whether that is acceptable depends on whether the stated task is inductive or explicitly transductive.

Hyperparameters should be selected without repeatedly checking the final test result. A nested validation scheme separates model selection from final evaluation. Permutation tests and null models should preserve the dependencies relevant to the null hypothesis rather than randomly shuffling everything simply because that produces an easy baseline.

Report the uncertainty at the level of the scientific inference. A narrow error bar across thousands of correlated time bins is not equivalent to replication across independent animals or preparations. Likewise, a striking result from one configuration may be an excellent observation without supporting a claim of species-wide universality.

These requirements are not obstacles added after the science. They define what the result means. “The representation generalises” is incomplete until we know what was genuinely withheld, what the method was allowed to learn and which population the uncertainty statement concerns.

18. When geometry predicts behaviour, what has actually been explained?

Suppose a task-related neural distance increases after training and behavioural accuracy also improves. That is informative, but several explanations remain possible. The representation may have become more discriminable. The downstream readout may have improved. Both may have changed because of another learning process. Or a change in state or measurement quality may influence the estimated distance and performance together.

A useful explanation therefore separates levels. At the descriptive level, we estimate how responses are arranged. At the predictive level, we ask whether that arrangement predicts held-out neural responses or behaviour. At the mechanistic level, we identify processes that create, transform or read the arrangement. At the causal level, interventions test whether changing the proposed process changes the consequence as predicted.

These levels can support one another without being identical. A geometrically successful decoder is not automatically a model of synaptic implementation. A biologically plausible mechanism is not automatically the one responsible for the observed geometry. An intervention that changes behaviour may affect many neural variables at once.

Our worked examples suggest discriminating predictions. If the problem is noise aligned with the task direction, then changing that covariance component should have a different effect from changing equally large orthogonal variability. If the problem is a rotated code with a fixed receiver, readout realignment should restore performance without requiring the original geometry to be rebuilt. If the problem is context-dependent reversal, a shared linear readout should fail across contexts even when within-context classification remains excellent.

In biological work, implementing those interventions selectively is difficult. The difficulty should shape the strength of the conclusion, not disappear from it. An observational study can establish a well-supported geometric association. A computational study can establish a consequence under specified dynamics and assumptions. Neither needs to claim more in order to be valuable.

A further subtlety is that immediately output-null activity can influence future state evolution. A perturbation that does not alter the present linear readout may still change a later trajectory. A causal test must therefore specify its time horizon. “No immediate effect” is not equivalent to “no causal role,” just as “decodable after the response” is not equivalent to “available before the response.”

The strongest geometric account earns its place by making different predictions from plausible alternatives. It tells us not merely that the brain has a shape, but what should happen when a relevant relationship is changed.

19. Design a geometry study around a claim, not a preferred method

Imagine planning a study of whether a population supports a rule across contexts. Begin by specifying the task variable s, the contextual variable c, the behavioural response and the time by which the response must be determined. Then design conditions that vary s and c independently enough to identify their separate and interactive effects.

Next, define the primary representational object. It might be the mean population response in a pre-response interval, the distribution of single-trial responses, or a time-resolved trajectory. Choose a primary distance or readout model appropriate to that object. Record in advance whether the analysis assumes common covariance, a fixed receiver, a linear boundary or a particular form of temporal integration.

Construct the validation split to match the claim. A cross-context claim requires cross-context testing. A claim about learning a novel rule requires a behavioural novelty manipulation, not simply withholding familiar conditions from an external classifier. Decide what controls would reveal confounding by movement, response preparation, sensory strength or global state.

Only then select the analysis tools. PCA may describe dominant variance. A supervised axis may estimate the task contrast. Crossvalidated distances may test condition separation. A classifier may measure readout performance. A manifold model may describe structured within-class variation. The tools should cooperate around the question rather than compete to produce the most dramatic picture.

Finally, specify the alternative outcomes. Good within-context but poor cross-context performance supports a different conclusion from poor performance everywhere. Preserved distances with failed fixed-decoder transfer suggests a different family of hypotheses from degraded within-session decoding. A geometric effect that vanishes after controlling movement should be interpreted differently from one that persists across matched behavioural states.

An informative study is allowed to return a complicated answer. Perhaps the task variable is abstract in one subspace and context-bound in another. Perhaps a representation is stable early and reconfigured later. Perhaps learning improves margin without substantially changing a global dimensionality score. Those results can be more scientifically useful than a universal sentence about “higher-dimensional brains.”

A compact study record should therefore preserve the population, preparation, response measure, time window, task labels, metric, model assumptions, train–test split, uncertainty unit and competing explanations. Without that record, geometric language becomes difficult to compare across studies because the same word may refer to different objects.

20. Advanced workshop: ten questions that expose a weak interpretation

Question 1: The neurons are farther apart after rescaling. Has information increased?

An analyst multiplies every recorded response by ten and reports a tenfold increase in Euclidean centre distance. The same operation also multiplies standard deviations by ten. The answer is no: this is a coordinate rescaling, not evidence of improved encoding. A consistently transformed noise-normalised separation remains unchanged. A claim of improved information would need something beyond a change of units.

Question 2: Two populations have the same dimensionality. Must a decoder perform equally well?

No. The aligned and twisted squares both have global covariance participation ratio two, but the task labels have different arrangements. One supports a single context-general linear rule; the other produces an exclusive-or problem. A dimensionality value cannot replace a task-labelled geometric analysis.

Question 3: Noise was reduced in a large-variance direction. Must classification improve substantially?

Not necessarily. For a fixed linear readout, the relevant projected variance is wᵀΣw. Reducing a large noise component orthogonal to w can have little immediate effect on that output. If the task can exploit the altered direction through a newly learned receiver, the consequences may differ. The statement must identify whether the readout is fixed or adaptable.

Question 4: An old decoder performs below chance after a context change. Was the code destroyed?

No. A reversed code can preserve the distinction while producing systematically wrong output through the old weights. Test a newly fitted decoder under a valid split and inspect the relationship between the old and new coding directions. Below-chance transfer can be evidence of a structured reversal rather than absence of information.

Question 5: A distance estimate is negative. Is the analysis physically impossible?

Not when it is a crossvalidated estimator of a non-negative underlying distance. Independent noisy estimates can yield a negative cross-product on a finite sample. The important questions are the estimator’s construction, its uncertainty and whether its assumptions are satisfied. Replacing negative estimates with zero may damage the statistical interpretation.

Question 6: A decoder transfers to withheld conditions. Did the animal solve an unfamiliar task?

Not automatically. The conditions may have been withheld only from the external decoder’s training set while remaining familiar to the animal. That result can support an abstract representational organisation. Demonstrating behavioural novelty requires a separate design in which the organism faces genuinely new task demands or combinations.

Question 7: A two-dimensional plot shows overlapping classes. Is the full code inseparable?

No. Projection may have discarded the separating direction. In the synthetic response (100c, s), a dominant-variance projection can hide the task variable. Evaluate the relevant classifier or distance in the declared full or independently selected feature space. Do not infer impossibility from a visualisation that was never designed to preserve the required distinction.

Question 8: A bounded class has margin 0.2. What perturbations are guaranteed harmless?

For the specified linear boundary and norm, every perturbation with norm strictly less than 0.2 is too small to cross it. This is a worst-case geometric guarantee for the stated model. It does not guarantee safety under a different norm, a changing boundary, an unbounded response distribution or a later dynamical amplification.

Question 9: The best task axis was selected using all trials before crossvalidation. Is the later score independent?

No. The held-out labels or responses have already influenced axis selection. The feature-selection operation belongs inside the training procedure, with model selection separated from final testing. Merely splitting the classifier stage does not undo leakage introduced upstream.

Question 10: Learning reduced manifold size. Is that necessarily beneficial?

No. Compression can remove irrelevant variation and improve margin, but it can also merge distinctions needed by another task. Examine what information survives, which receiver is being evaluated and whether transfer improves. A smaller representation is not a complete objective unless the retained task and acceptable losses have been specified.

The workshop has one recurring demand: attach every number to its object and every conclusion to its assumptions. “Distance increased,” “dimension decreased,” “accuracy improved” and “geometry stabilised” are beginnings of explanations, not their completion.

21. What can transfer to teaching without pretending to measure the brain?

The geometric perspective can sharpen instructional questions, but behavioural observations do not become neural measurements by metaphor. A teacher cannot infer a student’s neural dimensionality, population margin or communication subspace from an ordinary worksheet. Those terms belong to defined measurements and models.

What can transfer is the logic of testing across changes. Consider a student who solves 2x + 3 = 11 but hesitates at 11 = 3 + 2x. The algebraic solution is the same, while the surface arrangement differs. A teacher can test whether the student recognises the invariant relationship rather than relying on a familiar left-to-right layout. That is a behavioural transfer test, not evidence of a particular neural rotation.

Now change several things at once: notation, numbers, wording and time pressure. Failure becomes harder to interpret because the source of the difficulty is unclear. A more informative sequence first varies one surface feature while preserving the mathematical structure, then combines variations after the initial distinction is understood. This mirrors the experimental need to identify task variables and contexts rather than allowing them to remain confounded.

The same approach applies to English. A student may identify a character’s disappointment in a familiar story but fail when a new passage expresses it indirectly. Repeating the original answer tests memory for that instance. Presenting a new setting, different vocabulary and equivalent evidence tests whether the interpretation can transfer. The teacher should ask for the evidence and reasoning, not merely the final label.

A useful diagnostic comparison separates three possibilities. The student may not know the underlying concept. The concept may be available but tied to a narrow cue. Or the student may understand it but fail to express it in the required answer format. Those possibilities resemble, only at the level of reasoning, the distinction between missing information, context-bound access and an unsuitable readout.

No neural mechanism follows from that resemblance. The classroom evidence should remain classroom evidence: worked responses, explanations, transfer across questions and delayed independent performance. The value of Cognitive Art is to make those tests more discriminating, not to decorate ordinary teaching claims with neuroscience vocabulary.

The most useful question is therefore practical: what must stay the same for the idea to remain valid, what may change without changing the answer, and which change would require a different rule? That question can be taught directly, even when nothing is known about the learner’s neural geometry.

22. A stronger way to read the next paper—or the next impressive neural plot

When a paper presents separated neural clusters, begin with the experimental object. What does each point represent? Were points averaged across trials? Which variables were labelled, and which remained uncontrolled? Was the space chosen because it produced a convincing picture, or was the measure specified and evaluated independently?

Then identify the relationship being claimed. A within-area geometry describes local responses. A cross-area predictive subspace describes a relationship between populations. A behavioural readout describes a mapping into an output. A theory of learning additionally specifies how that mapping is acquired. These levels can connect, but a result at one level does not silently establish all the others.

Next inspect the metric and uncertainty. Are distances Euclidean, correlation-based or noise-normalised? Are covariances shared across conditions or estimated separately? How many independent observations support the result? Were negative crossvalidated distances retained appropriately? Were comparisons evaluated under matched sampling and signal quality?

Finally ask what would change your mind. A claim of abstraction should face a cross-condition test. A claim of robust separation should face relevant perturbations or held-out variation. A claim that learning improves geometry should face alternatives involving readout changes, state and measurement. A claim of biological use should identify a plausible receiver and evidence beyond offline decodability.

This does not make every conclusion tentative in the same way. Some statements are exact mathematical consequences of assumptions, such as the rotation invariance of Euclidean distances. Some are well-supported measurements within a preparation. Some are models that explain multiple observations while retaining unresolved alternatives. Good scientific writing lets the reader tell those categories apart.

That is also how the field can remain ambitious. Geometry is valuable because it generates precise questions that single-neuron labels or global activity averages often miss. It can expose hidden failures of transfer, distinguish a changing code from a failing receiver, and identify the conditions under which compression is helpful rather than destructive.

The ambition should be matched by a disciplined return to measurement. A code that looks abstract but does not generalise under the declared test needs re-examination. A supposedly safer margin that vanishes under realistic variation was not the margin the task required. A plot becomes explanatory only when its relationships survive contact with the job it was supposed to explain.

Advanced synthesis: when geometry becomes a theory of generalisation

The first half of this guide established a basic claim: neural population geometry matters because the same amount of neural activity, the same number of dimensions and even the same pairwise centre distance can support very different computations when signal, noise and readout are arranged differently. The advanced problem begins one step later. Suppose a representation is demonstrably informative. What determines whether a downstream learner can discover the useful variable quickly, whether the same rule transfers to a new context, whether a fixed biological receiver can use it, and whether learning should expand or compress the representation?

Those questions have become central in current work on neural population geometry, representational geometry, neural manifolds and latent structure. A 2026 Nature Neuroscience theory by Wakhloo, Slatton and Chung asks which geometric statistics govern linear-readout generalisation across tasks that share latent structure. A separate 2026 primate study by Wójcik and colleagues asks how prefrontal population geometry changes as a new rule is learned. Work in visual cortex, amygdala and olfactory networks asks related questions from different experimental angles. Together, these studies suggest that “information is present” is only the beginning of the reader job, not the end. See Wakhloo et al. (2026), Wójcik et al. (2026), O’Neill et al. (2026) and Srinath et al. (2026).

Information and accessibility are not the same quantity

Imagine two representations of the same binary variable. In representation A, the variable is carried by the sign of one coordinate. In representation B, the same variable is distributed through a complicated nonlinear pattern across many coordinates. If a sufficiently powerful decoder is allowed unlimited data and arbitrary nonlinear transformations, both may support near-perfect prediction. Yet a downstream neuron restricted to a weighted sum followed by a threshold may read A easily and fail on B.

This distinction can be stated cleanly. Statistical information asks whether a response contains dependence on a feature. Accessibility asks whether a specified receiver can extract the feature under realistic constraints. The receiver might be a linear decoder, a biologically plausible synaptic population, a recurrent circuit with limited time, or a learner with only a small number of labelled examples. A statement about one receiver is not automatically a statement about all of them.

The conceptual framework proposed by Pohl and colleagues in 2026 is useful here because it separates sensitivity, specificity, invariance and functionality rather than allowing the word “representation” to absorb every claim at once. A neural response can be statistically sensitive to a feature without being specific to it. It can be invariant to some nuisance variables without being used downstream. And a feature that is decodable by an external algorithm is not automatically the feature a biological circuit reads. See Clarifying the conceptual dimensions of representation in neuroscience.

Geometry enters because it describes accessibility relative to an operation. A linearly separable code makes a binary distinction available to a linear threshold. Parallel task axes can allow one set of weights to generalise across contexts. Orthogonal nuisance variation can leave one readout unchanged. A narrow margin can make a previously correct readout fragile. A representational rotation can preserve pairwise distances while destroying compatibility with an unchanged receiver.

This yields a hierarchy of claims that should not be collapsed:

  1. Dependence: neural response and task variable are statistically related.
  2. Decodability: a declared decoder can predict the variable on held-out data.
  3. Generalisation: the learned readout remains useful under declared changes of stimulus, context, task or action.
  4. Biological accessibility: a plausible receiver has the anatomy, timescale and operation required to use the code.
  5. Functionality: interventions or converging evidence support the claim that the represented information contributes to downstream computation or behaviour.

A longform owner needs all five because a short definition often jumps from step 1 or 2 directly to step 5. The result sounds confident while hiding the very problem neural population geometry is designed to expose.

Local geometry and global geometry answer different questions

Not every geometric problem is a classification problem between a few discrete clouds. Sometimes the represented variable is continuous: orientation, heading, curvature, position, elapsed time or a decision variable. Then the relevant question may be local sensitivity rather than global separation.

Suppose the mean population response depends smoothly on a scalar stimulus x, giving μ(x). For a small change dx, the mean response changes approximately by μ′(x)dx. If trial variability has covariance Σ and the usual regularity assumptions hold, a local quantity related to linear sensitivity is μ′(x)ᵀΣ⁻¹μ′(x). This expression resembles the Mahalanobis geometry used earlier, but now the signal vector is the local tangent of a continuous response curve rather than the difference between two distant class means.

The intuition is simple. A continuous code is locally informative when nearby values move the population in a direction that is large relative to noise. If the response curve bends through population space, the informative direction can change with x. A decoder trained around one part of the curve may therefore fail elsewhere even if local sensitivity is excellent throughout.

This is why local Fisher-information-like reasoning and global representational geometry should not be treated as substitutes. Local geometry asks how sharply nearby values can be distinguished around one operating point. Global geometry asks how the whole set of conditions is arranged, whether one readout can serve many of them, whether distant regions fold near one another, and whether context changes rotate the relevant direction.

A circle provides a clean example. Every point on a unit circle has the same local curvature and can be parameterised by one angle. Yet a one-dimensional linear projection can identify some distinctions and collapse others. Local precision does not tell us whether a global linear readout can recover the full represented variable. Conversely, a globally simple ordering can coexist with poor local sensitivity when neighbouring values produce almost indistinguishable responses.

For a research reader, the repair is to state the scale of the claim. Is the metric intended to capture infinitesimal sensitivity, discrimination between nearby values, separation between classes, or transfer of a common rule across a broad task family? The word “geometry” can cover all four, but the scientific object is different.

Signal geometry, noise geometry and task geometry must meet

The opening example compared one signal direction with two differently oriented covariance ellipses. Real tasks contain more structure. There may be several relevant variables, several nuisance variables, multiple behavioural outputs and a covariance matrix that itself changes with condition. The analyst therefore needs to keep at least three geometries separate.

  • Signal geometry: how condition means, trajectories or manifolds differ.
  • Noise geometry: how trial-to-trial variability is distributed around those signals.
  • Task geometry: which distinctions should be treated as equivalent, different, invariant or behaviourally actionable.

Suppose a task requires distinguishing four categories while ignoring colour. A representation can contain enormous colour variation. That variation is not automatically harmful. If colour moves responses mainly along directions orthogonal to the category readout, a stable category decision can survive. If colour rotates or translates the category axis so that the same category occupies incompatible regions in different colours, a fixed readout may fail even though category information remains present within each colour condition.

This is one reason a generic signal-to-noise ratio can be misleading. The task determines which signal counts. Noise matters through its alignment with that signal and with the receiver. A large source of variation can be irrelevant to one task and essential to another. In modern neural population analysis, “nuisance” should therefore be read as “irrelevant to the declared current readout,” not “biologically meaningless.”

The same caution applies to dimensionality reduction. If the first few principal components capture arousal and movement, discarding lower-variance dimensions may remove the task variable. If a supervised projection is fitted using all labels before crossvalidation, the apparent task geometry can become circular. If a nonlinear embedding separates classes beautifully but distorts global distances, the visual separation may reflect the method’s objective rather than a directly interpretable neural metric.

A defensible analysis declares which relationships should be preserved before choosing the transformation. This is the general principle behind every later section: geometry is meaningful only relative to a job.

Factorised geometry and conjunctive geometry solve different problems

One recurring debate in representation research concerns whether variables should be represented in separate, factorised directions or combined through mixed, conjunctive codes. The correct answer depends on what downstream tasks must be learned.

Take two binary variables: object identity s and context c. A simple factorised representation r = (s,c) makes each variable available along its own axis. One readout extracts s across both contexts; another extracts c across both objects. This geometry supports strong cross-condition generalisation because the direction corresponding to each variable stays stable while the other variable changes.

Now add an interaction coordinate sc. The representation r = (s,c,sc) is more expressive. A linear receiver can now extract not only s and c but also whether the two variables match. Mixed selectivity has increased the repertoire of functions available to simple downstream readouts. This is one reason high-dimensional mixed codes can support flexible cognition.

But a purely conjunctive representation can make transfer harder. If every combination occupies an unrelated direction, a learner may need examples from each combination before discovering the underlying factors. The representation has excellent capacity to distinguish individual states and poor sample efficiency for a task requiring abstraction across them.

The tension can be described as separability versus factor reuse. High-dimensional expansion can make many dichotomies linearly separable. Factorised structure can make common latent variables discoverable from fewer examples and reusable across tasks. A useful code may combine both: preserve low-dimensional axes for variables that should generalise while reserving additional dimensions for interactions that matter only in some tasks.

Wakhloo and colleagues’ 2026 work formalises this kind of trade-off for task families sharing latent structure. Their theory relates generalisation to geometric statistics including dimensionality, factorisation and correlation structure, and shows that the geometry optimal under limited training data need not match the geometry optimal later. The important point for this guide is not one universal prescription such as “low-dimensional early, high-dimensional late.” It is the conditional rule: optimal geometry depends on the task family, learner and amount of experience. See Nature Neuroscience, 2026.

Cross-condition generalisation turns abstraction into a testable geometric claim

Suppose a decoder can distinguish rewarded from unrewarded trials when trained and tested on randomly held-out trials drawn from all contexts. That result shows that reward information is decodable. It does not show that the code is abstract across contexts. The decoder may rely on a different cue in each context.

A stronger test trains on some condition combinations and tests on combinations withheld from decoder training. If the same decision rule transfers, the relevant coding directions are compatible across those conditions. This family of tests is often described as cross-condition generalisation performance, or CCGP, in work on abstract neural representations.

Consider four condition means corresponding to context A/rewarded, context A/unrewarded, context B/rewarded and context B/unrewarded. In an ideal factorised geometry, the rewarded-minus-unrewarded vector is parallel in both contexts. Train a linear decoder in context A and it can transfer to B. If the reward axis reverses in B, within-context decoding remains perfect while cross-context transfer fails. If the axes are merely tilted, transfer degrades gradually with alignment and noise.

The critical word is withheld. What exactly was withheld? A decoder can be denied condition B while the animal has experienced B thousands of times during training. Such a result tests whether the neural representation supports a context-general external readout; it does not establish zero-shot behavioural learning by the animal. Bernardi and colleagues make this representational use of CCGP explicit in their work on abstraction. See The Geometry of Abstraction in the Hippocampus and Prefrontal Cortex.

CCGP can also be misleading if the train and test sets share hidden confounds. Suppose all rewarded training conditions involve a leftward movement and all rewarded test conditions happen to share another correlated motor feature. The decoder may transfer through movement rather than the intended abstract variable. Good design therefore orthogonalises or models obvious alternatives and tests multiple cross-condition splits rather than selecting the one that gives the strongest result.

The broader lesson is powerful: abstraction can be expressed as a geometric relationship, but only after the conditions and transfer requirement are defined. “Abstract representation” without an explicit invariance or generalisation test is too loose for an advanced article.

Learning can reorganise geometry without simply adding information

A learner can improve even when the task variable was decodable before learning. The improvement may come from changing how that information is arranged relative to other variables and downstream readouts.

This is one of the most important implications of the 2026 primate prefrontal-cortex study by Wójcik and colleagues. The authors examined how neural representations evolved as monkeys learned a new rule. The study’s framing is geometric: prefrontal representations can support exploratory flexibility through high-dimensional mixed coding, but learning can reorganise activity toward more compact, task-relevant and abstract formats that generalise efficiently. See Learning shapes neural geometry in the primate prefrontal cortex.

The useful interpretation is not that learning always drives dimensionality downward. A different task could require adding distinctions and expanding dimensionality. Nor is “compact” automatically “better.” Compression is beneficial only when the discarded variation is irrelevant to the task family that must be supported.

Instead, think in terms of three possible learning operations:

  • Expansion: create new dimensions so previously entangled conditions become separable.
  • Rotation or alignment: preserve information while orienting relevant directions toward a stable receiver.
  • Compression or factorisation: reduce irrelevant variation so a common rule generalises across conditions.

These operations can coexist. Early learning may explore many mixed features. Later activity can retain useful interaction dimensions while strengthening one task axis. A downstream readout can also change at the same time. Behavioural improvement therefore does not identify which geometric layer changed unless the experiment measures representation and readout separately enough to distinguish alternatives.

A particularly informative comparison is old decoder versus new decoder. If a decoder trained before learning fails afterward but a newly trained decoder succeeds, information may remain while geometry has changed. If both fail, the information itself may have weakened or moved outside the measured population. If the old decoder improves without retraining, the population may have aligned itself with an existing receiver. These are schematic interpretations, not automatic conclusions, but they illustrate why geometry can make learning mechanisms testable.

One population can represent several variables without forcing them to interfere

Mixed selectivity sometimes sounds like a recipe for interference: if the same neurons respond to stimulus identity, value, movement and internal state, how can downstream systems isolate the variable they need?

Population geometry provides the answer. Variables can be mixed at the level of individual neurons yet organised into population directions that support specialised readouts. A single cell does not need to be a pure “valence neuron” for a valence-sensitive downstream projection to exist. What matters is whether the weighted population pattern contains a stable direction for that variable and whether nuisance variables lie in directions that the receiver discounts.

O’Neill and colleagues reported exactly this kind of population-level organisation in a 2026 Nature Neuroscience study of basolateral amygdala. Individual neurons commonly showed mixed selectivity for several variables. Yet population activity could realise geometries with two useful computational properties: generalisation across conditions and reduced interference between specialised readouts of different variables. See The representational geometry of emotional states in basolateral amygdala.

The result is important because it prevents a false dichotomy. A population need not choose between mixed selectivity and clean downstream control. Mixing at the unit level can coexist with structured separability at the population level.

A simple constructed example shows how. Let a three-dimensional population encode variables s and c as r = (s, c, sc). A neuron-like coordinate can mix the variables through sc. Yet the first coordinate still provides a clean s readout and the second a clean c readout. Now rotate the whole representation with an orthogonal transformation. Every recorded unit becomes a mixture of all three latent variables, but appropriately rotated readout weights still recover s, c and sc exactly. Unit selectivity has become messy while population accessibility remains clean.

This example also explains why searching only for “purely selective neurons” can miss a functional code. Conversely, finding a clean decodable axis does not prove the brain has a dedicated anatomical channel for it. Geometry describes a relation among population activity and a candidate operation; anatomy and causal use still require separate evidence.

Stable readout can support flexible behaviour even when individual responses vary

Generalisation is often imagined as a property of a representation that deletes irrelevant variation. But another possibility is that irrelevant variation remains while the task-relevant direction stays stable enough for one readout to work across it.

A 2026 Nature Communications study by Srinath, Czarnik and Cohen tested this idea in visual cortex. The authors recorded populations in V1 and V4 from two male monkeys performing a continuous visual-curvature estimation task while stimulus features and action mappings varied. Their analyses supported a relatively stable mapping from visual population activity to perceptual judgment across changes in task-irrelevant visual features. In V4, population activity could additionally reflect the upcoming saccade while maintaining the relevant perceptual representation. See Stable readout of visual representations mediates flexible generalization.

The key phrase is stable readout, not “unchanging neurons.” Individual responses can vary strongly with object identity, attention, action planning or other factors. Flexible behaviour can still be supported when the task-relevant population direction preserves a sufficiently consistent relationship to the receiver.

This creates an important alternative to the idea that every contextual change requires retuning downstream weights. If the representation itself organises relevant and irrelevant variation appropriately, one receiver can remain useful across many conditions. That can make rapid generalisation possible because the system does not need to relearn the mapping for every new object or response configuration.

Again, the claim is bounded. The experiment concerns particular visual-cortical populations, a particular continuous estimation task and two animals. It does not establish one universal readout mechanism for all cognition. Its value here is mechanistic: it shows how population geometry can be tested against a concrete generalisation problem rather than discussed only as an abstract shape.

A geometry-of-generalisation checklist

Before concluding that a neural representation supports generalisation, ask the following questions in order:

  1. What variable is supposed to generalise? Name it precisely.
  2. Across what changes? Stimulus identity, context, action, time, session, animal or task?
  3. What operation is allowed? Fixed linear readout, retrained linear readout, nonlinear decoder or recurrent receiver?
  4. What data are genuinely withheld? Trials, condition combinations, sessions or experiences?
  5. What nuisance variables covary with the target? Movement, arousal, reward, timing, sensory strength?
  6. Does the task axis preserve orientation? Or does it rotate, reverse or bend across conditions?
  7. How is noise oriented? Large orthogonal variability may be harmless; small aligned noise may not be.
  8. Is performance margin large enough? Perfect training separation can still be fragile.
  9. Can a plausible biological receiver access the code in time? Decodability after the behavioural deadline is not causal evidence.
  10. What alternative explanation predicts a different result? A useful geometric theory must be falsifiable.

This checklist is deliberately stricter than the phrase “the classes separate in PCA space.” The purpose of a 20,000+ word owner is to keep the reader from confusing an attractive geometric description with a completed explanation.

Representational drift: when the code moves but the capability survives

Record the same brain region today and next week. Individual neurons may change their apparent tuning. Population coordinates can shift. A decoder trained on the first session may lose accuracy on the second. It is tempting to say that the representation has become unstable or that the remembered information has disappeared. Neural population geometry gives us a more careful set of possibilities.

First, the represented information itself may weaken. If a fresh decoder trained and tested within the later session also performs poorly, this becomes plausible. Second, the information may remain but rotate or otherwise transform relative to the old coordinate system. A newly fitted decoder can then recover the distinction even while the old decoder fails. Third, the geometry relevant to a downstream biological receiver may remain stable even though single-neuron coordinates drift. Fourth, recording instability can imitate biological drift when different cells, signal qualities or spike-sorting decisions enter the two sessions.

These alternatives demand different tests. Within-session decoding measures whether the information remains available at each time. Cross-session decoding asks whether one set of weights transfers. Representational-similarity measures ask whether relationships among condition patterns are conserved. Alignment methods ask whether one low-dimensional transformation maps one session into another. Behaviour tells us whether the organism retained the capability. None of these measures alone is the whole story.

Suppose session 2 is an exact orthogonal rotation of session 1. Pairwise Euclidean distances among condition means remain identical. A representational dissimilarity matrix based on those distances is unchanged. A fixed decoder expressed in the old neural coordinates can nevertheless fail. If the downstream biological circuit changed its synaptic weights in concert, behaviour could remain stable. If it did not, behaviour might deteriorate despite preserved internal geometry. This one construction shows why “stable representation” needs a declared level: stable relative geometry, stable individual tuning, stable external decoder, stable biological readout or stable behaviour.

Now imagine a different transformation: all condition means collapse toward one another while the old orientation is preserved. Cross-session alignment can look excellent because the axes have not rotated, yet discriminability falls because the margin shrank relative to noise. Stability of direction is not stability of information. A serious longitudinal analysis therefore needs both alignment and scale relative to uncertainty.

Drift is especially important for brain–computer interfaces, chronic neural recordings and theories of durable memory because a practical receiver may need to operate for days or months. But even here one should avoid a simple moral that drift is bad. A system can preserve behaviour through redundancy, manifold-constrained change, recalibration or stable latent relationships while individual components move. The engineering question is which invariants a receiver can rely on.

For Cognitive Art, the deeper lesson is that stability is itself geometric. The relevant invariant may be an angle, a distance ordering, a subspace, a latent variable or a readout relation rather than the activity of any one neuron. This is why the series keeps Invariance, Robustness and Neural Population Geometry as separate owners.

The geometry you estimate depends on what your instrument can see

“Neural population geometry” can sound as though every experiment observes the same object. It does not. Electrophysiology, calcium imaging and fMRI measure different signals on different spatial and temporal scales. The geometry inferred from each modality is therefore a geometry of a particular measurement process, not a camera photograph of an instrument-independent neural state.

Extracellular electrophysiology can resolve spikes at millisecond timescales from selected neurons near electrodes. A population vector might contain spike counts in a 50-ms window or smoothed firing rates. Change the bin width and fast temporal structure may disappear. Include only well-isolated high-rate neurons and the sampled population can differ from the underlying circuit. Chronic recordings add the problem of matching units across days.

Calcium imaging observes fluorescence related to intracellular calcium, which is coupled to spiking but filtered through indicator kinetics, optical sampling and preprocessing. It can record many identified cells simultaneously and is powerful for manifold studies, but fast spike timing is blurred relative to electrophysiology. Two spike trains with different temporal codes can look similar after calcium dynamics integrate them. The resulting population geometry can therefore emphasise slower activity structure.

fMRI works at another level. A voxel aggregates vascular responses related indirectly to local neural activity over millimetres and seconds. Representational similarity analysis can still compare condition patterns across voxels, but a voxel dimension is not a neuron dimension. Fine temporal ordering and microscopic population structure are inaccessible. A geometry in fMRI pattern space is scientifically meaningful when interpreted at the level the signal supports, not when translated casually into single-neuron claims.

Even within one modality, preprocessing changes the object. Z-scoring each neuron changes relative scaling. Trial averaging removes within-condition variability from the displayed points. Baseline subtraction changes origins. Temporal smoothing changes covariance. Deconvolution changes calcium-derived estimates. Motion regression can remove genuine brain–behaviour covariance along with artefact if the model is poorly specified.

This does not make geometry arbitrary. It makes it measurement-bound. The repair is explicit provenance: define one coordinate, one observation, one time window, one normalisation procedure and one sampling population before interpreting distances or angles. Replication across measurement choices can then strengthen the conclusion.

A useful robustness analysis asks whether the qualitative geometric claim survives reasonable preprocessing alternatives. Does the context-general task axis remain when the bin width doubles? Does the class margin persist after controlling movement? Do crossvalidated distances stay positive when covariance regularisation changes? Do conclusions hold across animals rather than only across thousands of correlated trials from one animal? Those tests turn method sensitivity into evidence rather than embarrassment.

Nonlinear manifolds: when straight-line geometry stops being enough

Many of the examples in this guide use straight lines, planes and covariance ellipses because they make the logic inspectable. Real neural manifolds can be curved. A curved representation creates questions that Euclidean distance and linear separability cannot completely answer.

Imagine a one-dimensional ring embedded in fifty neural dimensions. The ring’s intrinsic coordinate is angle. Two states on opposite sides of the ring can be far apart along the manifold even if a projection makes them appear nearby. Conversely, two points can be close in Euclidean space because the manifold folds while requiring a long path along the represented variable to move between them.

This motivates the distinction between ambient distance and geodesic distance. Ambient distance measures the straight line through measurement space. A geodesic follows the manifold. Which one matters depends on the mechanism. A decoder receiving the raw population vector can sometimes exploit ambient separation. A dynamical system constrained to move along the manifold may care more about paths permitted by its dynamics.

Topology adds another layer. A ring and a line can both be one-dimensional locally, but one contains a loop and the other does not. A torus and a sheet can have the same intrinsic dimension while differing in global connectivity. In spatial-navigation systems, circular variables such as head direction naturally motivate loop-like population structure. The existence of such topology does not mean every neural manifold should be interpreted topologically; it means dimension alone does not specify global form.

Nonlinear embedding methods such as t-SNE or UMAP can make curved relationships visible, but their pictures require caution. Neighbourhood-preserving objectives can distort global distance. Parameter choices can change apparent cluster spacing. Axes often have no direct biological meaning. A beautiful two-dimensional embedding is exploratory evidence until claims are tested with metrics defined in the original or independently validated space.

Likewise, a nonlinear decoder can rescue classifications that a linear decoder cannot. That establishes nonlinear accessibility for the declared algorithm. It does not automatically establish a plausible biological implementation. The relevant question is what class of nonlinear operations the receiving circuit could perform at the required timescale and with the available training signal.

Work on high-dimensional visual responses and on nonlinear neural manifolds illustrates why one should resist equating “nonlinear” with “messy.” A population can have rich high-dimensional structure that remains smooth with respect to sensory inputs. See Stringer and colleagues’ High-dimensional geometry of population responses in visual cortex and De and Chaudhuri’s discussion of population codes and nonlinear manifolds in PNAS.

Local geometry is not the same as communication geometry

A brain area can possess many internal dimensions while sending only a small subset effectively to another area. This makes a crucial distinction between the geometry inside a population and the geometry between populations.

Suppose source population X has 100 effective dimensions. Target population Y covaries predictably with only five particular linear combinations of X. Those five combinations form a candidate communication subspace. A large source fluctuation outside that subspace can dominate local PCA yet have little relation to Y. A smaller source fluctuation aligned with the communication subspace can be more consequential downstream.

This is why the existing Cognitive Art article What Is a Communication Subspace? remains a separate owner. Neural Population Geometry asks how a representation is arranged for discrimination, generalisation and readout. Communication Subspace asks which source dimensions participate in a low-dimensional relation with another population.

The distinction also blocks a common inference error. If a task variable is beautifully separated inside area X, one cannot assume area Y receives that separation. The relevant task axis may be output-null for Y. Conversely, a low-variance dimension in X may align strongly with Y and carry behaviourally useful information despite being visually unimpressive in the source population’s dominant components.

A geometric route through the system therefore asks three successive questions. First, is the variable represented in the source? Second, is the relevant source dimension aligned with a source-to-target communication channel? Third, does the target’s own readout use the arriving variation? The route is:

Population geometry → communication geometry → receiver geometry.

This is an inspection path, not a universal three-stage pipeline. Recurrent loops, feedback, common input and bidirectional communication complicate the causal story. But the decomposition helps locate where an explanation has silently skipped a step.

It also clarifies perturbation logic. Perturbing a source direction that is task-informative but orthogonal to the target-linked subspace may change the local representation without changing the target immediately. Perturbing an aligned direction may have a larger downstream consequence. Such experiments are technically difficult, but they represent the causal prediction that a geometric communication account should aspire to make.

Comparing brain geometry with artificial neural networks without pretending they are the same system

Representational geometry is attractive for comparing biological brains and artificial neural networks because it can operate above the level of individual units. A model may have millions of artificial features while a brain recording contains hundreds of neurons or thousands of voxels. Their units do not correspond one-to-one, but one can still ask whether the relationships among stimuli have similar structure.

Representational similarity analysis does this by comparing dissimilarity matrices. Other approaches compare linear subspaces, canonical correlations, centred-kernel alignment or prediction between representational spaces. These tools can reveal useful correspondence. They can also invite overclaiming.

Suppose a visual model and visual cortex have similar RDMs for 100 images. This means that, under the chosen measure and stimulus set, pairs of images have related representational separations. It does not mean the model and cortex use the same neurons, learning rule, dynamics, causal mechanism or objective. Many distinct implementations can generate similar second-order geometry.

The reverse problem also matters. Two systems can implement the same behaviour through geometries that appear different under one metric. An invertible reparameterisation can preserve information while changing unit tuning and Euclidean distances. A nonlinear transformation can make one representation look unlike another while leaving a downstream task equally solvable. Similarity and functional equivalence are not identical.

Model comparison is strongest when it generates discriminating predictions. If model A and model B both fit current neural geometry, find stimuli on which their predicted representational relationships diverge. Recent neuroscience increasingly values experiments that make models disagree rather than merely rewarding another correlation with existing data. That principle belongs naturally inside a geometry article: the purpose of the metric is not only to say two spaces resemble each other, but to create tests capable of separating explanations.

For education readers, one additional boundary is essential. An artificial network’s embedding geometry can be inspected directly from the model. A student’s internal neural geometry cannot be inferred from essay marks or classroom behaviour without actual neuroscience measurement. Similar vocabulary across AI, cognition and education does not erase evidence boundaries.

Advanced failure signatures: how a persuasive geometry result can still be wrong

The more sophisticated the visualisation, the easier it becomes to mistake analytical flexibility for discovery. The following failure modes deserve explicit tests.

Mixed states create artificial dimensions

Resting trials occupy one low-dimensional subspace and movement trials another. Pool them and dimensionality rises. The new dimensions may describe state mixture rather than one computation. Repair: stratify by state, model state explicitly, then ask which geometry survives within states.

Class imbalance moves the apparent boundary

If one condition is much more common, raw accuracy can favour a biased classifier. Distances between means are unaffected by class frequency, but learned decision thresholds, covariance estimates and crossvalidation variance can be. Repair: report priors, balanced metrics and the decision objective.

Condition averaging hides overlap

Four mean points can form a perfect square while single-trial clouds overlap almost completely. An RDM of means then describes a clean arrangement that individual trials cannot support reliably. Repair: inspect within-condition distributions and crossvalidated performance, not only means.

Arbitrary normalisation changes which neurons dominate

Z-scoring each neuron gives a weakly firing but reliable neuron similar marginal scale to a highly variable neuron. Raw-rate geometry does the opposite. Neither is universally correct. Repair: connect scaling to the receiver or scientific objective and test sensitivity to reasonable alternatives.

Temporal leakage uses the answer before the decision

A decoder predicts choice from a 500-ms window that extends past the behavioural response. The geometry may describe consequences of the choice rather than information used to make it. Repair: constrain windows by causal timing and repeat the analysis before commitment.

Hyperparameter search turns the test set into training data

The analyst tries twenty smoothing windows, ten neuron subsets and eight embedding settings, then reports the combination with the strongest held-out separation. The test set has become part of model selection. Repair: nested validation or a final untouched test set.

Pseudo-replication turns neurons into animals

Ten thousand recorded neurons can create tiny error bars even when they came from one animal. If the scientific claim concerns a species or population of animals, neurons are not the independent replication unit. Repair: carry uncertainty at the experimental level matching the generalisation claim.

A confound becomes the most beautiful axis

Correct trials involve one movement and error trials another. A supervised axis cleanly separates correct from error. The geometry is real, but its interpretation as “confidence” may be wrong. Repair: orthogonalise, match or regress plausible behavioural and sensory alternatives, and report what remains.

These failures share one pattern: the geometry can be statistically real while the story attached to it is wrong. A mature article therefore audits both estimation and interpretation.

Full worked study: from raw population activity to a defensible generalisation claim

Consider a fictional experiment designed only to demonstrate the reasoning workflow. A monkey sees one of four shapes. Each shape can be red or blue. The task is to report whether the shape is rounded or angular. Colour is irrelevant. The animal reports its judgment with one of two saccades whose physical locations swap across blocks. We record 150 prefrontal neurons during the 400 ms before the response.

The scientific question is not simply whether rounded versus angular is decodable. We want to know whether the population geometry supports a shape-category readout that generalises across colour and action mapping.

Step 1: Declare the representational object

For each trial, count spikes from each neuron in a pre-response window that ends before saccade onset. This gives one 150-dimensional vector. We do not average trials before classification because we care about single-trial accessibility. We retain condition means separately for geometric visualisation and crossvalidated distance estimation.

Step 2: Create the right held-out test

Random trial-level crossvalidation within all conditions would answer whether category is decodable in the observed condition mixture. That is useful but insufficient. For colour generalisation, train the category decoder only on red trials and test it on blue trials, then reverse. For action generalisation, train under one saccade mapping and test under the swapped mapping. For the strongest test, train on red trials under mapping A and test on blue trials under mapping B.

Step 3: Measure axes rather than trusting the plot

Estimate category, colour and action contrast vectors using training data only. Compute their angles with uncertainty across sessions. If the category vector remains similar across colour and action contexts while colour and action lie in largely different directions, the geometry is compatible with a stable category readout. If category reverses when action mapping swaps, apparent category decoding may have been motor coding.

Step 4: Add the noise geometry

Estimate within-condition covariance on training data with regularisation selected inside the training folds. Project noise onto the category axis. A large Euclidean category difference with enormous aligned noise can yield poor single-trial generalisation. Large noise orthogonal to category can coexist with good performance.

Step 5: Separate representation from receiver

Compare three models: one fixed linear readout across all contexts, separate linear readouts for each context, and a nonlinear decoder. If separate linear decoders succeed but the fixed decoder fails, information is context-accessible but not factorised for the simple receiver. If only the nonlinear decoder succeeds, the category distinction is present in a more complex form. If all fail on genuine held-out contexts, the measured population may not support the desired generalisation.

Step 6: Test competing explanations

Match reaction time across categories. Control eye position before the response. Confirm the category axis appears before saccade initiation. Repeat the analysis after excluding neurons with the strongest premotor modulation. Examine whether arousal indicators covary with the classes. These controls cannot prove one unique mechanism, but they can remove obvious alternative routes.

Step 7: Link geometry to behaviour

Across sessions, ask whether category margin or cross-condition generalisation predicts behavioural transfer better than raw firing rate or within-context decoding. This is a model comparison, not proof of causality. If the relationship survives held-out sessions and animal-level replication, confidence increases.

Step 8: State the conclusion at the right strength

A careful conclusion might say: “Prefrontal population activity contained a category axis that remained linearly accessible across colour and action-mapping changes in the recorded pre-response interval, with cross-condition generalisation exceeding matched alternatives.” It should not leap immediately to: “The prefrontal cortex stores abstract shape concepts and causally controls category behaviour.” The second statement demands anatomy, perturbation and downstream-use evidence not supplied by the geometry alone.

This fictional study illustrates the reason for longform coverage. The reader needs to know not only the vocabulary but how to turn it into a sequence of decisions where each test eliminates a particular failure mode.

A reproducible blueprint for neural population geometry research

The following blueprint can be used to inspect a published paper, design an analysis, or review a claim. It is deliberately more detailed than a checklist because the order matters.

  1. Define the reader/scientific job. Discrimination, generalisation, dynamics, communication, learning or comparison across systems?
  2. Define one observation. Trial, time bin, condition mean, trajectory or session?
  3. Define the coordinate system. Spikes, rates, fluorescence, voxels, latent factors?
  4. Declare preprocessing. Binning, smoothing, centring, z-scoring, deconvolution, dimensionality reduction?
  5. Name the target variables and nuisance variables. Do not decide which is nuisance after seeing the strongest axis.
  6. Specify the metric. Euclidean, correlation distance, Mahalanobis/crossnobis, cosine, geodesic, classifier margin or another quantity?
  7. Specify the receiver class. Fixed linear, retrained linear, nonlinear, dynamic recurrent or biological target?
  8. Design the held-out set to match the claim. New trials, combinations, sessions, animals or tasks?
  9. Keep model selection inside training data. Including supervised axes, covariance regularisation and time-window choice.
  10. Estimate uncertainty at the correct experimental level. Trials are not animals; neurons are not independent organisms.
  11. Measure signal and noise geometry separately. Mean separation alone is not single-trial discriminability.
  12. Compare plausible alternative geometries. Rotation, scaling, state mixture, motor confound, common input.
  13. Check temporal causality. Was the signal available before the output it is claimed to explain?
  14. Check anatomy and communication if biological use is claimed. Who can receive this population dimension?
  15. Use perturbation where the claim requires mechanism. Prefer interventions that separate competing geometric predictions.
  16. Report non-claims. Say explicitly what the analysis does not establish.

A paper that satisfies only the first half of this blueprint can still make a valuable descriptive contribution. The purpose is not to demand causal certainty from every study. It is to stop descriptive, predictive, functional and causal claims from being blended together.

An evidence ladder for geometric claims

Different claims deserve different evidence. A practical ladder is:

  1. Visual observation: a projection suggests structure.
  2. Quantified description: a declared metric confirms distance, angle, dimension or margin.
  3. Crossvalidated prediction: the structure predicts held-out neural or task variables.
  4. Cross-condition generalisation: one rule transfers across declared changes.
  5. Cross-session or cross-subject replication: the result survives broader sampling.
  6. Downstream association: candidate receiver activity tracks the relevant source dimension.
  7. Temporal plausibility: source changes precede and can reach the receiver before behaviour.
  8. Pathway or dimension-specific intervention: perturbation changes the predicted downstream quantity.
  9. Behavioural consequence: the intervention changes behaviour in the predicted direction while alternatives are constrained.

Higher steps do not make lower steps useless. They answer stronger questions. A clean geometric description can be important even before causal tools exist. The problem begins only when the language of a higher rung is used for evidence from a lower one.

Where Neural Population Geometry sits in the Cognitive Art library

This article is a canonical owner because it performs a reader job none of the neighbouring pages can do alone. The corridor is best understood as a sequence of questions rather than a literal neural pipeline:

  • Population Code: what information exists jointly across neurons?
  • Neural Dimensionality: how many independent collective degrees of freedom are used?
  • Neural Population Geometry: how are task distinctions, nuisance variation, noise, manifolds and margins arranged relative to possible readouts?
  • Communication Subspace: which source-population directions covary with another region?
  • Neural Readout: which available dimensions does a receiver actually convert into downstream consequence?
  • Decision Making: how does evidence become commitment and action at the broader functional level?

Other owners remain essential around the edges. Manifold owns the structured set of states occupied in a larger space. Mixed Selectivity owns nonlinear combinations of task variables in neural responses. Nonlinearity owns the general failure of proportional superposition. Robustness owns survival under perturbation. Neural Population Geometry connects those ideas around one central question: what computations does the arrangement make easy, hard, transferable or fragile?

This ownership matters for both science and search. Instead of creating separate thin pages for “linear separability,” “representational angle,” “neural margin,” “RDM interpretation,” “cross-condition geometry” and “task-axis alignment,” the library places those closely dependent ideas inside one strong owner. The result is better conceptual navigation and less cannibalisation.

Additional current research used in the advanced synthesis

[13] Wakhloo, A. J., Slatton, W. & Chung, S. (2026). Neural population geometry and optimal coding of tasks with shared latent structure. Nature Neuroscience 29, 682–692. Theoretical generalisation framework with analyses of biological and artificial data; conclusions depend on the specified task family and learning assumptions.

[14] Wójcik, M. J. and colleagues (2026). Learning shapes neural geometry in the primate prefrontal cortex. Nature Neuroscience 29, 1966–1975. Primate prefrontal population activity during rule learning; relevant to the reorganisation of high-dimensional, task-relevant and abstract coding geometry.

[15] O’Neill and colleagues (2026). The representational geometry of emotional states in basolateral amygdala. Nature Neuroscience. Mouse basolateral-amygdala population activity, mixed selectivity and geometries supporting generalisation and specialised readouts.

[16] Srinath, R., Czarnik, M. M. & Cohen, M. R. (2026). Stable readout of visual representations mediates flexible generalization. Nature Communications 17, 7914. Population recordings in V1 and V4 of two male monkeys during continuous visual estimation.

[17] Pohl, S. and colleagues (2026). Clarifying the conceptual dimensions of representation in neuroscience. Nature Reviews Neuroscience 27, 357–372. Conceptual framework separating sensitivity, specificity, invariance and downstream functionality.

[18] Hu, B. and colleagues (2026). Representational learning by optimization of neural manifolds in an olfactory memory network. Nature Neuroscience. Zebrafish odour-discrimination learning and pDp population geometry; published 10 September 2026.

[19] Stringer, C. and colleagues (2019). High-dimensional geometry of population responses in visual cortex. Nature. Large-scale visual responses and power-law eigenspectra; useful for separating high-dimensional structure from randomness.

[20] De, A. & Chaudhuri, R. (2023). Common population codes produce extremely nonlinear neural manifolds. Proceedings of the National Academy of Sciences. Theoretical work on population coding and nonlinear manifold geometry.

World return: geometry earns its place when it changes what we can predict

The strongest reason to study neural population geometry is not aesthetic. It is predictive. Geometry can tell us that two codes with equal dimensionality should generalise differently. It can predict that noise aligned with a task axis will be more damaging than larger orthogonal noise. It can explain why an old decoder fails after a representational rotation while a newly trained decoder succeeds. It can separate information present in a population from information available to a fixed receiver.

Those predictions become scientifically valuable when they survive held-out data, alternative metrics, state controls, cross-session tests and—where the claim demands it—causal perturbation. They become weaker when a result exists only in one chosen projection, disappears under a reasonable preprocessing change, or relies on data that leaked into model selection.

For the wider Cognitive Art project, this is why the article has grown beyond a glossary definition. The reader should leave able to recognise the object, calculate simple consequences, diagnose bad interpretations, design stronger tests and navigate to the correct neighbouring owners without inventing a new URL for every geometric term.

The lasting idea is simple enough to carry into any complex system: what matters is not only which states exist, but how the distinctions that matter are arranged relative to variation, transformation and the receiver that must use them.

Mathematical appendix: the quantities that make geometric language precise

Advanced readers often encounter the same words—distance, angle, subspace, dimension, margin—used across papers with different definitions. The safest way to compare studies is to return each word to a mathematical object. This appendix is not intended as a complete textbook. It is a set of translation rules that prevent an intuitive geometric metaphor from outrunning the measurement behind it.

Distance between condition means

For two condition means μA and μB, Euclidean distance is ||μA − μB||. It treats one unit on every coordinate as comparable. If neurons have very different variances or physical units, that choice can make high-variance coordinates dominate. Standardising coordinates changes the metric and should therefore be reported as part of the analysis, not as an invisible cosmetic step.

A noise-normalised alternative uses a positive-definite covariance Σ and computes d² = ΔᵀΣ⁻¹Δ. This asks how large the mean difference is relative to variation described by Σ. The answer depends on which covariance is used, how it was estimated and whether regularisation was required. With limited trials and many neurons, the sample covariance can be singular or unstable, so naïve inversion can produce dramatic but unreliable distances.

Angles between coding directions

Given non-zero vectors a and b, their cosine similarity is aᵀb/(||a||||b||), and the angle is arccos of that value. Parallel vectors have cosine +1, orthogonal vectors 0 and antiparallel vectors −1. In cross-condition generalisation, the sign can matter: a reversed axis can carry the same binary distinction while producing below-chance transfer through an unchanged decoder.

Angles are meaningful only after defining the vectors. A difference of condition means, a linear-classifier weight vector, a principal component and a regression coefficient are not interchangeable “coding axes.” A classifier weight also depends on covariance and regularisation; it need not point in the same direction as the raw mean contrast.

Principal angles between subspaces

When two representations occupy multidimensional subspaces, one angle is no longer enough. Principal angles describe the sequence of best-aligned directions between two subspaces. The smallest principal angle finds the pair of unit vectors—one from each subspace—with greatest alignment. Subsequent angles repeat the process under orthogonality constraints.

This matters when asking whether two task epochs, sessions or brain regions share a subspace. Two five-dimensional spaces can share one almost identical direction while their remaining four dimensions differ substantially. Reporting only the smallest angle could therefore overstate global similarity. Conversely, an average angle can hide one highly conserved task-relevant direction.

A useful analysis links the angle spectrum to the reader job. If the hypothesis concerns one common readout, the best-aligned task direction may matter. If it concerns a reusable latent representation, broader subspace overlap may be relevant. Again, the metric follows the job.

Margin

For a linear boundary wᵀr + b = 0, geometric margin divides the signed score by ||w||. The normalisation removes the arbitrary scale of the classifier weights. A large margin means that labelled points sit farther from the decision boundary under the chosen norm. It provides a useful bridge from geometry to robustness because bounded perturbations smaller than the margin cannot cross that fixed boundary.

But margin is not an intrinsic property of the population alone. It depends on the labels, the boundary and often the training objective. Change class priors, misclassification costs or regularisation and the selected boundary can move. A high maximum-margin separator for one dichotomy says little about another.

Participation ratio and effective dimensionality

If covariance eigenvalues are λ₁,…,λN, the participation ratio D = (Σλᵢ)²/Σλᵢ² is one measure of effective dimensionality. Equal variance across k non-zero modes gives D = k. One dominant eigenvalue drives D toward one. The measure is smooth and avoids choosing an arbitrary explained-variance threshold, but it remains a variance statistic—not a direct measure of task information, abstraction or generalisation.

Two populations can have the same participation ratio and completely different task geometry, as the worked noise example earlier demonstrated. Dimensionality must therefore be interpreted alongside axis alignment, covariance orientation and the variables each mode carries.

Finite-sample geometry: the space can look cleaner than the evidence

Population neuroscience often operates in a difficult statistical regime: many neurons, relatively few independent trials. A 500-neuron recording with 40 trials per condition lives in a 500-dimensional measurement space, but the empirical covariance matrix cannot be estimated freely with high precision from that sample. Geometry must therefore be regularised, crossvalidated or reduced—and each choice adds assumptions.

Suppose an analyst estimates 500-dimensional condition means from 20 trials. Even if the true means are identical, the sample means differ in hundreds of coordinates. Their Euclidean distance will usually be positive. Add more neurons while keeping trial count fixed and the chance distance created by estimation noise can grow. This is why apparent high-dimensional separation is not automatically evidence of a high-capacity code.

Crossvalidated distances reduce one important bias by comparing independently estimated condition differences. But independence must be real enough for the expected cross-terms to vanish. Splitting adjacent time bins from the same trial into different folds does not create independent observations. Preprocessing fitted on all trials can couple the folds. Trial history and slow drift can also violate simple exchangeability assumptions.

Regularisation addresses a different problem. A covariance estimate may be too noisy to invert, so the analyst can shrink it toward a diagonal or identity matrix. Strong shrinkage improves stability at the cost of assuming more isotropic noise. Weak shrinkage preserves estimated correlations at the cost of variance. The shrinkage parameter belongs inside training or validation, not chosen after inspecting the final test performance.

Dimensionality itself is sample-dependent. With too few trials, weak dimensions disappear into estimation noise. With too few neurons, a larger latent structure cannot be observed. With many highly correlated neurons, adding cells may barely increase effective dimension. Good papers therefore examine scaling with neuron count and trial count rather than reporting one dimensionality number as though it were a permanent property of the region.

Finite samples also affect angles. Two independently estimated versions of the same weak task axis can appear poorly aligned because estimation error dominates the signal. A raw angle of forty degrees is hard to interpret without a reliability ceiling or bootstrap distribution. The scientific question is not only “how aligned are the estimated axes?” but “how aligned could they appear given the reliability of each estimate?”

This creates a practical rule: before interpreting geometric differences biologically, estimate how much geometric difference the measurement process itself can generate.

Class priors and decision costs: geometry does not choose the objective for you

Two class clouds can be geometrically identical while the optimal decision boundary changes because the world changes. Suppose class A and class B have equal covariance and separated means. Under equal priors and equal error costs, the Bayes-optimal boundary lies halfway in an appropriate noise-normalised coordinate. Now make A ninety-nine times more likely than B. The optimal decision threshold shifts toward B because a false B prediction is now encountered in a very different probabilistic environment.

Alternatively, keep equal priors but make one error ten times more costly. A medical triage system, predator-detection circuit or safety controller may rationally move the threshold even though the representation has not changed. The geometry supplies evidence; the objective supplies how evidence should be acted upon.

This is another reason classification accuracy cannot be the universal measure of representational quality. A geometry can support excellent ranking of evidence while a poorly chosen threshold produces bad accuracy. Receiver calibration, prior beliefs and asymmetric costs belong to the decision layer.

For neuroscience, the lesson is to distinguish encoding geometry from decision policy. A change in behaviour can arise because the neural evidence changed, because the readout weights changed, because the criterion moved, or because the subjective prior changed. All can be described mathematically but belong to different owners in the Cognitive Art map.

Regularisation changes the geometry a learner can exploit

A representation does not determine a learned decoder without also specifying how the decoder is trained. This becomes crucial in high-dimensional populations where many separating boundaries fit the training data.

Consider a task with 1,000 neural coordinates and 50 labelled training trials. If classes are separable, there may be many linear weight vectors that classify all 50 correctly. A minimum-norm, ridge-regularised, sparse or maximum-margin learner chooses among them differently. The resulting generalisation can differ even though the neural geometry and training labels are identical.

Ridge regularisation penalises large squared weights and tends to distribute the readout. Lasso-like penalties encourage sparse weights. Maximum-margin classification prefers a boundary with large separation under its norm. A biological synaptic learner may impose entirely different locality, sign, wiring and plasticity constraints. “Linearly decodable” therefore means only that some linear boundary exists or was learned by the specified algorithm—not that all plausible learners will find the same one.

This connects directly to sample-efficient generalisation. A geometry with one strong, stable task direction can make many simple learners agree. A geometry with the same information scattered among hundreds of weak dimensions may require more data or stronger inductive assumptions. The difference is not information content alone; it is information relative to the learner’s bias.

When comparing biological and artificial systems, this point becomes even more important. An artificial decoder can be trained with backpropagation, millions of examples and explicit labels. A downstream brain area learns through biological plasticity, local signals, recurrent dynamics and a finite lifetime. Equal asymptotic decodability does not imply equal learnability.

Advanced FAQ: neural population geometry, representational geometry and neural manifolds

Is neural population geometry the same as representational geometry?

They overlap strongly but are not always used identically. Representational geometry is a broad term for relationships among representation patterns, often studied with distances or similarity matrices across neural, behavioural or artificial systems. Neural population geometry emphasises the joint activity space of neural populations and often connects that geometry to readout, dynamics, manifolds, mixed selectivity and computation. In this article, neural population geometry is the narrower canonical owner because the reader job concerns population activity and its computational consequences.

Is a neural manifold the same thing as neural population geometry?

No. A neural manifold is the structured set of population states occupied within a higher-dimensional space, often with lower intrinsic dimension. Population geometry includes the arrangement of one or more manifolds, clouds, task axes, covariance directions, margins and readout relations. The Manifold owner asks what shape the states occupy. This owner asks what the arrangement means for computation.

Does higher neural dimensionality mean more information?

Not necessarily. Higher dimensionality can increase representational capacity and make more dichotomies linearly separable, but dimensions can also contain noise or task-irrelevant variation. A low-dimensional code can carry a task variable very precisely. Information, dimensionality and generalisation are related but distinct quantities.

Does better decoding prove the brain uses that representation?

No. External decoding demonstrates that information is statistically accessible to the declared decoder. Biological use requires a plausible receiver, appropriate timing and stronger evidence such as source-target relationships or perturbation. The Neural Readout article owns that distinction.

Why are noise correlations important?

Because trial-to-trial variability has direction. Correlated noise aligned with a task-relevant difference can blur conditions more strongly than equal or larger variance in orthogonal directions. The effect of correlation therefore depends on its geometry relative to signal and readout, not merely on whether correlation is positive or negative on average.

What does linear separability tell us?

It tells us whether a labelled distinction can be implemented by one hyperplane in the chosen feature space. This is computationally important because weighted sums followed by thresholds are simple readouts. But separability does not tell us margin, robustness, sample efficiency, biological plausibility or whether the brain actually uses that boundary.

Can two representations contain the same information but have different geometry?

Yes. An invertible transformation can preserve all information while changing distances, axes and compatibility with a fixed receiver. An orthogonal rotation even preserves Euclidean distances while changing the weights required for readout. This is one reason information content and accessibility should be separated.

Can two representations have the same geometry but different biological meaning?

Yes. The same abstract geometry can be implemented by different neurons, circuits or measurement modalities. Condition labels and task context matter. Similar RDMs do not establish identical mechanism, causality or subjective meaning. Geometry is a relational description, not a complete ontology.

Does learning always compress neural representations?

No. Learning can expand representations to separate previously confused conditions, rotate them to align with a receiver, compress irrelevant dimensions, factorise latent variables or combine these changes. Current 2026 research shows geometry can reorganise with learning, but the direction depends on the task and stage of learning.

Why is cross-condition generalisation stronger than ordinary decoding?

Ordinary decoding can exploit context-specific cues because training and test data include examples from the same contexts. Cross-condition generalisation withholds relevant combinations from decoder training and asks whether one rule transfers. It therefore tests whether the represented variable occupies a compatible direction across the declared contextual change.

Why can PCA be misleading?

PCA finds directions of large variance, not necessarily directions carrying the task variable or influencing a downstream target. A low-variance task axis can be discarded by a variance-preserving projection. PCA is valuable when its objective matches the scientific question; it is not a universal readout of meaning.

What is the single most important question to ask about a neural geometry plot?

Ask: what operation or prediction is this geometry supposed to support? Then verify that the metric, validation split, noise model, timing and receiver correspond to that job. Without that link, the plot may be descriptive but cannot carry the full explanatory claim.

Final longform audit: what this owner now covers

This article now covers the full reader job rather than stopping at a definition. It explains activity space, condition clouds, mean separation, covariance, noise-normalised distance, margins, readout alignment, linear separability, mixed selectivity, factorisation, cross-condition generalisation, RDMs, manifolds, dimensionality, nonlinear geometry, temporal representation, drift, communication subspaces, artificial-network comparisons, finite-sample estimation, regularisation, causal evidence, measurement modality and experimental design.

It also tells the reader what the concept does not establish. A geometry is not a literal anatomical object. Decodability is not biological use. Low dimension is not simplicity. High dimension is not superiority. Similar RDMs are not identical mechanisms. A clean projection is not proof that the full data have the same structure. Learning-related geometry changes in a particular animal and circuit are not classroom neuroscience.

That combination—mechanism, mathematics, diagnosis, evidence boundary, worked examples and navigation—is the reason the canonical owner deserves longform scale. The word count is a floor on functional coverage, not the objective itself.

The final distinction: a shape becomes useful only in relation to a question

Return to the two response clouds at the beginning. Their neurons did not become more numerous. Their mean separation did not increase. Their total noise did not shrink. The crucial change was that uncertainty pointed in a different direction relative to the task.

Then return to the square. The unlabelled points and dimensionality were identical, but one arrangement supported a common rule across contexts while another required a nonlinear combination. Finally, return to the rotated code. Its pairwise distances survived, yet an old fixed receiver could become blind to its signal.

These are not three unrelated curiosities. They reveal the same principle from different sides: a representation cannot be evaluated only by how much it contains. We must also ask how its relevant differences are arranged, how uncertainty crosses them, what transformations preserve them and what a receiving computation can actually do.

Neural population geometry is the study of those relationships. Its strongest claim is not that the brain draws beautiful shapes. It is that the arrangement of a code can make some operations easy, others fragile and others unavailable to a particular reader—even when the information is still there.

Research references and scope

The worked populations, numerical accuracies, square examples and algebraic arguments in this guide are original explanatory constructions. They are not participant data, simulations of a validated whole-brain model or independent replications of the papers below. The research summaries preserve the distinction between theoretical assumptions, experimental observations and possible interpretations.

[1] Chung, S. and Abbott, L. F. (2021). Neural population geometry: An approach for understanding biological and artificial neural networks. Current Opinion in Neurobiology. A conceptual framework for relating population representations to computation.

[2] Moreno-Bote, R. and colleagues (2014). Information-limiting correlations. Nature Neuroscience. A theoretical treatment of how particular correlation structures can limit population information.

[3] Chung, S., Lee, D. D. and Sompolinsky, H. (2018). Classification and Geometry of General Perceptual Manifolds. Physical Review X, 8, 031003. A theoretical framework for linear classification of manifolds; its effective geometric quantities should not be equated automatically with the simple disk radius used in this guide.

[4] Bernardi, S. and colleagues (2020). The Geometry of Abstraction in the Hippocampus and Prefrontal Cortex. Cell. Cross-condition decoding and population geometry provide operational tests of abstract representation; decoder-held-out conditions must be distinguished from novel animal experiences.

[5] Kriegeskorte, N. and Wei, X.-X. (2021). Neural tuning and representational geometry. Nature Reviews Neuroscience. A discussion of the relationship between single-neuron tuning and population-level representational structure.

[6] Diedrichsen, J. and Kriegeskorte, N. (2017). Representational models: A common framework for understanding encoding, pattern-component, and representational-similarity analysis. PLOS Computational Biology, 13, e1005508.

[7] Stringer, C. and colleagues (2019). High-dimensional geometry of population responses in visual cortex. Nature. Large-population visual responses illustrate that high-dimensional structure can coexist with constraints on representational smoothness.

[8] Wakhloo, A. J., Slatton, W. and Chung, S. (2026). Neural population geometry and optimal coding of tasks with shared latent structure. Nature Neuroscience, published 4 February 2026. Its optimality results depend on the specified task family, statistical assumptions, learning rule and sample regime.

[9] Hu, B. and colleagues (2026). Representational learning by optimization of neural manifolds in an olfactory memory network. Nature Neuroscience, published 10 September 2026. Zebrafish behaviour and post-training ex vivo population imaging provide the experimental setting, not human classroom instruction.

[10] Kobak, D. and colleagues (2016). Demixed principal component analysis of neural population data. eLife, 5, e10989. A population-analysis method that relates activity components to experimental variables.

[11] Walther, A. and colleagues (2016). Reliability of dissimilarity measures for multi-voxel pattern analysis. NeuroImage. A methodological study of dissimilarity estimation, crossvalidation and noise normalisation.

[12] Kriegeskorte, N., Simmons, W. K., Bellgowan, P. S. F. and Baker, C. I. (2009). Circular analysis in systems neuroscience: the dangers of double dipping. Nature Neuroscience. A methodological warning about using the same data to select and evaluate an analysis.

Discover more from eduKate Singapore

Subscribe now to keep reading and get access to the full archive.

Continue reading