Vector space is the geometry in which many Super Intelligence representations become computable. An embedding is one vector. A vector space is the larger coordinate system containing many vectors, where direction, distance, angle and neighbourhood can be used to compare representations.
This does not mean language literally lives on a two-dimensional map. Modern embedding spaces can have hundreds or thousands of dimensions. The geometry is learned or constructed for computation, and the relationships that matter depend on the model and training objective.
This guide explains vector space in AI from first principles. We will build vectors, dimensions, norms, dot products, cosine similarity, Euclidean distance, neighbourhoods, projections, centroids and high-dimensional search. We will also examine the limits of geometric metaphors, including why nearby vectors can still be contradictory or operationally different.
In the eduKateSG series, Super Intelligence, or SI, is our umbrella term for AI technologies. Standard mathematical terms remain visible because vector spaces connect embeddings, neural networks, attention, semantic search and many other machine-learning mechanisms.
Previous: 013 — Embeddings. Return to the How Super Intelligence Works hub for the complete series.
Vector Space at a Glance
- A vector is an ordered list of numbers; a vector space is the mathematical environment in which vectors can be added and scaled.
- Dimensions are coordinate directions, not automatically human-readable concepts.
- Distance measures separation; direction and angle can measure alignment.
- Cosine similarity compares direction; Euclidean distance compares straight-line separation; dot product combines alignment and magnitude.
- High-dimensional spaces behave less intuitively than ordinary 2D diagrams.
- Nearest-neighbour search retrieves geometrically close candidates, not necessarily truthful, authoritative or logically equivalent items.
- Projections and dimensionality reduction can help visualisation but do not preserve every property of the original space.
The Simplest Vector: A Point With Coordinates
A two-dimensional vector can be written [3,4]. Geometrically, we can draw an arrow from the origin [0,0] to the point [3,4]. The same vector can also be treated simply as an ordered pair used in computation.
A three-dimensional vector such as [2,-1,5] adds another coordinate. Human beings can still picture three dimensions. Once an embedding has 768 dimensions, visual intuition stops being literal, but the algebra remains well-defined.
This is why 2D diagrams are teaching tools. They show distance and direction cleanly, but a production embedding space can contain far richer high-dimensional structure.
Vector Addition
If a = [1,2] and b = [3,4], then a + b = [4,6]. We add corresponding coordinates. Geometrically, vector addition combines displacements.
Neural networks use vector and matrix addition constantly. Residual connections in transformers, for example, add one representation to another with matching shape.
The operation is simple but powerful because learned representations can be transformed and combined while preserving differentiability.
Scalar Multiplication
Multiplying vector [1,2] by scalar 3 produces [3,6]. Direction remains the same for a positive scalar while magnitude changes.
Scaling appears throughout machine learning: attention weights scale value vectors, normalization rescales vectors, optimisation updates parameter vectors and linear layers combine scaled inputs.
A vector space is closed under operations such as vector addition and scalar multiplication according to its mathematical definition.
Dimensions Are Coordinates, Not Labels
In a classroom graph, the x-axis might mean time and the y-axis might mean temperature. In a learned embedding space, dimensions usually do not come with such neat labels.
A model can distribute one semantic relationship across many coordinates. A rotation of the whole space can preserve distances and dot products while changing every individual coordinate.
Therefore, do not inspect embedding dimension 217 and casually label it “politeness” or “algebra”. Interpretation requires specific evidence.
Basis Vectors and Coordinate Systems
A basis is a set of directions used to express vectors in a space. In ordinary 2D Cartesian coordinates, [1,0] and [0,1] are familiar basis vectors.
The same geometric vector can be expressed using a different basis. Coordinates change while the underlying vector remains the same relative to that coordinate system.
This helps explain why individual embedding coordinates are not sacred. What often matters is relational geometry rather than one axis value.
Vector Magnitude: The Norm
The Euclidean norm of [3,4] is √(3²+4²) = 5. It measures vector length. In higher dimensions, the same formula extends by summing squared coordinates.
Magnitude can carry information in some learned spaces. In others, applications normalize vectors to unit length so direction becomes the primary similarity signal.
Never assume magnitude is irrelevant unless the model or task supports that assumption.
Dot Product
The dot product of a = [1,2] and b = [3,4] is 1×3 + 2×4 = 11. The operation increases when vectors are strongly aligned and/or have large magnitudes.
Matrix multiplication is built from dot products, so this operation appears throughout neural networks. In transformer attention, learned query and key vectors are compared through scaled dot products.
The dot product is therefore both a geometric measure and a computational primitive.
Angle and Alignment
Two vectors can point in similar directions even when one is much longer. Angle captures directional alignment. Cosine similarity converts the angle relationship into a convenient score.
Vectors pointing in the same direction have cosine similarity 1. Perpendicular vectors have similarity 0. Opposite directions have -1 in the standard mathematical definition.
Embedding models can produce score distributions that occupy only part of this theoretical range. Interpret scores empirically.
Cosine Similarity
Cosine similarity is dot(a,b) divided by ||a||×||b||. Because the denominator removes vector length, the metric focuses on direction.
For a=[1,0] and b=[0.6,0.8], dot product is 0.6, both lengths are one, so cosine similarity is 0.6.
Semantic-search libraries often use cosine similarity, but some embedding models are trained for raw dot product. The correct scoring function belongs to the model specification and evaluation.
Euclidean Distance
Euclidean distance between points a and b is the norm of a−b. For [1,0] and [0.6,0.8], the difference is [0.4,-0.8], whose length is √0.8, approximately 0.894.
Distance and cosine can rank neighbours differently when vector magnitudes vary. If vectors are normalized to unit length, the two measures become closely related.
Manhattan Distance and Other Metrics
Not every task uses Euclidean distance or cosine similarity. Manhattan distance sums absolute coordinate differences. Other metrics and learned similarity functions exist.
The word “nearest” is incomplete until the metric is specified. A vector database index is configured around a particular notion of similarity or distance.
A Worked Geometry Example
Place three toy points in 2D: algebra=[0.9,0.8], equation=[0.8,0.85], banana=[-0.7,0.2]. Algebra and equation are close in Euclidean distance: √((0.1)²+(-0.05)²), roughly 0.112.
Algebra and banana are much farther apart: difference [1.6,0.6], distance about 1.709. Under this invented geometry, the mathematical terms form a neighbourhood distinct from the fruit term.
The example illustrates geometry only. A real embedding space is learned from data and objectives, not manually drawn with human categories.
Neighbourhoods
A vector’s neighbourhood is the set of nearby points under a chosen metric. Semantic search asks for nearest neighbours of a query vector.
Neighbourhoods can reveal clusters, analogies or related items. They can also contain false friends—items that are close for broad topic but wrong for the operational question.
The quality of a neighbourhood therefore depends on the representation and task.
Centroids
A centroid is the average vector of a set of points. If three document embeddings are [1,0], [0.8,0.2] and [0.9,0.1], their centroid is [0.9,0.1].
Centroids can summarise clusters or classes. A classification system might compare a new point with class centroids as one simple strategy.
Averages can also blur multimodal groups. If a class contains two very different subclusters, the centroid may lie in a region containing no actual example.
Clusters
Clustering algorithms group points according to geometric relationships without necessarily having human labels. In an educational corpus, embeddings might form clusters around algebra, geometry, assessment logistics and school policy.
The resulting clusters require interpretation. A cluster is evidence of geometric grouping under the embedding model, not proof that one human concept cleanly defines the group.
Linear Interpolation
Between vectors a and b, an interpolation such as 0.5a + 0.5b produces the midpoint. In generative latent spaces, interpolations can sometimes create smooth transitions.
But a midpoint between two meaningful embeddings is not guaranteed to correspond to a valid real-world concept. Geometry is model-specific.
Vector Analogies
Classic word embeddings became famous for vector arithmetic where relationships could sometimes be approximated by differences between vectors. Examples such as king−man+woman ≈ queen illustrated regularities learned from text.
These demonstrations are interesting but should not be treated as universal symbolic rules. Results depend on model, vocabulary, training data and evaluation.
Modern contextual models also complicate the picture because representation changes with context.
Subspaces
A subspace is a subset of a vector space closed under vector addition and scalar multiplication. In machine learning, data may approximately occupy lower-dimensional structure within a much larger ambient space.
This motivates dimensionality-reduction methods and analysis of learned representations. However, learned semantic structure does not always form clean linear subspaces.
Projection
A projection maps a vector onto another direction or subspace. In simple 2D geometry, projecting onto the x-axis discards the y-coordinate.
Projection can deliberately keep some variation while discarding other variation. Neural networks use learned linear projections constantly, including query, key and value projections in transformers.
A projection is a transformation with a purpose; it is not automatically a visualization method.
Matrices Transform Vector Spaces
A matrix can scale, rotate, shear or project vectors. A neural-network linear layer multiplies an input vector by a weight matrix and often adds a bias.
This operation maps one representation space into another. If an input has 768 dimensions and the layer outputs 3,072 values, the matrix changes dimensionality as well as coordinates.
Deep networks repeatedly transform spaces so later representations become useful for the task objective.
A Worked Matrix Transformation
Let x=[2,1] and matrix W=[[1,0],[0,2]]. Multiplying W×x gives [2,2]. The first coordinate is unchanged; the second is doubled.
A learned weight matrix uses the same algebra but its values are optimised from data rather than chosen manually for a textbook transformation.
High-Dimensional Space Is Counterintuitive
In hundreds of dimensions, many intuitions from 2D fail. Volume concentrates differently, random vectors often become nearly orthogonal, and distance distributions can behave unexpectedly.
This is one reason visualising embeddings in 2D must be treated carefully. A plot is a projection of a much richer space.
The Curse of Dimensionality
As dimensionality increases, the volume of the space grows rapidly. Data becomes sparse relative to the ambient space, and some distance-based methods require more data or specialised indexing.
High dimension also increases memory and compute costs. Approximate nearest-neighbour search, dimensionality reduction and model-specific structure help manage the problem.
The phrase “curse of dimensionality” covers several related phenomena, not one universal failure law.
Concentration of Measure
In high-dimensional spaces, quantities such as norms and pairwise distances can concentrate around narrow ranges under certain distributions. This can make naive distance intuition less discriminative.
Learned embedding spaces are not random Gaussian clouds, but high-dimensional geometry still influences index design and analysis.
Anisotropy
Some learned embedding spaces are anisotropic: vectors occupy preferred directions rather than spreading uniformly. This can affect cosine-similarity distributions and retrieval thresholds.
Methods such as centering, whitening or task-specific training can sometimes improve representation geometry, but changes must be evaluated on the actual task.
Dimensionality Reduction
Dimensionality reduction maps high-dimensional vectors into fewer dimensions. Principal Component Analysis is a linear method that finds directions explaining substantial variance. t-SNE and UMAP are nonlinear methods often used for visualisation.
A 2D projection can reveal patterns but can also distort global distances or create apparent clusters. Never infer the full original geometry solely from the plot.
PCA Intuition
Imagine 3D points mostly lying near a tilted plane. PCA can find directions within that plane that capture most variation, allowing a 2D representation with relatively limited loss.
If important task distinctions lie in a low-variance direction, however, discarding that direction can hurt performance. Variance is not identical to task relevance.
A Worked Projection Warning
Suppose documents A and B are far apart in 768 dimensions but become neighbours after a 2D visualisation. The plot does not prove the original embeddings were close. The dimensionality-reduction algorithm changed geometry to display it.
Use original-space similarity for retrieval decisions unless the reduced representation was specifically trained and validated for that purpose.
Vector Space and Semantic Search
Semantic search encodes a query and corpus items into a vector space, then retrieves nearby items. The quality of search depends on whether that geometry aligns with relevance.
Sentence Transformers’ current semantic-search documentation describes embedding queries and corpus items into a shared space and finding nearest embeddings using similarity functions.
The database index accelerates this geometry; it does not define relevance by itself.
Vector Space and Classification
A classifier can learn a decision boundary in representation space. In a two-dimensional toy example, a line might separate two classes. In higher dimensions, the boundary becomes a hyperplane or nonlinear surface.
The network can also transform the space before classification so examples become easier to separate.
Vector Space and Attention
Transformer attention creates learned query, key and value vectors from hidden states. Query-key similarity determines attention weights, and weighted value vectors are combined.
The mechanism is therefore geometric: learned projections create spaces in which alignment controls information routing.
Article 017 will examine attention directly, but the vector-space foundation explains why dot products are central.
Vector Space and Neural Networks
Each neural-network layer maps one set of vectors into another. Activation functions add nonlinearity so a deep network can learn transformations far richer than one matrix multiplication.
Article 015 will build these transformations from neurons and layers. The vector-space view keeps the shapes and geometry visible.
Vector Space and Transformers
Transformers maintain a vector representation for each token position through many layers. Attention mixes information across positions, while feed-forward networks transform each position’s representation.
The sequence can therefore be viewed as a matrix: number of token positions × hidden dimension. Transformer blocks repeatedly update that matrix.
Vector Space Is Not Physical Space
Embedding distances are mathematical relationships learned for a computational objective. They do not imply that concepts literally occupy locations in the brain or world.
Spatial language—near, far, direction—is useful because the algebra is geometric. Keep the metaphor tied to the actual vector operations.
Similarity Is Not Causality
Two vectors can be close because their items appear in similar contexts. That does not show that one causes the other.
An embedding space built from news may place “rain” near “umbrella”. The relationship reflects patterns in data, not a causal model proving umbrellas produce rain.
Similarity Is Not Moral or Institutional Equivalence
A vector model can place two policies close because they discuss the same subject. One may be current and one superseded. One may apply to students and one to staff.
Operational systems need metadata and rules beyond geometry.
Nearest Neighbour Can Be the Wrong Answer
If a query asks for an exact identifier, nearest-neighbour search is unnecessary and potentially harmful. If the query asks for the current policy, semantic similarity alone cannot establish current status.
Choose exact lookup, filtering or vector retrieval according to the task.
A Complete Worked Vector-Retrieval Example
Corpus vectors represent four policy passages: P1 current camera rule, P2 archived camera rule, P3 laboratory hours and P4 field-trip transport. The query asks about camera borrowing.
Similarity ranks P2 first, P1 second, P4 third, P3 fourth. Pure vector ranking would return the archived rule first. Metadata says P2 status=archived and P1 status=current.
The application filters to current approved documents or reranks by authority, so P1 becomes the usable source. The model answers from P1 and can mention the field-project exception if retrieved.
The lesson is not that vector search failed. It correctly found semantically relevant camera policies. The complete task required a second dimension of authority that was not encoded in similarity alone.
Reality Check: High-Dimensional Geometry Is Useful but Not Self-Interpreting
- A 2D visualisation is a projection, not the original space.
- A close neighbour can contradict the query while remaining semantically related.
- A large similarity score has no universal meaning across embedding models.
- Dimensions are not automatically named concepts.
- Vector geometry should be evaluated against the downstream task rather than admired abstractly.
Vector-Space Failure Class 1: Wrong Metric
The embedding model expects dot-product retrieval, but the index uses Euclidean distance without normalisation. Relevant candidates move down the ranking.
Repair by following the model’s intended similarity function and re-evaluating retrieval.
Failure Class 2: Mixed Coordinate Systems
Half the corpus uses embedding model E1 and half E2. The vectors have equal dimension but incompatible geometry. Query distances become meaningless across groups.
Version and rebuild embeddings so compared vectors inhabit the same representation space.
Failure Class 3: Visualisation Overinterpretation
A 2D UMAP plot shows two clusters touching, and a team concludes the categories are indistinguishable. The original high-dimensional classifier performs well.
The conclusion came from the projection, not the task metric. Evaluate original-space performance.
Failure Class 4: Neighbourhood Without Metadata
The closest policy is archived. Add exact authority filters or reranking; do not ask geometry to infer an unavailable governance rule.
Failure Class 5: Threshold Copied Across Models
A similarity threshold of 0.75 works for E1. E2 has a different score distribution. Applying the same threshold causes many false rejections.
Calibrate thresholds per model and task.
Repair Pathway
When vector retrieval fails, inspect one query in original space. Confirm model version, vector dimension, metric, normalisation and metadata filters. Compare exact similarity scores for known relevant and irrelevant examples.
Only after the basic geometry is correct should you tune approximate indexes or add complex reranking.
Stabilisation Pathway
Build a labelled neighbourhood set containing obvious matches, hard negatives, exact-ID cases, old/current versions and multilingual queries. Track ranking quality through model and index changes.
Extension Pathway
Add hybrid lexical-vector search, learned reranking, multi-vector retrieval or dimensionality compression only when evaluation shows a bottleneck these techniques address.
Independent Exercise 1: Dot Product
What is the dot product of [2,3] and [4,-1]?
Answer
2×4 + 3×(-1) = 8−3 = 5.
Independent Exercise 2: Unit Vectors
Vectors A and B are both normalized to length one. What relationship exists between dot product and cosine similarity?
Answer
They are equal because the cosine denominator ||A||×||B|| equals one.
Independent Exercise 3: Similarity Versus Authority
An archived policy scores 0.91 and the current policy 0.88. Which should answer a current-policy question?
Answer
The current authoritative policy, assuming metadata confirms its status. Similarity ranking alone does not determine authority.
Independent Exercise 4: Projection
Two points overlap in a 2D PCA plot. Must their original embeddings be identical?
Answer
No. The projection discards dimensions. Different original vectors can map to similar or identical projected locations.
Independent Exercise 5: Exact Lookup
The user supplies document ID D913. Should the application search nearest vectors to find the document?
Answer
Use exact ID lookup. Vector search is for fuzzy similarity, not known identity.
Orthogonality: Perpendicular Directions in Representation Space
Two non-zero vectors are orthogonal when their dot product is zero. In 2D, [1,0] and [0,1] are the familiar perpendicular axes. In high dimensions, many directions can be mutually orthogonal even though we cannot draw them.
Orthogonality matters because a projection onto one direction can ignore components lying in orthogonal directions. It also appears in basis construction, numerical stability and analysis of representations.
Do not translate orthogonality directly into “unrelated meanings”. The mathematical relation is exact in the chosen coordinates; semantic interpretation depends on how the representation was trained.
Linear Independence
A set of vectors is linearly independent when none can be written as a linear combination of the others. Independent directions contribute new degrees of freedom.
If one vector is exactly twice another, the pair is dependent: [2,4] = 2×[1,2]. They lie on the same line through the origin.
Rank, basis size and dimensionality are connected to this idea. A representation matrix can have many columns while its data effectively varies in a smaller number of independent directions.
Rank: How Many Independent Directions Are Present?
The rank of a matrix is the number of linearly independent rows or columns. Consider three 2D vectors [1,0], [0,1] and [1,1]. There are three rows but only two independent directions, so the matrix rank is at most two.
Rank becomes relevant in neural networks because learned matrices can compress or expand representations, and because low-rank approximations can reduce parameter or compute cost.
A high-dimensional representation does not guarantee every dimension contributes independent information. Effective structure can be lower-dimensional.
Hyperplanes: Decision Boundaries in Higher Dimensions
A line can divide a 2D plane. A plane can divide 3D space. The general higher-dimensional analogue is a hyperplane, often defined by w·x + b = 0.
A linear classifier can assign one class when w·x+b is positive and another when it is negative. The vector w defines the boundary’s orientation; b shifts the boundary away from the origin.
Neural networks build nonlinear decision regions by composing many linear transformations with nonlinear activation functions. The linear hyperplane is the starting geometry.
A Worked Linear Classifier
Let w=[1,-1] and b=0. For point x=[3,1], score w·x=3−1=2, so it lies on the positive side. Point y=[1,3] gives 1−3=-2, so it lies on the negative side.
The boundary consists of points where x1=x2. In a learned classifier, training adjusts w and b so this boundary better separates labelled examples.
This example connects vector geometry to Article 015: a neuron can compute a weighted sum plus bias, then pass the result through an activation.
Affine Transformations: Linear Maps Plus a Shift
A neural-network layer typically computes Wx+b rather than only Wx. The bias b shifts the transformed space, so the operation is affine rather than strictly linear.
This matters because a pure linear map always sends the origin to the origin. Adding bias lets the model move decision boundaries and representation regions more flexibly.
Feature Directions Are Not the Same as Raw Dimensions
Researchers sometimes find directions in activation space correlated with a property. A direction can be a combination of many raw coordinates. This is different from claiming one individual dimension stores the property.
If direction v corresponds to a useful probe, moving along v changes the projection onto that learned feature direction. Whether the intervention causally changes model behaviour requires additional experiments.
Representation analysis should therefore distinguish correlation, decodability and causal control.
Eigenvectors and Principal Directions
For certain square matrices, an eigenvector keeps its direction under multiplication: Av = λv. The scalar λ is the eigenvalue. Eigenvectors identify directions transformed by simple scaling under that matrix.
This idea underlies Principal Component Analysis. PCA finds orthogonal directions that capture decreasing amounts of variance in a dataset after centering.
PCA does not know which variance matters for your task. It is a statistical geometry tool, not a semantic oracle.
A Worked PCA Intuition
Imagine student-performance vectors with coordinates [algebra score, geometry score, statistics score]. If all three scores tend to rise together, much of the dataset’s variance may lie along a direction resembling [1,1,1].
PCA could identify that broad “overall performance” direction as a principal component. A second component might capture algebra-versus-geometry differences.
Those interpretations are hypotheses drawn from loadings and data patterns. PCA itself returns directions of variance, not named educational constructs.
Covariance Connects Dimensions
Covariance measures how two variables vary together. A covariance matrix summarises pairwise relationships across dimensions. PCA can be derived from eigenvectors of a covariance matrix under standard formulations.
Learned embedding dimensions are not independent input features, but covariance and related tools can still reveal dominant geometric directions.
Whitening and Centering
Centering subtracts the dataset mean so the cloud is arranged around the origin. Whitening additionally transforms coordinates so selected dimensions have standardised variance and reduced correlation.
These operations can change similarity geometry. They are sometimes used in representation analysis or retrieval pipelines, but they should be validated because the embedding model may have been trained expecting its original distribution.
K-Means Clustering
K-means assigns points to k clusters by alternating between two steps: assign each point to the nearest centroid, then recompute each centroid as the average of its assigned points.
The objective minimises within-cluster squared Euclidean distance under standard k-means. This means the algorithm assumes a particular geometry and roughly compact clusters.
If semantic categories have irregular shapes or different densities, k-means can produce misleading partitions.
A Worked K-Means Toy Example
Points A=[0,0], B=[0.1,0], C=[5,5], D=[5.1,5]. With k=2 and sensible initial centroids, A/B form one cluster and C/D another. The centroids converge near [0.05,0] and [5.05,5].
Now add point E=[2.5,2.5]. It is equally distant from the two centroids. A tiny perturbation or initialization choice can determine its assignment. Clustering boundaries are algorithmic choices, not discovered natural laws.
Outliers
An outlier lies unusually far from the main data cloud under some metric. It can represent an error, rare but valid case, attack, new distribution or novel category.
Outlier detection therefore requires domain interpretation. Automatically deleting every distant point can erase exactly the cases a safety system needs to notice.
Local Versus Global Geometry
An embedding can preserve local neighbourhoods well without preserving large-scale global distances. A retrieval system may only care that the right passage appears among nearby candidates.
Visualisation methods such as t-SNE deliberately emphasize local neighbourhood structure and can distort distances between clusters. UMAP makes different assumptions and trade-offs.
Always match geometric analysis to the property the method preserves.
Manifold Intuition
Although a representation may live in a 1,024-dimensional ambient space, meaningful data can concentrate near a lower-dimensional curved structure, often described informally as a manifold.
This idea motivates dimensionality reduction and local representation analysis. Real learned spaces are complex and need not form one smooth, clean manifold.
Batch Similarity Is Matrix Multiplication
Suppose Q is a matrix whose rows are query embeddings and D a matrix whose rows are document embeddings. If vectors are suitably arranged, a batch of dot-product similarities can be computed with Q×Dᵀ.
This is one reason vector retrieval maps efficiently to modern accelerators: many similarity comparisons become large matrix operations.
Index structures further reduce how many candidates need exact scoring at large scale.
Approximate Nearest-Neighbour Search
Exact nearest-neighbour search compares a query with every stored vector. For millions or billions of vectors, this can be too slow or costly.
Approximate nearest-neighbour methods build indexes that search a much smaller candidate region, trading some recall for large speed gains. Graph-based indexes such as HNSW are a common family.
The application should evaluate recall and latency together. A fast index that consistently drops the only relevant passage is not useful.
HNSW Intuition
Hierarchical Navigable Small World indexes organise vectors into a graph with multiple navigation layers. Search starts from sparse high-level connections and descends toward denser local neighbourhoods.
The method avoids scanning every point. Parameters control construction cost, memory, query speed and recall.
The embedding model and HNSW index solve different problems: one creates geometry, the other navigates it efficiently.
Quantised Vector Search
Vector indexes can use compressed or quantised representations to reduce memory and speed search. Product quantisation and scalar quantisation are examples.
Compression changes distances approximately, so top candidates can shift. Many systems use compressed search to find candidates and then rescore a smaller set with full-precision vectors.
Vector Space and Hybrid Retrieval
Keyword retrieval and vector retrieval create different score spaces. Hybrid search needs a method to combine or rerank them rather than directly adding incomparable raw scores without calibration.
A system can use reciprocal-rank fusion, learned ranking or normalized score combinations. The correct method depends on evaluation.
Vector Arithmetic Is Coordinate-System Dependent
If two embedding models learn different rotated versions of the same conceptual geometry, raw vector coordinates cannot be compared across models even when neighbourhood structure is similar.
This reinforces the versioning rule: same dimension does not mean same space.
Alignment Between Spaces
Sometimes systems learn a mapping between two vector spaces—for example, aligning representations from two languages or modalities. A linear transformation can map one coordinate system toward another if enough paired examples support the relation.
Cross-modal systems can therefore place images and text into a shared or aligned space where similarity enables text-to-image retrieval.
Residual Streams as Evolving Vector Spaces
In decoder-only transformers, each token position carries a hidden-state vector through many blocks. Residual connections add outputs from attention and feed-forward sublayers back into that evolving representation.
The result is not one fixed semantic vector space. Each layer defines a different stage of representation with its own geometry.
Query, Key and Value Spaces
Attention applies learned linear projections to hidden states to create queries, keys and values. These vectors inhabit learned spaces with different operational roles.
Query-key dot products determine attention scores; value vectors carry information that is mixed according to those scores.
The fact that all are vectors does not make them interchangeable. Their learned projections define their function.
A Worked Attention Geometry Preview
Assume query q=[1,0]. Key k1=[0.9,0.1] gives dot product 0.9. Key k2=[0.1,0.9] gives 0.1. Before softmax scaling, k1 aligns much more strongly with q.
Attention converts these scores into weights and uses them to combine value vectors. Article 017 will derive the full mechanism.
Decision Boundaries Can Be Nonlinear After Layering
One affine layer plus a monotonic activation has limited geometry. Stacking layers lets networks bend and partition representation space into complex regions.
This is why deep networks can separate patterns that are not linearly separable in the original input space: earlier layers learn transformations that make later decisions easier.
XOR: Why One Linear Boundary Is Not Enough
The XOR pattern assigns one class to [0,0] and [1,1], and the other class to [0,1] and [1,0]. No single straight line separates the classes in 2D.
A small neural network with a hidden nonlinear layer can transform the points so the final class becomes linearly separable. This classic example motivates nonlinear representation learning.
Vector Norms and Numerical Stability
Very large or very small activation magnitudes can make optimisation difficult. Normalisation techniques help networks control activation statistics.
Layer Normalization, used prominently in transformers, normalizes features within a representation according to model-specific formulas and learned scale/shift parameters.
Normalization changes geometry and optimisation behaviour; it is not simply cosmetic rescaling.
Vector Space and Bias
Embedding geometry reflects patterns in training data and objectives, including social associations and representation imbalances. A vector space can therefore encode undesirable correlations.
Debiasing one geometric direction does not guarantee removal of all downstream bias. Evaluation must examine actual system behaviour on relevant populations and tasks.
Vector Space and Privacy
Representations can preserve information about source data. A vector should not be treated as anonymous merely because it is high-dimensional and not human-readable.
Protect vector stores according to the sensitivity of the represented material and the possibility of linking vectors to source objects.
A Vector-Space Audit Worksheet
- What objects are represented as vectors?
- Which model and version created the vectors?
- How many dimensions does the space have?
- Which similarity or distance metric is intended?
- Are vectors normalized?
- Are query and corpus vectors produced in compatible modes?
- Which metadata constraints sit outside the geometry?
- Is search exact or approximate, and what recall is measured?
- Are 2D visualisations being mistaken for original-space evidence?
- What happens after the embedding model or index version changes?
Independent Exercise 6: Hyperplane
For classifier score s=2×1−x2−1, what is the score at x=[2,1], and which side of the boundary is it on?
Answer
s=2×2−1−1=2. It lies on the positive side because the score is greater than zero.
Independent Exercise 7: Rank
Vectors [1,0], [2,0] and [0,1] live in 2D. How many independent directions do they contain?
Answer
Two. [2,0] is a scalar multiple of [1,0], while [0,1] adds a second independent direction.
Independent Exercise 8: Approximate Search
An ANN index answers in 5 ms but retrieves the known relevant document in only 85% of test queries. An exact scan answers in 500 ms and retrieves it 100%. Which is better?
Answer
The answer depends on the application’s latency and recall requirements. The trade-off must be evaluated; neither number alone determines the correct design.
Independent Exercise 9: Visualisation
A t-SNE plot places two documents far apart. Can you conclude their cosine similarity in the original embedding space is low?
Answer
No. t-SNE transforms geometry for visualisation and does not preserve every global distance. Compute similarity in the original representation space.
Metric Calibration: Similarity Scores Need Task-Specific Meaning
A vector-search score becomes useful only after it is connected to outcomes. Suppose relevant query-document pairs in one evaluation mostly produce cosine similarities between 0.72 and 0.91, while hard negatives frequently reach 0.76. A threshold of 0.70 would keep many relevant items but also admit difficult negatives. A threshold of 0.80 could discard legitimate matches.
The right operating point depends on whether the system retrieves top-k candidates for a later reranker, directly accepts matches, or escalates uncertain cases. This is analogous to Article 005: a numerical score is not a complete decision policy.
Record score distributions when changing embedding model, normalization or index. A threshold calibrated for one vector space should not be transferred blindly to another.
Recall at k and Precision at k
Retrieval evaluation often asks whether relevant material appears within the first k results. If a known relevant passage appears in the top five for 92 of 100 queries, recall-at-5 for a one-relevant-document setup is 92%.
Precision at k asks how many returned items are relevant. A system can achieve high recall by returning many candidates, but this can flood the language model with noise. Retrieval design balances candidate coverage against context quality and cost.
These metrics are simplifications when queries have multiple graded relevant documents, but they illustrate why vector quality must be measured through retrieval outcomes rather than geometric elegance alone.
Reranking Changes the Effective Geometry
A first-stage vector index may retrieve twenty semantically plausible passages. A cross-encoder reranker then reads each query-passage pair jointly and assigns a new relevance score. The final ranking is no longer determined solely by nearest-neighbour geometry.
This is often a strength. Fast embeddings provide broad recall; the more expensive reranker makes finer pairwise distinctions. When diagnosing a bad result, inspect both stages separately.
The Complete Geometry-to-Answer Chain
For a grounded answer, the full path is: text or object → embedding model → vector space → similarity metric → vector index → candidate set → metadata filters → optional reranker → context assembly → language model → verification. Every arrow can fail independently.
The vector space is therefore a middle layer, not the whole SI system. It is powerful because it makes learned similarity computationally accessible. It becomes dependable when exact identity, authority, permissions and task evaluation surround that geometry.
Independent Exercise 10: Threshold Transfer
Embedding model E1 uses a relevance threshold of 0.78. A new model E2 replaces it. Should 0.78 remain automatically?
Answer
No. Measure E2’s score distribution and retrieval outcomes. Different models can produce different geometries and similarity ranges even on the same corpus.
Frequently Asked Questions About Vector Space in AI
What is a vector space?
A mathematical set of vectors with defined addition and scalar multiplication. In AI, learned representations are often treated as points or directions in high-dimensional vector spaces.
What is a dimension?
One coordinate direction in the vector representation. In learned spaces it usually does not correspond to one simple human-readable feature.
What is a vector norm?
A measure of vector magnitude. The Euclidean norm is the square root of the sum of squared coordinates.
What is cosine similarity?
A measure of directional alignment between vectors, calculated from dot product and vector lengths.
What is Euclidean distance?
Straight-line distance between two points in the coordinate space.
Why use high dimensions?
They provide capacity for complex learned representations, though they increase compute/storage and create geometric challenges.
Can I understand embeddings by plotting them in 2D?
Plots can help explore patterns, but dimensionality reduction distorts some relationships. Validate conclusions in the original space.
Is nearest neighbour always the most relevant item?
No. It is nearest according to the chosen representation and metric. Authority, exact identity and task constraints can require additional logic.
Do vectors have meaning without a model?
Their coordinates and useful geometry are defined by how they were produced. A raw vector without model/version/context metadata is difficult to interpret reliably.
Selected Technical References
- Efficient Estimation of Word Representations in Vector Space.
- Sentence Transformers Semantic Search.
- Sentence Transformers Quickstart.
Vector Space Gives Learned Representations a Geometry
Embeddings become useful together because they inhabit a space where models and retrieval systems can compute relationships. Addition, scaling, dot products, angles, distance, projections and neighbourhoods form the mathematical language of that geometry.
The geometry is powerful precisely because it is continuous and learnable. It also has limits. Similarity is not authority, identity, truth or causality. High-dimensional spaces resist simple visual intuition, and every score belongs to a specific model and metric.
Next: 015 — Neural Networks, where we move from representing points in a vector space to learning functions that transform those vectors through layers of weights, biases and nonlinear activations.
How Super Intelligence Works Series Navigation
Previous: 013 — Embeddings · Series Hub · Next: 015 — Neural Networks.
