VIEW THIS AS

Auto mode follows the Route Engine until you choose a viewpoint.

YOU ARE HERE

ROUTE CHECK

CONNECTED TO

WHAT NEXT

Use the canonical route for this room, or HELP if you are unsure.

How Super Intelligence Works | Embeddings — Turning Tokens, Sentences and Documents Into Learned Vectors

eduKate Secondary students reviewing open books for How Super Intelligence Works: Embeddings.

Embeddings are learned numerical representations that let Super Intelligence systems compute with tokens, words, sentences, documents, images and other objects. A token ID is only a discrete label. An embedding turns that label—or a larger piece of information—into a vector of numbers whose position can be useful to a model.

This transition is fundamental. Neural networks cannot directly multiply, add or compare the human meaning of “algebra”, “equation” and “variable”. They operate on numbers. Embeddings provide a learned numerical interface between discrete identities and continuous computation.

This guide explains how embeddings work from first principles. We will follow a token ID into an embedding matrix, build small toy vectors by hand, compare dot product and cosine similarity, distinguish token embeddings from contextual representations and sentence embeddings, and show how embeddings power semantic search, retrieval, clustering and recommendation-like systems.

In the eduKateSG series, Super Intelligence, or SI, is our umbrella term for AI technologies. Standard terms such as embedding vector, embedding matrix, dimension, dot product, cosine similarity, semantic search, contextual representation and vector database remain visible so readers can connect this explanation to current technical work.

Previous: 012 — Tokenisation. For the broader representation stack, return to the How Super Intelligence Works hub.


Embeddings at a Glance

  • A token ID is an index; an embedding is a learned vector.
  • An embedding matrix contains one learned vector per vocabulary item in a standard token-embedding layer.
  • Vectors live in a multidimensional coordinate space where useful relationships can emerge through training.
  • Similarity is task-dependent: cosine similarity, dot product and Euclidean distance answer different geometric questions.
  • Token embeddings inside a language model are not the same thing as sentence or document embeddings built for retrieval.
  • Contextual transformer representations change with surrounding tokens; the same token can be represented differently in different sentences.
  • Vector similarity retrieves candidates; it does not prove equivalence, truth, authority or permission.

The Hidden Transition: Token ID to Learned Vector

Article 011 introduced token IDs as discrete vocabulary labels. Suppose our toy vocabulary assigns token ID 42 to “ algebra”. The integer 42 does not contain algebraic meaning. It simply identifies which vocabulary item is present.

The model needs a continuous numerical representation it can combine with other representations. An embedding layer performs a lookup: token ID 42 selects row 42 from an embedding matrix. That row might contain hundreds or thousands of learned floating-point values in a real model.

For teaching, imagine a three-dimensional embedding for “ algebra”: [0.8, -0.2, 0.5]. Another token, “ equation”, might be [0.7, -0.1, 0.6]. These numbers are invented. They let us see the mechanism without pretending a real model has only three dimensions or that any individual coordinate has a simple human label.

The vector becomes input to later neural-network computation. During training, gradients can update the embedding values together with many other model parameters so the representations become useful for the model’s objective.

The Embedding Matrix

If a model has a vocabulary of V tokens and an embedding dimension of D, a simple token-embedding matrix has shape V × D. Each row corresponds to one token ID. Looking up several token IDs returns a sequence of vectors.

A toy vocabulary with 10,000 tokens and 256-dimensional embeddings would have 2.56 million values in that matrix. Large language models can use much larger vocabularies and hidden dimensions. The exact architecture varies, but the lookup principle is straightforward.

The embedding matrix is learned, not manually filled with dictionary definitions. Training adjusts numerical values because certain configurations help reduce the model’s loss. Semantic-looking structure can emerge from that optimisation.

A Toy Embedding Table We Can Inspect

Imagine four words represented in two dimensions: cat = [0.9, 0.8], dog = [0.8, 0.9], banana = [-0.8, 0.2], apple = [-0.7, 0.3]. We deliberately choose these values so animal words are near each other and fruit words are near each other.

Plotting these points would place cat and dog in one region, apple and banana in another. A distance or similarity function can turn that visual intuition into a number.

This is a toy construction, not a claim about any production embedding model. Real learned spaces contain many dimensions, and relationships can be distributed across those dimensions in ways that do not map cleanly onto one human concept per axis.

Why Continuous Vectors Are Useful

Discrete IDs are exact but isolated. Token 42 and token 43 are simply different indices. Their numerical difference of one has no semantic significance. Continuous vectors allow graded relationships: two points can be close, distant, aligned or opposed according to a chosen geometry.

This makes it possible to generalise. A model can learn patterns involving related items without requiring a separate hard-coded rule for every token pair. Retrieval systems can find documents whose embeddings are close to a query even when exact keywords differ.

Continuous representation does not eliminate discrete identity. A database still needs exact document IDs. A tokenizer still needs vocabulary IDs. The vector layer adds a learned geometry on top of those identities.

The Word2Vec Milestone

The 2013 paper Efficient Estimation of Word Representations in Vector Space by Mikolov and colleagues helped popularise efficient learned word vectors at large scale. The work demonstrated that trained continuous word representations can capture useful syntactic and semantic regularities.

Word2vec-style models are not the same as modern transformer embeddings, but they provide a clean historical bridge: instead of representing every word as an isolated one-hot category, learn vectors whose geometry reflects statistical patterns in language.

The deeper idea survives across modern systems: useful representations can be learned from an objective rather than manually specifying every feature.

One-Hot Vectors Versus Dense Embeddings

A one-hot vector for a vocabulary of 50,000 items has 50,000 positions with one 1 and the rest 0. It preserves exact identity but gives every pair of different items the same simple orthogonality. “cat” is no closer to “dog” than to “spreadsheet”.

A dense embedding uses far fewer dimensions than the vocabulary size and stores learned real-valued coordinates. The representation sacrifices the trivial interpretability of one-hot identity to gain a geometry useful for computation.

One-hot vectors remain useful conceptually because they show that identity and similarity are different questions. Embeddings add similarity structure learned from data and objectives.

Embedding Dimensions Are Not Usually Human-Labeled Features

It is tempting to look at a vector like [0.8, -0.2, 0.5] and say dimension one means “mathematical”, dimension two means “difficulty”, and dimension three means “school subject”. Real learned embeddings generally should not be interpreted that way without specific evidence.

Meaning is distributed. A concept can depend on patterns across many dimensions, layers and contexts. Rotating a vector space can preserve geometric relationships while completely changing individual coordinates.

Therefore, treat the vector as a computational representation, not a spreadsheet of named psychological traits.

Dot Product: Alignment With Magnitude

For vectors a and b, the dot product multiplies corresponding coordinates and sums the results. For a = [1, 2] and b = [3, 4], the dot product is 1×3 + 2×4 = 11.

A larger dot product can reflect both directional alignment and vector magnitude. Neural networks use dot products constantly because matrix multiplication is built from the same operation.

Attention mechanisms also rely on dot-product-style comparisons between learned query and key vectors. The next articles will build that connection carefully.

Cosine Similarity: Compare Direction

Cosine similarity divides the dot product by the product of vector lengths. The result focuses on angle rather than absolute magnitude. Identical directions have cosine similarity 1, orthogonal directions 0, and opposite directions -1 in the mathematical definition.

Sentence Transformers’ current semantic search documentation describes embedding queries and corpus items into a shared vector space and finding nearby embeddings, with cosine similarity used by its semantic_search utility by default.

Cosine similarity is common in semantic search, but it is not universally best. Some embedding models are trained expecting dot product or another scoring function. Use the similarity measure appropriate to the model and task.

A Worked Cosine-Similarity Example

Let vector A = [1, 0] and vector B = [0.8, 0.6]. The dot product is 0.8. The length of A is 1. The length of B is √(0.64 + 0.36) = 1. Therefore their cosine similarity is 0.8.

Now let C = [-1, 0]. A · C = -1, and both have length 1, so cosine similarity is -1. In this toy geometry, C points exactly opposite A.

Real embedding scores need empirical interpretation. A cosine similarity of 0.8 does not universally mean “80% semantically equivalent”. Thresholds vary by model, domain and task.

Euclidean Distance: Straight-Line Separation

Euclidean distance asks how far apart two points are in the usual geometric sense. Between [1,0] and [0.8,0.6], the distance is √((0.2)² + (-0.6)²), about 0.632.

If vectors are normalised to unit length, cosine similarity and Euclidean distance become closely related. Without normalisation, magnitude matters differently.

The important lesson is that “nearest” depends on the metric and representation. Vector search is not one universal geometry.

Normalisation Changes the Geometry

Normalising a vector usually means scaling it to length one. Its direction remains while magnitude is removed. When embeddings are normalised, dot product becomes equivalent to cosine similarity.

Some retrieval pipelines store normalised embeddings because it simplifies similarity computation. Others preserve magnitude because the model uses it meaningfully.

Do not normalise automatically without checking the model’s intended scoring method.

Token Embeddings Are Context-Free at the Lookup Step

A basic token-embedding lookup assigns the same initial learned vector to token ID 42 whenever it appears. At that first lookup stage, “bank” starts from the same token embedding whether the sentence concerns money or a river.

Context enters through later network computation. Attention and feed-forward layers transform token representations using surrounding information. After several layers, the representation associated with “bank” can differ greatly between contexts.

This distinction—static lookup embedding versus contextual hidden state—is central to understanding modern transformers.

Contextual Representations

Take “The bank approved the loan” and “We sat on the river bank.” The token or subword for “bank” begins with the same vocabulary embedding in a standard setup, but transformer layers process different neighbours.

The hidden representation at a later layer can encode different task-relevant relationships. That is why modern language models are not simply bags of static word vectors.

Calling every hidden state an “embedding” is common but can blur important distinctions. Ask which layer, which pooling rule and which training objective produced the vector.

Positional Information Joins Token Information

Token identity alone does not tell the model whether one word came before another. Transformer architectures therefore include positional information through learned or fixed mechanisms, depending on the model.

The representation entering a transformer block reflects both token identity and position. Later articles will examine these mechanisms more closely.

For now, remember that the initial token embedding is one ingredient, not the entire model input representation.

Sentence Embeddings Compress Variable-Length Text

A sentence embedding represents a whole sentence with a fixed-size vector. The system must decide how to transform multiple token-level states into one vector, often through pooling, a special representation, a trained projection or an architecture designed specifically for sentence-level encoding.

Sentence Transformers’ current quickstart describes bi-encoder models that compute fixed-size embeddings and support semantic similarity, semantic search, clustering and related tasks.

This is different from taking one arbitrary token embedding and calling it the meaning of the whole sentence.

Document Embeddings

Documents can also be represented as vectors. A long document may first be chunked, encoded, pooled or represented with multiple vectors depending on the retrieval system.

A single vector is compact and fast to compare, but compression can lose detail. Multi-vector and late-interaction approaches preserve more token-level information at additional cost.

The right representation depends on the retrieval problem. A short FAQ search and a legal-document retrieval system do not necessarily need the same embedding architecture.

Semantic Search: Query and Corpus in One Space

Semantic search encodes a query and candidate documents into vectors, then retrieves nearby candidates. This allows matching beyond exact words. “How do I reset my password?” can retrieve a document titled “Recover account access” if the embedding model learned those meanings as related.

Sentence Transformers distinguishes symmetric tasks—such as finding paraphrased questions—from asymmetric tasks where a short query retrieves a longer passage. Its current documentation recommends task-appropriate query and document encoding methods for retrieval models.

This is a practical reminder that “an embedding” is not a universal representation independent of task. Training objective matters.

Worked Semantic Search Example

Corpus A: “Students may borrow cameras for three school days.” Corpus B: “The laboratory closes at 5 PM.” Corpus C: “A teacher can authorise a longer field-project camera loan.” Query: “Can I keep a camera longer for a project?”

A good retrieval embedding should place the query near C and possibly A, while B should be less similar. The exact scores depend on the model.

A retrieval system can return A and C for the language model to interpret together. Vector similarity helps discover evidence; the policy status and final conclusion still require metadata and reasoning.

Embedding Search Is Not Truth Search

If an archived policy and a current policy are semantically similar, both can be close to the query. Vector geometry does not know which document is authoritative unless the training or metadata pipeline provides that distinction.

Therefore, combine semantic retrieval with exact filters: approved = true, organisation = correct tenant, date within range, access permission satisfied.

This is the same lesson established in Article 009: meaning can be fuzzy while authority and identity remain exact.

Embedding Similarity Is Not Logical Equivalence

“Students may borrow cameras for three days” and “Students may borrow cameras for five days” can have very high semantic similarity because almost every word is shared. Yet the statements conflict on the fact that matters.

A similarity model is behaving reasonably if it recognizes that the sentences discuss the same topic. It is not a contradiction detector unless trained and evaluated for that task.

Use similarity to find candidates, then use task-appropriate reasoning or verification to judge agreement, contradiction or authority.

Embedding Similarity Is Not Identity

Two different records can have identical descriptions. Two copies of a policy can be semantically indistinguishable but have different version IDs. Do not substitute nearest-neighbour search for exact record identity when the identifier is known.

A vector database should usually store metadata or references that map the embedding back to an authoritative object.

Embedding Similarity Is Not Permission

A private student note can be semantically perfect for a query and still be unauthorised. Access control must operate independently of similarity.

The retrieval service should filter by the user’s permissions before or alongside vector ranking. A model should not see every candidate and be asked to “ignore the private ones”.

Training Makes the Geometry Task-Sensitive

Embeddings become useful because their training objective rewards certain relationships. Word2vec used local context prediction objectives. Sentence embedding systems can use pairs, triplets, contrastive losses or retrieval objectives.

A model trained to cluster paraphrases may not be optimal for product recommendations. A model trained on code may represent programming relationships better than a general sentence model.

Evaluate embeddings on the task they will actually support.

Positive and Negative Pairs

Contrastive training often uses examples that should be close and examples that should be farther apart. A query and its relevant passage form a positive pair. An irrelevant passage forms a negative example.

Hard negatives—items that look similar but are wrong—can be particularly valuable. For the camera policy, a superseded five-day rule is a harder negative than an unrelated cafeteria notice.

This connects representation learning directly to retrieval quality: the examples define what distinctions the geometry should learn to preserve.

Embedding Drift Across Model Versions

If an organisation changes embedding models, stored vectors from the old model generally should not be assumed comparable to new query vectors. The coordinate systems can differ even when dimensions match.

A migration may require re-embedding the corpus and rebuilding indexes. Version the embedding model alongside stored vectors.

A silent model change can create retrieval regressions that look like search problems rather than representation-version problems.

Dimension Count Does Not Equal Intelligence

A 3,072-dimensional embedding is not automatically more meaningful than a 768-dimensional one. Higher dimension increases representational capacity and storage cost, but task performance depends on architecture, data, training objective and evaluation.

Some embedding systems can truncate or project dimensions with controlled quality trade-offs. Compare real retrieval metrics rather than ranking models by vector length.

Storage Cost of Embeddings

A corpus of one million vectors with 1,024 float32 dimensions contains roughly 1.024 billion float values, requiring about four gigabytes just for raw vector values before index and metadata overhead. This is an illustrative arithmetic estimate.

Compression, lower-precision formats and specialised indexes can reduce memory. The relevant trade-off includes recall, latency, storage and update cost.

Embedding architecture therefore connects to infrastructure decisions later in the SI stack.

Approximate Nearest Neighbour Search

For small corpora, an application can compare a query vector with every candidate. At large scale, exact comparison becomes expensive. Approximate nearest-neighbour indexes trade some exactness for much faster search.

The vector database or search engine handles this index; the embedding model produces the vectors. These are separate components with separate failure modes.

A retrieval error can come from poor embeddings, a poorly configured index, metadata filtering, or simply asking for too few candidates.

Reranking After Embedding Retrieval

A common architecture uses embeddings to retrieve a candidate set quickly, then applies a more expensive reranker to those candidates. Sentence Transformers’ current documentation describes bi-encoder embeddings as an efficient first stage and cross-encoders as a possible reranking stage.

This architecture acknowledges that one fixed-size vector may be excellent for broad retrieval while a pairwise model can make finer distinctions among top candidates.

Embeddings in Recommendation-Like Systems

Users, products, lessons or content items can be represented in vector spaces so nearby items reflect learned affinity. A learner who performs similarly across certain skills might be matched with appropriate practice material.

Such systems require careful evaluation because similarity can reproduce biases in data or correlate with sensitive attributes. The geometry is learned from examples; it does not arrive value-neutral.

Embeddings Beyond Text

Modern systems can embed images, audio and other modalities. Some models map different modalities into related spaces, allowing text-to-image retrieval or image similarity.

The word “embedding” therefore refers to a general representation concept, not only word vectors. Always ask what object was embedded and what training objective makes distances meaningful.

A Complete Worked Retrieval Pipeline

Task: answer “How long can a camera be borrowed for a field project?” The system has ten thousand school documents. Step one: split approved documents into retrievable passages while preserving source IDs and status metadata.

Step two: encode each passage with the chosen document embedding model and store vectors with metadata. Step three: encode the user query using the model’s query route. Step four: retrieve the top candidates by similarity while filtering to current approved documents.

Step five: rerank or inspect the top passages. Step six: supply the relevant passages to the language model. Step seven: generate the answer and cite the authoritative passage. Step eight: verify that the cited text supports the conclusion.

Embeddings solve candidate discovery. They do not replace document governance, reasoning, citation support or human authority.

Failure Class 1: Wrong Embedding Model for the Task

A generic similarity model may group passages by topic but fail to distinguish the exact relevance needed for question answering. The solution is not to tune the vector database first if the representations themselves do not rank relevant examples well.

Create an evaluation set of queries and known relevant passages. Test the embedding model directly.

Failure Class 2: Query and Document Encoded Inconsistently

Some retrieval models use different prompts or task routes for queries and documents. If the application encodes everything with the wrong mode, retrieval quality can drop.

Sentence Transformers’ current API exposes encode_query and encode_document specifically for this reason in models that define different query/document prompts or routing.

Failure Class 3: Stale Embeddings

A document changes, but its stored embedding is not refreshed. Search can continue retrieving a vector representing old content while the displayed document is new.

Version embeddings with source content and re-embed after meaningful updates.

Failure Class 4: Metadata Filter Missing

The relevant archived policy scores highest and is returned because the vector search ignored approved/current status. This is not primarily a geometry failure; exact metadata filtering was missing.

Failure Class 5: Chunk Representation Too Broad

A long chunk contains several unrelated topics. Its embedding averages or compresses multiple meanings, making retrieval less precise. Smaller or structure-aware chunks can improve candidate quality.

Failure Class 6: Similarity Threshold Misused

A team chooses 0.8 as a universal relevance threshold because it sounds high. Another embedding model produces a different score distribution, and useful results disappear.

Thresholds must be evaluated for the specific model and task rather than copied across spaces.

Reality Check: Embeddings Do Not Contain Little Dictionary Entries

Embeddings are learned coordinates useful for computation. They are not miniature definitions stored inside dimensions, and similarity is not a universal semantic truth score.

  • A close vector can still correspond to a contradictory statement.
  • A distant vector can still be relevant under a specialised task the embedding model was not trained for.
  • Embedding quality depends on training objective and evaluation domain.
  • Vector search retrieves candidates; it does not determine authority or permission.
  • Changing embedding models changes the coordinate system and can require re-embedding stored content.

Repair Pathway

When semantic search fails, first isolate the representation. Use a small query set with known relevant passages. Compare similarity rankings before involving the language model. If the correct passage is not near the top, investigate model choice, chunking and query/document encoding.

If the right candidate ranks well in an exact scan but disappears in production, inspect the vector index, approximate-search settings and metadata filters.

Stabilisation Pathway

Build retrieval regression tests containing ordinary queries, synonyms, abbreviations, hard negatives, conflicting versions and multilingual examples relevant to your users. Track recall at k, ranking quality and downstream answer support.

Re-run the set after embedding-model, chunking, index or source changes.

Extension Pathway

Once first-stage retrieval is stable, add reranking, hybrid keyword-plus-vector search, query rewriting or multi-vector representations when measurements justify the complexity.

Every extension should preserve source identity and permissions so improved relevance does not widen access or weaken provenance.

Independent Exercise 1: Token ID Versus Embedding

Token ID 500 corresponds to “geometry”. Is 500 itself the learned semantic representation?

Answer

No. The ID is a vocabulary index. The embedding layer uses that ID to select a learned vector.

Independent Exercise 2: Similar but Contradictory

Two policy sentences differ only in “three days” versus “five days”. Their embeddings are very close. Has the embedding model failed?

Answer

Not necessarily. The sentences are topically and semantically similar. A retrieval model can reasonably place them close. A later task must inspect the factual difference and version authority.

Independent Exercise 3: Model Upgrade

A corpus was embedded with model E1. The application switches query encoding to E2 without rebuilding the corpus. Both output 768-dimensional vectors. Is comparison safe?

Answer

Not automatically. Matching dimension count does not mean the coordinate systems align. Re-embed the corpus or use a documented compatible migration path.

Independent Exercise 4: Exact Asset ID

The user asks for asset C07. Should the system select the nearest embedding among asset descriptions?

Answer

No when the exact identifier is known. Use exact lookup for identity. Embeddings are useful for fuzzy discovery such as “a camera suitable for field work”.

Independent Exercise 5: Query and Document Roles

An embedding model was trained with different query and document instructions. The application uses document encoding for both. What should be tested?

Answer

Re-encode queries using the intended query route and measure retrieval. The representation mismatch may be the first unstable point.


A Complete Embedding-Matrix Training Example

Imagine a vocabulary with four tokens: “cat”, “dog”, “apple” and “banana”. Give each token a two-dimensional trainable embedding. At the start of training the vectors may be essentially arbitrary: cat = [0.10, -0.30], dog = [-0.20, 0.05], apple = [0.40, 0.10], banana = [-0.15, -0.25]. These numbers are invented.

Now suppose the training task repeatedly rewards the model for predicting similar contexts around cat and dog, and different contexts around the animal words and fruit words. Gradients flow through the network and into the embedding matrix. The vectors move because changing them helps lower the training loss.

After many updates our toy table might look like cat = [0.85, 0.70], dog = [0.78, 0.74], apple = [-0.70, 0.25], banana = [-0.74, 0.20]. No human manually told dimension one to mean animal. The geometry emerged because the training objective rewarded representations useful for predicting the observed data.

The same principle scales. In a transformer, token embeddings are only one set of parameters among many. Their values are learned jointly with attention projections, feed-forward networks and other components. The final behaviour therefore cannot be understood by inspecting the input embedding table alone.

Embedding Lookup Is Differentiable Through the Selected Rows

The lookup operation itself selects rows by token ID, but the selected embedding values participate in differentiable computation. When the loss gradient reaches those values, optimisation can update the corresponding rows.

A token that appears frequently may receive many updates across varied contexts. A rare token may receive fewer direct updates. Subword tokenisation helps rare words because their pieces can share representations with many other strings.

This creates a direct bridge between Article 012 and embeddings: tokenisation determines which discrete IDs exist; the embedding matrix determines their initial learned continuous representations.

Static Word Vectors and Contextual Transformer States Are Different Products

Classic word-vector systems such as word2vec assign one learned vector to a word type. “Bank” has one vector even though it can mean a financial institution or the side of a river. The surrounding statistics shape that vector toward an average of its uses.

A transformer still begins with token embeddings, but later hidden states depend on context. The representation of “bank” after attention in “the bank approved the loan” can differ from the representation after attention in “the canoe reached the bank”.

This means a modern language model has many layers of representation. The input embedding is the starting coordinate. Contextual hidden states are transformed coordinates conditioned on the sequence. A sentence embedding is yet another representation, usually constructed for a particular downstream objective.

Pooling: How Several Token States Become One Sentence Vector

Suppose a sentence contains eight token representations, each with 768 dimensions. A retrieval system wants one 768-dimensional sentence vector. It needs a pooling rule.

Mean pooling averages token representations across the sequence, often excluding padding. A special-token strategy uses one designated representation. Some models add a learned pooling layer or projection. Different choices can change retrieval quality even when the underlying transformer is the same.

Therefore, “we use model X embeddings” is incomplete. Ask which layer, which pooling method, which normalisation step and which training objective produce the final vectors.

Worked Mean-Pooling Example

Use three toy contextual token vectors: t1 = [1,0], t2 = [0.5,0.5], t3 = [0,1]. Mean pooling adds them to [1.5,1.5] and divides by three, producing [0.5,0.5].

If t3 is only padding and should be excluded, the correct mean of t1 and t2 is [0.75,0.25]. Accidentally averaging padding changes the sentence representation. This is a small example of why masks and pooling implementation matter.

Query Embeddings and Document Embeddings May Use Different Instructions

Some retrieval models are trained asymmetrically: a short query should match a longer passage that answers it. The same semantic content can therefore be encoded differently depending on whether it is acting as a query or a document.

Sentence Transformers’ current documentation recommends encode_query and encode_document for models that define query/document prompts or routing. Using one generic encoding path can still work for some models, but it should not be assumed for all.

This is an important production detail. A retrieval system can have a high-quality model and a high-quality vector index but still perform poorly because it encodes queries with the wrong task route.

Hard Negatives Teach Fine Distinctions

Consider the query “How long is a normal camera loan?” Positive passage: “The normal camera loan is three school days.” Easy negative: “The canteen opens at 7 AM.” Hard negative: “Approved field-project loans may last five school days.”

The hard negative is topically very similar and contains the same asset type and duration concept. Training that distinguishes it from the normal-loan answer teaches the embedding model a finer retrieval boundary.

Hard negatives are powerful but must be constructed carefully. If a supposedly negative passage actually answers the query under some interpretation, the training signal becomes contradictory.

Embedding Evaluation Before RAG

Before connecting embeddings to a language model, evaluate retrieval by itself. Prepare queries with human-identified relevant passages. Encode the corpus and queries, retrieve top-k candidates and measure whether the correct evidence appears.

Useful metrics include recall at k, mean reciprocal rank and normalized discounted cumulative gain depending on the task. The exact metric matters less than isolating retrieval from generation so a poor answer is not automatically blamed on the LLM.

Once retrieval is strong enough, connect the candidates to generation and evaluate answer support. This two-stage discipline mirrors the model-versus-system method used throughout the series.

A Worked Retrieval Evaluation

Suppose five test queries each have one clearly relevant passage. At top-1, the system finds the relevant passage for three queries. At top-3, it finds the relevant passage for all five. Top-1 recall is 60%; top-3 recall is 100% for this tiny illustrative set.

A reranker may improve which candidate reaches rank one. Increasing k may also increase context noise. The evaluation helps choose the architecture based on measured behaviour rather than intuition.

Five queries are far too few for a production reliability claim. They are enough to demonstrate how retrieval evaluation is separated from generation evaluation.

Hybrid Search Combines Lexical and Vector Evidence

Vector search is strong when wording differs but meaning is related. Keyword or sparse retrieval is strong for exact names, rare codes and distinctive phrases. Hybrid search combines the signals.

A query for “C07” should preserve the exact asset code. A query for “camera for a field project” benefits from semantic similarity. A hybrid system can retrieve exact identifier matches while also finding conceptually relevant policy passages.

The architecture should expose which signal retrieved each result where diagnosis matters. Otherwise a vector-model change can be mistaken for a keyword-index regression.

Embedding Compression and Quantisation

Large vector collections consume memory. Systems can store embeddings with lower precision, quantise them or reduce dimensionality to save space and accelerate search.

Compression creates a trade-off: smaller representations can reduce storage and bandwidth while changing similarity rankings. The right configuration depends on acceptable retrieval quality and infrastructure constraints.

Measure recall after compression rather than assuming the effect is negligible.

Embeddings Can Leak Information About Their Training and Inputs

An embedding is not automatically safe merely because it is a vector rather than readable text. Research has shown that model representations can preserve information about inputs, and vector stores can contain sensitive organisational material.

Apply normal data governance: access control, retention rules, tenant separation, encryption where appropriate and careful decisions about which material should be embedded.

Do not use “it is only numbers” as a privacy argument.

Embedding Provenance Should Travel With the Vector

A production vector record should normally retain enough metadata to identify the source object, content version, embedding model version and access scope. The vector itself is insufficient for audit.

If a source document is withdrawn, the corresponding vectors should be removable. If the document changes, the system should know which vectors need re-embedding. If the embedding model changes, the index should know which coordinate system each stored vector belongs to.

A Practical Embedding Registry

  • Source ID and source version.
  • Chunk or item identity.
  • Embedding model and model version.
  • Embedding dimension and expected similarity metric.
  • Encoding mode or prompt, such as query versus document where relevant.
  • Creation time and re-embedding status.
  • Access-control or tenant metadata.
  • Optional checksum or source-content fingerprint for change detection.

This registry is not required for every toy experiment. It becomes increasingly valuable when embeddings are durable infrastructure rather than disposable notebook outputs.

Embedding Failure Matrix

  • Right source, poor vector: investigate embedding model and chunk representation.
  • Right vector, poor index result: investigate ANN index configuration and top-k search.
  • Relevant archived source retrieved: investigate metadata filters and source governance.
  • Good retrieval, bad answer: investigate generation and evidence use rather than re-embedding immediately.
  • New query model against old corpus vectors: investigate representation-version mismatch.
  • Exact ID not found semantically: route exact identity through keyword/database lookup instead of forcing vector search.

Independent Exercise 6: Pooling Error

A sentence has four token states, but the fourth position is padding. Mean pooling averages all four. What is the likely problem?

Answer

The padding representation contaminates the sentence vector. Use the attention mask or equivalent metadata so pooling includes only real tokens.

Independent Exercise 7: Hard Negative

Query: “What is the standard loan period?” Candidate A says “standard period is three days.” Candidate B says “field-project exception allows five days.” Why is B a useful hard negative?

Answer

It is highly related in topic and vocabulary but answers a different condition. Distinguishing A from B tests whether the embedding space preserves the operational distinction that matters.

Independent Exercise 8: Vector Privacy

A team says embeddings need no access controls because humans cannot read them directly. What is wrong with the argument?

Answer

Vectors can still encode sensitive information, can be searched and can be linked back to source records. Apply data governance based on the underlying information and use case, not visual readability.

Frequently Asked Questions About Embeddings

What is an embedding?

A learned numerical vector representing an item for a particular model and objective. The item may be a token, sentence, document, image or other object.

Is an embedding the same as a token?

No. A token is a discrete vocabulary unit. Its token ID can be mapped to an embedding vector before neural-network computation.

What does embedding dimension mean?

It is the number of numerical coordinates in the vector. More dimensions provide more capacity but do not automatically imply better quality.

What is cosine similarity?

A measure of directional similarity between vectors, calculated from their dot product and lengths. It is common in semantic search but must be interpreted within the model and task.

Are close embeddings synonyms?

Not necessarily. They may be related by topic, usage or task-specific training. Similarity needs empirical interpretation.

Can embeddings store current facts?

They can represent text containing facts, but a vector is not a substitute for authoritative current state. Preserve the source record and metadata.

What is a sentence embedding?

A fixed-size vector representing a sentence, produced by a model and pooling or encoding process designed for sentence-level tasks.

Do vector databases create embeddings?

They may integrate embedding services, but conceptually the embedding model produces vectors and the database stores/indexes them for retrieval.

Should I use cosine similarity or dot product?

Use the metric the embedding model and evaluation support. Do not assume one metric is universally best.

Why can an old policy rank highly?

Because it can be semantically similar to the current policy. Use exact metadata such as version and status alongside vector similarity.

Selected Technical References

Embeddings Turn Discrete Identity Into Learnable Geometry

The most important transition is now visible. Tokenisation produces discrete IDs. An embedding layer turns those IDs into continuous vectors. Transformer layers then transform those representations in context. Retrieval systems can also create sentence or document embeddings designed for similarity search.

The geometry is useful because training makes relationships computable. But geometry is not authority, truth or identity. The system must still preserve source records, permissions, exact IDs and verification.

Next: 014 — Vector Space, where we zoom out from individual embeddings and study the geometry that contains them: dimensions, distance, direction, neighbourhoods, projections and the limits of interpreting high-dimensional spaces.


How Super Intelligence Works Series Navigation

Previous: 012 — Tokenisation · Series Hub · Next: 014 — Vector Space.

Discover more from eduKateSG

Subscribe now to keep reading and get access to the full archive.

Continue reading