VIEW THIS AS

Auto mode follows the Route Engine until you choose a viewpoint.

YOU ARE HERE

ROUTE CHECK

CONNECTED TO

WHAT NEXT

Use the canonical route for this room, or HELP if you are unsure.

How Super Intelligence Works | Semantic Search — Retrieving Beyond Exact Words

eduKate Secondary students reviewing open books for How Super Intelligence Works: Attention.

Semantic search helps a Super Intelligence system find information when a question and a useful passage use different words. A person might ask, “Can I work somewhere quiet after dinner?” while a relevant record says, “Silent study room available until 21:00.” Exact-word overlap is limited, but the information need and the passage are related. A learned representation can make that relationship searchable.

The difficult part is deciding which relationship matters. “Quiet”, “study”, “evening” and “available” can point towards useful material, yet a result may still be closed on the requested day, require a booking, describe an old schedule or prohibit the activity the reader has in mind. Finding a semantically nearby passage is an invitation to inspect it, not a certificate that every condition has been met.

This article explains the mechanism behind that invitation: lexical matching, embeddings, similarity measures, candidate search, hybrid rankings and reranking. It also shows how to test exact constraints, multilingual queries, unfamiliar terminology and changes in the underlying index. The central question is practical: why did this passage appear, and what would make another passage deserve a higher position?

Here, Super Intelligence is the name of the learning series. The techniques discussed are used in present-day AI and information-retrieval systems; their existence does not establish that any current system has achieved artificial superintelligence. Semantic search is a bounded retrieval capability, with measurable strengths and failures.

Return to the How Super Intelligence Works guide. For the underlying numerical concepts, read Embeddings and Vector Space.

1. The search problem begins with an information need

An information need is what someone must learn, locate or decide. A query is the expression they happen to type. The two are not identical. “Quiet room tonight” might mean a place to revise, a space for a confidential call, a venue for an interview or somewhere to record music. A search system sees an incomplete description and has to rank the available records under that uncertainty.

Traditional search already does more than compare whole strings. It can split text into terms, normalise case, account for word forms, search fields and apply curated synonym rules. Semantic search adds learned relationships that can connect different expressions. “Recover my account” may connect with “reset access credentials”; “reduce waiting time” may connect with “shorten the queue”. The connection is useful when it reflects the intended task and harmful when it quietly changes that task.

Three different relationships often get called similarity. Topic similarity asks whether two texts discuss the same subject. Paraphrase similarity asks whether they express approximately the same meaning. Retrieval relevance asks whether a particular document helps with a particular information need. A paragraph about a building being closed is topically similar to a question about its opening hours. It can be a highly relevant answer even though it contradicts the user’s hope that the building is open.

This distinction explains why an embedding model trained to group similar sentences is not automatically the best model for answering short questions from long documents. A query may need a complementary passage rather than a paraphrase. “What should I bring?” and “A photo identity document is required at collection” serve different grammatical roles but can form a useful question–answer pair. The Sentence Transformers semantic-search documentation distinguishes symmetric tasks, such as finding similar questions, from asymmetric tasks, such as finding a passage that answers a question.

Before evaluating a search result, make the intended relationship explicit. Are you finding duplicates, supporting evidence, a specific record, alternative explanations or an answer-bearing passage? A system can succeed at one and fail at another. “The results look related” is too weak a success criterion when a person needs the exact current exception to a rule.

The rest of this article uses an invented collection of room notices. It lets us inspect the difference between related language and usable information without depending on an opaque commercial search engine or a changing real-world booking service.

Back to contents · Continue: 2. A small original corpus we can inspect

2. A small original corpus we can inspect

Imagine a fictional community learning centre called Harbour Learning Rooms. The centre is invented for this article. Its names, schedules, room capacities and document identifiers are teaching data, not statements about an actual venue. Each record has a short text plus structured fields such as room identifier, day, closing time and version status.

D1 — Cedar room, current notice. “Cedar is a silent study room. On Tuesdays it opens at 18:00 and closes at 21:00. No reservation is needed. Maximum occupancy: 12.” Fields: room CEDAR; day Tuesday; end 21:00; booking false; activity silent study; current true.

D2 — Cedar room, archived notice. “Cedar is a silent study room. On Tuesdays it opens at 18:00 and closes at 22:00. No reservation is needed. Maximum occupancy: 12.” The same room identifier is retained, but current is false and the notice is labelled archived.

D3 — Maple room, current notice. “Maple supports group discussion. Tuesday sessions run from 18:00 to 22:00. Reserve a place before arrival. Maximum occupancy: 20.” Fields distinguish a discussion room from a silent room and record booking true.

D4 — Birch room, current notice. “Birch offers silent reading on Tuesdays from 18:00 to 20:00. Walk-ins are welcome. Maximum occupancy: 8.” This is a plausible semantic match for quiet study, but its closing time can disqualify it for a late-evening visit.

D5 — Linden room, current notice. “Linden has silent study on Wednesdays from 18:00 to 22:00. No booking is necessary. Maximum occupancy: 10.” Its day is different even though most of its language resembles the desired answer.

D6 — Cedar room, activity restriction. “Cedar does not permit music rehearsals or recorded performances. Silent individual study is permitted.” This is about the right room and may be essential to an activity question, but it supplies no opening schedule.

D7 — Cedar room, translated notice. “星期二,Cedar 安静自习室开放至晚上九点,无需预约。” The teaching translation is: on Tuesday, Cedar’s quiet self-study room is open until nine in the evening and needs no reservation. It repeats D1’s relevant schedule facts in Chinese and retains the Latin-script room identifier.

D8 — Reference-code register. “Reference CED-021 identifies the replacement desk lamp specification. It is not a room timetable.” Its identifier happens to look similar to the Cedar room name. The record exists to test whether a search respects exact codes and separates furniture documentation from room information.

D9 — General study advice. “To concentrate after dinner, silence notifications, choose a calm environment and work in short sessions.” This may be useful advice for studying, but it does not establish that any particular room is available.

D10 — Birch accessibility note. “The entrance to Birch is step-free. Tuesday silent-reading hours end at 20:00.” This supplies a property that the other room notices do not establish. Unknown accessibility must not silently become accessible.

Our main query is: “Find a quiet place to study on Tuesday at 20:30 without booking.” We will treat 20:30 as the intended time of attendance, interpret quiet study as silent study or silent reading, and use current notices only. Under those explicit assumptions, D1 is a direct answer-bearing record; D7 is a translated duplicate of its key facts. D3 fails the activity and booking constraints, D4 and D10 fail the time constraint, D5 fails the day constraint, and D2 is obsolete for a current-availability question.

This answer key is deliberately richer than a list of similar passages. It tells us why a record is useful, which requirement it violates, and which facts remain unknown. Those distinctions will matter more than an attractive-looking similarity score.

Back to contents · Continue: 3. What lexical search preserves

3. What lexical search preserves

Lexical search represents text through words or other tokens and their occurrences. An inverted index can record which documents contain each term. For a query containing “Tuesday”, the system can efficiently find postings for that term instead of reading every document from the beginning. Field-specific indexes can distinguish a room name from a word that appears incidentally inside the body.

BM25 is a widely used lexical ranking family. It combines evidence such as query-term occurrence, the rarity of a term across the collection, and document-length normalisation. Repeating the same term does not ordinarily add unlimited proportional benefit because term frequency is saturated. The Stanford information-retrieval text’s BM25 discussion explains the formulation and its parameters. A BM25 score is a ranking quantity, not a percentage chance that a result is correct.

For the Harbour query, literal matching can preserve “Tuesday” and “booking”. But D1 uses “silent” and “reservation”, while the query uses “quiet” and “booking”. An unexpanded term-based system may therefore miss two useful connections. A well-designed lexical system could handle them through synonym rules or query expansion. The meaningful comparison is between actual configurations, not between a capable semantic model and an intentionally crippled keyword baseline.

Lexical evidence becomes particularly valuable when a rare string is the user’s target. A query for CED-021 should strongly favour the code register even if the surrounding sentence says, “I need the relevant document for this item.” A semantically related room notice cannot substitute for the exact identifier. Search designers may preserve such strings in a keyword field, while also indexing the prose in an analysed text field.

Tokenisation decisions matter here. Punctuation, spaces and letter case can change whether “CED-021”, “CED 021” and “ced021” are treated as identical, related or different. There is no universally correct answer. If codes are case-sensitive identifiers, normalising them carelessly can merge distinct items. If a user often inserts harmless spaces, refusing all variation may make the search unnecessarily brittle. The system needs a documented identity rule.

The same caution applies to phrases. Searching “no booking” as two independent terms is not equivalent to enforcing the condition booking equals false. D3 contains “Reserve a place”, while a longer document could include both “booking” and “not” in unrelated sentences. Lexical matching supplies evidence about wording; it does not by itself turn every natural-language constraint into executable logic.

Lexical search is therefore an essential source of precision rather than an obsolete stage of search history. Its strengths are inspectable: exact names, distinctive terms, phrases and field values. Its weaknesses are also inspectable: the system may require words that the relevant author never used. Semantic representations address part of that vocabulary mismatch.

Back to contents · Continue: 4. What an embedding actually represents

4. What an embedding actually represents

An embedding is an ordered collection of numbers produced by a model. In a common dense-retrieval design, an encoder turns a query into one vector and turns each passage into another. The passage vectors can be computed before a user searches. At query time, the system compares the query vector with stored passage vectors using a chosen scoring rule.

The numbers are learned features, not a neatly labelled dictionary of meanings. A coordinate does not usually mean “Tuesday”, another “quietness” and another “permission”. The useful information is distributed through many coordinates and their relationships. A map of points can illustrate neighbourhoods, but it should not be mistaken for a literal atlas in which every human concept has an obvious independent axis.

Training determines which relationships become easy to retrieve. A model trained with question–passage pairs can be encouraged to score an answer-bearing passage above distracting passages. Negative examples matter: if the model repeatedly sees the wrong room, wrong year or wrong language as distractors, it receives pressure to represent distinctions that separate those cases. Poorly chosen negatives can also teach shortcuts or label another valid answer as wrong.

Sentence-BERT established an influential approach to producing sentence representations suitable for similarity comparison without jointly processing every possible sentence pair. Dense Passage Retrieval studied separately encoded questions and passages for open-domain question answering. These are research contributions with particular training and evaluation settings, not evidence that every vector representation solves every search task.

In our corpus, a useful representation might place “no reservation is needed”, “walk-ins are welcome” and “无需预约” close enough to help retrieve related notices. That does not mean it has built a reliable Boolean database field for booking. Similarity can preserve a broad relationship while blurring the precise condition. An application that requires guaranteed compliance should use explicit fields or separately checked text evidence for the relevant rule.

The encoder’s input is also part of the representation. Embedding only “Closes at 21:00” omits the room, day and status. Embedding “Cedar | Current Tuesday timetable | Closes at 21:00” supplies far more discriminating context. A heading can make an otherwise ambiguous paragraph useful. Conversely, repeatedly attaching a very long generic heading may drown out the passage-specific content.

Finally, representations are model-specific. A query vector from encoder A is not safely comparable with a passage vector from encoder B just because both have the same number of dimensions. Their coordinate systems and training objectives may differ. A functioning search index depends on compatible query and document encoders, preprocessing, dimensions and similarity conventions. Matching array lengths is necessary in many implementations, but it is not a semantic compatibility test.

Back to contents · Continue: 5. A transparent geometry example

5. A transparent geometry example

We can calculate a similarity ranking without pretending to reproduce a real trained model. Consider three invented two-dimensional document vectors and one query vector. These coordinates are chosen solely for arithmetic practice. They are not measurements of the Harbour records, and their axes have no asserted linguistic meaning.

Let the query be q = (1, 0). Let A = (0.8, 0.6), B = (0.6, 0.8), and C = (4, 3). The length of q is 1. A and B both have length 1 because the square root of 0.64 + 0.36 is 1. C has length 5 because the square root of 16 + 9 is 5.

The dot product multiplies matching coordinates and adds them. Thus q · A = 0.8, q · B = 0.6, and q · C = 4. Under raw dot-product ranking, C comes first, A second and B third. C benefits from its larger magnitude even though it points in the same direction as A.

Cosine similarity divides the dot product by both vector lengths. The cosine scores are therefore 0.8 for A, 0.6 for B and 4 ÷ 5 = 0.8 for C. A and C tie. Normalising C to unit length produces (0.8, 0.6), exactly A. After unit normalisation, dot product and cosine produce the same ordering for these vectors.

This small example explains why a similarity measure cannot be selected independently of the model and its recommended representation. If magnitude carries useful learned information, normalising can discard it. If the intended score uses direction, allowing magnitude to dominate can distort the ranking. Neither measure is universally superior; compatibility and tested task performance matter.

For unit vectors, squared Euclidean distance equals 2 minus twice the cosine similarity. A has squared distance 0.4 from q: (1 − 0.8) squared plus (0 − 0.6) squared equals 0.04 + 0.36. B has squared distance 0.8. Thus minimising this distance agrees with maximising cosine in this unit-normalised example. The agreement does not justify casually switching metrics on unnormalised vectors.

Notice what none of these calculations tells us. A cosine of 0.8 does not mean 80 percent accuracy, an 80 percent probability of truth or that 80 percent of the query’s conditions are met. It describes a geometric relationship under a particular representation. Two passages that differ only in “permitted” versus “not permitted” may still be close because they discuss almost everything else in common.

If the display rounds one result’s 0.804 and another document’s 0.801 to 0.80, the visible tie can conceal a real ranking difference. If scores differ only slightly, small representation or indexing changes may reorder them. Inspect full scores when diagnosing ranking changes, but judge usefulness from the actual information need rather than from numerical precision alone.

Back to contents · Continue: 6. Hard constraints need a different kind of answer

6. Hard constraints need a different kind of answer

The Harbour query contains several requirements: silent study, Tuesday attendance, a time of 20:30, no reservation and a current notice. Some concern meaning; others can be expressed as exact predicates. Treating all of them as interchangeable hints is the fastest way to return a persuasive but unusable result.

Suppose the retrieval model puts D2 first because its archived description closely matches the query. A freshness boost that merely adds a little weight to current notices may still leave D2 above D1. If the task requires current availability, the correct action is to exclude archived notices from the eligible set, not hope that a soft ranking preference will usually suppress them.

Similarly, time comparisons belong in a reliable representation. In our teaching data, the room is usable at 20:30 only if its opening interval contains 20:30. Cedar’s 18:00–21:00 interval qualifies. Birch’s 18:00–20:00 interval does not. The system must decide whether a closing boundary is inclusive or exclusive and whether the user needs a minimum remaining duration. Arriving at 20:59 is a different task from finding a room for a two-hour study session.

Numbers need units and scope. “Maximum occupancy: 12” does not mean twelve places are currently free. “Closes at 21:00” is a local-time statement in our invented setting, not a timezone conversion. “At least ten seats” and “exactly ten seats” are different conditions. A similarity score may react to the presence of ten without reliably enforcing either comparison.

Structured filters are useful when trustworthy fields exist. A filter can require day equals Tuesday, booking equals false and current equals true before ranking eligible candidates. If a field is missing, the application must choose an explicit unknown-data policy. It should not infer that a missing booking flag means no booking is needed, or that an absent accessibility field means step-free access.

Filters themselves can be wrong. If D4’s closing time was extracted as 22:00 rather than 20:00, exact filtering will enforce the wrong value perfectly. The remedy is to preserve the source passage, validate important extraction and test updates. Precision in execution does not repair an inaccurate data model.

A helpful design separates three states: satisfied, violated and unknown. D1 satisfies the stated study-time conditions. D3 violates them. D1’s step-free status is unknown because the record does not say. This three-way distinction prevents the absence of contrary evidence from becoming evidence of compliance. It also makes clarification useful: “Do you also need step-free access?” changes the eligible answer set rather than merely adding a decorative preference.

Back to contents · Continue: 7. Similarity is not truth, permission or entailment

7. Similarity is not truth, permission or entailment

Consider these three statements: “Cedar permits music rehearsals”, “Cedar does not permit music rehearsals”, and “Maple requires reservations for group discussion”. The first two are extremely close in topic and vocabulary. Yet one directly contradicts the other. The third is more distant in wording and subject even though it may be a safer alternative for a group activity.

Semantic similarity measures relatedness under a learned representation. Entailment asks a directional question: if a source statement is true, does a proposed claim follow from it? Truth asks whether the claim corresponds to reality. Permission asks whether an authorised rule allows an action. These questions can interact, but none can be replaced by a high similarity score.

D6 is an excellent result for “Is music rehearsal allowed in Cedar?” because it directly addresses the permission. It is not evidence supporting an affirmative answer. A system that retrieves D6 and responds “Yes, Cedar supports music rehearsal” has found the right topic and failed to read the decisive negation. Improving candidate similarity alone may not repair that failure.

Negation also has scope. “No reservation is needed on Tuesdays” does not establish the rule for Wednesdays. “Not every room requires booking” does not identify a room that is free to enter. “The room is not unavailable” differs from “The room is unavailable”. A bag of important nouns cannot settle these differences, and a dense representation should not be assumed to handle them reliably without tests.

Research can examine such representation behaviour separately from broad task averages. ALIGN-SIM proposes tests of semantic alignment involving distinctions such as antonyms, paraphrase and sentence changes. Its relevance here is methodological: evaluate the specific distinction you need. A generally strong similarity benchmark score is not a guarantee of correct performance on every negated room rule.

For our collection, build paired tests. Keep a query fixed and replace “permitted” with “not permitted” in a candidate. Then keep the document fixed and change the question from “Is rehearsal allowed?” to “Which activities are prohibited?” Both tests involve the same vocabulary, but the desired interpretation changes. A useful search system may return the same document in both cases; the answer extraction must change accordingly.

Finally, a false statement can be a relevant search result. Someone investigating a rumour may need the document that made the false claim. The retrieval label should reflect the task, while truth assessment remains explicit. This is why relevance judgements must describe what counts as useful rather than silently equating useful, agreeable and true.

Back to contents · Continue: 8. Hybrid search combines evidence, not guarantees

8. Hybrid search combines evidence, not guarantees

Hybrid search combines different retrieval signals, commonly lexical and vector-based search. The motivation is complementarity. Literal matching can preserve CED-021. A learned representation can connect “walk-ins” with “without booking”. Neither signal needs to be discarded merely because the other succeeds on a different kind of query.

One combination method is a weighted score. For instance, a system might normalise two score ranges and compute a weighted sum. But raw BM25 and cosine scores do not share a natural common scale. Adding 12.4 from one method to 0.83 from another does not express a principled balance unless the application defines how those values are calibrated or transformed.

Another method is reciprocal rank fusion, or RRF. It uses positions in result lists rather than the original scores. For each list in which a document appears, add 1 divided by a constant plus that document’s rank, with ranks starting at one. The official Elasticsearch RRF reference describes this rule and the use of a finite ranking window. The constant and window still affect behaviour even though raw-score calibration is avoided.

Here is an entirely invented ranking example using two lists of four records each. Suppose the lexical list is D2, D1, D8, D3, while the semantic list is D1, D3, D2, D9. These lists are teaching assumptions, not outputs measured from a model. Choose a small constant of 10 to keep the arithmetic visible; this is not a recommended production setting.

D1 receives 1/12 from lexical rank two and 1/11 from semantic rank one: approximately 0.174242. D2 receives 1/11 plus 1/13: approximately 0.167832. D3 receives 1/14 plus 1/12: approximately 0.154762. D8 appears only at lexical rank three, so it receives 1/13, approximately 0.076923. D9 appears only at semantic rank four, so it receives 1/14, approximately 0.071429.

The fused order is D1, D2, D3, D8, D9. The outcome rewards agreement between the two lists. It does not make D2 current, make D3 a silent room or turn D8 into a timetable. Eligibility constraints should be handled explicitly. If archived records are forbidden for the task, remove them through an appropriate filter rather than interpreting a lower fused rank as adequate protection.

RRF also loses information. A large gap between first and second place in one input list becomes a one-rank difference, just as a tiny gap does. A document absent from a retrieved window contributes nothing from that list, even if it would have ranked immediately below the cutoff. Changing the window can therefore change the final order without any underlying text changing.

The responsible question is not “Is hybrid always better?” It is “Which failures does this combination reduce on our query set, and which new ones does it introduce?” Keep a lexical baseline, a semantic baseline and the combined system in the same evaluation. Otherwise complexity can be mistaken for improvement.

Back to contents · Continue: 9. Dense vectors are not the only semantic representation

9. Dense vectors are not the only semantic representation

It is convenient to contrast keyword search with one dense vector per passage, but that is not the whole design space. Semantic information can also be represented through learned sparse weights or multiple vectors. Different representations preserve different details and create different storage, indexing and computation costs.

A sparse lexical representation has many possible dimensions but relatively few nonzero entries. Traditional term-frequency representations are sparse. Learned sparse methods can assign weights to terms and expand beyond the words literally present in the original passage. SPLADE is a research example of learned sparse retrieval using lexical expansion and controlled sparsity. Its existence makes a simple “sparse means no semantics, dense means semantics” distinction inaccurate.

For Harbour, imagine a learned sparse representation that associates “reservation” with “booking”. The index can gain a useful bridge while retaining token-linked evidence. But expansion can also create false associations. A room document mentioning that reservations are unnecessary may acquire a strong booking-related feature; that feature alone does not encode whether booking is required, forbidden or optional.

Multiple-vector methods preserve finer-grained representations than one pooled vector. ColBERT independently encodes queries and documents and uses a later interaction between their token-level representations. The broad idea is to retain more local matching information while still precomputing document-side work. It is an architectural alternative with its own efficiency tradeoffs, not a promise that exact constraints become infallible.

An analogy helps. A single-vector representation is like a compact description of a room notice’s overall character. A multiple-vector representation keeps a richer collection of local signals. A learned sparse representation keeps weighted vocabulary-linked features, potentially including inferred associations. These analogies explain what is being retained; none implies that the system literally reads or remembers in the human way.

Choosing among them begins with a failure case. If search finds the right general topic but misses small decisive clauses, a richer interaction model may be worth testing. If identifiers are lost, a reliable lexical field may be simpler than replacing the entire retrieval stack. If the collection is small enough, a more expensive direct comparison of every candidate may be practical. The best architecture depends on task size, constraints and evidence from evaluation.

Keep representation claims separate from implementation claims. “Uses vectors” says little about whether an index is approximate, whether it applies filters safely, or whether it can distinguish Tuesday from Wednesday. “Uses a transformer” says little about whether the model was trained for paraphrase, classification or retrieval. Ask what inputs are encoded, what training relationship is rewarded, what score is used and which details are retained.

Back to contents · Continue: 10. Approximate nearest-neighbour search creates candidates

10. Approximate nearest-neighbour search creates candidates

After computing embeddings, a search system still has to find promising neighbours. With ten documents, it can compare the query against every stored vector. With a much larger collection, an index can reduce the search work by exploring a promising subset. Approximate nearest-neighbour search, usually abbreviated ANN, trades exactness in neighbour discovery for efficiency.

Approximate does not mean the system fabricates new documents. It means the returned set may omit a vector that would have appeared among the top neighbours in an exhaustive search under the same scoring rule. Even an exact vector search can return semantically wrong documents if the representation is poor. These are two different failure sources and should be measured separately.

Hierarchical Navigable Small World graphs are a well-known graph-based approach to approximate search. At a high level, graph connections guide exploration towards nearby points instead of requiring an exhaustive comparison. Other index families partition space or compress representations. The Faiss documentation describes a range of similarity-search methods and implementation choices. No single index setting is best for every collection and workload.

Use a second toy example. Suppose exhaustive vector search says the top five IDs are D1, D7, D2, D4 and D9. An approximate index returns D1, D2, D4, D9 and D3. Four of the exact top five are present, so the overlap-based ANN recall at five is 4/5, or 0.8, under this explicitly defined test.

That number is not task relevance recall. D7 is a translated duplicate; D2 is archived; D4 closes too early; D9 offers no venue. A system can have excellent agreement with an exact vector ranking while that ranking itself is poor for the user’s requirements. Conversely, an approximate result might omit a redundant neighbour without harming the practical answer. Report both kinds of measurement rather than merging them.

Filters complicate candidate generation. If a system retrieves a small global list and then removes ineligible records, it may be left with too few results even though qualifying records exist elsewhere. A filter-aware index or a sufficiently broad candidate search can help, but actual behaviour depends on the implementation. A post-filtered empty list does not prove that the corpus contains no eligible answer.

The diagnostic sequence is straightforward. First check whether the target record is indexed. Next check whether its vector is compatible with the query encoder. Then compare approximate results with exhaustive scoring on a manageable sample. Finally assess the relevance of the exact ranking. This sequence tells you whether to repair ingestion, representation, index search effort or task modelling.

Candidate budgets should be treated as design parameters. Increasing them can improve coverage while costing more computation downstream. The goal is enough useful candidates for the next stage, not the largest list the system can technically return.

Back to contents · Continue: 11. Reranking asks a more focused question

11. Reranking asks a more focused question

A first-stage retriever must search broadly and cheaply. A reranker can spend more work on a smaller candidate set. In a common cross-encoder design, the query and each candidate passage are processed together, allowing the model to compare their details directly rather than only comparing two independently computed summary vectors.

The Sentence Transformers retrieve-and-rerank guide illustrates this two-stage arrangement. The important distinction is computational: document vectors can be reused across many queries, while a joint query–document calculation ordinarily depends on the current pair. Spending that joint calculation on a shortlist is more manageable than applying it to every record in a large collection.

For the Harbour query, a reranker may notice that D1 covers Tuesday, the relevant time and no reservation, whereas D3 requires reservation and supports discussion. It may lower D9 because advice about concentration does not identify a venue. Whether it actually does so is an empirical question; the presence of a sophisticated model is not a substitute for inspecting the results.

Reranking has a strict ceiling: it cannot promote an answer that never entered its candidate set. If D1 and D7 were both omitted, the reranker can only arrange the remaining records. The top-ranked result may still be wrong. A user interface that always presents rank one as “the answer” hides this limitation, especially when the collection has no qualifying passage.

Scores from rerankers also require interpretation. Some models output logits; some output transformed values; some are trained for a particular relevance scale. A number between zero and one is not automatically a calibrated probability. Before using a threshold to accept or reject a result, test what that threshold means across your actual query categories and corpus versions.

Separate text relevance from business rules. If archived material is ineligible, the reranker should not be asked to override that prohibition because the archived prose is beautifully aligned with the query. Conversely, if the user asks for historical opening hours, an archived record becomes eligible and may be the best answer. The rule belongs to the task definition, not to a permanent preference for new documents.

A good reranking evaluation includes nearly identical distractors: wrong day, wrong closing time, wrong room identifier, negated permission and missing booking information. Easy negatives such as an unrelated recipe can make a reranker look excellent without testing the details that determine usefulness. The hard examples are where the model must earn its place in the system.

Back to contents · Continue: 12. Passage boundaries change the searchable meaning

12. Passage boundaries change the searchable meaning

Documents often need to be split into smaller retrieval units. A room handbook might contain booking rules, timetables, accessibility notes and equipment policies. Encoding the entire handbook as one vector can blur its individual topics. Splitting every sentence independently can remove the context that makes a sentence interpretable. Chunking is therefore a representation decision, not merely a file-size convenience.

Imagine D1 is split into “Cedar is a silent study room”, “On Tuesdays it opens at 18:00 and closes at 21:00”, and “No reservation is needed”. The third fragment might match the booking condition but no longer identify the room or day. A query could retrieve the right sentence while losing the information needed to apply it. Including a concise heading or parent identifier can preserve that connection.

Now imagine the opposite: D1, D2 and D3 are concatenated into one large passage. Its vector may mix current and archived schedules with different booking rules. A returned passage can contain all the query’s words while leaving the reader to disentangle conflicting records. More text is not automatically more helpful context.

Overlapping chunks can protect a sentence that falls across a boundary, but excessive overlap creates near-duplicate results. If the top five hits repeat the same Cedar notice, they do not provide five independent pieces of evidence. At the search interface, grouping by document or room can improve diversity. In evaluation, decide whether relevance is counted at passage level, document level or unique-answer level.

Tables require particular care. A cell containing “21:00” needs its column heading and row identity. Flattening a timetable into an arbitrary text sequence can attach a closing time to the wrong day. Preserve the relationship between headers and values, or store the table as structured data alongside a readable representation. Search should not depend on a model guessing which row a number belongs to.

Representations should also retain enough provenance to reopen the original. A passage ID, source document ID, section label and version marker make a hit inspectable. They do not themselves establish truth, but they allow a person or a later checking step to find the surrounding conditions. Losing those pointers makes debugging much harder when a result appears convincing but cannot be located again.

Test boundary changes with real questions. Ask whether the same target can still be found after a heading is removed, a table is extracted, or a long notice is split. If a small formatting change destroys retrieval, the representation may depend on fragile context. Repairing that context is often more useful than simply increasing the number of results.

Back to contents · Continue: 13. Multilingual search needs language-specific checks

13. Multilingual search needs language-specific checks

A multilingual embedding model may place related expressions from different languages into a shared representation space. That can allow an English query to retrieve D7’s Chinese notice without a separate English translation in the index. The useful capability is cross-language matching; it is not proof that every meaning, domain term or regional expression is preserved equally well.

The Sentence Transformers multilingual-training documentation describes training with parallel sentences. The MIRACL research dataset evaluates same-language query-to-corpus retrieval across 18 languages; it is not a cross-language query-to-document benchmark. These are distinct evaluation needs. Measure each language and cross-language direction explicitly rather than infer universal performance from a strong English result.

Our D7 notice is intentionally simple, but realistic collections contain mixed scripts, transliterated names, local abbreviations and incomplete translations. “Cedar” remains in Latin script inside Chinese prose. A user may type a transliteration instead, a building nickname or a shortened room label. The search system needs a tested relationship between those expressions and the authoritative room identifier.

Numbers and dates introduce additional ambiguity. Nine in the evening is 21:00 in the teaching notice, but date formats such as 03/04 can represent different dates under different conventions. A model may retrieve the right topic while leaving the date interpretation unresolved. Structured dates should use unambiguous stored forms, with display conventions handled separately.

Translated duplicates also affect rankings. D1 and D7 describe the same room availability. Returning both can reassure a multilingual reader, but counting them as independent confirmation exaggerates the evidence. A deduplication policy can group them under one source fact while allowing the reader to choose a language. Preserve the relationship rather than deleting useful translations indiscriminately.

When evaluating, create matched queries in the languages the audience actually uses. Include native phrasing rather than only literal translations of English questions. Measure whether the relevant document is found, whether the displayed language is useful, and whether exact constraints survive. A system can retrieve well across languages but present the answer in a form the user cannot comfortably read.

Finally, keep abstention available. If an important phrase is ambiguous, the application can show the source and ask for clarification rather than silently choose an interpretation. Cross-language confidence should be earned from representative tests and, for consequential uses, competent human review. Broad language coverage on a model description is a starting point for testing, not a service guarantee.

Back to contents · Continue: 14. Domain language can move the neighbourhood

14. Domain language can move the neighbourhood

Words acquire specialised meanings in particular settings. “Reservation” can describe a booking or a doubt. “Capacity” can mean seats, storage, production output or legal competence. “Recall” can mean retrieval coverage, memory or a product withdrawal. A model that retrieves sensibly in everyday conversation can make different mistakes in technical, organisational or educational material.

Within Harbour’s collection, “silent” is a room-use category. In another collection, “silent installation” describes software deployment. A query for “quiet setup after dinner” could connect poorly if the collection mixes building notices with technical manuals. Adding an explicit domain field or a source filter may be more effective than expecting one embedding to resolve every ambiguous phrase.

Domain adaptation can involve better text preparation, curated aliases, relevant training pairs, a different pretrained model or fine-tuning. These options have different costs and risks. Start with the failure you observed. If the real problem is that every passage loses its section heading during extraction, more training may leave the ingestion fault untouched.

Training examples should reflect the intended relevance relationship. For a room-search service, “quiet place without booking” paired with D1 is a useful positive. D3 is a challenging negative because it shares the Tuesday-evening topic but violates important conditions. Yet if the query changes to “group discussion room I can reserve”, D3 becomes a positive. Labels attach to query–document pairs, not to documents permanently.

The BEIR benchmark was designed to study retrieval across heterogeneous tasks and domains. Its broader lesson is that performance can change when the evaluation setting changes. The MTEB benchmark likewise evaluates text embeddings across multiple task types rather than reducing quality to one similarity demonstration. Neither benchmark replaces local evaluation on the user’s actual information needs.

Avoid adapting on the same examples used for the final quality claim. If the ten Harbour records are used repeatedly to tune synonyms and thresholds, success on those ten records becomes less informative. Keep unseen queries, later document versions and separately constructed distractors for evaluation. Otherwise the system may learn the test collection’s quirks rather than a transferable retrieval capability.

Also examine whose language the training data represents. Expert vocabulary and beginner vocabulary can describe the same need differently. A parent may ask for “somewhere my child can revise”, while an administrator writes “independent study provision”. Search quality should be tested across those perspectives without treating the less technical phrasing as a user error.

Back to contents · Continue: 15. Relevance judgements define what good means

15. Relevance judgements define what good means

You cannot meaningfully improve a ranking until you define a good result. For the main Harbour query, a useful judgement guide might require a current passage identifying a silent-study or silent-reading room open at 20:30 on Tuesday without a reservation. It should distinguish direct answers from background advice and label disqualifying conditions explicitly.

One possible graded scheme assigns three points to a direct, current answer-bearing record; one point to useful context that does not independently answer the question; and zero points to a non-answer or an ineligible record. D1 and D7 receive three under this scheme. D6 may receive one if the task includes understanding permitted activity. D9 can receive zero because generic study advice does not locate a venue. These are declared editorial judgements for this example, not universal labels.

Another team might give D9 one point for helpful background. That disagreement is not automatically a mistake. It reveals that the teams have different definitions of the search task. Resolve the guideline before celebrating a metric improvement. Otherwise one system can appear better simply because its results resemble the preferences of the person who wrote the labels.

Judges should see the query, the necessary task context and the relevant source content. They should not need to infer hidden user requirements. When possible, conceal which retrieval system produced a candidate so a familiar product name does not bias the judgement. Have a second judge review a sample, discuss disagreements and record any rule changes.

Unjudged is not the same as irrelevant. In large collections, evaluators often assess only a pooled set of candidates. A new system may retrieve a genuinely relevant document that no earlier system contributed to the pool. Treating every unseen item as certainly wrong can penalise discovery. Record coverage limits and expand the judged pool when comparing substantially different retrieval methods.

Include difficult query categories intentionally: exact identifiers, rare terminology, wrong-year distractors, numerical constraints, negation, multiple conditions, multilingual queries and requests with no answer in the collection. Ordinary paraphrases are necessary, but they do not reveal all the ways a plausible hit can fail a real person.

Finally, label the unit of usefulness. Two overlapping chunks from D1 may count as two relevant passages but only one useful room option. If the user wants alternatives, duplicate passages should not inflate the result count. The judgement guide must align with the interface and the decision the person is making.

Back to contents · Continue: 16. Metrics should expose the error you care about

16. Metrics should expose the error you care about

Precision at k measures how many of the first k returned items are relevant under a declared binary judgement. Recall at k measures how many of all known relevant items are present in those first k. These answer different questions: how clean is the visible shortlist, and how much of the relevant material did it cover?

Suppose the only binary-relevant passage IDs for our main query are D1 and D7. A system returns D3, D1 and D9 as its first three results. Precision at three is 1/3, approximately 0.333. Recall at three is 1/2, or 0.5. The first relevant result occurs at rank two, so this query’s reciprocal rank is 1/2. Mean reciprocal rank would average that quantity across a set of queries, using an explicit convention for queries with no relevant hit.

These numbers depend on the unit. If D1 and D7 are treated as one unique room answer, retrieving D1 gives complete coverage of that answer even though passage-level recall is only one half. Both statements can be correct. A report must identify whether it measures passages, documents, rooms or distinct facts.

Graded relevance can be evaluated with discounted cumulative gain, or DCG. Choose the common teaching convention gain = 2 to the power of relevance grade, minus 1, and discount each rank i by log base 2 of i + 1. This convention must be stated because variants exist. Normalised DCG divides a ranking’s DCG by the best possible DCG for the same judged set and cutoff.

For a tiny three-document set, suppose A has grade 3, B grade 2 and C grade 0. Ranking B, A, C gives DCG at three of 3/1 + 7/log2(3) + 0/2, approximately 7.4165. The ideal ranking A, B, C gives 7/1 + 3/log2(3), approximately 8.8928. The normalised score is therefore about 0.8340. This calculation rewards putting the more useful item earlier; it does not measure the truth of the underlying documents.

The Stanford treatment of ranked-retrieval evaluation discusses ranking measures and graded relevance. Our arithmetic above is original teaching data. No number here is a benchmark result for a product, a model or the Harbour collection running in a real engine.

Aggregate metrics need slices. A system could improve average ranking quality while becoming worse on exact codes or a minority language. Report the categories that matter, plus latency, failure rate and coverage where appropriate. A small overall gain is not a reason to accept a large regression in a safety-critical constraint.

For no-answer queries, evaluate the system’s ability to return no qualifying result or ask a useful clarification. Standard top-k ranking always has a first item when the index is nonempty. That structural fact does not imply that the first item satisfies the task. Calibrate acceptance and abstention separately from ordering.

Back to contents · Continue: 17. Diagnose a failure by finding its first broken step

17. Diagnose a failure by finding its first broken step

When a search misses D1, changing the embedding model immediately is premature. First establish whether the record exists in the searchable corpus. A connector may have skipped it, an extraction step may have produced empty text, or a document update may not have reached the index. No representation can retrieve a passage that was never indexed.

Next inspect the indexed text and fields. Did the Tuesday heading survive? Is “no reservation” still attached to Cedar? Is the current flag correct? Are time fields parsed consistently? An accurate encoder of a damaged passage can still produce an unusable representation. Compare the stored record with the original before tuning search parameters.

Then test the candidate generator. Run the same query with filters displayed. A wrong day filter can remove the answer before semantic scoring begins. An overly small post-filtered candidate list can produce an empty result. If possible, compare approximate and exhaustive vector search on a controlled subset to isolate index-search loss from model-ranking error.

If D1 enters the candidate set but appears too low, examine the ranking stage. Compare lexical, dense and fused lists separately. Check whether an exact room name was weakened, whether an obsolete duplicate dominates, or whether score transformations changed unexpectedly. The individual lists often reveal a fault that is invisible in the final combined output.

If D1 reaches the reranker and then drops, inspect the actual query–passage input. Truncation may have removed the closing time or the booking clause. The reranker may have been trained for a different task, or its output may be interpreted backwards. A score treated as higher-is-better when the interface returns a distance will invert the intended preference.

If the correct passage is visible but the answer is wrong, the failure is downstream interpretation. D6 might be retrieved accurately and its prohibition ignored. A search-quality metric can look healthy while the answer-generation stage fails. Keep the search test independent so the wrong component is not blamed for the wrong symptom.

Record one compact diagnostic trace: query text, parsed constraints, index version, candidate IDs, per-stage ranks, applied filters, selected passage text and judgement reason. Avoid logging private content unnecessarily, and use appropriate access controls. The purpose is to reproduce the failure, not to accumulate every user’s sensitive search indefinitely.

The most useful repair is the smallest one that addresses the demonstrated cause. Fix a broken date field before replacing the model. Expand candidates if ANN recall is the bottleneck. Improve a relevance model when the candidates are present but poorly ordered. Ask a clarification when the query itself does not specify which interpretation is intended.

Back to contents · Continue: 18. Version drift can change results without changing the question

18. Version drift can change results without changing the question

A search result depends on more than query text. It depends on the corpus snapshot, document processing, embedding model, indexing method, filter logic, fusion settings and reranker. Changing any of them can alter the ranking. A reproducible search test therefore needs a versioned configuration, not only a saved list of example questions.

Consider an embedding-model upgrade. Even if the new vectors have the same dimension, they may occupy a different learned space. Replacing only the query encoder while leaving old document vectors can make comparisons meaningless. A migration should preserve compatible query and document processing and validate the rebuilt index before switching live traffic.

Document edits create a different kind of drift. If Cedar’s closing time changes to 20:00, D1’s text, metadata and vector should be refreshed coherently. Updating the displayed source while leaving an old embedding and old closing-time field creates inconsistent representations of the same record. The ranking may still surface the old implication even though the original page has changed.

Deletion and permission changes also need propagation. A document removed from the source should not remain discoverable indefinitely through stale indexed content. Access controls must be enforced by the application and retrieval infrastructure, not inferred from semantic similarity. A passage’s relevance never grants a user permission to see it.

Use a small set of stable canary queries alongside a broader evaluation set. For Harbour, include the main quiet-study query, the exact code CED-021, a rehearsal-permission question and a Chinese schedule query. After an update, compare candidate presence, ranks and constraint satisfaction. A canary is an early warning, not proof that every other query is safe.

When results change, inspect the changed condition before calling it a regression. A new current timetable may correctly displace an older one. A newly added room may deserve first place. A better model may retrieve a valid passage that the old judgement set omitted. Version comparison should distinguish broken behaviour from a legitimate change in the world or in the available evidence.

Keep rollback feasible when a change materially degrades important query categories. Preserve the prior configuration and an auditable migration record, subject to retention and security requirements. Search improvements are easier to trust when they can be explained as specific changes to representation and ranking rather than as an unexplained replacement of everything at once.

Back to contents · Continue: 19. Changed-condition checks reveal brittle understanding

19. Changed-condition checks reveal brittle understanding

Return to the main query and change one condition at a time. This is a powerful test because the expected answer changes for a clear reason. A system that repeats the same favourite result across every variant may be responding to broad topic similarity while ignoring the condition that determines usefulness.

Change the time to 21:30 on Tuesday. D1 and D7 no longer qualify because Cedar closes at 21:00. D2 cannot rescue the answer because it is archived. D3 remains open, but it is a discussion room requiring reservation. Under the original silent-study and no-booking conditions, the collection establishes no qualifying room. “No qualifying result in these notices” is the correct conclusion, not a global claim that no room exists anywhere.

Change the day to Wednesday at 20:30. D5 now becomes the direct answer-bearing record. A model that continues to rank Cedar as an affirmative answer has failed to preserve the changed day. The test is especially informative because the rest of the query remains nearly identical.

Change the activity to group discussion and allow reservations. D3 becomes appropriate on Tuesday at 20:30. The result previously labelled ineligible is now useful. This shows why relevance labels depend on the whole query. A permanent “good document” list cannot represent changing user goals.

Keep Tuesday and no booking, but ask for 19:30. D4 now qualifies on time and offers silent reading. D1 also qualifies. If the user needs more than a short visit, ask for intended duration or apply an explicitly stated minimum. A room that closes in thirty minutes may satisfy attendance-at-a-time while failing a two-hour-session need.

Add step-free access at 20:30. D10 establishes accessibility for Birch but also says it closes at 20:00. D1 does not establish Cedar’s accessibility. The available evidence therefore supports no fully confirmed match. The system should preserve the unknown rather than fill it with an assumption based on the building type or room name.

Ask for the historical Tuesday schedule. D2 can become relevant, but “historical” is underspecified if several archived versions exist. The search should return the version information or ask which period matters. Recency is now less important than matching the requested time period.

These tests can be extended systematically. Swap a negation, alter one digit in an identifier, change a minimum to a maximum, change an inclusive boundary to an exclusive one, or replace a known field with an unknown value. Each variation should come with a reasoned expected outcome. That turns a plausible demo into a meaningful behavioural test.

Back to contents · Continue: 20. Practice: explain the ranking before trusting it

20. Practice: explain the ranking before trusting it

Practice 1: the convincing obsolete notice. A semantic system returns D2 above D1 for the original Tuesday-at-20:30 query. Should the team immediately retrain the encoder? Think about the role of the current flag before reading the answer.

Reasoned answer. Not immediately. The archived record is genuinely similar and would be useful for some historical questions. For a current-availability task, an explicit current-document eligibility rule is the direct repair. Check that the flag is trustworthy and that the filter reaches candidate generation. Retraining may improve ranking, but it should not be the only protection against a prohibited version.

Practice 2: a reassuring score. A result has cosine similarity 0.92 and violates the no-booking requirement. Can the system describe it as a 92 percent match to the user’s needs?

Reasoned answer. No. The number is a geometric score under a specific model and representation. It is not a percentage of satisfied conditions. The booking violation is decisive under the stated task. A clearer display would identify the related result and its failed condition, or exclude it from qualifying answers while showing it only as an explicitly labelled alternative if useful.

Practice 3: the missing translated result. Exhaustive vector search places D7 among the nearest five, but ANN search omits it. The English D1 remains first. Has multilingual relevance necessarily failed?

Reasoned answer. The observed fault is candidate-search disagreement, not necessarily multilingual representation failure. The exact vector ranking did include D7. Practical impact depends on the task: a Chinese-language display may suffer, while an English answer using D1 may remain complete. Measure ANN overlap and user-level usefulness separately, then test whether more search effort recovers the missing candidate.

Practice 4: fusion arithmetic. A record ranks first in one list and is absent from another. A second record ranks third in both lists. With the teaching RRF constant 10, which scores higher?

Reasoned answer. The first record receives 1/11, approximately 0.090909. The second receives 1/13 + 1/13, approximately 0.153846. The second wins because agreement across two lists contributes twice. This is a property of the chosen fusion rule, not proof that the second record is more truthful or satisfies more constraints.

Practice 5: duplicate evidence. The top three hits are D1, D7 and another overlapping chunk of D1. Does that establish three independently available rooms?

Reasoned answer. No. These hits refer to the same Cedar option, with translated or overlapping evidence. Group them by room and source relationship. The result set may offer useful language choices, but it contains one distinct room answer. An evaluation measuring alternative venues should not reward duplicates as extra coverage.

Practice 6: an exact code almost matches. A user asks for CED-021, but the highest vector result is a record labelled CED-012. The body discusses similar desk lamps. What should the application do?

Reasoned answer. Preserve the exact identifier requirement. Retrieve or filter on the authoritative code field, with only documented normalisation rules. Similar surrounding prose does not make two codes interchangeable. If CED-021 is absent, report that absence within the searched collection instead of silently substituting CED-012.

Practice 7: a changed score after an upgrade. An old accepted result scored 0.81; after a model migration a clearly irrelevant record scores 0.84. Does the old threshold of 0.80 still provide a reliable acceptance rule?

Reasoned answer. No such reliability follows. Scores belong to particular models, preprocessing and corpora. Re-evaluate the threshold on representative accepted and rejected cases after the migration, including no-answer queries. Also verify that all document vectors were generated with a compatible configuration. A numerically larger score is not evidence that the new result is better.

Back to contents · Continue: 21. Frequently asked questions

21. Frequently asked questions

Is semantic search the same as asking a chatbot?

No. Semantic search retrieves or ranks existing items using learned relationships. A chatbot may use those results, but it can also generate text without retrieval. A search interface can present passages directly without generating an answer. Keeping these functions separate helps diagnose whether a problem came from finding the material or from interpreting it.

Does every semantic search system use one embedding per document?

No. A system can use passage embeddings, several vectors per document, learned sparse features, reranking or combinations of these. “Semantic” describes the intended use of meaning-related signals, not one mandatory storage format. Ask which representation and scoring method the actual system uses.

Can lexical search understand synonyms?

It can use configured synonyms, stemming, normalisation and expansion. The useful distinction is whether relationships are explicitly engineered, learned or combined. Avoid comparing semantic search only with a bare exact-string search if the real alternative is a capable lexical engine with well-maintained domain rules.

Why does search return the opposite of what I asked?

A passage that denies a proposition may discuss the same topic in nearly identical language. It can also be the correct answer-bearing result for a yes-or-no question. Read whether it supports, contradicts or merely mentions the proposed claim. Similarity alone does not encode the direction of the answer reliably enough to assume.

Is a higher similarity score always better?

Within one compatible configuration it usually means a stronger match under that scoring rule, but it does not necessarily mean greater task usefulness. An archived notice can score higher than a current one. Scores from different models or query types may not be comparable. Evaluate the outcome against the actual requirement.

Should every search use hybrid retrieval and reranking?

No universal requirement follows. An exact code lookup may be best served by a deterministic field query. A tiny corpus may allow direct comparison of all passages. A broad, mixed-language collection may benefit from several stages. Use the simplest arrangement that meets measured quality, latency, access and maintenance requirements.

How many candidates should a system retrieve?

There is no context-free number. The useful choice depends on corpus size, filter selectivity, relevant-answer density, ranking quality and downstream cost. Increase the candidate budget when measured misses justify it, and check whether the additional candidates improve useful recall rather than merely adding duplicates or noise.

Can a good embedding model fix stale information?

No. It may retrieve stale information very effectively because that information closely matches the question. Source updates, version labels and index refreshes handle freshness. If the user asks a historical question, old material may be appropriate; the task must specify which time period counts.

Can embeddings safely replace access controls?

No. An embedding score estimates a relationship, not a permission. Eligibility to view a document must be enforced through authorised access rules. Private material should not become visible because it is relevant, and permission changes must propagate through the index and any caches or result stores that expose content.

What should a reader do with a suspicious result?

Identify the exact failed condition: wrong entity, date, number, polarity, language or source status. Open the underlying passage and inspect its context. If the result is merely incomplete, label the missing fact as unknown. This produces a specific diagnosis that a search team can reproduce and repair.

Back to contents · Continue: 22. The practical mental model

22. The practical mental model

Semantic search builds a searchable representation of relationships in language. It can bridge vocabulary differences, connect a question with an answer-bearing passage and make large collections easier to explore. Its usefulness comes from what the representation preserves; its failures often come from what the representation blurs or what the surrounding system fails to enforce.

Use five questions when inspecting a result. What information need was represented? Which words, fields and context entered the representation? How was the candidate found? Why was it ranked above alternatives? Which exact conditions and source claims still require checking? Each question points to a different responsibility and a different possible repair.

In the Harbour example, the meaningful achievement is not merely finding several notices about studying. It is identifying the current Cedar notice for Tuesday at 20:30 without booking, recognising the translated duplicate, rejecting obsolete or ineligible alternatives, and changing the conclusion when the time, day, activity or accessibility requirement changes. That is a much more demanding and useful standard than topical resemblance.

For Super Intelligence systems, semantic search is best understood as a selective access mechanism. It helps bring potentially useful information into view. Human-defined relevance, exact constraints, compatible representations and explicit evaluation make that mechanism dependable enough for its intended role. A nearby passage is a candidate to understand; a usable answer still has to satisfy the question.

Continue with the How Super Intelligence Works guide to connect this mechanism with context, retrieval and the wider system.

Back to contents · Continue: Sources and further reading

Sources and further reading

The room corpus, ranking lists, vectors, arithmetic examples and practice answers are original teaching constructions. They illustrate mechanisms and judgement choices; they are not measured product benchmarks or real venue information.

Back to contents · Return to the full guide

Discover more from eduKateSG

Subscribe now to keep reading and get access to the full archive.

Continue reading