VIEW THIS AS

Auto mode follows the Route Engine until you choose a viewpoint.

YOU ARE HERE

ROUTE CHECK

CONNECTED TO

WHAT NEXT

Use the canonical route for this room, or HELP if you are unsure.

Retrieval-Augmented Generation Vocabulary

THE CORE AIM OF VOCABULARY MASTERY · RETRIEVAL-AUGMENTED GENERATION VOCABULARY · QUERY → RETRIEVE → RANK → CONTEXT → GENERATE

Retrieval-augmented generation vocabulary describes systems that retrieve external information and supply it to a generative model before answering. Core terms include query, chunk, embedding, vector search, retrieval, ranking, context, grounding and citation.

The core aim of vocabulary mastery for retrieval-augmented generation vocabulary is evidence-flow clarity: explain what was searched, what was retrieved, why it was selected, what entered model context and whether the answer is supported by the source.

This page is the RAG Vocabulary owner inside the eduKateSG Vocabulary hub, linked to Generative AI Vocabulary and Large Language Model Vocabulary.

Central proposition: RAG vocabulary is mastered when a claim can be traced from query to retrieved evidence to context to generated answer.


The RAG Vocabulary Router

  • Prepare: document, chunk, metadata, index.
  • Represent: embedding, vector, similarity.
  • Retrieve: query, search, top-k, filter.
  • Rank: relevance, reranker, score.
  • Generate: context, grounding, answer.
  • Verify: citation, source, faithfulness, recall.

Retrieval and Generation Are Different

Retrieval finds existing information. Generation produces language conditioned on available context. Combining them does not guarantee the answer correctly uses the evidence.

Chunk and Embedding

A chunk is a segment indexed for retrieval. An embedding is a numerical representation used to compare semantic similarity. Chunk boundaries and representation quality strongly influence what can be found.

Ranking and Reranking

Initial retrieval may return many candidates. A reranker applies another relevance stage to reorder them before selected passages enter generation context.

Grounding and Faithfulness

Grounding connects generation to evidence. Faithfulness asks whether generated claims actually follow from that evidence. A citation is useful only when the cited material supports the associated claim.

How to Learn RAG Vocabulary

  • Index a small document set.
  • Experiment with chunk boundaries.
  • Compare keyword and semantic retrieval.
  • Inspect retrieved passages before generation.
  • Rerank candidates.
  • Check citations claim by claim.
  • Separate retrieval failure from generation failure.

Common Mistakes

Assuming retrieval guarantees correctness

Repair: verify how evidence was interpreted.

Confusing embedding with stored text

Repair: embeddings are numerical representations.

Treating citations as automatic proof

Repair: check whether each source supports the claim.

Frequently Asked Questions

What is RAG?

Retrieval-augmented generation combines external information retrieval with generative-model synthesis.

What is vector search?

It retrieves items based on similarity between vector representations.

What is a reranker?

It is a second-stage method that reorders retrieved candidates by relevance.

The RAG Vocabulary Standard

Mastery means diagnosing whether a weak answer came from retrieval, ranking, context construction or generation.

That is the standard: evidence language that keeps search and synthesis visibly connected.

Discover more from eduKate Singapore

Subscribe now to keep reading and get access to the full archive.

Continue reading