Glossary
Reference
intermediate
A quick-reference glossary of every significant term you will encounter when building RAG and knowledge-retrieval systems. Definitions are written for engineers: precise, brief, and grounded in how the concept behaves in code.
See also: Cheat Sheet, Vector DB Reference, Troubleshooting.
A
| Term | Definition |
|---|---|
| ANN (Approximate Nearest Neighbour) | A family of algorithms that find the k most similar vectors without scanning every entry in the index. Trades a small amount of recall for large speed gains. HNSW and IVF are the dominant ANN algorithms in production vector databases. |
| Answer Correctness | An end-to-end evaluation metric that compares the generated answer against a ground-truth reference answer, usually measured with an LLM judge or semantic similarity score. Distinct from faithfulness, which only measures agreement with the retrieved context. |
| Approximate Recall | The fraction of true nearest neighbours that an ANN index actually returns. A recall of 0.95 means 5 % of the true top-k results are missed. Tunable via index parameters (ef_search in HNSW, nprobe in IVF). |
| Augmentation | The step in RAG where retrieved document chunks are inserted into the LLM prompt as context before generation. The quality and placement of augmentation directly determines whether the model grounds its answer in the retrieved evidence. |
B
| Term | Definition |
|---|---|
| BGE | A family of open-source bi-encoder embedding models from Beijing Academy of AI. BGE models (e.g. BAAI/bge-large-en-v1.5) consistently rank near the top of the MTEB leaderboard and run locally via sentence-transformers. |
| Bi-encoder | An embedding architecture where query and document are each encoded independently into a vector, and similarity is measured by dot product or cosine distance. Fast at retrieval time because document vectors are pre-computed. Contrast with cross-encoder. |
| BM25 | Best Match 25 — a probabilistic sparse keyword-ranking algorithm. Scores documents by term frequency (how often a term appears) and inverse document frequency (how rare the term is across the corpus). The backbone of Elasticsearch and OpenSearch full-text search; the "sparse" half of hybrid search. |
C
| Term | Definition |
|---|---|
| Chunking | The process of splitting source documents into smaller text segments before embedding. Chunk boundaries determine what the retriever can surface; poor chunking is the single most common cause of low retrieval quality. |
| Chunk Overlap | A deliberate repetition of N tokens at the boundary between adjacent chunks to prevent a key sentence from being split across two chunks where neither chunk carries full context. Typical overlap is 10–20 % of chunk size. |
| Chunk Size | The target token or character length of each chunk. Smaller chunks (128–256 tokens) improve precision but lose surrounding context; larger chunks (512–1024 tokens) preserve context but dilute relevance scores. |
| CLIP | Contrastive Language–Image Pretraining — an embedding model that projects both images and text into the same vector space, enabling cross-modal similarity search (e.g. retrieve images with a text query). Foundation of multi-modal RAG. |
| Cohere Embed | A commercial embedding API from Cohere with multilingual support and an input_type parameter that produces separate query and document embedding spaces for asymmetric retrieval tasks. |
| Context Window | The maximum number of tokens an LLM can process in a single call (prompt + completion). RAG must fit retrieved chunks plus instructions inside this limit; exceeding it requires truncation or summarisation strategies. |
| Context Assembly | The step that formats retrieved chunks into the prompt: ordering chunks (by score, recency, or relevance), adding source citations, applying a system instruction, and fitting within the context window. |
| Cosine Similarity | A measure of the angle between two vectors, ranging from –1 (opposite) to 1 (identical direction). The standard distance metric for dense embedding search. `similarity = (A · B) / ( |
| Cross-encoder | A re-ranking model that takes a (query, document) pair as joint input and outputs a relevance score. Much more accurate than bi-encoder dot-product similarity but too slow to score every document in the index — used to re-rank the top-k ANN candidates. |
D
| Term | Definition |
|---|---|
| Dense Retrieval | Retrieval based on dense vector embeddings and ANN search. Finds semantically similar text even when exact words differ. Complement to sparse retrieval; the two are combined in hybrid search. |
| Document Loader | A component (in LlamaIndex: Reader; in LangChain: DocumentLoader) that reads raw content from a source (PDF, URL, S3, Notion, database) and returns normalised Document objects ready for chunking and embedding. |
| Dot Product | An alternative similarity metric: A · B = sum(a_i * b_i). Equivalent to cosine similarity when vectors are L2-normalised. Pinecone and many FAISS configs use inner-product (dot product) for speed. |
E
| Term | Definition |
|---|---|
| E5 | A family of open-source text embedding models from Microsoft Research (intfloat/e5-large-v2). Designed for asymmetric tasks — queries are prefixed "query: " and documents "passage: " to shift their embeddings appropriately. |
| Embedding | A dense vector of floating-point numbers that encodes the semantic meaning of a piece of text. Produced by an encoder model; texts with similar meanings produce vectors close together in high-dimensional space. A typical embedding has 384–3072 dimensions. |
| Embedding Dimension | The length of the embedding vector (e.g. 1536 for text-embedding-3-small, 768 for bge-base-en-v1.5). Higher dimensions can capture more nuance but cost more storage, index build time, and query latency. |
| Embedding Model | The neural network (usually a transformer encoder) that converts text into an embedding vector. Must be used consistently: the same model for indexing documents and encoding queries; mixing models produces nonsensical similarity scores. |
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("BAAI/bge-large-en-v1.5")
query_vec = model.encode("query: what is cosine similarity?", normalize_embeddings=True)
doc_vec = model.encode("passage: Cosine similarity measures ...", normalize_embeddings=True)
score = float(query_vec @ doc_vec) # dot product of unit vectors == cosine similarity
F
| Term | Definition |
|---|---|
| FAISS | Facebook AI Similarity Search — an open-source C++ library for efficient dense vector search. The inner engine behind many vector databases and the default local backend for LangChain and LlamaIndex dev workflows. |
| Faithfulness | A generation-quality metric: the fraction of claims in the generated answer that can be directly supported by the retrieved context. A faithfulness score of 1.0 means the model invented nothing; 0.0 means every claim was hallucinated. Measured by RAGAS and similar frameworks. |
| Fixed-size Chunking | Splitting text every N characters or tokens regardless of sentence or paragraph boundaries. Simple to implement but can break mid-sentence, degrading retrieval quality. Use recursive character splitting or semantic chunking instead. |
| Flat Index | A brute-force vector index that computes exact similarity against every stored vector. 100 % recall by definition, but O(n) query time. Practical only for collections under ~50 k vectors. |
G
| Term | Definition |
|---|---|
| Generation | The second stage of RAG: passing the assembled prompt (question + retrieved context) to an LLM and obtaining a grounded answer. The LLM should be instructed to answer using only the provided context and to indicate when the context is insufficient. |
| Golden Dataset | A curated set of (question, relevant_document_ids, expected_answer) triples used to benchmark retrieval and generation quality. Building even 50–100 golden examples is the single highest-leverage investment in RAG quality assurance. |
| Graph RAG | An architecture that stores knowledge as a graph of entities and relationships (rather than flat chunks), enabling multi-hop reasoning (e.g. "find all documents involving the same team that owns service X"). Often combines a graph database (Neo4j, Amazon Neptune) with a vector store. |
H
| Term | Definition |
|---|---|
| Hallucination | An LLM generating plausible-sounding but factually incorrect content not grounded in retrieved evidence. RAG reduces hallucination by constraining generation to retrieved context, but does not eliminate it — the model can still misread or ignore context. |
| HNSW (Hierarchical Navigable Small World) | The dominant ANN indexing algorithm. Organises vectors in a multi-layer proximity graph; queries navigate from coarse to fine layers, typically achieving >0.99 recall with sub-millisecond latency at millions of vectors. The default index type in Chroma, Qdrant, Weaviate, and pgvector. |
| Hybrid Search | A retrieval strategy that combines dense vector search (semantic similarity) with sparse keyword search (BM25/TF-IDF) and merges the ranked lists, typically with Reciprocal Rank Fusion (RRF). Recovers keywords that embeddings miss (product codes, names, acronyms) while still finding semantic matches. |
| HyDE (Hypothetical Document Embedding) | A query-expansion technique: ask the LLM to draft a hypothetical ideal answer to the question, embed that text, and use the resulting vector to search the index. Improves recall for short or vague queries because the hypothetical answer shares vocabulary with real documents. |
# HyDE in practice
hypothetical = llm.generate(f"Write a short passage that would answer: {query}")
search_vec = embed_model.encode(hypothetical)
results = vector_db.query(search_vec, top_k=10)
I
| Term | Definition |
|---|---|
| IVF (Inverted File Index) | An ANN algorithm that k-means clusters all vectors at index build time; a query vector is compared only to the nearest nprobe cluster centroids' members, reducing search space. Trades higher build time and configurable recall (nprobe) for good scalability on very large corpora. |
| Index | In vector databases: the data structure that enables fast ANN search (HNSW, IVF, Flat). In traditional databases: a B-tree structure for fast lookups. Choosing the right vector index type is covered in Lesson 3. |
| Ingestion Pipeline | The offline pipeline that loads, cleans, chunks, embeds, and stores documents in the vector database. Distinct from the query-time retrieval pipeline. Must be re-run (fully or incrementally) whenever source documents change. |
K
| Term | Definition |
|---|---|
| Knowledge Graph | A directed graph where nodes are entities (people, products, concepts) and edges are typed relationships (MANAGES, PART_OF, CAUSES). Enables structured queries and multi-hop reasoning that flat vector search cannot perform. |
| k-NN (k-Nearest Neighbours) | A query that retrieves the k most similar vectors to a query vector. The top_k / k parameter in every vector database query call. Typical values are 5–20 for RAG context assembly. |
L
| Term | Definition |
|---|---|
| LangChain | A Python (and JavaScript) framework for building LLM applications. Provides document loaders, text splitters, embedding wrappers, vector store integrations, and chain abstractions. Popular for rapid RAG prototyping. |
| LlamaIndex | A Python framework purpose-built for RAG and knowledge retrieval. Offers a richer set of index types, node post-processors, and evaluation tools than LangChain. Recommended for production RAG systems. |
| L2 Distance (Euclidean Distance) | The straight-line distance between two vectors. Less common than cosine similarity for text embeddings because it conflates semantic distance with vector magnitude. Some indexes (FAISS IndexFlatL2) use it by default — verify your distance metric matches what your embedding model was trained on. |
M
| Term | Definition |
|---|---|
| Metadata Filtering | Applying structured filters (date range, category, author, document_id) before or during vector search to restrict the candidate set. Supported by Qdrant, Pinecone, Weaviate, and pgvector. Reduces noise and enforces access control. |
| MRR (Mean Reciprocal Rank) | A retrieval metric: the average of 1/rank across queries, where rank is the position of the first relevant result. MRR = 1.0 if the top result is always relevant; MRR = 0.5 if the first relevant result is always rank 2. |
| MTEB (Massive Text Embedding Benchmark) | The standard public leaderboard for comparing embedding models across retrieval, classification, clustering, and similarity tasks. Use it to shortlist models, then evaluate the top candidates on your actual data. |
| Multi-modal Embedding | An embedding model that projects more than one data type (text, image, audio, code) into a shared vector space, enabling cross-modal similarity search. CLIP is the canonical example for text + images. |
| Multi-hop Reasoning | Answering a question that requires combining information from two or more distinct documents or facts. Pure vector search struggles with multi-hop queries; knowledge graphs and agentic retrieval loops handle them better. |
N
| Term | Definition |
|---|---|
| Node (LlamaIndex) | LlamaIndex's term for a chunk of text plus its metadata after splitting. A TextNode carries content, embedding, source document reference, and arbitrary metadata fields. |
| nprobe | The IVF index parameter controlling how many cluster partitions are scanned per query. Higher nprobe improves recall at the cost of latency. Analogous to ef_search in HNSW. |
P
| Term | Definition |
|---|---|
| Parent-Child Retrieval | A retrieval strategy that indexes small child chunks for high-precision matching but returns their larger parent chunks to the LLM for fuller context. Prevents the "relevant sentence, missing context" failure mode. |
| pgvector | A PostgreSQL extension that adds a vector column type and HNSW/IVF indexes. Lets existing Postgres users add vector search without a new database. Good choice when your metadata and content already live in Postgres. |
| Pinecone | A fully managed vector database service. No infrastructure to operate; supports metadata filtering, namespaces, and hybrid search (sparse-dense). Suited for teams that want to avoid operational overhead. |
| Precision@k | The fraction of returned results (top k) that are actually relevant. precision@5 = 0.8 means 4 of the 5 returned chunks were useful. High precision reduces noise in the LLM prompt. |
| Prompt Engineering (RAG) | Designing the system and user prompts that instruct the LLM to: use only the retrieved context, cite sources, and state when the context does not contain enough information. A weak RAG prompt wastes good retrieval. |
System: You are a helpful assistant. Answer the question using ONLY the
context below. If the context does not contain the answer, say
"I don't have that information." Cite the source filename after each claim.
Context:
{retrieved_chunks}
Question: {user_query}
Q
| Term | Definition |
|---|---|
| Qdrant | An open-source vector database written in Rust. Supports filtered search, named vector collections, on-disk indexing, and a REST + gRPC API. Popular for self-hosted production deployments. |
| Query Expansion | Generating multiple reformulations of the original query (synonyms, sub-questions, HyDE) and merging their retrieval results to improve recall. |
R
| Term | Definition |
|---|---|
| RAG (Retrieval-Augmented Generation) | An architecture that grounds LLM generation in retrieved documents: (1) embed the query, (2) retrieve relevant chunks, (3) augment the prompt with those chunks, (4) generate an answer. Addresses hallucination, knowledge cutoffs, and private data access. |
| RAGAS | An open-source evaluation framework for RAG pipelines. Measures faithfulness, answer relevance, context precision, and context recall using LLM-as-judge scoring. |
| Recall@k | The fraction of all relevant documents that appear in the top-k results. recall@10 = 0.9 means the retriever found 90 % of the relevant chunks within its top 10. Critical for ensuring the LLM has enough evidence to answer correctly. |
| Reciprocal Rank Fusion (RRF) | A score-fusion algorithm for combining ranked lists from different retrievers (e.g. dense + sparse). RRF_score(d) = sum(1 / (k + rank_i(d))). Parameter k (typically 60) prevents top-ranked items from dominating. |
| Re-ranking | A post-retrieval step where a cross-encoder model scores each (query, chunk) pair and reorders the ANN results by true relevance before passing them to the LLM. Typically improves precision@5 by 10–30 % over raw ANN ranking. |
| Recursive Character Splitting | A chunking strategy that tries to split on \n\n, then \n, then ., then , respecting natural text boundaries before falling back to hard character cuts. The default splitter in LangChain (RecursiveCharacterTextSplitter). |
| Retrieval | The first stage of RAG: given an embedded query, find the k most similar document chunks in the vector index. Retrieval quality is measured by recall@k and MRR against a golden dataset. |
S
| Term | Definition |
|---|---|
| Semantic Chunking | A chunking strategy that splits text at topic-boundary sentences — detected by measuring the cosine similarity drop between consecutive sentence embeddings. Produces more coherent chunks than fixed-size splitting, at higher compute cost. |
| Semantic Search | Retrieval based on meaning rather than exact keyword match. Enabled by embedding both the query and documents in the same vector space and measuring cosine similarity. |
| sentence-transformers | A Python library wrapping HuggingFace transformer encoders for fast, batched sentence embedding. The standard tool for loading and running open-source embedding models (BGE, E5, all-MiniLM-L6-v2, etc.) locally. |
| Sparse Retrieval | Retrieval using high-dimensional sparse vectors where most dimensions are zero, representing term frequency. BM25 is the dominant sparse retrieval algorithm. Fast on keyword matches but blind to synonyms. |
| Sparse-Dense Hybrid | See Hybrid Search. |
T
| Term | Definition |
|---|---|
| TF-IDF (Term Frequency – Inverse Document Frequency) | A classic sparse relevance score. TF measures how often a term appears in a document; IDF down-weights terms that appear in most documents (common words). BM25 is a probabilistic improvement on raw TF-IDF. |
| Top-k | The number of chunks returned by a retrieval query (k). Larger k improves recall but increases prompt size, latency, and the risk of including irrelevant context. Typical values: 5–20 before re-ranking, 3–5 after. |
V
| Term | Definition |
|---|---|
| Vector | An ordered list of floating-point numbers representing a point in high-dimensional space. In RAG, each chunk of text is represented as a vector; proximity in this space encodes semantic similarity. |
| Vector Database | A database optimised for storing, indexing, and querying high-dimensional embedding vectors. Provides ANN search, metadata filtering, and upsert/delete operations. Examples: Chroma, Qdrant, Pinecone, Weaviate, pgvector. |
| Vector Store | Often used interchangeably with vector database. In LangChain/LlamaIndex, a VectorStore is the abstraction class wrapping the actual backend (Chroma, Pinecone, etc.). |
import chromadb
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("BAAI/bge-base-en-v1.5")
client = chromadb.Client()
col = client.create_collection("docs")
col.add(
documents=["RAG grounds LLMs in retrieved context.", "BM25 is a sparse keyword ranker."],
ids=["doc1", "doc2"],
embeddings=model.encode(["...", "..."], normalize_embeddings=True).tolist(),
)
results = col.query(
query_embeddings=model.encode(["how does retrieval work?"], normalize_embeddings=True).tolist(),
n_results=2,
)
W
| Term | Definition |
|---|---|
| Weaviate | An open-source vector database with built-in hybrid search (BM25 + vector), a GraphQL API, and a module system for attaching embedding models. Supports multi-modal vectors natively. |