Skip to content

RAG: Retrieval-Augmented Generation ​

Group: 4 · Grounding (reading the world) | Previous group exit: write and validate input/output schemas, and deliver a cancellable, observable end-to-end interaction (groups 2–3) | This page exit: wire retrieval into the generation loop—answers cite sources, no-hit refuses, and a re-runnable rebuild pipeline exists for corpus updates Prerequisites: Embeddings and Retrieval | Next: Advanced Retrieval, Tool Execution Engineering

1. Overview ​

The model cannot memorize your documents, tickets, and code. RAG (Retrieval-Augmented Generation) first retrieves a few source-attributed passages, then has the model answer from them; if retrieval finds nothing, it refuses. The original paper (Lewis et al., NeurIPS 2020) defines it as a combination of "parametric + non-parametric memory": the parametric memory is a pre-trained generator, the non-parametric memory an external vector index; the paper explicitly lists "providing provenance for decisions" and "updating world knowledge" as open problems for pure parametric models—citations and updates are not RAG accessories but the reason it exists (arXiv:2005.11401, retrievedAt 2026-09-01).

A RAG chain splits in two halves: an offline pipeline (corpus into index) and an online path (question into answer), six stages chained together—any stage failing disguises itself as "the model answered wrong":

StageResponsibilityFailure modeSymptom
ingestfetch, clean, dedupe (multi-format checklist in 3.3)dirty/stale documents indexedanswers carry old versions
chunkcut into retrieval unitsboundaries split semantics; granularity too coarsemisses; citations mismatch
embedtext → vectorsquery and index use different modelsglobally broken results
retrievetop-k recallthreshold/lexical mismatchfalse refusals or noise
rerankprecise reordering and truncation (→ Advanced Retrieval)missingcontext crowded out by noise
generateanswer from passagesno citation constraint; guessing on no-hithallucination; unsourced answers

When to use / when not to ​

  • Use: answers need private corpora (product docs, tickets, policies, code), need citations, or the knowledge updates over time.
  • Do not use: a batch of corpus fits in the context—stuff it (Anthropic's reference line: up to about 200,000 tokens with prompt caching, put the whole knowledge base in the prompt, retrievedAt 2026-09-01); you need to change behavior/style rather than facts—that is fine-tuning territory (→ Learn LLM bridge); an assistant scenario wanting whole-repo context injection—see context engineering first.

Decision table: knowledge injection options ​

OptionDirectionControlStateTrust domainMinimum complexity
RAG (this page)Query → retrieve → contextChunking/index/filtering/thresholds all yoursIndex derived and rebuildable; corpus updates → rebuild indexOnly hit fragments reach the modelMedium (this page's fixture is the minimal shape)
Long-context stuffingCorpus → prompt bulk movePrompt-layer concatFull text resent per request; caching amortizesEntire corpus enters contextLowest (small single-batch corpus)
Fine-tuningCorpus → parametersTraining recipeBaked into weights; update = retrainCorpus enters the training pipelineHighest

Selection order: stuff → RAG → fine-tune. Move to RAG when any of "no longer fits / updates often / needs citations" appears; consider fine-tuning only to change behavior—patching facts with fine-tuning is an anti-pattern; when knowledge changes, rebuild the index, do not retrain.

Historical milestones ​

  • 2020: RAG paper (Lewis et al., NeurIPS 2020, submitted to arXiv 2020-05-22): two formulations (RAG-Sequence conditions on the same retrieved passages across the sequence / RAG-Token can switch passages per token); set then-state-of-the-art on three open-domain QA tasks (abstract re-verified, arXiv:2005.11401, retrievedAt 2026-09-01).
  • 2024-09-19: Anthropic published Contextual Retrieval (contextualized chunks + BM25 + reranking)—publish date cross-checked against multiple independent secondary sources (retrievedAt 2026-09-01); content in Advanced Retrieval.

2. Usage ​

Minimal hands-on: a zero-API-key complete RAG loop. The retrieval side reuses the exact functions from Embeddings and Retrieval; "generation" is a deterministic extractive mock—it picks the best-supported sentence from the retrieved passages and labels its source. Real systems swap generate() for a chat-model call (with a prompt requiring "answer only from the given passages, cite, say so if not found") and change nothing else.

Save as rag.ts (Node 22.18+ / 24 has built-in type stripping—run directly):

ts
// Teaching fixture: complete RAG loop with a deterministic extractive "generator".
// Zero API key: embed/retrieve are the exact functions from the embeddings-retrieval
// page; the "generation" step extracts the best-supported sentence instead of
// calling a chat model, so the whole loop runs offline and deterministically.
// Run: node rag.ts

// ---- Embedder: identical to embeddings-retrieval.ts (hashing bag-of-words) ----

const DIM = 512;

const STOPWORDS = new Set([
  "a", "an", "and", "are", "as", "at", "be", "but", "by", "do", "does", "for",
  "i", "if", "in", "is", "it", "its", "my", "of", "on", "or", "our", "that", "the",
  "this", "to", "we", "were", "what", "when", "where", "which", "who", "why", "you", "your",
]);

function tokenize(text: string): string[] {
  return (text.toLowerCase().match(/[a-z0-9]+/g) ?? []).filter((token) => !STOPWORDS.has(token));
}

function fnv1a(token: string): number {
  let hash = 0x811c9dc5;
  for (let i = 0; i < token.length; i++) {
    hash ^= token.charCodeAt(i);
    hash = Math.imul(hash, 0x01000193) >>> 0;
  }
  return hash >>> 0;
}

function embed(text: string): number[] {
  const vector = new Array<number>(DIM).fill(0);
  for (const token of tokenize(text)) {
    vector[fnv1a(token) % DIM] += 1;
  }
  const norm = Math.sqrt(vector.reduce((sum, value) => sum + value * value, 0));
  return norm === 0 ? vector : vector.map((value) => value / norm);
}

function cosine(a: number[], b: number[]): number {
  let dot = 0;
  for (let i = 0; i < a.length; i++) dot += a[i] * b[i];
  return dot;
}

// ---- Ingest: chunk each document (fixed size + overlap, deterministic) ----

type Chunk = { id: string; text: string; source: string; vector: number[] };

function chunkText(text: string, size = 12, overlap = 4): string[] {
  const words = text.split(/\s+/);
  const chunks: string[] = [];
  for (let start = 0; start < words.length; start += size - overlap) {
    const chunk = words.slice(start, start + size).join(" ");
    if (chunk.trim().length > 0) chunks.push(chunk);
    if (start + size >= words.length) break;
  }
  return chunks;
}

function ingest(source: string, text: string): Chunk[] {
  return chunkText(text).map((chunk, i) => ({
    id: `${source}#${i}`,
    text: chunk,
    source,
    vector: embed(chunk),
  }));
}

const knowledgeBase = [
  ...ingest(
    "runbook.md",
    "The checkout service returns error 502 when the payment gateway times out after ten seconds. " +
      "To mitigate, the gateway timeout must stay below the service timeout. " +
      "On-call runbook: check the payment gateway dashboard first, then restart the checkout pods. " +
      "Escalate to the payments team if the gateway error rate exceeds five percent for five minutes.",
  ),
  ...ingest(
    "oncall.md",
    "The on-call rotation changes every Monday morning. " +
      "Handover notes must list open incidents and pending follow-ups. " +
      "Acknowledge a page within five minutes or it escalates to the secondary on-call.",
  ),
];

// ---- Retrieve: top-k with a floor threshold ----

function retrieve(query: string, topK = 2, minScore = 0.2): Chunk[] {
  const queryVector = embed(query);
  return knowledgeBase
    .map((chunk) => ({ chunk, score: cosine(queryVector, chunk.vector) }))
    .filter((hit) => hit.score >= minScore)
    .sort((a, b) => b.score - a.score || a.chunk.id.localeCompare(b.chunk.id))
    .slice(0, topK)
    .map((hit) => hit.chunk);
}

// ---- Generate: extractive mock. A real system calls a chat model here. ----

type Answer = { text: string; sources: string[] };

function generate(query: string, chunks: Chunk[]): Answer {
  if (chunks.length === 0) {
    // Refusal tier 1: nothing retrieved -> say no instead of guessing.
    return { text: "I cannot answer: nothing in the knowledge base matches this question.", sources: [] };
  }
  const queryTokens = new Set(tokenize(query));
  let best = { sentence: "", overlap: 0, source: "" };
  for (const chunk of chunks) {
    for (const sentence of chunk.text.split(/(?<=\.)\s+/)) {
      const overlap = tokenize(sentence).filter((token) => queryTokens.has(token)).length;
      if (overlap > best.overlap) best = { sentence, overlap, source: chunk.source };
    }
  }
  if (best.overlap === 0) {
    // Refusal tier 2: retrieved chunks exist but none supports an answer.
    return {
      text: "I cannot answer: retrieved passages do not address this question.",
      sources: chunks.map((chunk) => chunk.source),
    };
  }
  return {
    text: `${best.sentence} [${best.source}]`,
    sources: [...new Set(chunks.map((chunk) => chunk.source))],
  };
}

function ask(query: string): void {
  const chunks = retrieve(query);
  const answer = generate(query, chunks);
  console.log(`Q: ${query}`);
  console.log(`A: ${answer.text}`);
  console.log(`sources: ${answer.sources.length > 0 ? answer.sources.join(", ") : "(none)"}`);
  console.log("");
}

// Normal case: a question the runbook can support, answered WITH its source
ask("what to do when the checkout service returns error 502");

// Negative case: off-corpus question -> refusal instead of hallucination
ask("where is the company summer party");

Actual output (deterministic—compare character by character):

text
Q: what to do when the checkout service returns error 502
A: The checkout service returns error 502 when the payment gateway times out [runbook.md]
sources: runbook.md

Q: where is the company summer party
A: I cannot answer: nothing in the knowledge base matches this question.
sources: (none)

Case-by-case reading:

  • Normal path: retrieval hits a runbook chunk; the extractive generator picks the sentence with the highest query overlap and pins the provenance runbook.md onto the answer. sources is a structured field the UI renders as clickable citations—not decorative text.
  • Negative path: an off-corpus question returns empty at the retrieval layer (below threshold) and the generator enters refusal tier 1. Note it does not fall back to parametric memory to invent an answer—that is "rather say no".
  • When you replace generate() with a real model call, the refusal requirement goes into the system prompt. The OpenAI guide's QA sample is worded exactly so: "If the answer cannot be found, say 'I don't know.'" (retrievedAt 2026-09-01).

Acceptance: node rag.ts matches the output above exactly (the supported question carries [runbook.md]; the off-corpus question refuses with sources: (none)). Cleanup: delete the script.

Scenario matrix ​

ScenarioInput / actionOutputFitsDoes not fit
Basic: document QAQuestion → retrieve → generateAnswer + source listProduct docs / policy FAQOpen-ended chat
Common: refusal pathOff-corpus / low-score question"Not in the knowledge base"Every RAG entryPadding answers from outside the threshold
Combined: rebuild after updateDocument change → fingerprint compare → incremental rebuildNew index entriesFrequently changing knowledgeFull re-embedding every time (wasteful)

3. Principles ​

3.1 Traceability is a product interface ​

A RAG answer's trustworthiness comes from every claim tracing back to a source document. In engineering terms: chunks enter the index with stable ids (source#index) and metadata (path, title, permissions, content fingerprint); retrieval passes chunk ids into generation; the answer's citations ([1], [runbook.md]) are rendered as clickable backlinks by the consumer. A RAG without citations is just "chat that read something"—it cannot be audited.

3.2 The no-hit path: rather say no ​

No-hit is not an exception branch but product behavior, handled in two tiers: retrieval returns nothing (refusal tier 1); retrieval returned results but none supports the question (refusal tier 2, still listing what was looked at). Both tiers need explicit user copy and instrumentation—the false-refusal rate is a core RAG operating metric (methodology at evals). The opposite is silently passing emptiness to the model and letting it improvise: the direct source of hallucination.

3.3 Multi-format ingestion: corpora are rarely plain text ​

The first lesson of ingest: real corpora are seldom plain text. Markdown, PDF, web pages, spreadsheets, and scans each have their own extraction path and pitfalls; the cost of extracting the wrong shape only surfaces at retrieval time—"should have hit but didn't", and it is hard to attribute. A multi-format checklist (coverage parallels the structure of Ch.1 of Huang Jia's RAG in Practice; tool ideas are generic approaches, not product endorsements):

FormatTool approachDominant failure mode
Markdown / HTMLparse into a heading tree, chunk on structural boundariesnavigation/footer boilerplate leaks into chunks; source and rendered text diverge
PDF (text layer)extract the text layer (e.g. pypdf / pdf-parse), keep page numbers for citation backlinkstwo-column reading order scrambles; tables shatter into fragments
Scans / imagesOCR first (e.g. Tesseract or a cloud OCR), store confidence in metadatalow-confidence garbage enters the index and pollutes it—retrieval "answers wrong" with no clear cause
Web pagesfetch + main-content extraction (e.g. readability), record fetch timestampsilent content loss on site redesign; dynamically rendered parts missed
CSV / tableschunk per row, prepend headers to every recorda whole table in one chunk loses column semantics; cross-row aggregation questions go unanswered
Office (docx / pptx)unpack the XML or convert to Markdown, then chunkcomments / revisions / hidden slides leak in; untitled slides lose structural cues

Two format-independent guardrails: normalize every ingest product into an internal representation of "plain text + metadata (source, page or line number, fetch time)", so downstream chunk / embed / index never sees the source format; and route low-confidence OCR and failed parses to a quarantine queue for human review, not into the index. Hosted reference points: OpenAI vector stores accept doc / docx / pdf / html / md / pptx / json / code formats directly and auto-chunk (official MIME table, retrievedAt 2026-09-01). Multimodal embedding is the other route: Cohere embed-v4.0 embeds mixed text+image content such as screenshots and slide decks directly (officially positioned as removing the text-extraction ETL, retrievedAt 2026-09-01), suited to fidelity-first corpora.

3.4 The update pipeline: rebuilding when data changes ​

RAG's unit of knowledge update is rebuilding the index, not retraining the model. A re-runnable incremental pipeline:

text
document change → compute a content fingerprint per chunk (e.g. SHA-256)
               → compare against the manifest's old fingerprints: same → skip, different → re-embed and upsert
               → document deleted → remove all its chunks by stable-id prefix
               → persist the manifest (id → fingerprint) as the next baseline

Stable ids (docId:chunk:i style) make updates and deletes deterministic; fingerprint comparison avoids paying to re-embed unchanged content. (Pattern distilled from the on-site semantic-search case study; the case itself lives in the appendices. For upsert semantics, see your vector DB's docs.)

3.5 Invariants ​

  • Index and query use the same embedding model: changing models requires an index rebuild; validate model name and dimensions at startup.
  • Permission filtering at retrieval time: ACLs enter retrieval predicates, not post-generation hiding (see Embeddings and Retrieval).
  • The answer exit uniformly goes through "cite or refuse": any path that bypasses retrieval and generates directly is a hallucination backdoor on this chain.

Spec vs. local measurement ​

ClaimOfficial/spec positionThis page's fixture
Chunk granularityAnthropic: usually a few hundred tokens per chunk12 words per chunk, 4-word overlap (demo)
Refusal promptingOpenAI sample: "if it cannot be found, say I don't know"Generator has built-in two-tier refusal, prompt-independent
Citation shapeOpenAI file search returns provenance via file_citation annotationssource field + inline [runbook.md]
Update semanticsVector DBs provide upsert (update if exists)Not implemented (pipeline described in 3.4)
Hosted RAGOpenAI file search: vector stores + built-in semantic/keyword searchThis page hand-writes the full chain, for teaching

4. Development ​

Integration, testing, rollback ​

  • Into the product interaction (group 2): after swapping in a real model, stream answer rendering reuses Streaming; the citation list arrives once after the stream ends, with sources.
  • Version pinning: the index records the embedding model version; upgrading = new-version index + alias switch + golden-set comparison, never in-place overwrite.
  • Testing: at least 20 real questions as fixtures—half "should hit" (answers must cite the right document), half "should refuse" (must refuse). Testing only "feels fluent" is not acceptance.
  • Rollback: keep the previous index version; the traffic switch lives on the alias—rollback is pointing it back.

Symptom → Evidence → Action → Done when ​

Symptom: the cited passage does not match the answer. Evidence: replay sampled answers against the text their chunk ids point to; check whether boundaries split "symptom" and "fix" into different chunks. Action: re-chunk on structural boundaries (headings/paragraphs) with overlap; rebuild with stable ids. Done when: 20 sampled answers all cite passages that support them.

Symptom → Evidence → Action → Done when ​

Symptom: many answerable questions get refused. Evidence: refusal list plus those questions' score distribution (separate "vacuum" from "borderline"). Action: borderline → adjust the threshold; vacuum → lexical/semantic mismatch—add hybrid search and query rewriting (→ Advanced Retrieval). Done when: "should hit" items all hit and "should refuse" items still all refuse on the golden set.

Symptom → Evidence → Action → Done when ​

Symptom: after a document update, answers still give the old version. Evidence: compare index-entry fingerprints against current document content; check whether the update pipeline runs and stale chunks are cleaned. Action: add fingerprint comparison and deletion per 3.4; hook index rebuilds into the document publishing flow. Done when: within N minutes of publishing, the same question's answer points to the new version.

Symptom → Evidence → Action → Done when ​

Symptom: an unauthorized user asked out restricted content. Evidence: run the unauthorized case set with a low-privilege account; check retrieval logs for ACL predicates. Action: fold ACLs into retrieval filtering; all retrieval entries share one filter builder. Done when: the unauthorized set is empty; normal-user recall unchanged.

Anti-patterns ​

  • Fine-tuning instead of updating documents. Knowledge changed → rebuild the index; fine-tuning changes behavior, not the fact source.
  • Embedding the whole repo as one chunk. Granularity too coarse—retrieval dilutes, citations blur.
  • Treating a hit as fact. High similarity only means "similar"; generation must still be constrained to the given passages.
  • Letting the model answer on no-hit. Empty results must short-circuit to refusal copy.
  • Accepting fluency only. A RAG shipped without a golden set has no acceptance at all.

5. Resource Library ​

Four-level reading route ​

  • Beginner: run this page's fixture → the OpenAI Retrieval guide (what hosted RAG looks like).
  • Builder: the OpenAI file search tool docs (vector stores, citation annotations, metadata filtering) → the Anthropic Contextual Retrieval post (where retrieval quality ceilings are).
  • Operator: this page's runbooks → evals (turn citation correctness / false-refusal rate into release gates).
  • Researcher: the RAG paper (Lewis et al. 2020) → Learn LLM chapters 11–12 (mechanisms and evaluation of chunking, recall, citation correctness).

Resource table ​

NameLevelcanonical URLUseSupported claimNext
RAG paperL0https://arxiv.org/abs/2005.11401Definition, parametric/non-parametric memory, provenance & update motivationRAG-Sequence/Token; SOTA on three QA tasksRead the experimental setup
Anthropic: Contextual RetrievalL1https://www.anthropic.com/news/contextual-retrievalRetrieval failure modes and the 200k-token stuffing reference lineContext loss on chunking; 49%/67% gains (expanded in Advanced Retrieval)Try its cookbook
OpenAI: RetrievalL1https://developers.openai.com/api/docs/guides/retrievalSemantic search, attribute filtering, chunk defaults, response synthesisFilter operators; 800/400 defaultsWire up file search
OpenAI: File searchL1https://developers.openai.com/api/docs/guides/tools-file-searchHosted RAG shape and citation annotationsfile_citation; semantic + keyword hybridCompare against a self-built chain
OpenAI: Vector embeddingsL1https://developers.openai.com/api/docs/guides/embeddingsRefusal wording in the QA sample"If it cannot be found, say I don't know"—

All pages retrievedAt 2026-09-01.

Active falsification and open questions ​

  • "Up to about 200,000 tokens can be stuffed" is Anthropic's reference line (with prompt caching); windows and prices change—recheck on the day of use.
  • This page's extractive generator cannot hallucinate, so it also cannot validate "how well prompt constraints suppress real-model hallucination"—that question belongs to controlled evaluation at evals.

Where learn-ai stops / where to go next ​

This page delivers the minimal RAG chain with citations and refusals. Retrieval quality upgrades (hybrid / rerank / contextualized chunking) → Advanced Retrieval; handing retrieval to an Agent as a tool → Tool Execution Engineering; mechanisms and evaluation of chunking and recall → Learn LLM chapters 11–12 and evals.

Built for frontend engineers · Powered by VitePress