Skip to content

Group 4 · Grounding ​

Group: 4 · Grounding (reading the world) | Previous group exit: write and validate input/output schemas, and deliver a cancellable, observable end-to-end interaction (groups 2–3) | This group exit: build a traceable retrieval chain and an update pipeline Prerequisites: Streaming | Next: Embeddings and Retrieval, Tool Execution Engineering

1. Overview ​

This group addresses one specific symptom: answers lack private or fresh facts. A model's parametric memory is frozen the moment training ends—it does not know your tickets, your code, or the document you changed last week. This group ties answers to external evidence: retrieve first, generate second, attach provenance to every answer, and rebuild when the data changes.

The original RAG paper (Lewis et al., NeurIPS 2020) named this structure "parametric memory + non-parametric memory": the parametric memory is a pre-trained generator, the non-parametric memory is an external vector index accessed by a retriever (in the paper, a dense vector index of Wikipedia). The paper also states that for pure parametric models, "providing provenance for their decisions" and "updating their world knowledge" remain open problems—exactly the two things this group delivers (a traceable retrieval chain, a re-runnable update pipeline) (arXiv:2005.11401, abstract re-verified, retrievedAt 2026-09-01).

This group shares the intuition of letting the model touch things outside itself with the Action group (05), but the direction is opposite: reading the world vs. writing the world. Everything here is a read—chunking, indexing, retrieving, citing—with no side effects by default; the failure mode is "wrong answer" or "should have refused but didn't", never destruction. The Action group writes to the world (executing actions); its failure mode is "broken state". Hence the different acceptance criteria: this group audits citation correctness and refusal correctness; the Action group audits restricted permissions and recoverability.

When to enter this group / when not to ​

  • Enter: answers need private corpora (product docs, tickets, code), need citations, or the knowledge updates frequently.
  • Do not enter: answers are unstable or unparseable—go back to Group 3 Context (the output-shape landing point: structured output); not yet wired into a product interaction—start Group 2 Inference & Interface; you need the system to take actions—go to Group 5 Action.

Symptom → topic navigation ​

SymptomGo toAfter reading you can
Want search by meaning; keywords missEmbeddings and RetrievalBuild a vector retrieval entry with metadata filtering and a no-hit path
Building "answer from our docs" featuresRAG: Retrieval-Augmented GenerationRun the full ingest→chunk→embed→retrieve→generate chain with citations and refusals
Baseline RAG recall not enough / error codes not foundAdvanced RetrievalRaise recall with hybrid search, query rewriting, reranking; pin permissions at the retrieval layer

Decision table: knowledge injection options ​

OptionDirectionControlStateTrust domainMinimum complexity
Long-context stuffingCorpus → prompt (bulk move)Prompt-layer concat, no retrieval controlFull text resent per requestEntire corpus enters model contextLowest (small single-batch corpus)
RAG (this group)Query → retrieve → context (selective move)Index, chunking, filtering all yoursIndex is derived, rebuildableOnly hit fragments reach the modelMedium (embedding + index + pipeline)
Fine-tuningCorpus → parameters (rewrite the model)Training recipe and data mixBaked into weights; update = retrainCorpus enters the training pipelineHighest (see Learn LLM bridge)

Selection principle: start at the lowest complexity. If a batch of corpus fits in the context (Anthropic's reference line is about 200,000 tokens, with prompt caching), stuff it first; move to RAG when the corpus is large, updates often, or citations matter; consider fine-tuning only to change behavior style, not facts.

Historical milestones ​

  • 2016: HNSW approximate nearest-neighbor index published (Malkov & Yashunin, arXiv:1603.09320); became the mainstream index foundation of vector databases (retrievedAt 2026-09-01).
  • 2020: RAG paper published (Lewis et al., NeurIPS 2020, arXiv:2005.11401), establishing the "parametric + non-parametric memory" framing (retrievedAt 2026-09-01).
  • 2024-09-19: Anthropic published Contextual Retrieval (contextualized chunks + BM25 + reranking); experiment data in Advanced Retrieval. The publish date was cross-checked against multiple independent secondary sources (retrievedAt 2026-09-01).
  • 2026-09: the v6 restructure (Issue #116) replaced the v5 six-layer pyramid with ten dependency-ordered groups; this group went from "Layer 3 · Knowledge Grounding" to "Group 4 · Grounding (reading the world)". The read/write boundary and acceptance criteria are unchanged.

2. Usage ​

This group's minimal hands-on is the zero-key example in Embeddings and Retrieval: a pure-TypeScript deterministic vector retriever that runs in a clean environment within 15 minutes.

bash
# Save the full example from the embeddings-retrieval page as embeddings-retrieval.ts, then:
node embeddings-retrieval.ts

Acceptance: three output groups present—a lexically overlapping query hits the right document with its source; a paraphrased query (zero lexical overlap) correctly takes the no-hit path; an off-corpus question correctly refuses. Once it runs you have seen this group's two core behaviors: a score is not truth (high similarity only means "similar"), and no-hit is a path, not an exception.

Every topic's Usage section follows the same constraints: zero API keys, single self-contained file, deterministic output, positive and negative paths in pairs.

3. Principles ​

On the capability chain, this group upgrades Group 2's "one interaction" into "one grounded interaction":

text
Offline: corpus → chunk → embed → index (text + vectors + metadata + content fingerprint)
Online:  question → retrieve top-k → filter (permissions/tenant) → rerank → generate (with citations)
Update:  document change → fingerprint compare → incrementally rebuild entries → delete stale ones
  • Invariant 1: cite, or say there is nothing. Every answer must trace back to a source document; if retrieval finds nothing, refuse—never let the model fill the gap from parametric memory.
  • Invariant 2: permission filtering happens at retrieval time. Documents a user cannot see must not be retrievable by them; filtering happens during recall, not by hiding after generation.
  • Invariant 3: the index is a derived artifact. It is determined by "corpus + chunking strategy + embedding model version"; any change requires a rebuild; the index can always be re-derived from the corpus.

Exit criteria: build a traceable retrieval chain and an update pipeline—any answer traces back to its source document and chunk; questions that should be refused are refused; after corpus updates, a re-runnable rebuild pipeline exists.

Per-topic details: embeddings-retrieval, rag, advanced-retrieval.

4. Development ​

Before entering the layer, use three diagnostics to locate the right page.

Symptom → Evidence → Action → Done when ​

Symptom: the model fabricates private facts, confidently. Evidence: ask the same question to the bare model and to the retrieval-backed product; check retrieval logs for hits. Action: confirm the question actually goes through the retrieval chain; add no-hit refusal and citation constraints per rag. Done when: every "should refuse" item in the golden set is refused; every "should hit" answer carries a correct citation.

Symptom → Evidence → Action → Done when ​

Symptom: the answer cites a passage that does not match what it says. Evidence: replay sampled answers against the text their chunk ids point to; check whether chunk boundaries cut semantics apart. Action: re-chunk per embeddings-retrieval (structure-first boundaries, with overlap); rebuild the index with stable ids. Done when: 20 sampled answers all cite passages that support them.

Symptom → Evidence → Action → Done when ​

Symptom: the document changed; the answer still gives the old version. Evidence: compare index entry timestamps/content fingerprints against the current document version. Action: fix the indexing pipeline per rag (content fingerprint + incremental rebuild + stale-entry deletion). Done when: after the rebuild triggered by the document update, the same question's answer points to the new version.

5. Resource Library ​

Four-level reading route ​

Active falsification and open questions ​

  • "About 200,000 tokens fits by stuffing" comes from the Anthropic Contextual Retrieval post (retrievedAt 2026-09-01); actual windows and cache prices change—recheck the official docs on the day you use them.
  • The hybrid-search and reranking gains are Anthropic's experimental results on their corpus and configuration; run your own golden-set comparison before porting the numbers.

Where learn-ai stops / where to go next ​

This group answers "how to ground results". It does not answer "embedding training objectives and vector geometry" (→ Learn LLM), "evaluation methodology and benchmarks" (→ evals), or "executing actions and permissions" (→ Group 5 Action).

Built for frontend engineers · Powered by VitePress