Retrieval & Data Entry

Retrieval augmented generation, explained

Reviewed October 2026

TL;DR: RAG (retrieval augmented generation) fetches the most relevant pieces of your data at question time and hands them to a language model along with the question. The model answers from those sources instead of from memory alone - which means current information, private data, and citations, without retraining anything.

How it works

A RAG system has two phases. At indexing time, your documents are split into chunks, each chunk is converted into an embedding (a list of numbers capturing its meaning), and the embeddings are stored in a vector database. At query time, the user's question is embedded the same way, the database returns the chunks whose embeddings sit closest to the question's, and those chunks are pasted into the model's prompt as context. The model then writes its answer grounded in what it just read.

The quality levers live in the details: how documents are chunked, which embedding model is used, how many chunks are retrieved, and whether a reranker re-sorts the candidates before the model sees them. A RAG pipeline that answers badly is almost always retrieving badly - the model can only be as good as the context it is given.

RAG exists because stuffing everything into the prompt does not work. Models know nothing after their training cutoff, and while context windows have grown to a million tokens and beyond, accuracy and recall degrade as one fills - the effect commonly called context rot - so more context is not automatically better. Retrieval also costs far less per answer than re-reading a corpus on every question, and it is where document-level permissions get enforced. Retrieval turns an open-book exam into a short one - find the right page first, then read it.

A second pattern has grown up next to this one. Rather than chunking and embedding a corpus in advance, an agent is handed search tools and left to look things up as it works: grep across a repository, a query against a database, a file read. Anthropic calls this just-in-time context and uses it in its own coding agent. Its guidance is that teams are augmenting embedding-based retrieval with the technique rather than replacing it, because runtime exploration is slower than retrieving pre-computed data, and that the most effective agents often combine the two. The practical division: agentic search suits a corpus that is already navigable and changes constantly, such as a codebase or a ticket system, while chunk-and-embed retrieval wins where you need low latency, predictable cost, permission filtering, and recall across more documents than any agent could read.

Where it sits in the AI stack

RAG is the bridge between your data layer and the model. It consumes embeddings, lives alongside the vector database, and feeds the model's context window:

Key tools and implementations

  • pgvector

    Vector search inside Postgres - the boring-but-proven default when your data already lives there.

  • LlamaIndex

    A Python framework for ingestion, indexing and retrieval pipelines. Still maintained, though the company behind it has shifted its focus to document parsing.

  • LangChain

    General LLM orchestration with a large ecosystem of retrievers, loaders, and RAG chains.

  • Managed retrieval APIs

    Hosted stores and file-search tools - OpenAI's vector stores, Google's File Search, Bedrock Knowledge Bases, Azure AI Search - that bundle chunking, embedding and retrieval behind one API. Anthropic is the exception: it ships citation-ready search results and a Files API but no hosted vector store, and no embedding model of its own.