What Is RAG? Retrieval-Augmented Generation Explained
RAG pairs large language models with live document retrieval so answers stay grounded in real sources instead of memorized training data. Here's how the pipeline works.

RAG, short for Retrieval-Augmented Generation, is a technique that connects a large language model to an external knowledge source — a document store, a company wiki, a live search index — so it can look up relevant facts before writing an answer, instead of relying only on what it memorized during training. The result is an AI response that is grounded in retrieved, citable text rather than guessed from the model's internal parameters alone, which is why RAG has become the default pattern for building enterprise AI assistants, support bots and search-augmented chat products.
- What it stands for: Retrieval-Augmented Generation
- What it does: retrieves relevant external text, then feeds it to an LLM as extra context before generation
- Where it came from: a 2020 paper by Patrick Lewis and colleagues at Facebook AI Research (now Meta)
- Core components: an embedding model, a vector database, a retriever, and a generative LLM
- Main benefit: fewer hallucinations and more current answers, without retraining the model
- Main alternative: fine-tuning, which updates the model's own weights instead of feeding it fresh context
What Is RAG, in Plain English?
Large language models learn everything they know during training, then freeze. Ask a plain LLM about something that happened after its training cutoff, or about a private company policy it never saw, and it has two options: say it doesn't know, or guess — and guessing convincingly is exactly how hallucinations happen. AWS defines RAG as the process of optimizing an LLM's output so it references an authoritative knowledge base outside its training data before generating a response, which sidesteps both problems at once: the model gets access to information it never trained on, and it can point to where that information came from.
The idea was formalized in the 2020 paper "Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks," written by Patrick Lewis and co-authors at Facebook AI Research. The paper described combining a pre-trained parametric model (the LLM's learned weights) with a non-parametric memory (a searchable index of documents), and showed that models built this way produced more specific, more diverse and more factual text than generation models that relied on parametric memory alone.
How the Retrieve-Then-Generate Pipeline Works
Strip away the branding and every RAG system follows roughly the same sequence, which Microsoft's Azure Machine Learning documentation and Google Cloud's RAG overview both describe in similar terms:
- Prepare the knowledge base. Source documents — PDFs, wiki pages, support tickets, product manuals — are split into smaller chunks of text.
- Embed the chunks. An embedding model converts each chunk into a vector: a long list of numbers that represents its meaning. These vectors are stored in a vector database alongside metadata for citations.
- Embed the query. When a user asks a question, that same embedding model converts the question into a vector too.
- Retrieve. The system searches the vector database for the stored chunks whose vectors are mathematically closest to the query's vector — a semantic similarity search rather than a keyword match.
- Augment the prompt. The retrieved chunks are inserted into the prompt sent to the LLM, alongside the original question, via prompt engineering.
- Generate. The LLM writes its answer using both the question and the retrieved context, typically citing or quoting the source material it was given.
NVIDIA's developer documentation adds a practical detail that matters in production: document chunk size is a real design decision, with the company's own tutorials recommending chunks in roughly the 100–600 token range to balance how much context a chunk retains against how well it fits inside the model's prompt window.
Vector Embeddings and Vector Databases, Explained
The piece of this pipeline that trips people up is the vector database, so it's worth slowing down on it. An embedding model doesn't store words — it stores meaning. Two sentences that use completely different vocabulary but mean similar things end up with vectors that sit close together in that numerical space. That's what makes retrieval "semantic" instead of a simple keyword search: a question like "how do I get a refund" can still retrieve a document chunk titled "Returns and Reimbursement Policy" even though it shares almost no words with the query.
A vector database is simply a data store built to make that kind of nearest-neighbor search fast across millions or billions of vectors. AWS's own explainer walks through this as a four-step cycle: create external data as embeddings, retrieve the relevant matches for a query, augment the prompt with what was retrieved, and periodically refresh the underlying data so the index doesn't go stale. Cloud vendors each ship a managed version of this: Google Cloud offers Vertex AI Vector Search and the broader Vertex AI Search product; AWS offers Amazon Bedrock Knowledge Bases, which can write vectors into Amazon Aurora, OpenSearch, Neptune, MongoDB, Pinecone or Redis depending on what a team already uses; and Microsoft supports Azure AI Search (formerly Cognitive Search) and open-source options like FAISS inside Azure Machine Learning's RAG tooling.
Why RAG Reduces Hallucination
A plain LLM answering from memory is, structurally, always guessing — it's predicting the statistically likely next words based on patterns learned during training, with no built-in way to check those words against a source. RAG changes the task: instead of "recall this from memory," the model's job becomes "summarize and reason over this specific retrieved text." Wikipedia's summary of the technique frames this as enabling "more factual consistency" and improving "reliability of the generated responses" specifically by grounding generation in retrieved documents rather than parametric memory alone, which helps mitigate hallucination on knowledge-intensive tasks.
Grounding also buys something hallucination-reduction alone doesn't: verifiability. Because the retrieved chunks carry metadata back to their source, a RAG system can show users exactly which document, page or passage an answer came from. Google builds this directly into the Gemini API's Grounding with Google Search feature, which connects Gemini models to real-time web content specifically so responses can "cite verifiable sources beyond its knowledge cutoff" — the model decides when a query needs a live search, runs it, and writes an answer with inline citations back to the pages it used.
It's worth being precise about the limits here too. RAG reduces hallucination; it doesn't eliminate it. Wikipedia's entry notes that a model can still misinterpret retrieved content or state something confidently without fully respecting the context it was given — bad retrieval (wrong or irrelevant chunks) or bad synthesis (misreading correct chunks) can both still produce wrong answers. RAG narrows the problem to those two failure modes instead of leaving the model to free-associate from training data.
RAG vs. Fine-Tuning vs. Plain Prompting
Teams building an AI product usually have three ways to make a general-purpose model behave like a specialist: prompt it carefully, retrieve context for it, or retrain it. These aren't mutually exclusive — many production systems use RAG and a lightly fine-tuned model together — but they solve different problems.
| Approach | What changes | Best for | Cost & speed to update |
|---|---|---|---|
| Plain prompting | Nothing in the model; only the instructions in the prompt | Simple tasks, no external data needed | Free, instant, but limited to what's in the prompt |
| RAG | Nothing in the model; retrieved context is added at query time | Fast-changing data, source attribution, large or proprietary knowledge bases | Cheap; update the index any time, no retraining |
| Fine-tuning | The model's own weights, via additional training | Teaching a consistent style, tone, format, or a narrow skill | Expensive and slower; needs retraining whenever the domain shifts |
Microsoft's own framing of this trade-off is blunt about where each one wins: its Azure Machine Learning documentation describes fine-tuning as "suitable for continuous domain adaptation," enabling real quality improvements but "often incurring higher costs," while RAG "allows the use of the same model as a reasoning engine over new data provided in a prompt," enabling in-context learning "without the need for expensive fine-tuning." AWS makes the same point from the other direction, describing RAG as a way to extend an LLM's capabilities to an organization's internal knowledge base "all without the need to retrain the model." In practice, the deciding factor is usually how often the underlying information changes: a support knowledge base updated weekly is a RAG problem; teaching a model to always respond in a brand's exact voice is more of a fine-tuning problem.
Where RAG Shows Up in Real Products Today
RAG isn't a lab technique anymore — it's shipped infrastructure at every major cloud and AI vendor:
- Amazon Bedrock Knowledge Bases is AWS's fully managed RAG service: it ingests company data from sources like Amazon S3, Confluence, Salesforce and SharePoint, automatically chunks and embeds it, writes the vectors into a supported store such as Aurora, OpenSearch or Pinecone, and handles retrieval and prompt augmentation so developers don't have to wire the pipeline together themselves, per AWS's own announcement of the service.
- Google's Gemini API ships Grounding with Google Search as a built-in tool, letting any application connect Gemini to live web results for up-to-date, cited answers, while Google Cloud's Vertex AI offers both a do-it-yourself path (Document AI plus Vertex AI Vector Search plus Gemini) and a managed one (Vertex AI Search) for enterprise RAG.
- Azure AI Search paired with Azure OpenAI models is Microsoft's reference architecture for enterprise RAG chatbots, with Azure Machine Learning's prompt flow tooling built specifically to wire retrieval, chunking and vectorization together for business use cases.
- NVIDIA's enterprise RAG Blueprint, built on its NeMo Retriever embedding and reranking models, is a production reference pipeline the company publishes so enterprises can stand up high-accuracy retrieval systems rather than build the chunking-and-embedding plumbing from scratch, as described on NVIDIA's developer blog.
These aren't niche integrations — they're the default architecture behind most "chat with your documents" features, internal support assistants, and search-grounded AI chat products shipping today.
How RAG Relates to Agents, MCP and Skills
RAG is often confused with some adjacent AI terms, so it's worth drawing clear lines. RAG is specifically about retrieving text to ground a single generation step. It's a different layer from an AI agent, which is a system that can plan and take multiple actions over time — an agent might use RAG as one tool among several. It's also distinct from the Model Context Protocol, which standardizes how an AI application connects to external tools and data sources (a RAG pipeline's retriever could be exposed to a model over MCP, but MCP itself isn't a retrieval technique). And it's not the same as Claude Skills, which package reusable instructions and scripts for a specific task rather than retrieving documents from a knowledge base. All three can be combined with RAG in a real system, but none of them is RAG.
What RAG Doesn't Fix
RAG is not a magic fix for bad data. If the underlying document store is outdated, contradictory or poorly organized, retrieval will faithfully surface outdated, contradictory or poorly organized chunks, and the model will generate a fluent answer built on them. Chunking strategy matters too: split documents too aggressively and a retrieved chunk loses the surrounding context needed to interpret it correctly; split them too coarsely and irrelevant text dilutes what the model sees. And because retrieval relies on an embedding model's notion of semantic similarity, a RAG system can still retrieve the wrong document for an ambiguous or unusually phrased question, feeding the generator confident-sounding but mismatched context. Keeping the knowledge base current, well-chunked and access-controlled is ongoing operational work, not a one-time setup step.
Quick FAQ
Is RAG the same as fine-tuning? No. Fine-tuning changes a model's own weights through additional training; RAG leaves the model untouched and instead feeds it retrieved context at the moment it answers a question.
Do I need a vector database to do RAG? You need some way to search your documents semantically, and a vector database is the standard way to do that at scale, but the core idea — retrieve, then generate — doesn't require any specific product.
Does RAG completely stop hallucinations? No. It grounds answers in retrieved text, which sharply reduces fabricated claims, but a model can still misread or overstate what it retrieved.
Bottom Line
RAG answers a simple, practical problem: large language models are frozen in time and limited to what they memorized, but most useful questions depend on information that changes constantly or lives behind a company's own walls. By retrieving relevant text from an external source and handing it to the model at the moment it answers, RAG keeps responses current, grounded and traceable to a source — without the cost and delay of retraining the model itself. That combination is why it has become default infrastructure at AWS, Google Cloud, Microsoft Azure and NVIDIA alike, and why it's likely to remain the backbone of enterprise AI assistants and search-grounded chat products for the foreseeable future.
Frequently asked questions
What does RAG stand for?
RAG stands for Retrieval-Augmented Generation, a technique that retrieves relevant external text and feeds it to a large language model as context before it generates an answer.
Is RAG the same as fine-tuning?
No. Fine-tuning retrains a model's own weights on new data, which is slower and more expensive. RAG leaves the model unchanged and instead retrieves fresh context at the moment it answers a question.
How does RAG reduce hallucination?
By grounding the model's answer in retrieved, citable text instead of relying only on what it memorized during training, RAG narrows the model's job to summarizing real sources rather than guessing from memory.
What is a vector database used for in RAG?
It stores documents as numerical embeddings so a system can find text that is semantically similar to a user's question, even when the wording doesn't match exactly.
Who invented RAG?
The technique was introduced in a 2020 paper by Patrick Lewis and colleagues at Facebook AI Research (now Meta), titled 'Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks.'
Which companies use RAG in their products?
AWS (Amazon Bedrock Knowledge Bases), Google (Gemini API grounding and Vertex AI Search), Microsoft (Azure AI Search with Azure OpenAI), and NVIDIA (its enterprise RAG Blueprint) all ship managed RAG infrastructure.
Sources
- AWS - What is Retrieval-Augmented Generation (RAG)?aws.amazon.com
- Google Cloud - Retrieval-augmented generation (RAG) use casecloud.google.com
- Microsoft Learn - Retrieval Augmented Generation using Azure Machine Learninglearn.microsoft.com
- NVIDIA Developer Blog - Tips for Building a RAG Pipelinedeveloper.nvidia.com
- Wikipedia - Retrieval-augmented generationen.wikipedia.org
- Meta AI Research - Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (Lewis et al., 2020)ai.meta.com
Theo Park runs the AI desk at Pandromeda. He follows model launches from the frontier labs and the open-weight community, tracks the assistants and developer tools built on them, and explains what each release changes on pricing, capability and safety. His reporting leans on primary sources: model cards, technical reports, API documentation and the companies' own announcements.

