The GenAI Field Guide

RAG and knowledge systems

Find trustworthy passages before asking the model to answer.

A language model only knows what was in its training data, frozen at a cutoff, plus whatever you put in the prompt. Most useful business questions depend on neither: they depend on this quarter's policy, a customer's contract, last night's incident report. Retrieval-augmented generation (RAG) closes that gap by searching your own content at question time and handing the best passages to the model as evidence. The model stops being the source of facts and becomes a reader and writer over sources you control, which is also what makes answers citable, permission-aware and fixable without retraining.

The mental model is a search engine bolted to a writer, and the search engine does most of the work. A typical pipeline has an offline half (ingest documents, split them into chunks, embed and index them with metadata) and an online half (rewrite the query, retrieve candidates with keyword and vector search, rerank, assemble a prompt, generate an answer with citations). When a RAG system answers wrongly, the cause is usually upstream of the model: the right passage was never extracted, was split away from its context, was outranked, or was filtered out. Teams that treat RAG as a retrieval problem first and a prompting problem second ship better systems.

The questions follow the pipeline. The Basic questions define RAG and walk through each stage: ingestion, chunking, embeddings, vector, keyword and hybrid search, reranking and citations. The Advanced questions cover the decisions that separate a demo from a production system: how to size chunks, how to measure retrieval with recall@K and precision@K, when relationship-aware Graph RAG helps, how to enforce document permissions, when to fine-tune instead, how to keep an index current, how to rewrite queries, how to debug a wrong answer stage by stage, and when to let the model drive retrieval itself.

What is RAG?

Retrieval-augmented generation (RAG) is a pattern where an application searches a knowledge source for passages relevant to a question, puts them in the prompt, and asks the model to answer from that evidence rather than from memory.

Retrieval-augmented generation splits answering into two jobs. A retriever finds a handful of relevant passages from a corpus you control: policy documents, tickets, product manuals, a database. A generator, the language model, reads those passages along with the question and writes the answer. The term comes from a 2020 research paper that combined a neural retriever with a sequence-to-sequence model, but in practice today it means any system that fetches context at query time and grounds a model's answer in it.

The reason it works is that models are much better at reading and synthesising text placed in front of them than at recalling specific facts from their weights. Recall from weights is lossy and frozen at the training cutoff; the model may half-remember a policy and fill the gap with something plausible, which is a hallucination. Text in the context window is exact and current. RAG turns a memory problem into a reading-comprehension problem, which is the one models are good at.

A production pipeline has an offline half and an online half. Offline, documents are parsed, split into chunks, embedded into vectors and stored in an index along with metadata such as source, date and access rights. Online, the user's question is optionally rewritten, used to search the index, the top results are reranked, and the best few are inserted into a prompt that instructs the model to answer only from them and cite which passage supports each claim.

The trade-off is that you now own a search system. Answer quality is capped by retrieval quality: if the right passage is not in the top results, the model either says it does not know or, worse, answers anyway from memory. RAG also adds latency (an embedding call, a search and often a reranking step before generation starts) and adds tokens to every request. Those costs are usually worth it for any question whose answer lives in your documents rather than in general knowledge.

Why use RAG?

Use RAG when answers depend on information the model was not trained on or that changes: private documents, recent updates, per-customer data. It adds that knowledge at query time, with citations and access control, without retraining anything.

A model's built-in knowledge has three gaps that matter in business settings. It is stale, ending at a training cutoff. It is public, containing nothing from your intranet, contracts or tickets. And it is unattributable: the model cannot tell you which document a fact came from, so nobody can verify it. RAG addresses all three by supplying the relevant text at the moment of the question.

Updating knowledge becomes a data operation rather than a training run. When the travel policy changes on Monday, you re-index one document and the assistant answers from the new version within minutes. Fine-tuning cannot do this reliably: it is slow, expensive to repeat, and research consistently shows that training is a poor way to inject new facts that a model will recall precisely. Retrieval also lets one model serve many customers, because each request retrieves only from that tenant's documents.

RAG also makes answers auditable. Because the evidence is explicit, you can show citations to users, log which passages supported which answer, and evaluate whether the answer actually follows from them. When an answer is wrong, you can see whether the source was wrong, missing or misread, and fix the right thing.

The main alternative today is long context: putting entire documents into the prompt. When the whole corpus fits comfortably, say a 40-page contract, that is often simpler and more accurate than retrieval. It stops working when the corpus is thousands of documents, when per-request cost and latency matter, or when users have different permissions. Even very long context windows tend to attend less reliably to material buried in the middle, so curating fewer, better passages still pays. RAG is not needed for tasks that do not depend on specific facts, such as rewriting an email.

What is document ingestion?

Document ingestion is the offline pipeline that turns raw sources (PDFs, web pages, wikis, tickets) into clean, structured, searchable chunks with metadata, ready for indexing. Its quality sets the ceiling for everything retrieval can do later.

Ingestion covers everything between a file landing in a folder and a chunk sitting in an index: connecting to the source, extracting text, recovering structure, cleaning, chunking, attaching metadata, embedding and writing to the store. It is the least glamorous part of RAG and the one that most often decides whether the system works.

Extraction is harder than it looks. PDFs store positioned glyphs, not paragraphs, so a naive extractor can merge two columns line by line, split words with hyphens, or flatten a table into a stream of numbers with no headers. Scanned documents need optical character recognition. Slides, spreadsheets and HTML each have their own traps, such as navigation menus repeated on every page. Layout-aware parsers and, increasingly, vision-capable models can recover headings, lists and tables, and converting everything to a consistent format such as Markdown makes later chunking much easier.

Metadata is the other half. Every chunk should carry its source document id, title, section path, page number, URL or file path, last-modified date, version, language and the access-control groups allowed to read it. Metadata powers filtering ("only 2026 policies"), permissions, citations, freshness and deletion. A chunk without a document id cannot be removed when its source is deleted.

Treat ingestion as a repeatable, versioned pipeline rather than a one-off script. Store a content hash per document so unchanged files are skipped, keep the parsed intermediate output so you can re-chunk without re-parsing, and record which parser and embedding model version produced each chunk. When you later change chunk size or embedding model, you will need to reprocess everything, and a pipeline you can rerun saves days.

What is chunking?

Chunking is splitting documents into smaller passages that are indexed and retrieved individually. Good chunks are self-contained units of meaning, small enough to match a specific question and large enough to answer it.

Retrieval returns chunks, not documents, so the chunk is the unit of everything downstream: what gets embedded, what gets matched, what the model reads and what gets cited. Chunking exists for two reasons. Embeddings compress a passage into one vector, so a passage covering ten topics produces a blurry vector that matches none of them well. And context is limited, so you want to send the model the three paragraphs that matter, not the 80-page manual.

The common strategies range from simple to structural. Fixed-size chunking cuts every N tokens, often with an overlap of 10 to 20 percent so a sentence on the boundary appears in both chunks. It is easy but cuts through the middle of ideas. Recursive splitting tries paragraph breaks first, then sentences, then words, to stay under a size limit. Structure-aware chunking follows headings, sections, list items or table boundaries, which usually matches how authors grouped ideas. Semantic chunking places boundaries where the embedding similarity between consecutive sentences drops, though its benefit over structure-aware splitting is inconsistent in published comparisons.

A chunk often loses context that the full document had. A paragraph saying "this limit does not apply to contractors" means little without knowing which limit. Two cheap fixes help a lot: prepend the document title and section path to each chunk before embedding, and optionally add a one-sentence model-generated summary of where the chunk sits in the document. Another pattern is small-to-big retrieval: match on small chunks for precision, then pass the larger parent section to the model for context.

There is no universally right size. It depends on the documents, the embedding model and the kind of questions, which is why chunk size should be chosen by testing retrieval on real questions rather than copied from a tutorial.

What is an embedding model?

An embedding model converts a piece of text into a fixed-length vector of numbers, positioned so that texts with similar meaning land close together. In RAG it encodes both chunks and queries so they can be compared by meaning.

An embedding is a list of numbers, typically a few hundred to a few thousand long, that represents a text's meaning as a point in space. An embedding model is a neural network, usually a transformer encoder, trained so that related texts get nearby vectors. The training is mostly contrastive: the model sees pairs that should match (a question and the passage that answers it) and many that should not, and learns to pull matches together and push the rest apart. That is why "How many days off do I get?" can land close to a paragraph about "annual leave entitlement" with no word in common.

Embedding models are different from the chat models that generate answers. They are smaller, faster and much cheaper per token, and they output a vector rather than text. Some are asymmetric, expecting a different prefix or mode for queries than for documents, because a short question and a long passage look different. Using the wrong mode quietly degrades retrieval.

Choosing one involves a handful of trade-offs. Domain and language matter most: a model trained on general web text may be weak on legal clauses, source code or Hindi. Dimension affects storage and speed; some models are trained so the first part of the vector works on its own (often called Matryoshka embeddings), letting you truncate to save space at a small accuracy cost. Maximum input length limits chunk size. Public leaderboards are a starting point, but they measure other people's data, so test the top few candidates on your own labelled questions.

One operational rule matters above the others: queries and documents must be embedded by the same model and version. Vectors from different models live in unrelated spaces, so switching models means re-embedding the entire corpus, which is why the model version belongs in your chunk metadata.

What is vector search?

Vector search finds the stored vectors closest to a query vector, using a similarity measure such as cosine similarity. In RAG it retrieves chunks whose meaning matches the question even when they share no keywords.

Once every chunk has an embedding, retrieval becomes a geometry problem: embed the query, then find the k chunk vectors nearest to it. Nearness is measured by cosine similarity (the angle between vectors), dot product, or Euclidean distance. For vectors normalised to unit length, cosine similarity and dot product give the same ranking, which is why many systems normalise at write time. Use the measure the embedding model was trained with.

Exact search compares the query with every vector. That is fine for tens of thousands of chunks but too slow for tens of millions, so production systems use approximate nearest neighbour (ANN) indexes. The most common is HNSW, a layered graph where search hops from node to closer node, giving sub-linear search with high recall. Others partition vectors into clusters and only search the nearest clusters (IVF), or compress vectors (product quantisation) to fit more in memory. All of them trade a little recall for a lot of speed, controlled by parameters such as how many candidates the search explores.

Vector search is strong at paraphrase and intent: "can I bring my dog to work" finds the pet policy. It is weak at exact identifiers, rare names, codes and numbers, because embeddings blur surface form. It also always returns k results, even when nothing relevant exists, so a similarity threshold or a reranker is needed to recognise "no good match".

You do not always need a dedicated vector database. Many relational and search engines now support vector columns and ANN indexes, and keeping vectors next to your metadata and permissions simplifies filtering. A specialised store earns its place at large scale or when you need features such as fast filtered search over very many tenants.

What is keyword search?

Keyword search, also called lexical or sparse search, ranks documents by the query words they contain, weighting rare words more heavily. It excels at exact terms such as codes, names and jargon that vector search often blurs.

Keyword search builds an inverted index: for every word, a list of the documents that contain it. At query time it looks up each query word and scores documents. The standard scoring function is BM25, which rewards a document for containing query terms, gives more weight to rare terms (a term in 3 of 100,000 documents is far more informative than "policy"), and applies diminishing returns so repeating a word twenty times does not dominate. It also normalises for document length so long documents do not win by size alone.

Its strengths are the mirror image of vector search. It matches exact identifiers (invoice INV-20931, error E4012, a product SKU), proper names, acronyms and new jargon that an embedding model never learned. It is transparent: you can see which words matched. It is fast and cheap, needs no model, and has decades of tooling behind it, including stemming, synonyms, phrase queries and field boosting.

Its weakness is vocabulary mismatch. A question about "time off" will not find a document that only says "annual leave" unless you add synonyms or expand the query. It also has no notion of meaning, so "bank" the river and "bank" the lender look the same, and analysis choices such as tokenisation and stemming matter, especially for languages without spaces between words or with heavy inflection.

There are also learned sparse models that use a neural network to decide which terms, including related terms not in the text, to put into a sparse representation. They keep inverted-index efficiency while closing some of the vocabulary gap, and they are a reasonable middle ground where available.

What is hybrid search?

Hybrid search runs keyword and vector search on the same query and merges their results into one ranking. It catches exact terms and paraphrases alike, and is the usual default for production RAG.

Keyword and vector search fail on different queries: keyword misses paraphrases, vector misses identifiers and rare terms. Hybrid search runs both, usually in parallel, and fuses the two candidate lists. Because their errors are only partly correlated, the union recovers relevant chunks that either method alone would miss, and published retrieval benchmarks and practitioner reports consistently show hybrid matching or beating either method on its own.

The fusion step is the subtle part, because the two scores are on unrelated scales: a BM25 score of 14.2 and a cosine similarity of 0.81 cannot be added. Reciprocal rank fusion (RRF) sidesteps this by using only ranks: each document gets the sum of 1 / (k + rank) across lists, with k commonly set to 60. Documents that rank well in both lists rise to the top. The alternative is to normalise each score list to 0 to 1 and take a weighted sum, which lets you tune the balance but is sensitive to score distributions that shift between queries.

Hybrid search is a candidate generation step. A typical configuration retrieves the top 20 to 50 from each method, fuses them, and passes the top 20 to 50 fused results to a reranker that picks the final handful for the prompt. Metadata filters, such as permissions, date or product, should apply to both searches before fusion, so neither list fills up with chunks that will be discarded.

Many search engines and databases now offer hybrid queries natively. Whether you use a built-in or write the fusion yourself, measure it: on some corpora one method dominates and a weighted fusion favouring it beats equal weighting.

What is reranking?

Reranking is a second, more expensive relevance pass over the candidates that first-stage search returned. A model reads the query and each candidate together, scores how well the candidate answers the query, and reorders the list before the top few go to the generator.

First-stage retrieval has to search millions of chunks in milliseconds, so it uses representations computed in advance: an embedding per chunk, or an inverted index. The query and the document never meet inside a model; they are compared only through their precomputed summaries. That is fast but coarse. A reranker takes the 20 to 100 candidates that survived the first stage and scores each one with the query and the chunk visible together.

The most common reranker is a cross-encoder: a transformer that takes the query and a passage as one input and outputs a relevance score. Because attention runs across both texts, it can notice that a passage mentions annual leave but only for contractors, or that it answers a different question with similar words. That precision costs one model call per candidate, which is why it is applied only to a short list. Language models can also act as rerankers, scoring or ordering candidates from a prompt, which is more flexible and more expensive.

Reranking matters because of what the generator does with its context. Models attend unevenly across a long prompt and are easily distracted by plausible but irrelevant passages, so sending the best 5 chunks usually beats sending the top 20 unreranked ones. A good reranker raises precision in the top few positions, which is exactly where it counts.

The trade-off is latency, typically tens to a few hundred milliseconds depending on model size, candidate count and hardware. Reranker scores can also serve as a cut-off: if even the best candidate scores low, the system can say it found nothing rather than answer from weak evidence.

What are citations in RAG?

Citations link each claim in a generated answer to the specific retrieved passage that supports it. They let users verify answers, let teams audit and evaluate groundedness, and make unsupported claims visible.

A citation in RAG is a pointer from part of the answer back to a source chunk, rendered as a footnote, a link to the document and section, or a highlighted quote. The model produces them because the prompt gives each retrieved passage an identifier and instructs it to tag claims with the identifiers that support them. Some provider APIs also offer built-in citation features that return exact source spans; the mechanism and format vary by provider.

Citations do three jobs. For users, they turn "trust me" into "check here", which matters in legal, medical, financial and policy settings where people must be able to verify. For evaluation, they let you test claim by claim whether the cited passage actually supports the claim, which is a direct groundedness check. For operations, logged citations show which documents drive answers, which ones are never used, and which outdated ones keep getting cited.

A model can cite incorrectly. It may attach [2] to a claim that came from [3], cite a passage that only loosely relates, or invent an identifier. So the application should validate citations: check every cited id exists among the passages actually supplied, and for high-stakes uses run a support check (a judge model or an entailment model) on each claim and its cited passage. Requiring short verbatim quotes alongside ids makes verification mechanical, because you can check the quote appears in the source.

Good citation design points at the narrowest useful unit (section and page, not just a 200-page PDF), carries the document date or version so readers can spot stale sources, and degrades gracefully: if the answer cannot be supported, the system should say so rather than produce an uncited paragraph.

How do you choose chunk size?

Choose it empirically: follow document structure, then compare a few sizes on labelled questions using retrieval recall and answer quality. Small chunks match precisely but lose context; large chunks keep context but dilute the embedding and waste prompt space.

Chunk size is a tension between two needs. Matching wants small chunks: an embedding of one focused paragraph sits close to questions about that paragraph, while an embedding of a whole section averages several topics and sits moderately close to many questions without being the best match for any. Answering wants larger chunks: the model needs the rule together with its exception, the definition the clause refers to, and the table with its headers. Too small and the retrieved text is precise but incomplete; too large and the right sentence is buried in noise and you can fit fewer distinct sources in the prompt.

The right size depends on three things. The documents: dense contracts with cross-referencing clauses behave differently from FAQs where every answer is a paragraph. The questions: factoid lookups ("what is the notice period") favour small chunks; synthesis questions ("how do our parental leave and sabbatical policies interact") favour larger ones or more of them. The embedding model: each has a maximum input length and is usually trained on passages of a certain length, and performance can degrade well before the hard limit. Typical working ranges are a few hundred tokens, with 200 to 800 common, but published comparisons disagree on a single best value because it varies by corpus.

You can partly escape the trade-off by decoupling what you match from what you send. Small-to-big (parent document) retrieval indexes small chunks but returns their enclosing section to the model. Sentence-window retrieval returns the matched sentence plus its neighbours. Contextual enrichment prepends the document title, section path or a short generated summary to each small chunk before embedding, so it carries context without growing. These often beat any single fixed size.

The method is an experiment, not a guess. Build a set of 100 to 300 real questions labelled with the passages that answer them, run each candidate configuration through ingestion and retrieval, and compare recall@K at the K you will actually send. Then check end-to-end answer quality on the top two configurations, because a configuration can retrieve the right chunk while still cutting it badly. Watch cost too: smaller chunks mean more vectors to store and more items to rerank.

What are recall@K and precision@K?

Recall@K is the share of all relevant passages that appear in the top K results; precision@K is the share of the top K results that are relevant. Recall tells you whether retrieval found the evidence, precision how much noise came with it.

Both metrics need relevance labels: for each test question, the set of chunks (or documents) that contain the answer. Given the top K retrieved items, recall@K = relevant items in the top K divided by all relevant items; precision@K = relevant items in the top K divided by K. If a question has 6 relevant chunks and the top 5 results contain 4 of them, recall@5 is 4/6, about 67 percent, and precision@5 is 4/5, 80 percent. Many RAG questions have exactly one relevant chunk, in which case recall@K becomes a hit rate: was the answer anywhere in the top K, yes or no.

For RAG, recall@K at the K you actually send is usually the metric that matters most, because a passage that is not retrieved cannot be used and no prompt can fix that. Precision matters second: irrelevant passages cost tokens, and plausible distractors can mislead the model. Rank-aware metrics add order. Mean reciprocal rank (MRR) averages 1 divided by the position of the first relevant item, rewarding systems that put the answer first. nDCG handles graded relevance and discounts relevant items lower in the list.

Measure at each stage separately. Recall@50 of first-stage hybrid search tells you whether the candidates contain the answer at all; recall@5 after reranking tells you whether it made the cut. If recall@50 is high but recall@5 is low, the reranker or K is the problem. If recall@50 is low, look at parsing, chunking, the embedding model or query rewriting. This decomposition is the fastest way to know which part of the pipeline to work on.

Labels are the expensive part. Start with 100 to 200 real user questions and have domain experts mark the supporting passages; a model can draft labels for humans to confirm. Label at document or section level if chunk boundaries will change, so the set survives re-chunking. Report the sample size with every number, because a 3-point change on 100 questions is within noise.

What is Graph RAG?

Graph RAG retrieves through a knowledge graph of entities and their relationships, alongside or instead of text chunks. It helps with multi-hop questions and corpus-wide questions that no single passage answers, at a significant cost in extraction and maintenance.

Standard RAG retrieves passages that look similar to the question. That fails in two situations. Multi-hop questions need facts from several places joined by relationships: "which customers are affected if supplier X stops shipping?" requires supplier to component, component to product, and product to customer, and those facts sit in different documents that do not resemble the question. Global questions ask about the corpus as a whole, such as "what are the main themes in these 2,000 incident reports?", and no top-K set of chunks represents the whole.

Graph RAG adds a graph whose nodes are entities (people, products, suppliers, clauses) and whose edges are typed relationships (supplies, contains, owns, supersedes). The graph can come from structured data you already have, which is the most reliable source, or be extracted from text by a language model that reads each chunk and emits entities and relations. At query time the system identifies entities in the question, traverses their neighbourhood, and gives the model the connected facts plus the text chunks they came from.

One widely discussed variant, published in 2024, clusters the extracted graph into communities of densely connected entities and pre-generates a summary for each. Global questions are answered by combining those community summaries, which addresses the "summarise everything" case that plain retrieval cannot.

The costs are real. LLM extraction over a large corpus means many model calls at indexing time, and extracted graphs contain duplicate entities ("IBM", "International Business Machines"), missed relations and hallucinated ones, so entity resolution and quality checks are needed. Keeping the graph current when documents change is harder than re-embedding a chunk. And for simple factual lookups, a graph adds latency without improving answers. Many teams get most of the benefit more cheaply by putting structured relationships in a database the model can query through a tool, or by letting an agent do iterative retrieval.

How do document permissions work in RAG?

Store each chunk's access-control list as metadata at ingestion, then filter retrieval by the requesting user's identity and groups inside the search query itself, so the model never receives text the user could not open in the source system.

A RAG index copies content out of systems that have their own permissions: a file share, a wiki, a CRM. If the index ignores those permissions, the assistant becomes a way around them, and a question such as "what is the CEO's salary" retrieves the compensation spreadsheet for anyone. The rule is simple: a user may only retrieve what they could open in the source. The model cannot be the enforcement point, because once text is in the prompt, no instruction reliably stops it from being repeated or paraphrased, and prompt injection can undo such instructions anyway.

Enforcement happens in two places. At ingestion, each chunk inherits the access control list (ACL) of its source document: the users and groups allowed to read it, plus tenant id in multi-tenant systems. At query time, the application resolves the authenticated user's identity and group memberships and passes them as a filter in the search request, so both keyword and vector search only consider permitted chunks. Filtering inside the query (pre-filtering) is preferred over retrieving the top 50 and dropping forbidden ones afterwards (post-filtering), which can leave too few results and makes it easy to forget a code path.

Permissions change, and the index must keep up. When someone leaves a group or a document is restricted, the old ACL on the chunk is now wrong. Options include syncing ACL changes from the source on a short schedule, storing group ids rather than expanded user lists so membership changes resolve at query time, and, for the most sensitive sources, re-checking each retrieved document against the source system's permission API before use. Decide an acceptable staleness window explicitly, such as minutes for HR data.

Other leaks are worth closing too. Caches, including semantic caches, must be keyed by permission scope or one user's answer will be served to another. Logs and traces contain retrieved text and need the same protection. Search suggestions and "no results" messages can reveal that a document exists. For strict multi-tenancy, a separate index or namespace per tenant gives stronger isolation than a filter field.

RAG or fine-tuning?

Use RAG to give a model knowledge that changes, is private, or must be cited. Use fine-tuning to change behaviour: format, style, tone, domain-specific task skill or efficiency. They solve different problems and are often combined.

The two techniques act on different things. RAG changes what the model sees at inference time: it inserts current, specific evidence into the prompt. Fine-tuning changes what the model is: it adjusts weights using examples so the model behaves differently on every request, even with no extra context. A useful question is whether your problem is "the model does not have the information" or "the model has what it needs but does the task badly".

For knowledge, RAG is almost always the right first answer. Facts placed in context are used precisely and can be updated by re-indexing in minutes. Facts trained into weights are recalled unreliably, cannot be cited, cannot be removed cleanly when they become wrong, and cannot be filtered per user. Research on fine-tuning for knowledge injection has generally found it less effective than retrieval for new facts, and it can increase confident hallucination about nearby facts the model did not learn well.

For behaviour, fine-tuning shines. If you need a strict output schema followed every time, a house writing style, a classification taxonomy with subtle boundaries, or a smaller and cheaper model to match a larger one on a narrow task, examples in training teach that more reliably than long instructions in every prompt. It can also shorten prompts, which reduces cost and latency at high volume.

Many production systems combine them: retrieve the evidence, and use a model fine-tuned to read retrieved passages well, cite in a fixed format and say "not found" when appropriate. Before either, though, try prompting with good examples. Fine-tuning brings a data pipeline, an evaluation set, retraining when the base model changes, and the risk of degrading general abilities, so it should be earned by measured shortfalls that prompting cannot fix.

How do you keep an index current?

Run incremental sync: detect added, changed and deleted source documents, re-process only those, replace all of a changed document's chunks atomically, remove deleted ones, and monitor index lag. Plan separately for full rebuilds when the pipeline itself changes.

A RAG system is only as current as its index. If a policy changed yesterday and the index still holds last week's version, the assistant will cite the old rule with full confidence. Worse, if the old and new versions both remain, retrieval may return either. Keeping the index current is a data-engineering problem with three parts: detecting change, applying it correctly, and knowing how far behind you are.

Detection depends on the source. Many systems offer change feeds, webhooks or a modified-since query; use those where they exist. Otherwise, crawl on a schedule and compare a content hash of each document with the stored one, which also avoids reprocessing files whose timestamp changed but content did not. Deletions are the easy thing to miss: a crawler that only lists current files never sees what disappeared, so compare the full list of source ids against the index and remove the orphans. Permission changes are changes too.

Apply updates per document, not per chunk. When a document changes, its chunk boundaries usually shift, so updating chunks in place leaves fragments of the old version behind. Instead, process the new version, write all its new chunks, then delete every chunk with that document id and the old version number, ideally in one transaction or with a version flag that switches atomically. Superseded documents (a 2025 policy replaced by 2026) may need explicit handling, either removal or a metadata flag that ranks them down.

Some changes require a full rebuild: a new embedding model, a new chunking strategy, or a parser fix. Build the new index alongside the old one, evaluate it on your labelled set, then switch traffic, a blue-green approach that avoids serving a half-migrated index. Finally, monitor freshness: track the lag between a source change and its index update, the count of documents per source compared with the source itself, and failed ingestion jobs, and alert when they drift.

How do you rewrite or expand a query before retrieval?

Use a model to turn the user's raw message into better search queries: resolve follow-up references, add synonyms and identifiers, split compound questions, or generate a hypothetical answer to search with. It fixes the vocabulary gap between how users ask and how documents are written.

The literal user message is often a poor search query. In a chat, "what about for contractors?" means nothing without the previous turn. Users write "my laptop won't charge" while the manual says "power adapter fault". A single message may contain two questions that live in different documents. Query transformation puts a fast model call between the user and the retriever to produce queries that match how the corpus is written.

The main techniques address different gaps. Conversational condensation rewrites a follow-up into a standalone question using the chat history, and is close to mandatory for chat interfaces. Expansion adds synonyms, acronyms and likely terms, which particularly helps keyword search. Multi-query generates several phrasings or sub-questions, retrieves for each, and fuses the results, raising recall for broad or compound questions. Hypothetical document embeddings (HyDE, from a 2022 paper) asks the model to write a plausible answer and embeds that instead of the question, on the logic that an answer looks more like the target passage than a question does. Metadata extraction pulls filters such as product, region or date range out of the question so they can be applied exactly.

Each technique costs a model call before retrieval starts, adding latency to every request, and each can hurt. Expansion can drift from the user's intent; HyDE can confidently invent the wrong domain and retrieve the wrong documents; multi-query multiplies search load. Keep the original query in the mix, so the rewrite can only add candidates, and use a small fast model with a tight prompt.

Measure rewriting like any retrieval change: run the labelled set with and without it and compare recall@K, sliced by query type. It usually helps follow-ups and short, vague queries a lot, and does little for long, specific ones, so routing only some queries through rewriting is a reasonable optimisation.

Why does a RAG system give a wrong answer, and how do you debug it?

Trace the failing question through each stage and find the first one that went wrong: the answer was not in the corpus, was mangled at ingestion, was not retrieved, was ranked out, was dropped by context assembly, or was misread by the model. Fix that stage, not the prompt.

A wrong RAG answer has a cause somewhere in a chain, and the visible symptom (a wrong sentence from the model) usually appears far from it. The useful habit is to stop treating the system as one black box and walk the question through each stage, asking a yes-or-no question at each: does the corpus contain the answer; does the parsed chunk contain it intact; is that chunk in the first-stage candidates; did it survive reranking and the cut to K; did it make it into the final prompt; did the model use it correctly. The first no is the defect.

The early stages fail more often than teams expect. Content is missing (the answer lives in a system that was never indexed, or in a document version not yet synced). Content is damaged (a table flattened by the parser, a scanned page with OCR errors, a chunk boundary that split a rule from its exception). Content is not retrieved (vocabulary mismatch, an identifier embeddings blur, a metadata filter that excluded it). Content is outranked (a newer duplicate, a general overview page that looks more similar than the specific clause). Each has a different fix: connectors, parsing, chunking, hybrid search, query rewriting, reranking or deduplication.

Generation failures are real but should be confirmed, not assumed. The model may ignore the right passage when it is buried among distractors, merge two passages that disagree, answer from training knowledge despite instructions, or fail a reasoning step such as comparing dates. Test this by placing only the correct passage in the prompt: if the answer becomes correct, the problem is upstream context quality; if it stays wrong, the issue is the prompt, the model or the task.

Make this cheap to do by logging, for every request, the rewritten queries, candidate ids and scores at each stage, the final prompt and the citations. Then error analysis on 50 failing questions can be tallied by stage, which tells you where to spend effort. Teams often find that most failures sit in ingestion and retrieval, which is why prompt tweaking alone rarely moves the numbers.

What is agentic RAG, and when should the model drive retrieval?

Agentic RAG gives the model a search tool and lets it decide whether, what and how many times to retrieve, reading results and searching again until it has enough evidence. It helps multi-step research questions at the cost of latency, cost and predictability.

Classic RAG is a fixed pipeline: retrieve once with the user's question, then generate. That works for lookups but breaks on questions whose next search depends on what the first one found. "Which of our enterprise customers are on contracts that renew before the new pricing takes effect?" needs the date the pricing takes effect, then a search for contracts, then a filter. In agentic RAG, retrieval becomes a tool the model can call: it plans a query, reads the results, decides whether it has enough, and issues follow-up searches, possibly across several tools such as a document search, a SQL query and a ticket lookup.

This adds real capabilities. The model can decompose a compound question, reformulate after a miss, choose the right source per sub-question, skip retrieval entirely when the question needs none, and stop when evidence is sufficient. Models trained for tool use and reasoning do this reasonably well, and it is the pattern behind most "deep research" style features.

The costs are the usual agent costs. Each iteration adds a model call and a search, so latency grows from one or two seconds to tens of seconds, and token cost multiplies. Behaviour becomes less predictable: the model may search too little and answer from memory, or loop on near-identical queries. Every retrieved result is also untrusted input entering the loop, so indirect prompt injection in a document can steer later tool calls. And evaluation must look at the trajectory (which searches, in what order) as well as the final answer.

Use it where questions genuinely need multiple dependent lookups or several sources, and keep a fixed pipeline for the high-volume simple questions. Bound the loop with a maximum number of searches, require citations for the final answer, apply permissions inside the search tool rather than trusting the model to, and log every step. A common hybrid is a router: a cheap classifier sends simple questions through single-shot RAG and complex ones to the agent.