Naive RAG — embed documents, store in a vector database, retrieve top-K chunks, stuff into prompt — works for demos. In production, it fails in predictable ways: irrelevant chunks ranked high, relevant context split across chunk boundaries, duplicate information drowning out unique facts, and the LLM hallucinating when retrieval misses.
Advanced RAG techniques address each failure mode systematically.
Split every N tokens with overlap. Simple, fast, terrible for structured documents. A paragraph split mid-sentence loses semantic coherence. Use only for unstructured text with no clear section boundaries.
Split at natural boundaries: paragraphs, section headers, topic shifts. Embed sentences sequentially and split where cosine similarity between adjacent sentences drops below a threshold. Produces chunks that are semantically self-contained.
Start with large chunks (entire sections), if too large, split into subsections, paragraphs, then sentences. LangChain's RecursiveCharacterTextSplitter does this. Better than fixed-size but still somewhat mechanical.
Parse the document structure (markdown headers, HTML tags, PDF sections) and chunk at structural boundaries. A legal contract should chunk at clause boundaries. A research paper should chunk at section/subsection boundaries. This produces the highest quality chunks but requires per-format parsers.
Vector search finds semantically similar content but misses exact matches. Keyword search (BM25) catches exact terms but misses paraphrases. Combine both:
score = sum(1 / (k + rank)) for each result across both listsHybrid search consistently outperforms either method alone. I've measured 15-25% improvement in retrieval accuracy on every project where I've tested it.
Initial retrieval is fast but imprecise. A re-ranker is a more powerful (slower) model that scores each retrieved chunk against the query for relevance:
The pattern: retrieve 50-100 chunks with fast vector search, then re-rank to find the top 5-10 truly relevant ones. This two-stage approach gives you both speed and accuracy.
User queries are often vague, misspelled, or use different terminology than your documents. Transform the query before retrieval:
You must measure retrieval quality separately from generation quality:
Tools: RAGAS, TruLens, DeepEval provide automated evaluation pipelines. Build a test set of 50-100 question-answer pairs and measure after every change.