RAG · Retrieval · Production · 15 min read · October 8, 2026

Advanced RAG: Chunking Strategies, Re-Ranking, and Production Retrieval Pipelines

Why Basic RAG Falls Short

Naive RAG — embed documents, store in a vector database, retrieve top-K chunks, stuff into prompt — works for demos. In production, it fails in predictable ways: irrelevant chunks ranked high, relevant context split across chunk boundaries, duplicate information drowning out unique facts, and the LLM hallucinating when retrieval misses.

Advanced RAG techniques address each failure mode systematically.

Chunking Strategies

Fixed-size chunking

Split every N tokens with overlap. Simple, fast, terrible for structured documents. A paragraph split mid-sentence loses semantic coherence. Use only for unstructured text with no clear section boundaries.

Semantic chunking

Split at natural boundaries: paragraphs, section headers, topic shifts. Embed sentences sequentially and split where cosine similarity between adjacent sentences drops below a threshold. Produces chunks that are semantically self-contained.

Recursive chunking

Start with large chunks (entire sections), if too large, split into subsections, paragraphs, then sentences. LangChain's RecursiveCharacterTextSplitter does this. Better than fixed-size but still somewhat mechanical.

Document-aware chunking

Parse the document structure (markdown headers, HTML tags, PDF sections) and chunk at structural boundaries. A legal contract should chunk at clause boundaries. A research paper should chunk at section/subsection boundaries. This produces the highest quality chunks but requires per-format parsers.

Hybrid Search: Vector + Keyword

Vector search finds semantically similar content but misses exact matches. Keyword search (BM25) catches exact terms but misses paraphrases. Combine both:

  1. Run vector search (cosine similarity) — returns semantically relevant chunks
  2. Run BM25 keyword search — returns exact-match chunks
  3. Merge results using Reciprocal Rank Fusion (RRF): score = sum(1 / (k + rank)) for each result across both lists

Hybrid search consistently outperforms either method alone. I've measured 15-25% improvement in retrieval accuracy on every project where I've tested it.

Re-Ranking

Initial retrieval is fast but imprecise. A re-ranker is a more powerful (slower) model that scores each retrieved chunk against the query for relevance:

The pattern: retrieve 50-100 chunks with fast vector search, then re-rank to find the top 5-10 truly relevant ones. This two-stage approach gives you both speed and accuracy.

Query Expansion and Transformation

User queries are often vague, misspelled, or use different terminology than your documents. Transform the query before retrieval:

Evaluation Metrics

You must measure retrieval quality separately from generation quality:

Tools: RAGAS, TruLens, DeepEval provide automated evaluation pipelines. Build a test set of 50-100 question-answer pairs and measure after every change.