Deep-Dives / Retrieval & RAG
Retrieval & RAG
Retrieval-augmented generation past the basics — advanced architectures, graph RAG, retrieval as an agent tool, RAG security.
- Advanced RAG ArchitecturesThe naive→modular→agentic RAG spectrum and the levers that matter — CRAG, Self-RAG, query transformation, fusion, reranking — all attacking the same garbage-in/confident-wrong-out failure.
- GraphRAG & Multi-Hop RetrievalWhy flat top-k RAG cannot answer thematic or relational multi-hop queries, how Microsoft GraphRAG and iterative retrieve-reason loops solve it, and the cost/staleness heuristic for when not to.
- Hybrid Search & RerankingWhy one retriever is structurally not enough, reciprocal rank fusion across BM25 + dense, the two-stage retrieve-then-cross-encoder pattern, and ColBERT-style late interaction when cross-encoders are too slow.
- Document Parsing & Ingestion QualityIngestion is the half of RAG that doesn't get dashboards and decides whether the answer was ever indexable — layout-aware parsing, tables, OCR, structural chunking, and vision-RAG as a parsing escape hatch.
- Query Understanding & TransformationFixing the question before you search — rewriting, decomposition, multi-query, HyDE caveats, step-back prompting, and routing — the cheapest place to add intelligence to a RAG pipeline.
- Agentic Retrieval: Search as a ToolRetrieval as a tool the model calls iteratively in a ReAct-style loop, with budget, stopping criteria, and the new failure modes (looping, drift, premature stop) that come with handing the model the steering wheel.
- Evaluating RAGScore retrieval, grounding, and answer quality as three separate things — recall@k, faithfulness, answer relevance — plus how to build a small living eval set and use LLM judges without lying to yourself.
- Choosing a Vector DatabaseA constraint-first selection guide — what a vector DB actually is, the axes that genuinely differ between products (ANN algorithm, filter quality, freshness, hybrid, multi-tenancy, ops, cost), just enough HNSW/IVF/DiskANN internals to read a vendor pitch, and why most teams end at pgvector.
- Local-First RetrievalBuilding a knowledge base on hardware you own — starting with whether you need an index at all, since the leading coding agents removed theirs in favour of grep. Ingest quality, embedding-model sizing, the four store architectures, hybrid retrieval, serving it to an agent over MCP, and when to run a finished platform instead.
- Index Freshness & InvalidationA stale index returns a confident answer with a citation, and the three ways a corpus goes stale have different blast radii — an out-of-date paragraph is embarrassing, a deleted document that still answers is an incident. Why the nightly re-crawl is a bill rather than a guarantee, how to drive invalidation from a change stream that is reliably bad at deletions, tombstones before compaction, the derived artefacts that inherit no deletion at all, and change-to-queryable as a per-source p99 with a name on it.
- Contextual RetrievalThe largest source of recall failure in a competent RAG stack is the moment you cut the document up: the chunk that says revenue grew 3% names no company, no quarter and no year. Prepending a generated preamble before embedding cuts top-20 retrieval failure by 35%, 49% when the same preamble feeds BM25, and 67% with reranking on top (5.7% to 1.9%), at roughly $1.02 per million document tokens under prompt caching — but your index is now a function of a second model and a prompt you will want to change, so budget three rebuilds and try heading-path prefixes, a reranker and hybrid search first.
- Late-Interaction RetrievalThe standard objection — one vector per token is a hundred times the storage — has been obsolete for four years and still decides architectures: ColBERTv2's residual compression takes MS MARCO from 154 GiB to 25 GiB at two bits per dimension, the same order as a plain float32 single-vector index over the same 8.8 million passages, and token pooling removes half the vectors again at virtually no measured cost. So the real decision is diagnostic — a better first stage only fixes recall failures, and a 2026 separation result says the queries where single vectors structurally fail are the multi-constraint conjunctive ones an agent actually issues. For page images the comparison is not against your retriever at all; it is against the parsing pipeline ColPali-style retrieval deletes.