AI ArchitectureBuild

Designing a Production-Ready RAG Pipeline

Ingestion, chunking, embedding, retrieval, reranking, generation and evaluation: a reference architecture for RAG that survives real users.

By The AI Build Room2 min read89 views

Two pipelines, not one

A RAG system is really two pipelines with different requirements:

  • Ingestion (offline): load documents, clean, chunk, embed and index. Throughput matters; latency does not.
  • Query (online): embed the question, retrieve, rerank, build the prompt and generate. Latency matters a lot.

Design, deploy and monitor them separately.

Ingestion

  1. Load and normalize: extract text, preserve headings and tables, drop navigation and boilerplate.
  2. Chunk by structure: split on headings and paragraphs; attach the document title and section path to each chunk.
  3. Embed with a versioned model: store the embedding model name and version with every vector.
  4. Index with metadata: source, permissions, updated time and document id for filtering and deletion.
  5. Make it idempotent: re-ingesting a document replaces its chunks instead of duplicating them.

Query

  1. Rewrite the question when needed: resolve follow-ups like "what about the second one?" into standalone queries.
  2. Retrieve broadly: hybrid search for the top 20 to 50 candidates, filtered by user permissions.
  3. Rerank narrowly: a cross-encoder reranker selects the best 5 to 8 chunks.
  4. Generate with citations: instruct the model to answer only from the provided context and cite chunk ids.
  5. Handle "I don't know": if nothing relevant is retrieved, say so instead of letting the model guess.

Observability

Trace every request end to end: rewritten query, retrieved ids and scores, reranked ids, final prompt size, model, latency per stage and cost. Without this trace, a wrong answer is a mystery; with it, it is a five-minute investigation.

Evaluation as a gate

Keep a versioned evaluation set and run it on every change to chunking, embedding model, retrieval parameters or prompts. Track retrieval recall and answer faithfulness separately. Ship changes that improve the numbers, not changes that improve one demo.