rag-architecture

v2026.09.24

RAG system architecture and design decisions. Covers naive vs advanced vs agentic RAG, decision trees for RAG vs fine-tuning vs long context, production topology, latency budgets, and component sequencing. USE WHEN: user mentions "RAG architecture", "RAG design", "naive RAG", "advanced RAG", "agentic RAG", "RAG vs fine-tuning", "RAG vs long context", "production RAG" DO NOT USE FOR: chunking details - use `chunking-strategies`; query rewriting - use `query-transformations`; retrieval algorithms - use `advanced-retrieval`; evaluation - use `rag-evaluation`; agent loops - use `agentic-rag`

GitHub
安装命令
npx skhub add claude-dev-suite/rag-architecture
Markdown
SKILL.md

RAG Architecture

Three Architectural Tiers

TierComponentsBest ForComplexity
Naive RAGChunk + embed + top-K + stuffPrototype, < 10k docs, homogeneous contentLow
Advanced RAG+ query rewriting, hybrid search, reranking, metadata filtersProduction, heterogeneous content, > 10k docsMedium
Agentic RAG+ self-reflection, retrieval-as-tool, multi-hop, corrective fallbackComplex research, multi-source synthesis, high-stakes answersHigh

Naive RAG Pipeline

[Docs] -> [Splitter] -> [Embedder] -> [Vector DB]
                                          |
[Query] -> [Embedder] -> [Top-K Search] --+-> [Prompt Stuffer] -> [LLM] -> [Answer]

Single failure mode: bad retrieval = bad answer. No recovery path. Works for well-scoped FAQ bots on small corpora.

Advanced RAG Pipeline

                          pre-retrieval         retrieval              post-retrieval
[Query] -> [Router] -> [Rewrite/HyDE/Multi-Q] -> [Hybrid Search] -> [Rerank] -> [Compress] -> [LLM]
                                                      |
                           [BM25] + [Dense Vector] + [Metadata Filter]

Each stage is independently replaceable and measurable. See query-transformations, hybrid-search, reranking.

Agentic RAG Pipeline

[Query] -> [Planner Agent]
             |
             v
       +-----+-----+-------------------+
       |           |                   |
   [Retrieve]  [Web Search]      [Code Tool]
       |           |                   |
       +-----+-----+-------------------+
             |
        [Critic / Self-reflection]
             |
       +-----+-----+
       |           |
    [Answer]  [Replan / More retrieval]

Dynamic step count, dynamic tool selection, self-correction. See agentic-rag.

Decision Tree: RAG vs Fine-Tuning vs Long Context

Is the knowledge dynamic (changes > monthly)?
  Yes -> RAG
  No  -> Continue
    |
    Is it style/format/persona (not facts)?
      Yes -> Fine-tuning (SFT or DPO)
      No  -> Continue
        |
        Does total corpus fit in 200k-1M tokens?
          Yes -> Long context with prompt caching (cheaper than RAG at small scale)
          No  -> RAG
            |
            Need factual grounding with citations?
              Yes -> RAG (mandatory for auditability)
              No  -> Hybrid: long context for recent + RAG for archive

Rules of thumb:

  • Under 500 KB of content: long context with prompt caching beats RAG on latency and quality.
  • Over 10 MB or changes weekly: RAG wins on cost and freshness.
  • Between: measure both.

Latency Budget (Production Target: < 3s P95)

StageBudgetOptimization
Query embedding50-150msBatch + local model for simple queries
Query rewriting (optional)300-800msSkip for short factual queries
Vector search (top-50)20-100msHNSW with ef_search tuned
BM25 search10-50msParallel with vector search
Fusion (RRF)< 5msIn-memory
Reranking (top-50 -> top-5)100-400msCohere/Voyage API or local BGE
LLM generation1000-2000msStreaming, prompt caching
Total1500-3500ms

Parallelize embedding + BM25. Skip rewriting for short queries. Cache query embeddings for hot terms.

Production Component Diagram

                    +------------------+
                    |   Ingestion API  |
                    +--------+---------+
                             |
                 +-----------v------------+
                 | Chunker + Embedder     |
                 | (batch worker, queue)  |
                 +-----------+------------+
                             |
         +-------------------+---------------------+
         |                                         |
  +------v------+                          +-------v-------+
  |  Vector DB  |                          | Document Store|
  |  (HNSW)     |                          | (Postgres/S3) |
  +------+------+                          +-------+-------+
         |                                         |
         |    +------------------+                 |
         +----> Retrieval Service <----------------+
              | (hybrid + rerank)|
              +--------+---------+
                       |
              +--------v---------+
              |  LLM Gateway     |
              | (Claude/GPT)     |
              +--------+---------+
                       |
              +--------v---------+
              |   API / UI       |
              +------------------+

Separate the ingestion path from the query path. Never block queries on indexing.

Minimal Python Scaffold (Advanced RAG)

from dataclasses import dataclass
from langchain_anthropic import ChatAnthropic
from langchain_openai import OpenAIEmbeddings
from langchain_qdrant import QdrantVectorStore
from langchain_community.retrievers import BM25Retriever
from langchain.retrievers import EnsembleRetriever, ContextualCompressionRetriever
from langchain_cohere import CohereRerank

@dataclass
class RAGConfig:
    top_k_retrieve: int = 50
    top_n_rerank: int = 5
    alpha: float = 0.5  # dense weight in hybrid
    rerank_model: str = "rerank-english-v3.0"
    llm_model: str = "claude-sonnet-4-5-20250929"

def build_pipeline(docs, cfg: RAGConfig):
    emb = OpenAIEmbeddings(model="text-embedding-3-small")
    vstore = QdrantVectorStore.from_documents(docs, emb, collection_name="kb")
    dense = vstore.as_retriever(search_kwargs={"k": cfg.top_k_retrieve})
    sparse = BM25Retriever.from_documents(docs); sparse.k = cfg.top_k_retrieve
    hybrid = EnsembleRetriever(retrievers=[sparse, dense], weights=[1 - cfg.alpha, cfg.alpha])
    reranker = CohereRerank(model=cfg.rerank_model, top_n=cfg.top_n_rerank)
    retriever = ContextualCompressionRetriever(base_compressor=reranker, base_retriever=hybrid)
    llm = ChatAnthropic(model=cfg.llm_model, max_tokens=1024)
    return retriever, llm

Scaling Patterns

RegimeIndexStrategy
< 100k chunksHNSW in-memory (FAISS, Chroma)Single node
100k-10M chunksManaged HNSW (Qdrant, Pinecone, Weaviate)Replicate reads
10M-1B chunksSharded HNSW + IVF coarse filterPartition by tenant/namespace
> 1B chunksDisk-ANN / SPTAG / hierarchicalCustom; hire infra

Namespace per tenant avoids noisy-neighbor retrieval in multi-tenant SaaS.

Anti-Patterns

Anti-PatternFix
Starting with agentic RAGStart naive, measure, add complexity only when recall < 70%
Treating RAG as a solved boxEvery stage needs its own eval; see rag-evaluation
Single index for heterogeneous contentSeparate indexes per content type with a router
Synchronous indexing in query pathQueue ingestion; queries never block on embedding
No reranking above 10k docsReranking recovers 10-30% recall@5 vs raw vector search
Hardcoded top_k everywhereParametrize; tune with eval set
Storing chunks only in vector DBKeep source of truth in document DB; vector DB holds refs

Production Checklist

  • Ingestion and query paths are separate services
  • Document store is source of truth; vector DB holds IDs + embeddings only
  • Latency budget allocated per stage and measured in traces
  • Query router distinguishes factual, analytical, and conversational intents
  • Hybrid search baseline before any advanced technique
  • Reranker in place for corpora > 10k chunks
  • Fallback path when retrieval returns nothing (admit ignorance, offer web search)
  • Namespace isolation for multi-tenant deployments
  • Graceful degradation: naive path if advanced stage fails
  • Tracing across ingestion -> retrieval -> generation (LangSmith / OpenTelemetry)
发现
标签

此技能尚未发布标签。

版本
最新版本元数据

版本

v2026.09.24

发布时间

2026年9月24日

分类

未分类

许可证

MIT

源路径

skills/rag/rag-architecture

默认分支

main

最新提交

9496306

Tree SHA

fe4e2f1