chunking-strategies

v2026.09.24

Document chunking techniques for RAG. Fixed-size, recursive, semantic, token-based, document-aware, proposition, parent-child, sliding window, and Anthropic contextual retrieval. Tradeoff tables, LangChain and LlamaIndex code. USE WHEN: user mentions "chunking", "text splitter", "split documents", "semantic chunking", "contextual retrieval", "parent-child chunks", "proposition chunking" DO NOT USE FOR: retrieval after chunking - use `advanced-retrieval`; query-side transforms - use `query-transformations`; overall design - use `rag-architecture`

GitHub
安装命令
npx skhub add claude-dev-suite/chunking-strategies
Markdown
SKILL.md

Chunking Strategies

Strategy Tradeoff Matrix

StrategyPreserves SemanticsCostBest ForChunk Unit
Fixed-sizeNoFreeHomogeneous proseChars
Recursive characterPartialFreeGeneral textChars with hierarchy
Token-basedNoFreeExact token budgetingTokens
Document-aware (Markdown/HTML)YesFreeTechnical docs, wikisHeaders/sections
Code-awareYesFreeSource codeFunctions/classes
Semantic (embedding breakpoints)Yes$$Long narrative, research papersMeaning shifts
Proposition-basedYes$$$High-precision Q&A, legalAtomic facts
Parent-childYesFreeNeed small match + big contextHierarchy
Sliding windowPartialFreeDialogues, timelinesOverlap + stride
Contextual retrieval (Anthropic)Yes$$Production RAG > 5k chunksChunk + LLM context

Fixed-Size

def fixed_chunks(text: str, size: int = 800, overlap: int = 200) -> list[str]:
    return [text[i:i + size] for i in range(0, len(text), size - overlap)]

Use only as a baseline. Breaks mid-sentence and mid-token.

Recursive Character (default for most projects)

from langchain_text_splitters import RecursiveCharacterTextSplitter

splitter = RecursiveCharacterTextSplitter(
    chunk_size=800,
    chunk_overlap=200,
    separators=["\n\n", "\n", ". ", "? ", "! ", " ", ""],
    length_function=len,
    is_separator_regex=False,
)
chunks = splitter.split_documents(docs)

Tries separators in order so paragraph boundaries are preferred over arbitrary cuts.

Token-Based (exact LLM budgeting)

from langchain_text_splitters import TokenTextSplitter

splitter = TokenTextSplitter(
    encoding_name="cl100k_base",  # GPT-4, text-embedding-3
    chunk_size=512,
    chunk_overlap=64,
)
chunks = splitter.split_text(text)

For Claude, use anthropic.Anthropic().messages.count_tokens to measure; approximate ratio is ~3.5 chars per token for English.

Document-Aware: Markdown

from langchain_text_splitters import MarkdownHeaderTextSplitter, RecursiveCharacterTextSplitter

headers = [("#", "h1"), ("##", "h2"), ("###", "h3")]
md_splitter = MarkdownHeaderTextSplitter(headers_to_split_on=headers, strip_headers=False)
header_chunks = md_splitter.split_text(markdown_text)

# Secondary splitter for oversized sections
char_splitter = RecursiveCharacterTextSplitter(chunk_size=800, chunk_overlap=150)
chunks = char_splitter.split_documents(header_chunks)  # metadata carries h1/h2/h3

Header metadata enables section-level filtering at query time.

Document-Aware: Code

from langchain_text_splitters import RecursiveCharacterTextSplitter, Language

py_splitter = RecursiveCharacterTextSplitter.from_language(
    language=Language.PYTHON, chunk_size=1500, chunk_overlap=200
)
ts_splitter = RecursiveCharacterTextSplitter.from_language(
    language=Language.TS, chunk_size=1500, chunk_overlap=200
)

Splits on class, def, function boundaries. Larger chunks because code lines are shorter than prose.

Semantic Chunking (embedding breakpoints)

Splits where adjacent sentences diverge in meaning. Expensive (embeds every sentence) but yields coherent chunks on long-form content.

from langchain_experimental.text_splitter import SemanticChunker
from langchain_openai import OpenAIEmbeddings

splitter = SemanticChunker(
    OpenAIEmbeddings(model="text-embedding-3-small"),
    breakpoint_threshold_type="percentile",   # or "standard_deviation", "interquartile"
    breakpoint_threshold_amount=95,
    buffer_size=1,
)
chunks = splitter.create_documents([long_text])

Manual variant for fine control (break where cosine distance between adjacent sentence embeddings exceeds the 95th percentile):

import numpy as np
from sentence_transformers import SentenceTransformer

def semantic_chunks(text: str, pct: float = 95) -> list[str]:
    sents = [s.strip() for s in text.split(". ") if s.strip()]
    embs = SentenceTransformer("all-MiniLM-L6-v2").encode(sents)
    sims = [np.dot(embs[i], embs[i+1]) / (np.linalg.norm(embs[i]) * np.linalg.norm(embs[i+1]))
            for i in range(len(sents) - 1)]
    dists = 1 - np.array(sims)
    breaks = [i + 1 for i, d in enumerate(dists) if d > np.percentile(dists, pct)]
    out, start = [], 0
    for b in breaks + [len(sents)]:
        out.append(". ".join(sents[start:b])); start = b
    return out

Proposition-Based Chunking

Decompose text into atomic factual propositions using an LLM. Each proposition becomes one chunk.

from anthropic import Anthropic
import json

client = Anthropic()

PROMPT = """Decompose the passage into atomic propositions. Each proposition:
- Expresses exactly one fact
- Is self-contained (no pronouns without antecedents)
- Resolves coreferences inline

Return JSON array of strings. Passage:
{passage}"""

def propositionize(passage: str) -> list[str]:
    msg = client.messages.create(
        model="claude-sonnet-4-5-20250929",
        max_tokens=2048,
        messages=[{"role": "user", "content": PROMPT.format(passage=passage)}],
    )
    return json.loads(msg.content[0].text)

Highest retrieval precision; highest ingestion cost. Use for legal, medical, compliance.

Parent-Child Chunking

Index small chunks (for precise retrieval) but return parent chunks (for context).

from langchain.retrievers import ParentDocumentRetriever
from langchain.storage import InMemoryStore
from langchain_chroma import Chroma
from langchain_text_splitters import RecursiveCharacterTextSplitter
from langchain_openai import OpenAIEmbeddings

parent_splitter = RecursiveCharacterTextSplitter(chunk_size=2000)
child_splitter = RecursiveCharacterTextSplitter(chunk_size=400)
vectorstore = Chroma(collection_name="children", embedding_function=OpenAIEmbeddings())
docstore = InMemoryStore()

retriever = ParentDocumentRetriever(
    vectorstore=vectorstore,
    docstore=docstore,
    child_splitter=child_splitter,
    parent_splitter=parent_splitter,
)
retriever.add_documents(docs)

See also advanced-retrieval for multi-vector and sentence-window patterns that generalize this idea.

Sliding Window with Stride

def sliding_window(tokens: list[str], window: int = 512, stride: int = 256) -> list[list[str]]:
    return [tokens[i:i + window] for i in range(0, len(tokens), stride) if i + window <= len(tokens)]

Use for dialogue, legal contracts, timelines where context before and after each point matters.

Anthropic Contextual Retrieval

Prepend LLM-generated context to each chunk before embedding. Reduces retrieval failure rate by ~35% on Anthropic's benchmark.

from anthropic import Anthropic

client = Anthropic()

CONTEXT_PROMPT = """<document>
{whole_document}
</document>

Here is the chunk we want to situate within the whole document:
<chunk>
{chunk}
</chunk>

Give a short (1-2 sentence) context to situate this chunk within the overall
document for the purposes of improving search retrieval of the chunk.
Answer only with the succinct context and nothing else."""

def contextualize(whole_doc: str, chunk: str) -> str:
    msg = client.messages.create(
        model="claude-haiku-4-5-20250929",
        max_tokens=200,
        messages=[{"role": "user", "content": CONTEXT_PROMPT.format(
            whole_document=whole_doc, chunk=chunk)}],
        extra_headers={"anthropic-beta": "prompt-caching-2024-07-31"},
    )
    return msg.content[0].text

def contextual_chunks(doc: str, base_chunks: list[str]) -> list[str]:
    return [f"{contextualize(doc, c)}\n\n{c}" for c in base_chunks]

Use prompt caching on whole_document to drop cost by ~10x. Combine with BM25 + dense for best results.

LlamaIndex Equivalents

from llama_index.core.node_parser import (
    SentenceSplitter, SemanticSplitterNodeParser, MarkdownNodeParser,
    HierarchicalNodeParser,
)
from llama_index.embeddings.openai import OpenAIEmbedding

sentence = SentenceSplitter(chunk_size=512, chunk_overlap=64)
semantic = SemanticSplitterNodeParser(
    buffer_size=1, breakpoint_percentile_threshold=95, embed_model=OpenAIEmbedding()
)
markdown = MarkdownNodeParser()
hierarchical = HierarchicalNodeParser.from_defaults(chunk_sizes=[2048, 512, 128])

Sizing Heuristics by Content Type

ContentChunk SizeOverlapSplitter
Blog posts / articles800 chars150Recursive character
Technical docs (markdown)1000 chars200Markdown headers
Source code1500 chars200Language-aware
Legal / contracts400 chars100Proposition or sliding
Research papersVariableN/ASemantic
Customer support tickets300 chars50Per-turn
Book / long narrative1500 chars300Semantic + parent-child

Anti-Patterns

Anti-PatternFix
One chunk size for all content typesSplit pipeline per content type
Overlap = 0Always overlap 10-25% to preserve boundary context
Splitting code by character countUse Language.* language-aware splitters
Stripping headers before chunkingKeep or re-prepend headers for retrieval signal
Chunks larger than embedding model contextMeasure: text-embedding-3-* max 8191 tokens
No metadata on chunksAttach source, section, chunk_index, parent_id
Re-embedding whole corpus on tweakIncremental pipeline keyed by content hash
Semantic chunking on short docsNot worth the cost below ~10 pages

Production Checklist

  • Content-type detection routes to the right splitter
  • Chunk size tuned with a retrieval eval set (rag-evaluation)
  • Metadata stamped on every chunk (source, section, position, hash)
  • Chunk hash stored for idempotent re-ingestion
  • Token counts measured against embedding model limit
  • Overlap configured (default 20%)
  • Parent-child or multi-vector used when context window matters
  • Contextual retrieval considered for corpora > 5k chunks
  • Ingestion pipeline monitored (chunks/sec, embedding cost/doc)
  • Re-chunking pathway documented (rolling re-index)
发现
标签

此技能尚未发布标签。

版本
最新版本元数据

版本

v2026.09.24

发布时间

Sep 24, 2026

分类

未分类

许可证

MIT

源路径

skills/rag/chunking-strategies

默认分支

main

最新提交

9496306

Tree SHA

fe4e2f1