Blog/rag/RAG chunking, measured: heading-aware chunks doubled hit@1 on the FastAPI docs
RAG chunking, measured: heading-aware chunks doubled hit@1 on the FastAPI docs
I tested 5 chunking strategies on 151 real docs and 773 queries. Fixed-size chunks crossed section boundaries 59% of the time. Heading-aware chunks never did.
What this post covers
How this was tested
FastAPI English docs (151 Markdown files, 1.45M characters), 773 queries, BM25 retrieval (k1=1.5, b=0.75), Python 3.11. Full script at the end.
Chunking is the first decision in a RAG pipeline and the one people tune least. Most projects take the default of their framework, often fixed-size chunks with overlap, and move on. I wanted numbers, so I ran five strategies over a real documentation set and measured how often retrieval finds the right section.
Result: splitting at headings first, then packing paragraphs, raised hit@1 from 0.39 to 0.71. Adding the heading path to each chunk raised it again to 0.79.
The test
- Corpus: the English FastAPI docs, 151 Markdown files with 1.45 million characters. Real technical writing, with headings, code blocks and admonitions.
- Queries: every h2/h3 heading with at least 300 characters of body and 2+ words. That gave 773 queries, such as “Declare request example data” or “Dependencies with yield and except”.
- Correct answer: a retrieved chunk counts as a hit when at least half of its text lies inside the section that heading introduces.
- Retriever: BM25. It is deterministic and needs no API key, so anyone can reproduce the numbers exactly.
Strategies
- Fixed 800/100: 800-character windows with 100 characters of overlap. The classic default.
- Fixed 530/100: the same, but sized to match the heading-aware chunks, as a control.
- Paragraph-packed 800: greedily pack blank-line-separated paragraphs up to 800 characters. This is roughly what “recursive” splitters do.
- Heading-aware: split at every heading first, so a chunk never contains two sections, then paragraph-pack within the section.
- Heading-aware + heading path: the same chunks, with
Page title > Section > Subsectionprepended to the text that gets indexed.
Results
| Strategy | Chunks | Avg chars | Crosses a section | Splits a code block | hit@1 | hit@5 | MRR |
|---|---|---|---|---|---|---|---|
| Fixed 800/100 | 2,142 | 770 | 58.8% | 7.3% | 0.389 | 0.611 | 0.477 |
| Fixed 530/100 (size-matched) | 3,446 | 516 | 47.5% | 6.1% | 0.481 | 0.766 | 0.598 |
| Paragraph-packed 800 | 1,737 | 835 | 57.2% | 3.6% | 0.375 | 0.576 | 0.455 |
| Heading-aware | 2,713 | 534 | 0.0% | 2.1% | 0.712 | 0.933 | 0.804 |
| Heading-aware + heading path | 2,713 | 534 | 0.0% | 2.1% | 0.785 | 0.966 | 0.859 |
What the numbers say
More than half of fixed-size chunks straddle two sections. A chunk that is half “Install” and half “Configuration” is a mediocre match for both. That shows up directly as lower retrieval scores.
Smaller chunks help, but not as much as structure. The size-matched control improved on the 800-character default (hit@5 0.61 → 0.77), which confirms part of the gain is simply size. But heading-aware chunks of the same average size scored 0.93. The remaining gap comes from respecting section boundaries.
Paragraph packing alone didn’t help. It halves the number of broken code blocks, which is nice for answer quality, but its chunks cross sections just as often as fixed windows do.
The heading path is cheap recall. A chunk from the middle of a long section often doesn’t repeat the words of its title. Prepending Tutorial > Security > OAuth2 scopes gives every chunk that context. It cost nothing and added 7 points of hit@1.
Honest caveats
- Heading-as-query favors heading-aware chunking. Real user questions are phrased differently from section titles. The direction of the result should hold, but expect smaller gaps with real queries.
- The hit rule rewards staying inside one section. That is the point I’m testing, but a chunk that crosses sections and still contains the answer counts as a miss here.
- BM25, not embeddings. Dense retrievers are more forgiving of wording. Section-crossing chunks still dilute an embedding across two topics, so I expect the same ranking, but run it on your own stack.
- Markdown makes this easy. For PDFs you first need to recover headings (font size, bold, numbering). That extraction step is worth the effort.
A heading-aware chunker in 30 lines
import re
HEADING = re.compile(r"^(#{1,3}) (.+?)\s*(\{.*\})?$", re.M)
SIZE = 800
def sections(text):
"""Yield (start, end, 'Page > Section > Sub') for each heading."""
ms = list(HEADING.finditer(text))
stack = []
if not ms or ms[0].start() > 0:
yield 0, ms[0].start() if ms else len(text), ""
for i, m in enumerate(ms):
level, title = len(m.group(1)), m.group(2).strip()
stack = [s for s in stack if s[0] < level] + [(level, title)]
end = ms[i + 1].start() if i + 1 < len(ms) else len(text)
yield m.start(), end, " > ".join(t for _, t in stack)
def pack(text):
"""Greedily pack blank-line paragraphs up to SIZE characters."""
out, start, cur = [], 0, 0
for m in re.finditer(r"\n\s*\n", text):
if m.end() - start > SIZE and cur > start:
out.append(text[start:cur]); start = cur
cur = m.end()
out.append(text[start:])
return [c for c in out if c.strip()]
def chunk(text):
for s, e, path in sections(text):
for body in pack(text[s:e]):
yield {"text": body, "embed_text": f"{path}\n{body}" if path else body, "section": path}
Store text for display and citations. Embed and index embed_text.
What I’d do in your pipeline
- Never let a chunk cross a heading. It is the single biggest win in this test.
- Prepend the heading path to what you embed and index, not to what you show the user.
- Keep chunks around 500 characters for technical docs, then tune with your own eval set.
- Measure it. Generate queries from your own headings like I did here, and you have a free retrieval regression test in under an hour.
Full benchmark script
The complete script (BM25, all five strategies, metrics) is about 90 lines of standard-library Python. Point it at a folder of Markdown files:
# Same chunkers as above, plus:
import collections, math
TOK = re.compile(r"[a-z0-9_]+")
toks = lambda s: TOK.findall(s.lower())
def bm25_index(chunks, k1=1.5, b=0.75):
tf = [collections.Counter(c) for c in chunks]
N, L = len(chunks), [len(c) for c in chunks]
avg = sum(L) / N
df = collections.Counter(t for c in tf for t in c)
idf = {t: math.log(1 + (N - n + .5) / (n + .5)) for t, n in df.items()}
inv = collections.defaultdict(list)
for i, c in enumerate(tf):
for t, f in c.items():
inv[t].append((i, f))
def search(q, k=5):
sc = collections.defaultdict(float)
for t in set(toks(q)):
for i, f in inv.get(t, ()):
sc[i] += idf[t] * f * (k1 + 1) / (f + k1 * (1 - b + b * L[i] / avg))
return sorted(sc, key=sc.get, reverse=True)[:k]
return search
For each h2/h3 heading, search with the heading text, and count a hit when a top-k chunk falls mostly inside that heading’s section.
Want the production version of this?
A production RAG + MCP starter kit for FastAPI: hybrid search, validated citations, evals in CI, Docker, 32 tests. It runs offline in about 60 seconds.
See ShipRAG or get the free RAG checklist first.