Blog/rag/The RAG production checklist: 25 checks before real users see it
The RAG production checklist: 25 checks before real users see it
25 checks across ingestion, retrieval, generation, evaluation and operations that catch the RAG failures demos hide. Free printable PDF included.
What this post covers
A RAG demo only has to work on the questions you tried. Production has to survive the questions you didn’t. These are the 25 checks I run before a RAG feature goes to real users, grouped by where things break.
Want it as a printable 2-page PDF? Get the free checklist (enter $0).
Ingestion
- Chunks never cross a heading. In my chunking benchmark, fixed-size chunks crossed a section boundary 59% of the time, and retrieval suffered for it.
- Each chunk knows where it came from. Store document ID, title, section path and URL with every chunk.
- The heading path is embedded with the chunk. Cheap context that raised hit@1 by 7 points in the same test.
- Re-ingesting a document replaces it. Uploading v2 of a file must delete v1’s chunks, not duplicate them.
- Embedding model and dimension are recorded. If someone changes the model, the system forces a re-index instead of silently mixing vectors from two models.
- File size and type limits are enforced before parsing, not after.
Retrieval
- Exact tokens are findable. Error codes, SKUs and function names often fail with pure vector search. Add BM25 and fuse the rankings with Reciprocal Rank Fusion — measured here: pure vectors lost 49 MRR points to BM25 on exact identifiers.
- You know your corpus size and latency budget. Under ~100k chunks, exact search takes a few milliseconds, so you may not need a vector database yet.
- If you use an ANN index, you measured its recall on your own embeddings, not just the defaults.
- Empty retrieval is handled. If nothing relevant comes back, the system says so without calling the LLM.
- Permissions are applied at retrieval time. Users can never retrieve chunks from documents they can’t open.
Generation
- Citations are validated on the server. Numbers outside the retrieved set are removed. Here’s a 25-line validator with tests.
- There is an explicit refusal sentence, and it is tested on questions the docs can’t answer.
- Retrieved text is treated as data, not instructions. A document saying “ignore previous instructions” must not change behavior.
- No write tools on the answer path. The model answering questions can’t send emails, delete records or call APIs with side effects.
- Context size is bounded. A cap on passages and tokens per request, so one huge document can’t blow up cost or latency.
- Streaming sends sources first, so the UI can render citations while the answer streams.
Evaluation
- A golden set exists: 20 to 50 real questions with the document that should answer each.
- Retrieval is measured separately from generation (hit@k and MRR for retrieval; citation rate and refusal accuracy for answers).
- Evals run in CI, and the build fails when retrieval quality drops below a threshold.
- Every prompt or model change is compared on the golden set before it ships.
Operations
- Each request logs a request ID, retrieved chunk IDs, latency per stage and token counts. Without these, you can’t debug a bad answer after the fact.
- API keys, rate limits and request size limits are in place on every endpoint.
- Health and readiness checks distinguish “the process is up” from “the index is loaded”.
- Cost per query is known. Tokens in, tokens out, times price. Put it on a dashboard before finance asks.
How to use this list
Don’t try to fix all 25 at once. Go through it once and mark each item as done, not needed or missing. Then fix the missing ones in this order: 12, 10, 14, 7, 18, 20. Those catch the failures users notice first: invented sources, confident answers with no basis, and quality that slips quietly after a change.
Want the production version of this?
A production RAG + MCP starter kit for FastAPI: hybrid search, validated citations, evals in CI, Docker, 32 tests. It runs offline in about 60 seconds.
See ShipRAG or get the free RAG checklist first.