Saerosense새로센스 · applied AI

Writing · 2026-09-19

The production RAG checklist: 12 things to verify before you ship

A RAG demo takes a weekend. A RAG system that survives real users takes discipline. This is the checklist we run before any retrieval-augmented generation system we build goes live.

Retrieval quality (checks 1–4)

1. Measure retrieval separately from generation. If the right passage isn't in the context, the model cannot be right. Track recall@k on a labeled set of question–passage pairs before you look at any answer metric.

2. Audit your chunks by reading them. Open fifty random chunks. If a human can't tell what document a chunk came from or what it's about, neither can your retriever. Fix chunking (structure-aware, with headings carried in) before touching embeddings.

3. Test hybrid retrieval. Dense embeddings miss exact identifiers — part numbers, error codes, statute numbers. A BM25 + dense hybrid with reranking is the boring default that usually wins, especially on technical corpora and mixed-language (e.g., Korean–English) document sets.

4. Check the index against the source of truth. Documents change. Verify you have a re-indexing path and a freshness SLA, or your system will confidently cite last year's policy.

Answer quality (checks 5–8)

5. Build the golden set nobody wants to build. Fifty to two hundred real questions with reference answers, written with the people who will use the system. This is the single highest-leverage artifact in the whole project.

6. Require citations, then verify them. Every claim should trace to a retrieved passage. Automate a spot-check: does the cited passage actually support the sentence? Uncited generation is where hallucinations hide.

7. Test the "no answer" path. Ask questions the corpus cannot answer. A production system must say "I don't know, here's who to ask" — not improvise. Measure the refusal rate on unanswerable questions explicitly.

8. Red-team with real users' phrasing. Users write typos, half-questions, and mixed languages. Sample real queries early; your eval set should look like production traffic, not like documentation.

Operations (checks 9–12)

9. Log everything end to end. Query, retrieved chunks, prompt, answer, latency, cost, feedback — one trace ID across all of it. You cannot debug what you didn't record.

10. Model the cost at real traffic. Multiply expected queries by tokens per query before launch. Reranking and long contexts are where budgets die quietly.

11. Decide the fallback before the outage. Model APIs fail and rate-limit. Define what the user sees when they do: cached answers, a smaller local model, or an honest error — anything but a spinner.

12. Rerun the eval on every change. Prompt edits, model upgrades, index rebuilds — each one reruns the golden set. If the number drops, you know before your users do.


This is the discipline we apply in our AI agents & RAG practice. If you're between the demo and production right now, we're happy to take a look.