Your RAG demo works. Here's why production will break it.
Retrieval-augmented generation fails quietly: stale indexes, permission leaks and answers that sound right but aren't. An evaluation-first approach catches it before customers do.

A retrieval-augmented assistant over 200 hand-picked documents is a weekend project. The same assistant over 2 million documents, with permissions, updates and real users asking ambiguous questions, is an engineering discipline.
Five ways production RAG fails
- Retrieval misses. The right passage exists but never reaches the model, often because of chunking that splits a table from its header.
- Confident synthesis. The model fills a gap in the retrieved context with something plausible and wrong.
- Permission leakage. An answer quotes a document the user isn't allowed to open.
- Index drift. The policy changed last week; the index didn't.
- Silent regression. A prompt tweak improves one question type and quietly breaks three others.
Build the evaluation harness before the assistant
Before tuning anything, we build a golden dataset: 150–300 real questions with expected answers and the source passages that support them, written with subject-matter experts. Every change to chunking, retrieval, prompts or models runs against it automatically.
- Retrieval recall@k: did the right passages come back?
- Faithfulness: is every claim supported by a retrieved passage?
- Answer correctness, graded against the expert answer
- Refusal quality: does it say 'I don't know' when it should?
Enforce permissions at retrieval, not at display
Filtering answers after generation is too late; the model has already read the restricted document. Access control lists must be indexed alongside content and applied as a hard filter in the retrieval query.
The takeaway
If you can't say what your assistant's accuracy is today, as a number on a known dataset, you're not ready for production. The good news: building that harness takes days, not months, and it turns every future decision from opinion into measurement.
- RAG
- Evaluation
- GenAI



