Most retrieval-augmented pilots demo well and fail in month two. The difference is evaluation, chunking discipline, and honest failure states.
A retrieval demo is easy. A retrieval system that answers correctly on the twelve-thousandth document, in three languages, with permissions applied, is an engineering problem.
Start with the corpus, not the model. Document quality, versioning, and access control determine ceiling accuracy far more than the choice of embedding model. If two contradictory policy versions live in the index, no prompt will rescue the answer.
Build an evaluation set before the first deployment: a hundred real questions with reference answers, scored on every release. Without it, you cannot tell a model upgrade from a regression.
Finally, design the failure state. A system that says 'I could not find this in the approved sources' earns trust. A system that guesses confidently loses the room permanently.
Applied AI & Data
The AI practice designs, evaluates, and operates applied AI systems — retrieval, agents, document intelligence, and forecasting — with measurable business baselines.