Retrieval quality is the whole product.
Everything else in a retrieval-augmented system is plumbing. Notes from running one in production over Ethiopian economic data.
The first version of Ask Shega answered fluently and was often wrong. The model was fine; the retrieval was not. It returned passages about the right company and missed the one with the figure the question was about. A language model will write a confident paragraph over whatever you hand it, so the answer is settled before generation begins.
Chunking decides more than any prompt. Splitting on a fixed token count cuts tables from their captions and figures from their units, and a chunk that says something went up, without naming what, is worse than none. We split on meaning: sections, then paragraphs, with the heading path carried into every piece so a fragment knows where it came from.
Vector search alone is not enough. Embeddings are good at paraphrase and bad at exact things: a ticker, a birr figure, a name. Hybrid retrieval, vector and keyword over the same Azure AI Search index, catches both, and a semantic rerank over the union does the ordering neither can do alone.
Then measure the retrieval, not the answers: for a fixed set of real questions, did the passage containing the answer land near the top? It moved every time we changed chunking and almost never when we changed the prompt, which told us where the work was.
The last part is showing the working. Every answer streams its sources before its sentences, over plain Server-Sent Events, so a reader sees what was retrieved while the model is still writing. When the sources are wrong, they can tell. When they are right, they can trust the rest. Retrieval quality is the product; the interface exists to make it visible.