Retrieval is the piece of this work I have come back to most often. It turns up in client projects, in work systems, and in my own assistant, and it is the same shape every time: something has to find the right item out of a very large set, and the tempting move is always to ask a model to remember it.
It will answer. That is the problem. A model asked to recall an identifier from memory hands you one that is well formed, plausible, and wrong, and nothing downstream can tell the difference. So in any system where the answer has to be correct rather than merely convincing, I make the stage that finds it a search rather than a generation. On an architecture diagram it looks like a step backwards, and it is the single decision I would defend hardest.
In practice that means indexing the whole vocabulary, tens of thousands of entries, and doing a real lookup against it. Hybrid where it needs to be. Keyword search catches the entries that are lexically obvious and semantically unremarkable, vector search catches the ones worded nothing like the query, and the two get fused rather than either being trusted alone. Embeddings chosen for the subject matter rather than a general purpose encoder, because in a controlled vocabulary the exact wording is the entire problem.
What comes back is typed rather than prose, and anything accepted has to carry the evidence from the source that justifies it. A result with nothing behind it does not get through, which also means somebody reviewing the output can check it without rerunning anything.
And there is a floor. Below a certain confidence the system returns nothing instead of its best guess. A pipeline that quietly guesses is worse than one that stops, because afterwards you cannot tell which answers were found and which were invented.
This is the same shape as the fairy that fires the portals and the validator that decides whether a command may run. Something proposes, something deterministic decides. I did not plan that. I just kept arriving at it.
Most of this is client work, so it is described here by the tools it was built with and nothing else. No client, no product, no repository, no sector.
Built with
- Python
- FAISS
- BM25
- Reciprocal rank fusion
- sentence-transformers
- Domain-specific embeddings
- Cross-encoder reranking
- Confidence thresholds
- Pydantic
- FastAPI
- LangGraph
- pytest
This is client work. It is described by stack only, with no client, product, repository or sector named, here or anywhere else on this site.