Case 01Personal buildRetrieval & LLM systems
A multi-tenant RAG platform that has to show its sources
Answers stream back with citations, fit a strict context budget, and pass regression-tracked evaluation before anything changes.
PythonFastAPIPostgreSQL + pgvectorRedisLangChainDockerPrometheus
- Hybrid retrieval: BM25 keyword search plus vector search, reranked by a cross-encoder.
- Citation-enforced, streaming answers with strict context budgeting.
- Evaluation for faithfulness, relevance and completeness, tracked for regressions.
- Structured logging, tracing and Prometheus metrics from day one.
- 01IngestAsync document ingestion per tenant
- 02Indexpgvector embeddings + BM25
- 03RetrieveHybrid keyword + vector search
- 04RerankCross-encoder orders candidates
- 05AnswerStreamed, cited, within budget
- 06EvaluateFaithfulness · relevance · completeness
The problem
RAG chat is easy to demo and hard to trust. When several customers share one system, three things break first: data leaking between tenants, answers that sound right but aren't grounded, and quality drifting quietly each time a prompt or index changes.
What I built
A FastAPI backend with async ingestion into PostgreSQL and pgvector, retrieval scoped to each tenant, and an LLM layer that streams answers and ties every claim to a retrieved passage. An evaluation harness and observability stack sit around it so changes can be measured.
Where it stands
The retrieval, generation, evaluation and observability layers all run end to end. The evaluation suite is the part I'd point a reviewer to first: it's what turns "seems better" into a number.
Key decisions
- Hybrid over pure vectorsKeyword search catches exact terms such as IDs and names that embeddings blur. The cross-encoder then decides the final order.
- Citations are enforced, not requestedResponses must reference retrieved passages, so a reader can check every claim against its source.
- A hard context budgetRetrieved text is fitted to a fixed token budget, which keeps cost and latency predictable as documents grow.
- Evaluation as a regression testQuality scores are tracked across changes, so a change that makes answers worse shows up before it ships.