A retrieval-augmented generation pipeline whose retrieval quality is measured. Document ingestion, sentence-aware chunking, hybrid dense + BM25 retrieval with reciprocal-rank fusion, reranking, MMR diversification, cited answer generation, and a FastAPI service — plus the evaluation harness that says whether any of it is working.
Most RAG repositories demonstrate a pipeline. Very few can tell you whether the pipeline retrieves the right documents, and that gap produces a specific set of failures:
- Stages are added on faith. A reranker and an MMR pass get bolted on because they are good practice, with no measurement of whether either helps on this corpus. One of them usually doesn't.
- The baseline is never run. BM25 is twenty-five years old, costs nothing, and is frequently within a point of an embedding pipeline. If you never measure it you cannot know whether the vector database is earning its operational cost.
- Metric conventions are mixed silently. Retrieval returns chunks; relevance is labelled on documents. Recall computed one way and Precision the other, reported side by side without saying which is which, is the norm.
- Citations are unverified. A model that cites
[doc-14#2]when no such chunk was retrieved looks exactly like a model that cited correctly, unless something checks.
This repository takes the position that retrieval is an evaluation problem first and an engineering problem second.
| Problem | Response |
|---|---|
| Unmeasured stages | ragkit ablate evaluates dense, BM25, hybrid, +rerank and +MMR over one labelled query set and prints the table. CI runs it on every push. |
| Missing baseline | BM25 is a first-class retriever, not a fallback, and it appears in every ablation. |
| Mixed conventions | Recall and MRR are document-level; Precision is chunk-level; nDCG's ideal ranking is capped by the relevant chunks that actually exist. All three are stated in the output and in docs/evaluation.md. |
| Unverified citations | ragkit ground checks that every cited chunk was retrieved and that every quote occurs in the chunk it cites. Deterministic — no second model. |
Everything below came from a command in this repository, on the 23-document corpus in
data/. Reproduce with ragkit ingest then ragkit ablate (exact commands under
Usage).
Corpus: 23 documents → 34 chunks. Embedder hashing-tfidf-8192, 1112 distinct
tokens, 13.6% hash-collision rate. 24 labelled queries.
| Configuration | Recall@1 | MRR | nDCG@1 | Hit rate | p50 ms | p95 ms |
|---|---|---|---|---|---|---|
dense |
0.938 | 0.958 | 0.958 | 95.8% | 0.04 | 0.08 |
bm25 |
0.938 | 0.958 | 0.958 | 95.8% | 0.04 | 0.06 |
hybrid |
0.938 | 0.958 | 0.958 | 95.8% | 0.09 | 0.11 |
hybrid+rerank |
0.979 | 1.000 | 1.000 | 100.0% | 0.76 | 0.86 |
hybrid+rerank+mmr |
0.979 | 1.000 | 1.000 | 100.0% | 0.99 | 1.20 |
| Configuration | Recall@5 | Precision@5 | MRR | nDCG@5 | p50 ms | p95 ms |
|---|---|---|---|---|---|---|
dense |
1.000 | 0.250 | 0.979 | 0.864 | 0.05 | 0.07 |
bm25 |
1.000 | 0.250 | 0.979 | 0.869 | 0.04 | 0.06 |
hybrid |
1.000 | 0.250 | 0.979 | 0.874 | 0.10 | 0.14 |
hybrid+rerank |
1.000 | 0.233 | 1.000 | 0.848 | 0.76 | 0.83 |
hybrid+rerank+mmr |
1.000 | 0.225 | 1.000 | 0.831 | 1.23 | 1.30 |
Four findings, including the unflattering ones:
-
Recall@5 is saturated and therefore useless here. Every configuration scores 1.000. On a 23-document corpus, five slots are enough for anything to succeed. The discriminating measurement is at k=1, which is why both tables are published. A README reporting only Recall@5 would be reporting nothing.
-
The reranker perfects rank 1 and costs ranking quality below it. MRR goes to 1.000 — it puts a relevant document first on all 24 queries — while nDCG@5 falls from 0.874 to 0.848 and Precision@5 from 0.250 to 0.233. It promotes the single best chunk and reorders positions 2–5 worse. If you feed one chunk to a generator, enable it; if you feed five, it is a net loss on this corpus. That is a tradeoff, not a win, and the table is what makes it visible.
-
MMR does not pay for itself here. It costs nDCG at every k and roughly doubles p50 latency. The corpus has little redundancy for it to remove, so it has nothing useful to do. Shipped and defaulted off.
-
BM25 alone is within half a point of the full hybrid pipeline (nDCG@5 0.869 vs 0.874). With this embedder the honest summary is that hybrid retrieval buys a small, real improvement for roughly double the latency — and that anyone deploying a vector database on this evidence alone should first check what BM25 gives them for free.
| Metric | Result |
|---|---|
| Answers checked | 24 |
| Citations | 64 |
| Grounded | 64 (100.0%) |
| Citing a chunk that was never retrieved | 0 |
| Quote absent from the cited chunk | 0 |
This 100% is trivially true and should not impress anyone. The default generator is
extractive: it quotes retrieved sentences verbatim, so it cannot fabricate a citation.
The check exists for hosted LLM generators, which can and do — and no hosted-model
grounding rate is published here, because that would mean paying for a run and I have
not. --generator openai and --generator anthropic are the commands that produce it.
| Metric | Result |
|---|---|
| Tests | 198 passed |
| Statement coverage | 94% |
| ruff / ruff format / mypy --strict | clean |
| Retrieval latency, p95, hybrid, k=5 | 0.14 ms |
Latency is for retrieval only, in-process, over 34 chunks. It is a floor, not a production figure: it excludes network, generation and any corpus large enough to matter.
documents (JSONL, .md, .txt)
|
load_documents
|
chunk_documents (sentence-aware, overlapping,
| exact character offsets retained)
|
+----------+-----------+
| |
embedder.fit BM25Index.fit LexicalOverlapReranker.fit
embedder.encode | |
| | |
VectorStore.add | |
| | |
+----------+-----------+------------------------+
|
RAGSystem ---- save() / load() --> config.json
| embedder.json (fitted IDF)
| documents.jsonl
| store/{vectors.npy,chunks.jsonl}
|
============ query time ============
|
dense ----+---- bm25 (each retrieves `candidates`)
\ | /
reciprocal_rank_fusion (rank is the only common currency
| between a cosine and a BM25 score)
v
rerank? (IDF-weighted query-term coverage + phrase adjacency)
|
mmr? (relevance vs redundancy, lambda-weighted)
|
truncate to top_k
|
+----------+-----------+
| |
Generator evaluate_retrieval
(extractive / Recall / Precision / MRR / nDCG
openai / latency p50, p95, max
anthropic) |
| v
Answer + citations ablation table
|
check_grounding (does every citation resolve, and does every quote exist?)
Reciprocal-rank fusion, not a weighted score sum. A BM25 score and a cosine similarity have different ranges and different distributions, so any weighted sum of them has an arbitrary constant in it. RRF combines rankings, and rank is the one thing the two retrievers agree on the meaning of.
Chunks keep their character offsets. chunk.start and chunk.end index into the
source document, which is what makes a citation checkable against the original rather
than against a reconstruction of it.
Token buckets come from blake2b, not hash(). Python randomises hash() per
interpreter unless PYTHONHASHSEED is set. A hashing vectoriser built on it produces
different vectors after a restart, so a persisted index would embed queries into a
different space than its documents and return quietly wrong results forever. There is a
test that runs two interpreters under different seeds and requires the same bucket.
- Sentence-aware chunking with character-accurate offsets, configurable overlap, and a hard-split fallback for text with no sentence boundaries
- Hybrid retrieval: TF-IDF-weighted dense vectors + Okapi BM25, fused by RRF
- Reranking on a different signal from the one that retrieved (IDF-weighted query coverage plus surviving phrase adjacency)
- MMR diversification, with relevance normalised so
lambdameans the same thing whatever the incoming scores were - Cited generation: an extractive generator that cannot hallucinate, plus OpenAI and Anthropic generators prompted to cite chunk ids
- Retrieval evaluation: Recall@k, Precision@k, MRR, nDCG@k, hit rate, latency percentiles, with the chunk/document conventions stated
- Ablation harness over retrieval configurations, with a config digest per run
- Deterministic grounding check for citations
- Persistence that saves the fitted embedder state with the vectors, and refuses to load an index built at a different dimension
- FastAPI service with
/query,/retrieve,/healthz,/statsand per-request retrieval overrides
Python 3.10+ · pydantic 2 · numpy · FastAPI · uvicorn · pytest · ruff · mypy (strict) · Docker · GitHub Actions
Three runtime dependencies. Chroma, sentence-transformers and httpx are optional
extras, each imported lazily, so the core and the entire test suite run without them.
src/ragkit/
types.py validated data model (Document, Chunk, ScoredChunk, Answer, metrics)
config.py chunking / retrieval / system config, with a digest for reports
text.py normalisation, tokenisation, sentence spans, stable hashing
chunking.py sentence-aware chunking with offsets and overlap
embeddings.py Embedder protocol, hashed TF-IDF, sentence-transformers adapter
stores.py VectorStore protocol, exact numpy store with persistence, Chroma
retrieval.py BM25, RRF, MMR, the reranker, and the assembled Retriever
generation.py extractive generator, OpenAI and Anthropic, citation parsing
ingest.py JSONL and directory loaders with line-accurate errors
rag.py RAGSystem: ingest, retrieve, answer, save, load
evaluation.py metrics, retrieval evaluation, grounding, ablation
report.py JSON and markdown rendering
api.py FastAPI app
cli.py ingest / query / evaluate / ablate / ground / stats / serve
data/
corpus.jsonl 20 documents in 5 topic clusters
docs/ 3 Markdown documents, exercising the directory loader
eval/queries.jsonl 24 labelled queries
tests/ 198 tests, no network access
docs/
evaluation.md what each metric means here and why the conventions differ
architecture.md module boundaries and extension points
git clone https://github.com/technikky/production-rag-system.git
cd production-rag-system
python -m venv .venv && source .venv/bin/activate # Windows: .venv\Scripts\activate
pip install -e ".[dev,serve]"No API key is needed for the default pipeline, for any test, or for any number in this README. The hashing embedder and the extractive generator are entirely local.
| Variable | Required for | Purpose |
|---|---|---|
OPENAI_API_KEY |
--generator openai |
API key |
OPENAI_BASE_URL |
optional | compatible-endpoint override |
ANTHROPIC_API_KEY |
--generator anthropic |
API key |
ANTHROPIC_BASE_URL |
optional | endpoint override |
RAGKIT_INDEX_DIR |
ragkit serve, Docker |
where the API loads its index from |
cp .env.example .env to start. .env is gitignored, no key literal exists anywhere in
the repository, and CI runs gitleaks on every push.
Build an index:
ragkit ingest data/corpus.jsonl data/docs --out indexAsk a question — no key, no network:
ragkit query "prevent double charging when a payment request times out" --index indexAn idempotency key solves this: the client generates a unique key per logical operation
and sends it with every attempt, and the server stores the key with the result of the
first successful execution. [payments/idempotency-keys#0] ...
retrieval 0.9 ms | generation 0.2 ms | hybrid
[1] payments/idempotency-keys#0 - Idempotency keys (score 0.0328)
...
Measure retrieval, compare configurations, and check citations:
ragkit evaluate --index index --queries data/eval/queries.jsonl --out reports/eval
ragkit ablate --index index --queries data/eval/queries.jsonl --top-k 1 --out reports/k1
ragkit ablate --index index --queries data/eval/queries.jsonl --top-k 5 --out reports/k5
ragkit ground --index index --queries data/eval/queries.jsonl --out reports/groundingTry retrieval variants without rebuilding:
ragkit query "how often should we rebuild images" --index index --mode bm25 --top-k 3
ragkit query "how often should we rebuild images" --index index --rerank --top-k 1Use a hosted model for the answer, keeping the same retrieval:
export OPENAI_API_KEY=...
ragkit query "what should a postmortem contain" --index index --generator openaiServe it:
ragkit serve --index index --port 8000
curl -X POST localhost:8000/query -H 'content-type: application/json' \
-d '{"query":"why did adding a user id label blow up our metrics bill","top_k":3}'As a library:
from ragkit import RAGSystem, evaluate_retrieval, load_documents
from ragkit.cli import load_queries
system = RAGSystem()
system.ingest(load_documents(["data/corpus.jsonl", "data/docs"]))
answer = system.answer("how do idempotency keys prevent double charges")
print(answer.text, answer.cited_doc_ids, answer.retrieval_ms)
metrics = evaluate_retrieval(
system.retriever, load_queries("data/eval/queries.jsonl"), system.chunks
)
print(metrics.recall_at_k, metrics.ndcg_at_k, metrics.latency_p95_ms)pytest # 198 tests
pytest --cov=ragkit --cov-report=term-missing
ruff check . && ruff format --check .
mypyNo test touches the network. The whole pipeline — chunking, embedding, both retrievers, fusion, reranking, MMR, generation, the API and the CLI — is exercised against real in-process indexes rather than mocks, which is why the suite is a usable gate and why the measured numbers above reproduce.
tests/test_cli_and_corpus.py also validates the shipped corpus: every label points at
an indexed document, every query is answerable by something, and retrieval quality
meets a deliberately loose floor so ordinary tuning does not fail CI while a genuine
regression does.
docker compose up --build api # serves on :8000 with a prebuilt index
docker compose run --rm ablate # the ablation table
docker compose run --rm ground # citation groundingThe index is built into the image, so a container answers its first request with no
ingest step. Rebuilding the image is how the corpus changes, which keeps the served index
and the image tag in step. Multi-stage build, non-root user (uid 10001), and a
HEALTHCHECK that hits /healthz.
push / PR
|
+-- test (Python 3.10, 3.11, 3.12)
| ruff check -> ruff format --check -> mypy --strict -> pytest --cov
|
+-- retrieval evaluation
| ragkit ingest
| ragkit ablate (k=1, k=3, k=5) <- the README's tables
| ragkit ground <- non-zero exit on an ungrounded citation
|
+-- secret scan (gitleaks, full history)
|
+-- docker build -> start the container -> query it -> assert the answer
The evaluation job costs nothing to run — no key, no model download — which is why it
runs per-commit instead of nightly. mypy targets 3.12 because numpy 2.5's stubs require
it; the 3.10 floor is enforced by the test matrix and by ruff's target-version.
Stated plainly, because a RAG system that overstates itself is worse than none:
- The default embedder is lexical, not semantic. Hashed TF-IDF cannot match a
paraphrase that shares no vocabulary with the document. Every retrieval number above
is a lexical-retrieval number. A
sentence-transformersembedder is wired through the same interface and is not benchmarked here, because that would mean publishing numbers from a model I have not measured. - 13.6% of tokens share a hash bucket on this corpus at 8192 dimensions. Collisions create small spurious similarities. Measured and reported rather than waved away; raising the dimension reduces it (41.3% at 2048).
- The corpus is 23 documents. It demonstrates the harness and makes the numbers reproducible. It does not establish how any configuration behaves at 100k documents, and nothing here should be read as a claim about that scale.
- Exact search only, in the default store.
NumpyVectorStorescores every chunk. That is the right default — trading recall for latency before measuring recall is how RAG systems end up worse than the baseline they replaced — but it is O(corpus) per query and will not hold at scale. Chroma is wired in for that; it is not benchmarked. - Latency figures are in-process retrieval only, over 34 chunks, excluding network and generation. Treat them as a floor.
- Relevance labels are binary and single-annotator. No graded relevance, no inter-annotator agreement.
- Ingest is a full rebuild, because the embedder's IDF is fitted on the corpus. Correct, and O(corpus) per ingest. A corpus-independent embedder would not need it.
- No hosted-model evaluation. No generation quality, faithfulness or answer-accuracy numbers are published. Only retrieval and citation resolution are measured.
ChromaVectorStoreandSentenceTransformerEmbedderare not covered by tests, deliberately: a suite that silently skips is not a gate.
- Benchmark a sentence-transformers embedder against the hashing baseline on the same query set, so the lexical-vs-semantic gap is a number rather than an assertion
- A larger corpus, at a scale where approximate search and Recall@5 both become informative
- Graded relevance labels and a second annotator
- Cross-encoder reranking, compared against the lexical reranker
- Paired significance testing between two ablation rows, not just point estimates
- Answer-level evaluation with an LLM judge, reusing the rubric harness from llm-evaluation-framework
- Streaming responses and a
/query/streamendpoint - Incremental ingest for corpus-independent embedders
MIT — see LICENSE.