A reproducible, rubric-based evaluation harness for LLM outputs. Define what "good" means in a versioned JSON rubric, score it with deterministic checks and LLM judges, and get back a report that says how confident the numbers are and where the judge gave up.
Most LLM evaluation code produces a single number that nobody can reproduce or defend. Four failure modes account for most of it:
- A judge that cannot be parsed is scored 0.0. That is indistinguishable from a genuine failure and silently drags every average down.
- Means are reported without intervals. 0.82 over 12 samples and 0.82 over 1200 samples get printed the same way, and only one of them means anything.
- Thresholds and weights live in code. Nobody can tell which rubric produced a number that was published three weeks ago.
- Judges are used for things a regex could check, which costs money and adds variance for no gain in signal.
This harness takes a position on all four.
| Problem | What this does |
|---|---|
| Unparseable judge output | Abstains. CriterionScore enforces that a value and an abstention are mutually exclusive, and the report prints abstention counts per criterion. |
| Bare means | Every mean carries a seeded percentile-bootstrap 95% interval and its sample count. |
| Untracked config | The rubric is JSON, canonically hashed into a config_hash alongside the dataset digest, provider, model and parameters. |
| Judge overuse | Seven deterministic scorers ship first; the judge is for what genuinely needs one. |
dataset.jsonl rubric.json
| |
sha256 digest canonical sha256
| |
+-----------+------------+
|
EvaluationRunner
|
+-------------------+--------------------+
| |
predict(sample) score(criterion)
| |
ResponseCache (content-addressed) Scorer registry
| |
Model provider +------------+------------+
(deterministic / OpenAI / | |
Anthropic / your own) deterministic llm_judge
exact_match |
contains_all judge Model
regex_match (abstains on a
json_valid parse failure)
token_f1
word_count_within
no_forbidden_terms
| |
+-------------------+--------------------+
|
weighted roll-up per sample
(abstentions excluded, weights
renormalised, never zeroed)
|
aggregate + bootstrap intervals
|
RunManifest + Aggregate + results
|
report.json / report.md
- Rubrics as versioned data — weighted criteria, per-criterion pass thresholds, canonical hashing
- Seven deterministic scorers —
exact_match,contains_all,regex_match,json_valid,token_f1,word_count_within,no_forbidden_terms - LLM judge with a 0–4 anchored scale — normalised to
[0, 1], abstains on a parse failure or an off-scale score, and does not see the reference answer unless a criterion opts in - Bootstrap confidence intervals — seeded, so a report is byte-stable for a fixed config
- Content-addressed response cache — keyed on provider, model, prompt and parameters; atomic writes; a corrupt entry degrades to a miss instead of failing a run
- Fail-fast configuration — an unknown scorer or a judge-requiring rubric with no judge raises at construction, not on sample 400 of 500
- Failure isolation — a provider error or a broken scorer costs one abstention, not the whole run
- A built-in baseline — lead-N extractive summarisation, so a model's score always has a floor to beat
- CI regression gate —
--fail-underexits non-zero when the mean drops
Python 3.10+ · pydantic 2 · argparse · pytest · ruff · mypy (strict) · Docker · GitHub Actions
One runtime dependency (pydantic). httpx is an optional extra needed only for the
hosted providers, imported lazily, so the core and the whole test suite run without it.
src/llmeval/
types.py validated data model (Sample, Prediction, scores, manifest, report)
rubric.py Criterion, Rubric, canonical hashing, JSON loading
dataset.py JSONL loading with line-accurate errors; file digests
cache.py content-addressed response cache
aggregate.py weighted roll-ups, pass rates, bootstrap intervals
runner.py orchestration, retries, concurrency, config hash
report.py report.json and report.md rendering
cli.py run / validate / scorers
providers/
base.py the Model protocol
deterministic.py ExtractiveBaseline, ScriptedModel
http_providers.py OpenAI- and Anthropic-compatible clients
scorers/
base.py Scorer protocol and registry
deterministic.py the seven mechanical scorers
judge.py prompt construction, parsing, abstention
tests/ 139 tests, no network access
examples/
datasets/ summarisation set (text written for this repository)
rubrics/ deterministic and judge-augmented rubrics
docs/
architecture.md
methodology.md
git clone https://github.com/technikky/llm-evaluation-framework.git
cd llm-evaluation-framework
python -m venv .venv && source .venv/bin/activate # Windows: .venv\Scripts\activate
pip install -e ".[dev,providers]"No key is needed for the deterministic provider or for any test. For hosted providers, copy the template and fill it in:
cp .env.example .env| Variable | Required for | Purpose |
|---|---|---|
OPENAI_API_KEY |
--provider openai |
API key |
OPENAI_BASE_URL |
optional | override for a compatible endpoint |
ANTHROPIC_API_KEY |
--provider anthropic |
API key |
ANTHROPIC_BASE_URL |
optional | endpoint override |
LLMEVAL_CACHE_DIR |
optional | cache location (default .llmeval_cache) |
.env is gitignored. Keys are read from the environment only; there is no key
literal anywhere in this repository, and CI runs gitleaks on every push.
Check a dataset and rubric before spending anything:
llmeval validate --dataset examples/datasets/summarization.jsonl \
--rubric examples/rubrics/summarization.jsonRun the deterministic baseline — no keys, no network:
llmeval run --dataset examples/datasets/summarization.jsonl \
--rubric examples/rubrics/summarization.json \
--provider deterministic \
--out runs/baselineEvaluate a hosted model with judge-scored criteria:
llmeval run --dataset examples/datasets/summarization.jsonl \
--rubric examples/rubrics/summarization-with-judge.json \
--provider openai --model gpt-4o-mini \
--judge-provider anthropic --judge-model claude-sonnet-5 \
--out runs/gpt-4o-miniList the scorers a rubric can name:
llmeval scorersAs a library:
from llmeval import EvaluationRunner, Rubric, load_dataset
from llmeval.providers import build_provider
samples = load_dataset("examples/datasets/summarization.jsonl")
rubric = Rubric.from_json_file("examples/rubrics/summarization.json")
report = EvaluationRunner(
build_provider("deterministic"),
rubric,
provider="deterministic",
).run(samples, dataset_path="examples/datasets/summarization.jsonl")
print(report.aggregate.mean_weighted_score, report.manifest.config_hash)A rubric is plain JSON:
{
"name": "summarization-deterministic",
"version": "1",
"criteria": [
{
"name": "conciseness",
"scorer": "word_count_within",
"weight": 2.0,
"pass_threshold": 0.9,
"args": { "min_words": 8, "max_words": 60 }
}
]
}pytest # 139 tests
pytest --cov=llmeval --cov-report=term-missing # coverage
ruff check . && ruff format --check .
mypy # strictNo test makes a network call. Every test runs a real evaluation against
ExtractiveBaseline or ScriptedModel, so the suite exercises the whole pipeline
rather than mocking it away — which is why it is a usable gate in CI.
Dataset. examples/datasets/summarization.jsonl — 8 short passages with human
reference summaries, written for this repository. No third-party or proprietary text.
Methodology. Each passage is summarised, then scored on three deterministic criteria: unigram F1 against the reference (weight 3), a word-count band of 8–60 (weight 2) and the absence of model meta-commentary (weight 1). A sample's score is the weight-normalised mean over criteria that produced a value; a sample passes only if every criterion that produced a value passed. Intervals are seeded percentile bootstraps over 2000 resamples.
ExtractiveBaseline (lead-2), 8 samples, seed 0, config hash 76950c7469507125:
| Metric | Result |
|---|---|
| Mean weighted score | 0.691 (95% CI [0.651, 0.731]) |
| Pass rate | 62.5% |
content_overlap (token F1) |
0.381 (95% CI [0.302, 0.462]) |
conciseness |
1.000 (95% CI [1.000, 1.000]) |
no_meta_commentary |
1.000 (95% CI [1.000, 1.000]) |
| Prediction errors | 0 |
| Abstentions | 0 |
| Test suite | 139 passed |
| Statement coverage | 91% |
Reproduce with:
llmeval run --dataset examples/datasets/summarization.jsonl \
--rubric examples/rubrics/summarization.json \
--provider deterministic --no-cacheNo hosted-model numbers are published here. Running gpt-4o-mini or
claude-sonnet-5 against this rubric costs money and I have not paid for a run whose
results I would then be asserting as fact, so those rows read not measured rather
than carrying a plausible-looking figure. The command above is the one that produces
them.
- The dataset is 8 samples. The intervals are wide on purpose and should be read as "this pipeline works end to end", not "lead-2 scores 0.691 at summarisation".
token_f1is lexical overlap, not semantics. A correct paraphrase scores poorly. It is here because it is free and catches gross content drift.concisenessandno_meta_commentaryboth score 1.000 because a lead-2 extract is structurally incapable of failing them. That is a finding about the rubric, not a result about the model: two of three criteria carry no discriminative signal against this baseline. A rubric that cannot separate a real model from lead-2 is not yet measuring quality, which is exactly what shipping a baseline is for.
docker compose up --build # runs the deterministic baseline into ./runsOr directly:
docker build -t llmeval .
docker run --rm -v "$PWD/runs:/app/runs" llmeval \
run --dataset=examples/datasets/summarization.jsonl \
--rubric=examples/rubrics/summarization.json \
--provider=deterministic --out=runs/dockerMulti-stage build: a builder stage produces a wheel, the runtime stage installs it and
drops to a non-root user (uid 10001). Keys are supplied through env_file, never baked
into an image layer.
.github/workflows/ci.yml runs on every push and pull request:
push / PR
|
+-- test (Python 3.10, 3.11, 3.12)
| ruff check -> ruff format --check -> mypy --strict -> pytest --cov
|
+-- secret scan (gitleaks, full history)
|
+-- docker build (needs: test)
build runtime image -> --help -> a real evaluation inside the container
The Docker job runs an actual evaluation in the built image, so a broken entrypoint or a missing example file fails CI rather than being discovered by whoever clones the repository next.
Stated plainly, because a harness that overstates itself is worse than none:
- The example dataset is a demonstration, not a benchmark. 8 samples, single domain, written by one person.
token_f1is not a semantic metric. Treat it as a drift detector.- Single-judge scoring is not calibrated. No inter-judge agreement, no human-agreement measurement, no position-bias controls. Judge scores here are indicative. Calibration is on the roadmap and is not claimed.
- No multi-turn or tool-use evaluation. Single prompt in, single completion out.
- Concurrency is thread-based, which suits HTTP-bound work and nothing else.
- The cache never expires. Delete the directory when a provider changes a model behind a stable name.
- Judge calibration: inter-judge agreement and agreement against human labels
- Paired significance testing between two runs, not just per-run intervals
- Position-swap and self-preference controls for pairwise judging
- Multi-turn and tool-use evaluation
-
llmeval comparefor tworeport.jsonfiles, for CI regression diffs - Optional semantic similarity scorer behind an extra
MIT — see LICENSE.