Skip to content

About

Rubric-based LLM evaluation harness - an unparseable judge response abstains instead of scoring zero, every mean carries a seeded bootstrap interval, and rubrics are hashed into a config_hash. Python, pytest, Docker, GitHub Actions.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

llm-evaluation-framework

A reproducible, rubric-based evaluation harness for LLM outputs. Define what "good" means in a versioned JSON rubric, score it with deterministic checks and LLM judges, and get back a report that says how confident the numbers are and where the judge gave up.

CI License: MIT Python 3.10+

Problem

Most LLM evaluation code produces a single number that nobody can reproduce or defend. Four failure modes account for most of it:

  1. A judge that cannot be parsed is scored 0.0. That is indistinguishable from a genuine failure and silently drags every average down.
  2. Means are reported without intervals. 0.82 over 12 samples and 0.82 over 1200 samples get printed the same way, and only one of them means anything.
  3. Thresholds and weights live in code. Nobody can tell which rubric produced a number that was published three weeks ago.
  4. Judges are used for things a regex could check, which costs money and adds variance for no gain in signal.

This harness takes a position on all four.

How it answers them

Problem What this does
Unparseable judge output Abstains. CriterionScore enforces that a value and an abstention are mutually exclusive, and the report prints abstention counts per criterion.
Bare means Every mean carries a seeded percentile-bootstrap 95% interval and its sample count.
Untracked config The rubric is JSON, canonically hashed into a config_hash alongside the dataset digest, provider, model and parameters.
Judge overuse Seven deterministic scorers ship first; the judge is for what genuinely needs one.

Architecture

                 dataset.jsonl            rubric.json
                      |                        |
                sha256 digest          canonical sha256
                      |                        |
                      +-----------+------------+
                                  |
                          EvaluationRunner
                                  |
              +-------------------+--------------------+
              |                                        |
        predict(sample)                          score(criterion)
              |                                        |
      ResponseCache (content-addressed)       Scorer registry
              |                                        |
        Model provider                    +------------+------------+
     (deterministic / OpenAI /            |                         |
      Anthropic / your own)         deterministic              llm_judge
                                    exact_match                  |
                                    contains_all           judge Model
                                    regex_match          (abstains on a
                                    json_valid            parse failure)
                                    token_f1
                                    word_count_within
                                    no_forbidden_terms
              |                                        |
              +-------------------+--------------------+
                                  |
                      weighted roll-up per sample
                   (abstentions excluded, weights
                       renormalised, never zeroed)
                                  |
                   aggregate + bootstrap intervals
                                  |
                    RunManifest + Aggregate + results
                                  |
                    report.json  /  report.md

Features

  • Rubrics as versioned data — weighted criteria, per-criterion pass thresholds, canonical hashing
  • Seven deterministic scorers — exact_match, contains_all, regex_match, json_valid, token_f1, word_count_within, no_forbidden_terms
  • LLM judge with a 0–4 anchored scale — normalised to [0, 1], abstains on a parse failure or an off-scale score, and does not see the reference answer unless a criterion opts in
  • Bootstrap confidence intervals — seeded, so a report is byte-stable for a fixed config
  • Content-addressed response cache — keyed on provider, model, prompt and parameters; atomic writes; a corrupt entry degrades to a miss instead of failing a run
  • Fail-fast configuration — an unknown scorer or a judge-requiring rubric with no judge raises at construction, not on sample 400 of 500
  • Failure isolation — a provider error or a broken scorer costs one abstention, not the whole run
  • A built-in baseline — lead-N extractive summarisation, so a model's score always has a floor to beat
  • CI regression gate — --fail-under exits non-zero when the mean drops

Tech stack

Python 3.10+ · pydantic 2 · argparse · pytest · ruff · mypy (strict) · Docker · GitHub Actions

One runtime dependency (pydantic). httpx is an optional extra needed only for the hosted providers, imported lazily, so the core and the whole test suite run without it.

Project structure

src/llmeval/
    types.py                  validated data model (Sample, Prediction, scores, manifest, report)
    rubric.py                 Criterion, Rubric, canonical hashing, JSON loading
    dataset.py                JSONL loading with line-accurate errors; file digests
    cache.py                  content-addressed response cache
    aggregate.py              weighted roll-ups, pass rates, bootstrap intervals
    runner.py                 orchestration, retries, concurrency, config hash
    report.py                 report.json and report.md rendering
    cli.py                    run / validate / scorers
    providers/
        base.py               the Model protocol
        deterministic.py      ExtractiveBaseline, ScriptedModel
        http_providers.py     OpenAI- and Anthropic-compatible clients
    scorers/
        base.py               Scorer protocol and registry
        deterministic.py      the seven mechanical scorers
        judge.py              prompt construction, parsing, abstention
tests/                        139 tests, no network access
examples/
    datasets/                 summarisation set (text written for this repository)
    rubrics/                  deterministic and judge-augmented rubrics
docs/
    architecture.md
    methodology.md

Installation

git clone https://github.com/technikky/llm-evaluation-framework.git
cd llm-evaluation-framework
python -m venv .venv && source .venv/bin/activate    # Windows: .venv\Scripts\activate
pip install -e ".[dev,providers]"

Configuration

No key is needed for the deterministic provider or for any test. For hosted providers, copy the template and fill it in:

cp .env.example .env
Variable Required for Purpose
OPENAI_API_KEY --provider openai API key
OPENAI_BASE_URL optional override for a compatible endpoint
ANTHROPIC_API_KEY --provider anthropic API key
ANTHROPIC_BASE_URL optional endpoint override
LLMEVAL_CACHE_DIR optional cache location (default .llmeval_cache)

.env is gitignored. Keys are read from the environment only; there is no key literal anywhere in this repository, and CI runs gitleaks on every push.

Usage

Check a dataset and rubric before spending anything:

llmeval validate --dataset examples/datasets/summarization.jsonl \
                 --rubric examples/rubrics/summarization.json

Run the deterministic baseline — no keys, no network:

llmeval run --dataset examples/datasets/summarization.jsonl \
            --rubric examples/rubrics/summarization.json \
            --provider deterministic \
            --out runs/baseline

Evaluate a hosted model with judge-scored criteria:

llmeval run --dataset examples/datasets/summarization.jsonl \
            --rubric examples/rubrics/summarization-with-judge.json \
            --provider openai --model gpt-4o-mini \
            --judge-provider anthropic --judge-model claude-sonnet-5 \
            --out runs/gpt-4o-mini

List the scorers a rubric can name:

llmeval scorers

As a library:

from llmeval import EvaluationRunner, Rubric, load_dataset
from llmeval.providers import build_provider

samples = load_dataset("examples/datasets/summarization.jsonl")
rubric = Rubric.from_json_file("examples/rubrics/summarization.json")

report = EvaluationRunner(
    build_provider("deterministic"),
    rubric,
    provider="deterministic",
).run(samples, dataset_path="examples/datasets/summarization.jsonl")

print(report.aggregate.mean_weighted_score, report.manifest.config_hash)

A rubric is plain JSON:

{
  "name": "summarization-deterministic",
  "version": "1",
  "criteria": [
    {
      "name": "conciseness",
      "scorer": "word_count_within",
      "weight": 2.0,
      "pass_threshold": 0.9,
      "args": { "min_words": 8, "max_words": 60 }
    }
  ]
}

Testing

pytest                                              # 139 tests
pytest --cov=llmeval --cov-report=term-missing       # coverage
ruff check . && ruff format --check .
mypy                                                 # strict

No test makes a network call. Every test runs a real evaluation against ExtractiveBaseline or ScriptedModel, so the suite exercises the whole pipeline rather than mocking it away — which is why it is a usable gate in CI.

Evaluation

Dataset. examples/datasets/summarization.jsonl — 8 short passages with human reference summaries, written for this repository. No third-party or proprietary text.

Methodology. Each passage is summarised, then scored on three deterministic criteria: unigram F1 against the reference (weight 3), a word-count band of 8–60 (weight 2) and the absence of model meta-commentary (weight 1). A sample's score is the weight-normalised mean over criteria that produced a value; a sample passes only if every criterion that produced a value passed. Intervals are seeded percentile bootstraps over 2000 resamples.

Measured baseline

ExtractiveBaseline (lead-2), 8 samples, seed 0, config hash 76950c7469507125:

Metric Result
Mean weighted score 0.691 (95% CI [0.651, 0.731])
Pass rate 62.5%
content_overlap (token F1) 0.381 (95% CI [0.302, 0.462])
conciseness 1.000 (95% CI [1.000, 1.000])
no_meta_commentary 1.000 (95% CI [1.000, 1.000])
Prediction errors 0
Abstentions 0
Test suite 139 passed
Statement coverage 91%

Reproduce with:

llmeval run --dataset examples/datasets/summarization.jsonl \
            --rubric examples/rubrics/summarization.json \
            --provider deterministic --no-cache

No hosted-model numbers are published here. Running gpt-4o-mini or claude-sonnet-5 against this rubric costs money and I have not paid for a run whose results I would then be asserting as fact, so those rows read not measured rather than carrying a plausible-looking figure. The command above is the one that produces them.

What these numbers do and do not mean

  • The dataset is 8 samples. The intervals are wide on purpose and should be read as "this pipeline works end to end", not "lead-2 scores 0.691 at summarisation".
  • token_f1 is lexical overlap, not semantics. A correct paraphrase scores poorly. It is here because it is free and catches gross content drift.
  • conciseness and no_meta_commentary both score 1.000 because a lead-2 extract is structurally incapable of failing them. That is a finding about the rubric, not a result about the model: two of three criteria carry no discriminative signal against this baseline. A rubric that cannot separate a real model from lead-2 is not yet measuring quality, which is exactly what shipping a baseline is for.

Docker

docker compose up --build            # runs the deterministic baseline into ./runs

Or directly:

docker build -t llmeval .
docker run --rm -v "$PWD/runs:/app/runs" llmeval \
  run --dataset=examples/datasets/summarization.jsonl \
      --rubric=examples/rubrics/summarization.json \
      --provider=deterministic --out=runs/docker

Multi-stage build: a builder stage produces a wheel, the runtime stage installs it and drops to a non-root user (uid 10001). Keys are supplied through env_file, never baked into an image layer.

CI/CD

.github/workflows/ci.yml runs on every push and pull request:

push / PR
   |
   +-- test (Python 3.10, 3.11, 3.12)
   |     ruff check -> ruff format --check -> mypy --strict -> pytest --cov
   |
   +-- secret scan (gitleaks, full history)
   |
   +-- docker build (needs: test)
         build runtime image -> --help -> a real evaluation inside the container

The Docker job runs an actual evaluation in the built image, so a broken entrypoint or a missing example file fails CI rather than being discovered by whoever clones the repository next.

Limitations

Stated plainly, because a harness that overstates itself is worse than none:

  • The example dataset is a demonstration, not a benchmark. 8 samples, single domain, written by one person.
  • token_f1 is not a semantic metric. Treat it as a drift detector.
  • Single-judge scoring is not calibrated. No inter-judge agreement, no human-agreement measurement, no position-bias controls. Judge scores here are indicative. Calibration is on the roadmap and is not claimed.
  • No multi-turn or tool-use evaluation. Single prompt in, single completion out.
  • Concurrency is thread-based, which suits HTTP-bound work and nothing else.
  • The cache never expires. Delete the directory when a provider changes a model behind a stable name.

Roadmap

  • Judge calibration: inter-judge agreement and agreement against human labels
  • Paired significance testing between two runs, not just per-run intervals
  • Position-swap and self-preference controls for pairwise judging
  • Multi-turn and tool-use evaluation
  • llmeval compare for two report.json files, for CI regression diffs
  • Optional semantic similarity scorer behind an extra

License

MIT — see LICENSE.

About

Rubric-based LLM evaluation harness - an unparseable judge response abstains instead of scoring zero, every mean carries a seeded bootstrap interval, and rubrics are hashed into a config_hash. Python, pytest, Docker, GitHub Actions.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages