| description | Drive Coder Eval from inside Claude Code — install the plugin, scaffold a task directory, author and review a task, then read the run it produces. |
|---|
By the end you'll have installed the Coder Eval plugin and driven a full loop from slash commands: scaffold, author, review, run, analyze. ~15 minutes.
Cost: steps 1–2 are free; budget one paid agent run for steps 3–4. The optional step 5 is one run per row — 16 for an 8/8 suite — so budget it separately.
-
Claude Code installed and working.
-
The
coder-evalCLI. Installing the plugin does not install it — a plugin ships skills, not packages. You can let the skills handle it (init,taskandcheck-skillcheck for it and offer to install it, asking first), or do it now:uv tool install coder-eval # or: pip install coder-eval coder-eval --version -
An API key for whichever agent your tasks use (
ANTHROPIC_API_KEYfor the defaultclaude-codeagent).
Steps 2 onward assume you are in your own repository — not the coder_eval
clone used by Tutorials 01 and 04.
This repository is itself the marketplace. Both of these are typed at the Claude Code prompt, not in your shell:
/plugin marketplace add UiPath/coder_eval
/plugin install coder-eval@coder-eval
Verify by typing /coder-eval:. You should see six commands: init,
check-skill, task, lint-tasks, analyze, ci. Three of them — init,
check-skill and task — drive the same coder-eval CLI you would type by hand,
so what they write is a normal file you can commit, diff and run in CI. The other
three never invoke it: analyze reads a finished run directory, ci writes a
workflow, and lint-tasks only reads task files and reports.
/coder-eval:init
It scans for what is worth evaluating (Claude Code skills, an MCP server, a CLI), reports what it found, then scaffolds a task directory with one runnable task.
Note the directory it reports — the layout varies by repository (tasks/,
tests/tasks/, …) and later steps need that path:
ls tasks/ # or whichever path init reported
cat tasks/*.yaml | head -40Read that task before moving on. It is the shape every later task here gets modeled on, and step 3 is easier to follow once you have seen one.
/coder-eval:task a task that checks the CLI can list processes as JSON
It designs criteria against the bundled task-quality rubric, then re-checks the files it wrote against the rubric's framing question: what is the cheapest thing an agent could do that scores full marks?
Then it validates with coder-eval plan and offers to run the task, asking
first. Take the offer — step 4 needs a run. Read the score the way the skill
does: a 1.000 on a first attempt means re-read the criteria, not celebrate, and
a failure is a question about which layer is wrong (the prompt, or the skill or
tool it depends on) before it is a prompt edit.
Now review what you already have. Point the read-only linter at the directory from step 2:
/coder-eval:lint-tasks tasks/
Same rubric, applied to files on disk. Per task you get a severity, a line
reference and a concrete fix, covering criteria that cannot fail, prompts that give
away the answer, fixtures with no cleanup and near-duplicates. Gameability findings
name the weight at risk, e.g. "A single --file call satisfies 14.0 of 33.0
weight". It never edits a file, and it scores test design only, so it closes by
suggesting coder-eval plan for the schema half. Watch for ⚠ notices there: an
unknown top-level key warns rather than fails.
If you accepted the run in step 3, you already have a run directory: the skill invoked the CLI through Bash on your behalf. By hand it is the same command, which is the whole point of the plugin being a driver rather than a separate product:
coder-eval run # discovers tasks recursively; or pass explicit paths
ls runs/latest/ # the run that was just writtenHand it back to the agent:
/coder-eval:analyze runs/latest
It writes analysis.md into the run directory, containing:
- a TL;DR and a score breakdown,
- per-task findings with concrete fixes, ranked by estimated score recovery,
- on suites over 20 tasks, failures clustered into systemic patterns instead of repeated per task.
If this repository has Claude Code skills, the plugin can measure whether the model reaches for one at the right moment.
Export the skill location first — the evaluated agent runs in a fresh sandbox holding none of your files, so it is offered no skills unless the task says where they live. Point at the directory containing the skill's own directory:
export SKILL_SOURCE_PATH="$(pwd)/.claude/skills"/coder-eval:check-skill pdf-forms
You get recall, precision and F1 over a labeled suite of requests the skill should
win plus distractors it should decline, gated by suite_thresholds on
recall.yes / precision.yes. Each row is a full agent run, so the skill states
the count and asks before starting.
Low recall has three possible causes, not one: truncation and listing-budget
eviction look identical to bad wording.
The plugin page
covers telling them apart with /doctor and /context.
| Symptom | Cause |
|---|---|
No /coder-eval: commands after installing |
Check /plugin; re-run the install |
A skill offers to install the CLI, or Bash reports command not found |
The CLI isn't installed or isn't on PATH — accept the offer, or install it yourself |
coder-eval run matches nothing |
Wrong directory — use the path init reported in step 2 |
| Every positive row in step 5 scores 0 | SKILL_SOURCE_PATH is unset, so the skill was never offered |
To update after the marketplace moves, /plugin marketplace update coder-eval; to
remove it, /plugin uninstall.
/coder-eval:ciemits the CI workflow from Tutorial 02 for you — least-privilege by default, and it provides the agent runtime (Node plus the Claude CLI) that the Action deliberately does not install.- Claude Code plugin — the reference page: every skill, what ships in the plugin, and the activation-budget mechanics in full.
- Writing a task — the same authoring loop by hand, worth doing once to see what the skill is producing.
- Task Definition Guide — the complete task schema.