Evaluating workflows
A workflow that runs is not a workflow that works. Once a prompt, a model pin, or
a node boundary changes, the only honest answer to “did that help?” is a set of
cases you run before and after, graded the same way both times. keelson eval
is that loop: a YAML case set, a grader per case, a results file with the
statistics you need to tell signal from noise, and a compare command that says
keep or revert.
Why evals
Section titled “Why evals”Four properties make an eval worth trusting. The summary warns when any of them slips.
- Production-aligned cases. Each case is an input the workflow really
receives: the free text that arrives as
$ARGUMENTS, the named inputs that arrive asKEELSON_INPUTS_<key>. Synthetic prompts measure a workflow nobody runs. - Headroom. A set the workflow already passes at 100% cannot show
improvement. The summary prints a
HEADROOMwarning when the test split passes above 95%: add harder cases. - Low variance. A pass rate is an estimate. The summary carries a 95%
Wilson interval for every split and prints a
NOISEwarning when that interval is wider than 20 points. More reps per case, or more cases, narrow it. - Enough cases. Compare counts cases, not reps: at least six cases have to
change before any difference can be called, so a set smaller than that prints
a
POWERwarning. Twenty to thirty cases is a working size. - A held-out split. Tuning a prompt against the cases you stare at makes those cases pass. The train/test split exists so that overfitting shows up as a gap: train improves, test does not, and the verdict says revert.
Graders are either programmatic or an LLM judge with checkable claims. There is
no 1-to-5 scale anywhere: a claim is met or it is not, and the judge grades the
same output repeatedly (grader_reps, default 2) so its own disagreement is
measured and reported as grader noise rather than hidden in the pass rate.
The case file
Section titled “The case file”Scaffold one with keelson eval init <workflow>, which writes
<home>/evals/<workflow>.eval.yaml and refuses to overwrite an existing file.
The repository ships a working example at .keelson/evals/smoke-test.eval.yaml.
name: beads-groom-smokeworkflow: beads-groom # resolved like `workflow run`project: my-project # optional, as `workflow run --project`reps: 3 # default 1
split: # optional; default puts every case in test train: [c1, c2] test: [c3, c4] # or a seeded draw: # seed: 42 # train_fraction: 0.5
grader: # the default grader; a case may override type: contains
cases: - id: c1 arguments: "free text passed as $ARGUMENTS" inputs: { key: value } # KEELSON_INPUTS_key expect: # the grader's reference; never shown to the workflow strings: ["PASS"] - id: c2 node: bash-json-node # grade one node's output instead of the final output grader: { type: json_schema } expect: schema: { type: object, required: [status] }| Field | Means |
|---|---|
name | Names the results directory, <home>/evals/<name>/. |
workflow | The workflow to run, resolved exactly as workflow run resolves it. |
project | A named project to run against. Requires the server. |
reps | How many times each case runs. --reps overrides it. |
split | Explicit train and test lists covering every case once, or seed plus train_fraction for a deterministic draw. |
grader | The default grader. Each case may carry its own grader: block. |
node | Which node’s output to grade. Default is the run’s final output: the last node that succeeded. A case or grader may override it. |
cases[].expect | The grader’s reference. It never reaches the workflow, and the judge grader sees only expect.claims. |
Every case must land in exactly one split. Overlapping lists, an unknown id, or a case left out of both lists is a load error.
Graders
Section titled “Graders”| Type | Reference | Passes when |
|---|---|---|
exact | expect.text | The graded output equals the text, both trimmed. |
contains | expect.strings | Every string appears in the output. |
regex | expect.pattern, optional expect.flags | The pattern matches the output. |
json_schema | expect.schema | The output parses as JSON (a fenced block is fine) and validates against the schema, the same type / required / properties / items subset output_schema accepts on a node. |
bash | expect.script | The script exits 0. Exit 1 is a fail; any other exit, or a timeout, is an error. |
judge | expect.claims | An agent turn reports every claim met. |
The bash grader runs with EVAL_OUTPUT, EVAL_OUTPUT_FILE (the full text on
disk), EVAL_CASE_ID, EVAL_RUN_ID, and one EVAL_EXPECT_<KEY> per scalar in
expect, so a script can compare against its own reference values without
parsing YAML.
The judge grader builds a rubric from expect.claims, asks the provider the
in-process chat uses for a JSON verdict per claim, and passes only when every
claim is met. It asks grader_reps times (default 2) about the same output and
takes the majority; a tie reads as fail, so with the default two reps any
disagreement fails the case. The disagreement is reported as grader noise
either way. A rep that errors (no JSON, a provider failure) makes the case an
error, never a verdict from the reps that survived. The judge turn runs with no
tools, because it reads text the workflow produced and must not be talked into
touching the checkout; only claude, copilot, and stub enforce that rail,
so the grader rejects any other provider:. Set provider: and model: on
the grader to pin the judge; leave the rubric to things a reader could check by
quoting the output.
Running
Section titled “Running”keelson eval run .keelson/evals/smoke-test.eval.yamlkeelson eval run cases.eval.yaml --reps 3 --split test --jsonRuns go through the server when it is up and in-process when it is down, one
case at a time, exactly as workflow run does. --watch streams node events
per case (the default on a TTY). --provider, --base-url, and
--no-preflight mean what they mean on workflow run.
Every case executes the entire workflow. node selects which output to grade,
not which nodes run. Use a replayable, side-effect-safe copy before evaluating
workflows that change a backlog, edit a checkout, create PRs, or publish reviews.
| Exit code | Means |
|---|---|
0 | The eval ran and every case graded. |
1 | At least one case errored. |
2 | Bad arguments or an invalid case file. |
3 | The file names a project and the server is down. |
4 | The workflow was not found. |
Reading the summary
Section titled “Reading the summary”Each run writes <home>/evals/<name>/<timestamp>.json, a Markdown summary
beside it, and a <timestamp>.outputs/ directory holding every case’s full
output (the JSON inlines the first 16 KiB). --out <path> relocates all three.
| Statistic | What it tells you |
|---|---|
| Pass rate per split, with its Wilson interval | The estimate and how much to trust it. A narrow interval means more cases or reps. |
| Errors | Runs that never produced a gradable output. Fix the infrastructure before reading the pass rate. |
| Grader noise | The share of judged outputs where the judge’s reps disagreed. A high rate means the rubric, not the workflow, needs work. |
| Duration mean and p95 | What a change cost in wall time. |
| Cost | Summed from the server’s usage ledger when it prices every event of a run, null otherwise. A zero is a real zero, never a placeholder for unpriced. |
| Definition hash | The workflow definition each run executed. Two hashes in one eval means the workflow changed mid-run; compare refuses that eval. |
The warnings line names what to fix next: HEADROOM when the test split passes
above 95%, NOISE when an interval is wider than 20 points, POWER when the
set has fewer than six cases, HASH when more than one definition was observed,
ERRORS when any case errored.
The compare verdict
Section titled “The compare verdict”keelson eval compare before.json after.jsonBoth files ran the same cases, so compare pairs them: each case’s pass rate
before against its pass rate after, with reps averaged inside the case. A
sign-flip permutation test on those per-case differences gives each split a
verdict: improved or regressed when p is below 0.05, within-noise
otherwise. Only cases that moved carry evidence, which is why six must move
before any verdict can leave within-noise, and why reps steady a case without
standing in for more cases.
It then applies the hillclimb rule. Keep when the pooled result improved and the train and test splits each moved up. Revert when the pooled result is within noise, when either split regressed, or when the pooled gain came from train while test stayed flat: that is overfitting. A set with no train split decides on test alone and says so, which is a reason to declare one. The cost delta prints alongside, with the mean duration per case run and the token totals, so an improvement that doubled the spend is visible in the same place even on a provider the price table does not cover.
overall 60.0% [38.7%, 78.1%] → 90.0% [69.9%, 97.2%] +30.0 pts improved (6/20 cases changed, p=0.031)Compare refuses two files that did not run the same cases, project, reps, and
split filter (the results file carries a fingerprint of the case set), and reverts
when either side errored: an infrastructure failure is not a measurement.
Multiple workflow definition hashes in either file also make the comparison
invalid, even when the pooled quality improved. The CLI exits 2 with
NOT_COMPARABLE, the same error as mismatched case sets. The check reads both
the summary and per-case hashes, so a stale summary cannot
hide mixed definitions. A different single definition on each side is expected
for a workflow edit; missing hashes in older files make no claim about the
executed definition.
For comparable files, the command exits 0 for both keep and revert decisions.
Automation must inspect decision, not treat exit 0 as permission to keep a
candidate. Cost and duration are descriptive metrics, not acceptance gates.
It warns when both runs executed the same workflow definition, since any difference
then came from outside the workflow or from run-to-run variation, and when the
two runs used a different --provider or, for in-process runs, a different
keelson version. Comparing a results
file with itself is always within-noise.
The hillclimb workflow runs this loop for you: one root-cause prompt change per round from the train split, the eval re-run, and the compare verdict deciding keep or revert.
Related
Section titled “Related”- Authoring workflows: the loop the eval closes.
- Workflow nodes: the
output_schemasubset thejson_schemagrader validates against. - CLI reference: every
evaloption and exit code.