Skip to content

Evaluating workflows

A workflow that runs is not a workflow that works. Once a prompt, a model pin, or a node boundary changes, the only honest answer to “did that help?” is a set of cases you run before and after, graded the same way both times. keelson eval is that loop: a YAML case set, a grader per case, a results file with the statistics you need to tell signal from noise, and a compare command that says keep or revert.

Four properties make an eval worth trusting. The summary warns when any of them slips.

  • Production-aligned cases. Each case is an input the workflow really receives: the free text that arrives as $ARGUMENTS, the named inputs that arrive as KEELSON_INPUTS_<key>. Synthetic prompts measure a workflow nobody runs.
  • Headroom. A set the workflow already passes at 100% cannot show improvement. The summary prints a HEADROOM warning when the test split passes above 95%: add harder cases.
  • Low variance. A pass rate is an estimate. The summary carries a 95% Wilson interval for every split and prints a NOISE warning when that interval is wider than 20 points. More reps per case, or more cases, narrow it.
  • Enough cases. Compare counts cases, not reps: at least six cases have to change before any difference can be called, so a set smaller than that prints a POWER warning. Twenty to thirty cases is a working size.
  • A held-out split. Tuning a prompt against the cases you stare at makes those cases pass. The train/test split exists so that overfitting shows up as a gap: train improves, test does not, and the verdict says revert.

Graders are either programmatic or an LLM judge with checkable claims. There is no 1-to-5 scale anywhere: a claim is met or it is not, and the judge grades the same output repeatedly (grader_reps, default 2) so its own disagreement is measured and reported as grader noise rather than hidden in the pass rate.

Scaffold one with keelson eval init <workflow>, which writes <home>/evals/<workflow>.eval.yaml and refuses to overwrite an existing file. The repository ships a working example at .keelson/evals/smoke-test.eval.yaml.

name: beads-groom-smoke
workflow: beads-groom # resolved like `workflow run`
project: my-project # optional, as `workflow run --project`
reps: 3 # default 1
split: # optional; default puts every case in test
train: [c1, c2]
test: [c3, c4]
# or a seeded draw:
# seed: 42
# train_fraction: 0.5
grader: # the default grader; a case may override
type: contains
cases:
- id: c1
arguments: "free text passed as $ARGUMENTS"
inputs: { key: value } # KEELSON_INPUTS_key
expect: # the grader's reference; never shown to the workflow
strings: ["PASS"]
- id: c2
node: bash-json-node # grade one node's output instead of the final output
grader: { type: json_schema }
expect:
schema: { type: object, required: [status] }
FieldMeans
nameNames the results directory, <home>/evals/<name>/.
workflowThe workflow to run, resolved exactly as workflow run resolves it.
projectA named project to run against. Requires the server.
repsHow many times each case runs. --reps overrides it.
splitExplicit train and test lists covering every case once, or seed plus train_fraction for a deterministic draw.
graderThe default grader. Each case may carry its own grader: block.
nodeWhich node’s output to grade. Default is the run’s final output: the last node that succeeded. A case or grader may override it.
cases[].expectThe grader’s reference. It never reaches the workflow, and the judge grader sees only expect.claims.

Every case must land in exactly one split. Overlapping lists, an unknown id, or a case left out of both lists is a load error.

TypeReferencePasses when
exactexpect.textThe graded output equals the text, both trimmed.
containsexpect.stringsEvery string appears in the output.
regexexpect.pattern, optional expect.flagsThe pattern matches the output.
json_schemaexpect.schemaThe output parses as JSON (a fenced block is fine) and validates against the schema, the same type / required / properties / items subset output_schema accepts on a node.
bashexpect.scriptThe script exits 0. Exit 1 is a fail; any other exit, or a timeout, is an error.
judgeexpect.claimsAn agent turn reports every claim met.

The bash grader runs with EVAL_OUTPUT, EVAL_OUTPUT_FILE (the full text on disk), EVAL_CASE_ID, EVAL_RUN_ID, and one EVAL_EXPECT_<KEY> per scalar in expect, so a script can compare against its own reference values without parsing YAML.

The judge grader builds a rubric from expect.claims, asks the provider the in-process chat uses for a JSON verdict per claim, and passes only when every claim is met. It asks grader_reps times (default 2) about the same output and takes the majority; a tie reads as fail, so with the default two reps any disagreement fails the case. The disagreement is reported as grader noise either way. A rep that errors (no JSON, a provider failure) makes the case an error, never a verdict from the reps that survived. The judge turn runs with no tools, because it reads text the workflow produced and must not be talked into touching the checkout; only claude, copilot, and stub enforce that rail, so the grader rejects any other provider:. Set provider: and model: on the grader to pin the judge; leave the rubric to things a reader could check by quoting the output.

Terminal window
keelson eval run .keelson/evals/smoke-test.eval.yaml
keelson eval run cases.eval.yaml --reps 3 --split test --json

Runs go through the server when it is up and in-process when it is down, one case at a time, exactly as workflow run does. --watch streams node events per case (the default on a TTY). --provider, --base-url, and --no-preflight mean what they mean on workflow run.

Every case executes the entire workflow. node selects which output to grade, not which nodes run. Use a replayable, side-effect-safe copy before evaluating workflows that change a backlog, edit a checkout, create PRs, or publish reviews.

Exit codeMeans
0The eval ran and every case graded.
1At least one case errored.
2Bad arguments or an invalid case file.
3The file names a project and the server is down.
4The workflow was not found.

Each run writes <home>/evals/<name>/<timestamp>.json, a Markdown summary beside it, and a <timestamp>.outputs/ directory holding every case’s full output (the JSON inlines the first 16 KiB). --out <path> relocates all three.

StatisticWhat it tells you
Pass rate per split, with its Wilson intervalThe estimate and how much to trust it. A narrow interval means more cases or reps.
ErrorsRuns that never produced a gradable output. Fix the infrastructure before reading the pass rate.
Grader noiseThe share of judged outputs where the judge’s reps disagreed. A high rate means the rubric, not the workflow, needs work.
Duration mean and p95What a change cost in wall time.
CostSummed from the server’s usage ledger when it prices every event of a run, null otherwise. A zero is a real zero, never a placeholder for unpriced.
Definition hashThe workflow definition each run executed. Two hashes in one eval means the workflow changed mid-run; compare refuses that eval.

The warnings line names what to fix next: HEADROOM when the test split passes above 95%, NOISE when an interval is wider than 20 points, POWER when the set has fewer than six cases, HASH when more than one definition was observed, ERRORS when any case errored.

Terminal window
keelson eval compare before.json after.json

Both files ran the same cases, so compare pairs them: each case’s pass rate before against its pass rate after, with reps averaged inside the case. A sign-flip permutation test on those per-case differences gives each split a verdict: improved or regressed when p is below 0.05, within-noise otherwise. Only cases that moved carry evidence, which is why six must move before any verdict can leave within-noise, and why reps steady a case without standing in for more cases.

It then applies the hillclimb rule. Keep when the pooled result improved and the train and test splits each moved up. Revert when the pooled result is within noise, when either split regressed, or when the pooled gain came from train while test stayed flat: that is overfitting. A set with no train split decides on test alone and says so, which is a reason to declare one. The cost delta prints alongside, with the mean duration per case run and the token totals, so an improvement that doubled the spend is visible in the same place even on a provider the price table does not cover.

overall 60.0% [38.7%, 78.1%] → 90.0% [69.9%, 97.2%] +30.0 pts improved (6/20 cases changed, p=0.031)

Compare refuses two files that did not run the same cases, project, reps, and split filter (the results file carries a fingerprint of the case set), and reverts when either side errored: an infrastructure failure is not a measurement. Multiple workflow definition hashes in either file also make the comparison invalid, even when the pooled quality improved. The CLI exits 2 with NOT_COMPARABLE, the same error as mismatched case sets. The check reads both the summary and per-case hashes, so a stale summary cannot hide mixed definitions. A different single definition on each side is expected for a workflow edit; missing hashes in older files make no claim about the executed definition.

For comparable files, the command exits 0 for both keep and revert decisions. Automation must inspect decision, not treat exit 0 as permission to keep a candidate. Cost and duration are descriptive metrics, not acceptance gates.

It warns when both runs executed the same workflow definition, since any difference then came from outside the workflow or from run-to-run variation, and when the two runs used a different --provider or, for in-process runs, a different keelson version. Comparing a results file with itself is always within-noise.

The hillclimb workflow runs this loop for you: one root-cause prompt change per round from the train split, the eval re-run, and the compare verdict deciding keep or revert.