Skip to content

hillclimb

hillclimb closes the loop that evaluating workflows opens. You hand it an eval case set; it baselines the target workflow, then runs up to three rounds in which a proposer reads the train split’s failures and makes one root-cause change to the target YAML, a deterministic round node validates the edit, re-runs the eval, and asks keelson eval compare whether to keep it. Keep means a commit on a keelson/hillclimb/<workflow>-<stamp> branch; revert means the file is restored byte for byte. After two flat rounds it stops, a bucket node groups what still fails by root cause, and a report says whether to merge.

It optimizes existing node prompt, loop.prompt, and description text only. A structural check rejects changes to every other YAML field, including unknown fields. Node ids, order, edges, gates, bash bodies, models, and tools stay as they are. Adding or removing a node or an editable field is also out of scope. The proposer is instructed to make one logical wording change; the structural check enforces the editing surface, not the number of ideas in the rewrite.

Terminal window
# The case file is the only required argument:
keelson workflow run hillclimb --arguments .keelson/evals/triage.eval.yaml --watch
# One round, three reps per case, against an explicit copy of the workflow:
keelson workflow run hillclimb --arguments .keelson/evals/triage.eval.yaml \
--inputs rounds=1 --inputs reps=3 --inputs target=.keelson/workflows/triage.yaml
InputDefaultMeans
$ARGUMENTSrequiredPath to the eval case file (keelson eval init <workflow> writes one).
rounds3How many propose/evaluate rounds to allow, 1 to 3.
repsthe case file’s repsRepetitions per case for every eval the run makes.
targetresolved from the case fileThe workflow YAML to edit.

The target must be an editable copy: the .keelson/workflows/<name>.yaml of the project you run from, or the copy in your keelson home. Bundled assets are read-only, so a workflow that exists only as a bundled starter fails preflight with the copy command to run first. The eval resolves the same name through the same layers, and both the server catalog and the in-process runner re-read the file on every run, so each round evaluates the edit it just made without a restart.

The run needs keelson on PATH (it falls back to bun apps/cli/bin/keelson.ts inside a keelson checkout), plus git, jq, and bun. When the target is tracked by git it must be clean, and the kept changes land as one commit per round on a fresh branch; a target outside a repository is edited in place and restored from a snapshot on revert. Commits include only the target file, preserving any unrelated staged changes. If a commit fails, the round restores the last kept target, leaves the kept results and train view unchanged, and stops. Git hooks still run, so a round also stops when the committed file or the working copy no longer matches the bytes the eval ran: the commit is undone and the last kept target restored. An uncommitted round in a git repository never earns a merge recommendation.

Every other file in the repository is fingerprinted by content at preflight, uncommitted and untracked work included. A round that finds any of them changed, even a file that was already dirty, is a scope violation.

A straight line with three gated detours. preflight and baseline set the starting point; each propose-N runs only while the previous round said continue, and each round-N always settles so the chain never stalls on a failed proposer; collect joins the three rounds under all_done, bucket runs only when something still fails, and report writes the verdict.

NodeKindWhat it does
preflightbashParses the case file, resolves and validates the target, refuses a dirty target, creates the branch, snapshots the file and the rest of the repository, and writes the eval, scope-check, snapshot, and train-view helpers.
baselinebashRuns the case set, fails on any errored case or mixed workflow definitions, copies the results beside the run, and builds the train-only view.
propose-1..3prompt (deep)Reads the train view and makes one root-cause change to the target. Emits {changed, root_cause, change_summary}.
round-1..3bashValidates the YAML, checks protected fields against the last kept target, rejects out-of-scope or comment-only edits before evaluating, then evaluates and compares. Advances kept state only after a commit that holds the evaluated bytes when git is used.
collectbashAssembles the ledger: before/after table per split, the rounds and their reasons, cost delta, remaining failures. Copies the ledger and each round’s change and decision to hillclimb-<run id>/ beside the eval results, so they outlive the run. Returns to the original branch when nothing was kept.
bucketprompt (deep)Groups the remaining failures by root cause, flags ambiguous cases and grader contradictions, and recommends more reps or cases where the interval is wide.
reportprompt (balanced)Formats the ledger and the buckets, and ends with the one-line merge recommendation collect computed.

Train-only evidence is materialized. baseline and every kept round build train-view/ from the results file: the train rows, their output files, and the train cases’ inputs and expectations, nothing else. The prompt names that directory as the only evidence it may read and forbids opening the case file or the home’s evals/ directory. This is a prompt-level read restriction, not a filesystem sandbox. The test split participates in every keep/revert decision, so it is a reusable validation set rather than an untouched final holdout. Confirm the selected candidate on fresh acceptance cases before claiming generalization.

One change, never a quote. The proposer is told to find the one root cause shared by the most failures and fix it with a general rule. Pasting an output, an expected string, or a case id into the prompt is forbidden outright: a prompt that quotes its own eval passes that eval and measures nothing. The round node stores the diff in round-N-change.md so the operator can check the rule reads as a rule. It compares parsed YAML against the last kept target with only the existing editable text masked. A protected-field change is saved as round-N-rejected.yaml, restored, and stops the loop without an eval. An edit that changes only YAML formatting or comments is restored and stops as no-change.

Keep is the compare decision, not the agent’s opinion. The round node is deterministic bash. It runs keelson eval compare <kept> <round> and keeps only on keep, which the eval runner grants only when the paired per-case test calls the pooled result improved and train and test each moved up. Within-noise, a regression on either split, a flat test split, or any errored case run all restore the file. So does a round whose eval ran the same definition as the kept results: the edit never reached the copy the catalog runs, usually because another copy of the workflow shadows the target. Multiple definition hashes within the baseline stop the run before proposing; within a candidate they restore the target and stop the loop. A measurement that mixes definitions cannot justify keeping a change.

Flat rounds end the loop. Two consecutive non-keep rounds, a proposer that declines to change anything, a scope violation, or an eval error sets continue to false, and the next proposer skips. Rounds are unrolled rather than a loop node because a loop is one agent prompt and the keep/revert step must run outside the agent.

Nothing is held. The workflow declares mutates_checkout: false even though it edits a file, because a held project lock would refuse every nested keelson eval run the rounds depend on. Do not run two hillclimbs on the same project at once.

  • Classify, then branch: each propose-N gates on $round-(N-1).output.continue.
  • Feed agent output into a script safely: the round reads the proposer’s JSON from KEELSON_NODE_propose_N_OUTPUT and the decision never leaves bash.
  • Fail-closed structured output: every node a gate reads from carries an output_schema, the prompt nodes an output_format too, so a malformed verdict fails the node instead of steering the loop.
  • More rounds. Copy the file and add a propose-4 / round-4 pair; the round body is a YAML anchor, so the new pair is six lines.
  • A different surface. The proposer’s “only file you may edit” and “edit prompt text only” rules and the deterministic scope check must agree. Changing a model, a command file, or a skill needs a separate experiment with an explicit editing boundary; changing the proposer instructions alone does not allow those edits.
  • Reps per round. Pass reps higher than the case file’s default when cases flip between runs of an unchanged workflow. Reps steady each case; they do not replace cases, and the compare needs six cases to move before it can say improved.