Skip to content

Workflows

An agent in chat decides what happens next, every turn. That is the point of chat, and it is exactly wrong for the other half of real work: the release chore, the review pass, the data refresh you want to run the same way every time, reviewable before it runs and explainable after. A keelson workflow takes the control flow away from the agent and gives it to a YAML file. The file fixes what runs, in what order, under what conditions. Inside that fixed frame, individual nodes can still do agent work, with the leash declared in the same file.

What that file buys you, and what it costs, plays out below: how the graph runs, how data crosses between nodes, what a run leaves behind, and where the agency is bounded.

A workflow is a set of nodes and depends_on edges, a directed acyclic graph. The executor starts every node with no unmet dependencies at once, and starts each remaining node the moment its own dependencies settle, so a slow branch never holds up an independent one. A workflow can set scheduling: layered to wait for whole topological layers instead. The first-workflow tutorial walks this shape end to end, with two independent nodes starting together.

Whether a node runs at all is decided by two declarations on the node, both evaluated against upstream results:

  • when: is a data gate: a condition over upstream output, like $check.output.status == 'ok', with ==, !=, and numeric comparisons, joined by && / ||. A false gate skips the node.
  • trigger_rule: is a dependency-merge rule: what pattern of upstream outcomes lets the node fire.
trigger_ruleThe node runs when
all_success (default)Every dependency succeeded.
one_successAt least one dependency succeeded, so a branch can rescue a failing sibling.
none_failed_min_one_successNothing failed and at least one succeeded.
all_doneEvery dependency reached a terminal state, for collector nodes that summarize whatever happened.

The details fail closed. A depends_on reference the validator somehow never saw counts as failed, a malformed when: expression warns and skips its node, and a node that declares an output_schema fails rather than passing a mismatched payload downstream.

Seven node kinds, three jobs:

FamilyKindsCharacter
Agentprompt, command, loopRun agent turns. Take model, provider, and tool gates. command runs a named, reusable prompt file from the home; loop repeats a prompt until a completion token appears, hard-bounded by max_iterations.
Deterministicbash, scriptRun code with no model in the path: bash for shell, script for a declared runtime such as Bun. Timeouts and exit codes decide success.
Controlapproval, cancelShape the run itself. approval pauses for a human decision; cancel ends a branch deliberately, recording why.

The taxonomy is the determinism dial. A workflow of bash and script nodes is a build system. A workflow of prompt nodes with gates between them is a supervised agent. Most useful workflows sit in the middle: deterministic nodes establish facts, agent nodes do judgment work, and gates check the judgment before anything irreversible happens.

Upstream output reaches a downstream node through one of two channels, and the split is deliberate.

The workflow layer substitutes text. In prompts, when: conditions, and cancel reasons, $nodeId.output expands to the upstream node’s output, and $nodeId.output.field reaches into JSON output. In prompt and cancel bodies, workflow inputs arrive the same way, as $inputs.<name> and $ARGUMENTS. A when: condition resolves only $nodeId.output refs, and the loader rejects $inputs.*, $ARGUMENTS, or $ARTIFACTS_DIR inside one outright: encode the input through a producer node and compare its output. Substitution is a single atomic pass: if an upstream output happens to contain $something.output, the marker stays literal text instead of expanding again.

Shell nodes read the environment instead. A bash or script node runs its body exactly as written in the file, with upstream output delivered as KEELSON_NODE_<id>_OUTPUT environment variables. Nothing produced by a model is ever spliced into shell source, so upstream text containing $(...) or backticks is data, never code.

Starting a workflow creates a run, and the run is durable: a row for the run, a row per node with its output, timing, error, and the provider and model that node actually ran on, all in the harness’s SQLite store. That last pair makes an otherwise-unobservable auto node’s resolved served model visible in the run detail whenever the provider reports it. The run also records a definitionHash, the sha256 of the parsed workflow definition it executed, so an evaluation score or a cost figure can be pinned to one exact version of that definition. Edit the YAML in any way that changes a node, even whitespace inside an inline prompt, and the next run carries a different hash; a resumed run is re-stamped with whatever definition the resume ran. Content a node reads at execution time, such as the Markdown file behind a command node, is outside the hash. From those per-node timestamps, keelson workflow status and the run header in the browser derive a critical-path figure: the longest dependency chain by node duration, shown beside the wall clock, so the gap between the two reads as scheduling wait rather than node work. The watch stream you see in the CLI and the browser is the same event feed, emitted as each node starts, streams, and settles. The figure shows every state a run can occupy:

Workflow run states. From start, a run is running; an approval node moves it to paused and a resume moves it back. Running ends at succeeded, failed, or cancelled. A dashed edge from paused to failed is labeled: server restart, boot reconcile marks it failed.

Figure 1. The run lifecycle. Approval nodes are the only path into paused, and the dashed edge is the honest one: a pause does not survive a server restart.

The dashed edge is the honest one. An approval pause is held by the running server process, not by the database. If the server restarts while a run is paused or mid-flight, boot reconciliation finds the stale run and marks it failed rather than pretending it can continue. Runs are durable as records; failed and cancelled runs are resumable: the harness re-enters the executor from the first incomplete node, seeding completed node outputs so finished work is not repeated. A node marked always_run is the exception, it re-runs on resume even though it succeeded, so a gate or validation re-checks rather than replaying a stale pass.

Every agent node runs inside limits that are declared, not improvised, and they layer from the file up to the operator:

  • The file picks the provider and model per node, gates tools with allowed_tools and denied_tools, bounds loops with max_iterations, bounds stalls with idle_timeout, and pins output shape with output_schema.
  • The operator floor sits above the file. KEELSON_WORKFLOW_PROVIDER pins which provider every prompt node uses, and KEELSON_WORKFLOW_TOOL_DENYLIST subtracts tools no workflow may use, applied on top of whatever a node allows.
  • The human gate is a node. An approval node holds the run before the irreversible step, and the reply becomes that node’s output, so downstream nodes can branch on what the human said.

Provider support for the finer controls varies, and the loader is honest about it at load time: fields only one provider honors, like per-node hooks, are flagged as warnings rather than silently accepted or rejected.

The catalog draws from three layers, lowest precedence first: the bundled starters that ship with keelson (read-only), your global <home>/workflows/, and a project’s .keelson/workflows/. A later layer shadows an earlier one by name, and installed ribs contribute their own at activation, either built in code through the rib contract or shipped as plain YAML files in the rib package’s workflows/ folder. On first run the bundled starters are also copied into your home, so they list as global once seeded, while the bundled layer keeps them discoverable in a fresh checkout. The Workflows catalog documents the ones that ship. A rib-contributed workflow can additionally bind its run output to one of the rib’s snapshot keys, which is how a browser board’s refresh button becomes a workflow run; that pipeline is the subject of Snapshots and surfaces.