Starter workflows
The bundled catalog contains twelve workflows. They are real, runnable, and the authoring recipes show up composed: a three-line approval-gate recipe is one thing, the same gate sitting inside a thirty-node investigate-and-implement pipeline is where the shape clicks. This catalog groups what ships by the job it does, and each one gets a full node-by-node walkthrough.
Each is also readable in a single screen, so opening the file is its own documentation.
Where these come from
Section titled “Where these come from”All twelve are bundled starters: they ship with the package and seed into
your home (<home>/workflows/) on first run, so keelson workflow list shows
them as global once seeded. Copy and edit one freely; the bundled original
stays as a clean fallback. See scopes for
how the layers merge.
Run any of them the same way, streaming the DAG as it executes:
keelson workflow run smoke-test --watchkeelson workflow run fix-issue --inputs ARGUMENTS="42"Ship code
Section titled “Ship code”The delivery flagships: an issue taken to a draft PR, a scoped PR review, a plan-first lifecycle, and a PR convergence loop that drives a PR to a clean finish. Each composes gates, loops, and branches into a real pipeline, and each has a full walkthrough.
User wants to fix, resolve, or implement an issue end-to-end and open a draft PR.
Fetches the issue, classifies it, investigates / plans, pauses for human approval of the plan, implements the change, validates against the project's OWN checks (discovered from its docs/config — works in any repo), self-fixes failures, pushes the branch, creates a draft PR, runs a multi-lens review loop (independent correctness / conventions / coverage reviewers → adversarial triage → fix → re-review), and reports back.
"fix issue", "fix #123", "resolve issue", "implement issue 123", "fix this bug", "fix github issue".
PR review (use pr-review). Brainstorming or PRD-style exploration.
User wants a smart, scoped code review of an existing PR or MR that adapts to the PR's complexity and posts the synthesis as line-anchored inline review comments as a single batched review.
Fetches PR scope -> classifies complexity and which review lanes are relevant -> runs only those lanes (code review always, plus error handling / test coverage / docs impact when applicable) -> runs an opus triage gate to de-dupe and filter noise -> posts findings as a batched PR/MR review with line-anchored inline comments and suggestion blocks.
"review pr", "review #123", "smart pr review", "review this pr", "code review pr".
Fixing an issue (use fix-issue). Mutating the PR branch.
User wants to build a feature or land a change in ANY repo under the Plan→Act→Evaluate lifecycle — plan recorded as an issue, a human gate before any code, then a choice to implement locally OR delegate to the GitHub Copilot coding agent (GitHub only), with the PR/MR's CI as the authoritative check.
Normalizes the request (issue # / free text / plan file), drafts a plan under a read-only capability boundary, records it as an issue, pauses for human approval, then either implements locally (AI review loop on the project's own checks → draft PR) or assigns the issue to the Copilot coding agent and reviews its PR. Reads back remote CI, and reports the audit trail. Assumes nothing about the project's stack.
"plan act evaluate", "pae loop", "plan first", "lifecycle loop", "plan-act-evaluate", "build with a plan issue", "delegate to copilot".
Quick single-file or known-bug fixes (use fix-issue). PR review (use pr-review). Architectural sweeps (use architect). PRD / brainstorming.
User wants an existing same-repo PR driven to a mergeable state — review threads handled, fixes pushed, CI green.
Fetches unresolved review threads each round via GraphQL -> skips threads already answered (by this run, a prior run, or a human — any thread whose last comment is not the reviewer's) -> triages new threads (fix / reply with evidence, flagging rebuttals whose cited evidence this PR itself changes) -> pauses for approval ONLY when a round would publicly push back on a reviewer (wontfix / invalid / question) -> fixes actionable threads -> reviews its own fixes before anyone else does -> validates and pushes -> replies on every handled thread -> resolves fixed threads plus approved invalid / already-addressed rebuttals (question stays open; wontfix stays open unless explicitly enabled with `resolve_wontfix=true` and approved) -> waits for CI -> repeats until CI passes and a clean round observes no new threads, or escalates at the round cap.
"resolve PR #123", "finish pr", "get this PR green", "handle the review comments", "drive PR to merge".
Fork PRs, which cannot be pushed safely by this workflow. Creating a PR from an issue (use fix-issue). Reviewing a PR without changing it (use pr-review).
Plan and shape
Section titled “Plan and shape”Exploration workflows: turn a fuzzy idea into a validated requirements doc, or stress-test any artifact (a claim, design, or plan) to a consensus verdict.
User wants to create a product requirements doc through guided, problem-first conversation.
Asks foundation questions -> researches market + codebase -> asks deep-dive questions -> assesses technical feasibility from what already exists -> asks scope questions -> writes a PRD -> validates every technical claim against the real codebase.
"create a prd", "new prd", "interactive prd", "plan a feature", "product requirements", "write a prd".
Autonomous PRD generation with no human input, or implementing a plan.
User needs competing design proposals compared, fact-checked, and converged into one decision-ready draft.
Bundles a brief and optional criteria/evidence, runs three isolated proposal seats, blind cross-critique, an independent fact-check, and one synthesis with explicit dissent. Writes the panel's eight artifacts and pauses for a person to accept the draft or send it back for another round. Inputs: `brief` is a required non-empty file; `out` is the required artifact directory. Optional `criteria` is a file, `evidence` is a directory, `sources` is a comma- or newline-delimited list of local checkout paths, and `round` is a positive integer (default 1).
"converge on a design", "compare design options", "run a design panel", "challenge these approaches", "produce a consensus draft".
Implementing the draft or validating a finished artifact. After a design is accepted, use adversarial-review to stress-test it.
User wants an artifact stress-tested by independent agents that debate it and reach a consensus verdict — a claim, design doc, root-cause analysis, plan, spec, or diff.
Captures the artifact (inline text, a file path, or a directory of evidence files bundled verbatim), runs three adversarial reviewers in parallel (logic, evidence, risk lenses), independently re-verifies their load-bearing claims against source, then synthesizes one verdict — CONFIRMED / CONFIRMED-WITH-CHANGES / REFUTED — with the surviving claims, the dissent, and what is still unverified. Inputs: optional `subject` — a git ref of the code the artifact concerns; intake snapshots it read-only so the verifier checks the exact tree under review. For a PR, use `refs/pull/N/head`; a bare SHA resolves only when its object is already local or advertised by the remote. Set `subject_required=1` to fail instead of falling back when a provided subject cannot be snapshotted. Optional `author` names the model that wrote the artifact so it is excluded from every review, verification, and synthesis seat. Optional `out` names a directory where the five review artifacts are written.
"adversarial review", "pressure-test this", "debate this", "poke holes in", "review this and reach consensus", "have agents argue about".
Building or fixing code (use fix-issue), posting a GitHub PR review (use pr-review), or generating something new.
User needs a focused factual question answered from source and read-only runtime evidence, with an independent verification pass.
Gathers cited evidence, asks a second model vendor to re-check every load-bearing claim, updates exact verdicts, then deterministically writes the tally and effective model attribution to the requested output file. Inputs: `ARGUMENTS` is the required question; `out` is the required evidence file; optional `context` is a file or one-level directory; optional `access` is guidance text or a readable file; optional `fixtures` is a directory for saved raw probe responses; `tier` is deep|std; `verifier` is claude|grok.
"investigate this", "establish the facts", "verify this behavior", "research this claim", "produce an evidence file".
Changing source, enforcing a read-only sandbox, or publishing results to an external system.
Author and tune workflows
Section titled “Author and tune workflows”Write a workflow from a description, or improve an existing one’s prompts against an eval case set by measurement.
User wants to scaffold a new custom workflow for their project from a plain-language description.
Scans the repo for context -> extracts structured intent (JSON) -> generates a Keelson workflow YAML -> structurally validates it (Bun YAML parse + required fields) -> installs it under .keelson/workflows/ and reports how to validate and run it.
"build me a workflow", "create a workflow", "generate a workflow", "new workflow", "make a workflow for", "workflow builder".
Editing an existing workflow, or generating non-workflow files.
A workflow has an eval case set and you want its prompts improved by measurement instead of by hand: each round proposes one root-cause change, re-runs the eval, and keeps the change only when the paired compare says keep.
Baselines the case set, then up to three rounds of propose (one root-cause prompt change, train split only) -> validate -> evaluate -> keep or revert by `keelson eval compare`, committing kept changes on a `keelson/hillclimb/<name>-<stamp>` branch. Stops after two flat rounds, buckets the remaining failures by root cause, and reports before/after per split with a merge recommendation.
"hillclimb <workflow>", "improve the prompts against the eval", "tune this workflow's prompts", "run the eval loop".
writing the case set (keelson eval init), grading one run (keelson eval run), changing a workflow's structure, or chasing noise: a NOISE warning calls for more reps or cases, not more rounds.
Produce artifacts
Section titled “Produce artifacts”Author and publish a designed, self-contained HTML page to the canvas, a report or briefing or chart, validated fail-closed on the way out.
Verify the engine
Section titled “Verify the engine”Related
Section titled “Related”- Authoring workflows: the recipes these compose, and the loop you follow to write one.
- Workflow nodes: the full schema every node here is built from.
- Workflows: why deterministic YAML sits next to an agent, and the run lifecycle a workflow leaves behind.