Skip to content

fix-issue

fix-issue takes a GitHub issue from a number to a draft pull request: it fetches and classifies the issue, plans, pauses for a human to approve the plan, implements, validates against the repository’s own checks, self-fixes failures, opens a draft PR, then runs a multi-lens review loop where independent reviewers for correctness, conventions, and coverage feed a triage judge that adversarially verifies each candidate against the code and keeps only confirmed defects, which a fixer applies and a re-review confirms. Finally it watches the pushed PR’s CI and promotes the draft to ready for review the moment the checks go green. It is the catalog’s most complete pipeline, and the clearest example of putting a human gate in front of irreversible work.

It assumes nothing about your stack. The implement step discovers how this project checks itself by reading its docs and build config, writes those commands to a script, and every later validation runs that script. So the same workflow works in a Bun repo, a Cargo repo, or a Go repo without edits. For reviewing a PR rather than creating one, use pr-review.

Terminal window
keelson workflow run fix-issue --inputs ARGUMENTS="42"
keelson workflow run fix-issue --inputs ARGUMENTS="fix the SQLite timestamp bug"

It needs gh or glab authenticated (the bundled forge shim resolves whichever is present), a Copilot subscription, and a running server (keelson start) because the approval gate pauses the run. It mutates the working tree (branch, commits, push, PR), so it sets worktree.enabled: every run executes in an isolated git worktree and cannot trample your editor state. Worktree setup is a launch gate. If keelson cannot establish and record the checkout, the run fails before extract-issue instead of continuing in the source directory. --no-worktree is a deliberate override that makes the workflow mutate the supplied directory in place.

Ten phases, top to bottom. The spine is mostly linear, with three structures worth naming: a branch early (a bug is investigated, anything else is planned, both writing one file), a self-correction sequence after implementation (validate, fix, then a hard gate that refuses to open a PR over broken checks), and a CI gate at the end that promotes the draft to ready only when the pushed PR’s checks pass.

The fix-issue pipeline, top to bottom, enclosed in a dashed frame labeled isolated git worktree. Stages: fetch and classify; a branch where investigate (when bug) and plan (when not bug) both feed plan-ready; an approval gate shown in brass; implement; a validate, fix, revalidate self-correction group ending in a hard gate; create draft PR; the review loop; post-fix-validate; a CI gate and mark-ready phase (await-ci, triage-ci, fix-ci, finalize-pr); and report.

Figure 1. The fix-issue pipeline. The whole run executes in an isolated worktree (dashed frame). The brass node is the human approval gate; nothing downstream of it runs until the plan is approved. The hard gate after self-correction fails the run rather than open a PR over broken checks, and the closing CI gate marks the PR ready only when the pushed checks pass.

PhaseNodeKindWhat it does
Fetch & classifyextract-issuepromptParses the issue number from $ARGUMENTS, or searches forge issue list for the best match.
fetch-issuebashforge issue view to JSON, stashes the number in the scratch dir.
detect-basebashResolves the repository’s default branch with forge repo view --json defaultBranchRef, falling back to origin/HEAD then main, and writes it to $ARTIFACTS_DIR/.default-branch. Never assumes main.
extract-briefbashExtracts bullets under acceptance criteria, requirements, definition of done, success criteria, “What did you expect?”, or expected-behavior headings into $ARTIFACTS_DIR/brief.json. Heading suffixes such as (must all hold) are accepted.
classifypromptTyped JSON: issue_type (bug, feature, …), a title, and reasoning.
Planinvestigatepromptwhen: issue_type == 'bug', model: claude-opus-5.5, effort: medium: root-cause analysis into plan.md.
planpromptwhen: issue_type != 'bug', model: claude-opus-5.5, effort: medium: an implementation plan into the same plan.md.
plan-readybashFunnel (trigger_rule: one_success): asserts plan.md exists and has tasks, then prints it.
criteria-count-initialbashCounts criteria found by the structured heading scan.
extract-brief-llmpromptwhen: criteria-count-initial == 0, allowed_tools: []: extracts concrete, testable outcomes from prose-only issue bodies into typed JSON. A structured brief skips this call.
brief-readybashFunnel (trigger_rule: all_done): merges fallback criteria into brief.json. It still runs when the fallback is skipped, so structured criteria continue through the DAG.
criteria-countbashRecounts the merged brief so the coverage gate sees either extraction path.
coverage-checkpromptwhen: criteria-count > 0, model: claude-opus-4.8, effort: xhigh: semantic judge that maps each brief criterion to the plan step that covers it or MISSING. Returns typed JSON only; skipped when both extraction paths find no criteria.
coverage-readybashFunnel (trigger_rule: all_done): writes the judge’s output to coverage.json when criteria exist. Otherwise it removes the artifact and reports the skip in the run log.
Approveapprove-planapprovalPauses the run. capture_response makes the reply, approval or change requests, available downstream. The callout includes the criteria-coverage checklist, or a visible note that the divergence check did not run.
Implementimplementpromptmodel: gpt-6.1-sol, effort: xhigh: Writes verify.sh (the project’s own checks), branches, implements task by task, commits each.
ValidatevalidatebashRuns verify.sh fail-fast, saves the full transcript to $ARTIFACTS_DIR/validate.log, and prints a bounded summary. Exits 0 and reports VALIDATION_STATUS so the fixer can read it.
fix-validationpromptmodel: gpt-6.1-sol, effort: high: if it failed, fixes only what broke and commits; if it passed, stops.
revalidatebashHard gate: re-runs the checks, saves the full transcript to $ARTIFACTS_DIR/revalidate.log, and prints a bounded summary. Exits non-zero on failure, skipping the PR.
Draft PRcreate-prpromptPushes the branch, opens a draft PR with Fixes #N, and captures the PR number and URL. It follows a single standalone pull_request_template.md when one exists; repositories with only named templates under PULL_REQUEST_TEMPLATE/ use the fallback body. Its test plan records commands actually executed and their results, with pending CI reserved for checks that only CI can confirm.
enforce-draftbashVerifies the new PR is a draft; if it came up ready, runs forge pr ready --undo to force it back to draft, and fails the run if it cannot confirm the draft state.
Review loopcapture-diffbashFreezes the branch diff to $ARTIFACTS_DIR/diff.patch / diff-stat.txt so the lenses can run without shell access.
review-correctnesspromptmodel: claude-opus-5.5, effort: medium, allowed_tools: [Read, Glob, Grep]: one of three independent, parallel reviewers over diff.patch. Hunts by cost of failure (trust boundaries, irreversible state, idempotency gaps), follows each changed function to every entry path that reaches it (resume, cancel, the opt-out configuration, data written before the change), checks that docs and reference tables still agree with the code, labels unverified external-API claims INFERRED, and emits a confidence and a repro (the command, test, or input that shows the failure, or none plus the reason) per finding.
review-conventionspromptmodel: gpt-6-luna, effort: xhigh, same read-only rail: checks the diff against CLAUDE.md / CONTRIBUTING.md (comment policy, conventions), quoting the rule each finding violates.
review-coveragepromptmodel: gpt-6-luna, effort: xhigh, same read-only rail: flags new behavior left untested, and tests authored but never wired in.
triagepromptmodel: claude-opus-4.8, effort: xhigh, allowed_tools: [Read, Glob, Grep], typed JSON: de-dupes the findings, then adversarially refutes each against diff.patch and the source, starting from its repro. Survivors are CONFIRMED (85-100) or PLAUSIBLE (capped at 79, below the ≥80 keep line), so only CONFIRMED CRITICAL/HIGH findings reach the fixer.
apply-fixespromptmodel: gpt-6.1-sol, effort: high: applies only the triaged must-fix issues, running each one’s repro before and after the change, commits, pushes. No-ops when triage is clean.
re-reviewpromptmodel: gpt-6-luna, effort: high: closes the loop with a model independent of the fixer: confirms each must-fix is resolved with no new defect; fixes any straggler.
Re-validatepost-fix-validatebashRe-runs the full checks after the review loop (exits 0), saves the full transcript to $ARTIFACTS_DIR/post-fix-validate.log, and prints a bounded summary with the real status.
CI gateawait-cibashWatches the pushed PR’s CI; reruns failed jobs once to shake out flakes. Exits 0 so the run continues whether CI passed or not.
triage-cipromptmodel: claude-opus-4.8, effort: xhigh, typed JSON: classifies each CI failure as actionable (caused by this PR) or unrelated/flaky.
fix-cipromptmodel: gpt-6.1-sol, effort: high: fixes only the actionable failures on the PR branch and pushes. No-ops when triage finds none.
finalize-prbashRe-checks CI and runs forge pr ready to promote the draft only when every check is green; a red or check-less PR stays a draft.
Reportreport-statusbashReads the status header (CI conflict, CI red, CI unverified when no final CI status was recorded, or complete), the validation verdict, and the final PR state from what earlier nodes recorded. Names the full validation log in VALIDATION_LOG.
reportpromptThe operator summary: issue, branch, PR, what changed, and the review loop. Reads the verdict from report-status and names the full log without inlining the validation transcript.

The human gate is the spine’s hinge. Everything before approve-plan is read-and-think (fetch, classify, investigate or plan); everything after is write (implement, commit, push, PR). The gate sits exactly on that seam. When the issue has acceptance criteria, the gate shows how each criterion maps to the plan before the reviewer decides whether implementation should begin. capture_response: true means the reviewer’s reply is not just yes or no: if you send changes, implement is instructed to fold them into the plan before writing any code.

Two paths, one artifact. classify splits the flow with a when: gate, a bug goes to investigate, everything else to plan, but both write the same $ARTIFACTS_DIR/plan.md. plan-ready then joins them with trigger_rule: one_success, so nothing downstream has to know which path ran. Branching to converge on a shared artifact keeps the rest of the DAG linear.

Soft checks report, hard checks gate. validate and post-fix-validate exit 0 on purpose: their job is to report a status that an agent or the final report reads and acts on. revalidate exits non-zero on failure: its job is to gate, failing the node skips create-pr, so a PR is never opened over checks the run could not make pass. The exit code is the contract.

Each validation node saves its full transcript to $ARTIFACTS_DIR/<node>.log and prints only a bounded summary: the log path, exit code, failure lines when checks fail, and a capped tail, followed by the VALIDATION_STATUS verdict. The fixer can read the full log when the summary is not enough. The final report uses the verdict from report-status, not the validation transcript, so large test logs do not consume its model context.

It learns the project’s checks. implement reads CLAUDE.md, AGENTS.md, package.json, a Makefile, pyproject.toml, whatever exists, and writes the type-check, test, and lint commands to verify.sh with set -euo pipefail so the first failure aborts. Every validation node runs that exact script. The workflow carries no assumption about the stack, which is what makes it portable.

Operational guardrails live in the prompts. The implement and fix prompts forbid git add -A (stage by name) and forbid committing anything under $ARTIFACTS_DIR, so scratch files never leak into the PR. They also commit each task’s code and its tests together, immediately, never carrying uncommitted work forward, where a timeout would erase it. These are the kind of hard-won rules a generic “implement the plan” prompt omits.

Review is a panel, not a self-check. Phase 7 does not ask one agent to grade its own homework. capture-diff freezes the branch diff as an artifact, and three independent reviewers (correctness, conventions, and coverage) read it in parallel (they share a DAG layer, so the panel costs about one review’s wall-clock) under a read-only harness rail: allowed_tools: [Read, Glob, Grep], the same rail pr-review puts on its lanes, so a reviewer cannot edit or run shell no matter what its prompt says. Correctness and conventions findings carry a required repro: the command, test, or input that shows the failure, or none plus the reason. The coverage lens declares the field but leaves it optional, since it emits a test to add rather than a fix. The triage judge then holds the same tools and adversarially refutes each candidate against diff.patch and the source, starting from that repro (one that does not reproduce refutes the finding): a survivor is CONFIRMED when the failure was traced in real code (confidence 85-100) or PLAUSIBLE when it could not be refuted but not traced either (capped at 79, below the ≥80 keep line), so only CONFIRMED CRITICAL/HIGH findings reach apply-fixes, and re-review confirms the loop closed.

Four models, four roles, on purpose. The planning nodes (investigate, plan) are pinned to claude-opus-5.5; the judge nodes (coverage-check, triage, triage-ci) to claude-opus-4.8; the edit nodes (implement, fix-validation, apply-fixes, fix-ci) to gpt-6.1-sol. The review panel is mixed: review-correctness runs claude-opus-5.5, while review-conventions, review-coverage, and re-review run gpt-6-luna. The implementer never reviews its own work, no review lens shares the implementer’s model, and the judge that checks the plan and rules on the review findings is a different vendor from the planner, the implementer, and every reviewer. Each pin runs a chosen, deterministic model rather than auto, and each carries an explicit effort: tier; a model pin alone would still leave the node on that model’s default tier. The remaining prompt nodes stay on auto and let Copilot route them.

The CI gate fixes only its own mess. Once the PR is pushed, its real CI runs on every target the project covers, not just the local box. await-ci watches it and reruns failed jobs once to rule out a flake; triage-ci then classifies each failure as caused by this PR or pre-existing/flaky. fix-ci touches only the actionable ones. Auto-editing someone else’s failing test to force green would be worse than failing loud, so an unrelated red leaves the PR a draft for a human. finalize-pr promotes the draft to ready with forge pr ready only when every check passes, and never claims ready if the promotion itself fails.

  • Approval gate: approve-plan with capture_response, sitting between plan and implementation.
  • Classify, then branch: investigate vs plan, gated on classify.
  • Loop until done, unrolled: the validate, fix, revalidate triad is a self-correction sequence with a deterministic hard gate instead of a loop node.
  • Feed agent output into a script safely: every bash gate reads the agent’s work from verify.sh and the scratch dir, never from spliced text.
  • Worktree isolation: worktree.enabled so a mutating run cannot touch your checkout.
  • Change the bar. The whole validation story is verify.sh. To enforce extra checks (coverage, a security scan), have implement add them to that script; every gate picks them up.
  • Tune the review loop. Phase 7’s panel (capture-diff, review-correctness, review-conventions, review-coverage, triage, apply-fixes, re-review) is the run’s quality tail. Collapse it to a single reviewer plus apply-fixes for a cheaper run, or drop it so create-pr is the last step. The pins come in four roles (claude-opus-5.5 for planning and the correctness lens, gpt-6-luna for the other review nodes, gpt-6.1-sol for edits, claude-opus-4.8 for judgment); retarget any to whatever your Copilot account exposes, keeping each role on a distinct model if you can. All four ids are account-dependent: swap or drop any to auto if it does not appear in your Copilot model list, and drop that node’s effort: tier with it, since auto has no reasoning tier.
  • Swap the gate for autonomy. Remove approve-plan (and point implement at plan-ready) for an unattended run, at the cost of the safety interlock. Keep the gate for anything that pushes to a shared remote.
  • Drop the CI gate. To stop at a draft PR without waiting on CI, remove the await-ci, triage-ci, fix-ci, and finalize-pr phase and let report depend on post-fix-validate. Keep it when you want the PR promoted to ready automatically on green.