fix-issue
fix-issue takes a GitHub issue from a number to a draft pull request: it
fetches and classifies the issue, plans, pauses for a human to approve the
plan, implements, validates against the repository’s own checks, self-fixes
failures, opens a draft PR, then runs a multi-lens review loop where
independent reviewers for correctness, conventions, and coverage feed a triage
judge that adversarially verifies each candidate against the code and keeps
only confirmed defects, which a fixer applies and a re-review confirms. Finally it watches the pushed PR’s CI and promotes the draft to
ready for review the moment the checks go green. It is the catalog’s most
complete pipeline, and the clearest example of putting a human gate in front of
irreversible work.
It assumes nothing about your stack. The implement step discovers how this project checks itself by reading its docs and build config, writes those commands to a script, and every later validation runs that script. So the same workflow works in a Bun repo, a Cargo repo, or a Go repo without edits. For reviewing a PR rather than creating one, use pr-review.
Invoke it
Section titled “Invoke it”keelson workflow run fix-issue --inputs ARGUMENTS="42"keelson workflow run fix-issue --inputs ARGUMENTS="fix the SQLite timestamp bug"It needs gh or glab authenticated (the bundled forge shim resolves
whichever is present), a Copilot subscription, and a running
server (keelson start) because the approval gate pauses the run. It mutates
the working tree (branch, commits, push, PR), so it sets worktree.enabled:
every run executes in an isolated git worktree and cannot trample your editor
state. Worktree setup is a launch gate. If keelson cannot establish and record
the checkout, the run fails before extract-issue instead of continuing in the
source directory. --no-worktree is a deliberate override that makes the
workflow mutate the supplied directory in place.
The shape
Section titled “The shape”Ten phases, top to bottom. The spine is mostly linear, with three structures worth naming: a branch early (a bug is investigated, anything else is planned, both writing one file), a self-correction sequence after implementation (validate, fix, then a hard gate that refuses to open a PR over broken checks), and a CI gate at the end that promotes the draft to ready only when the pushed PR’s checks pass.
Figure 1. The fix-issue pipeline. The whole run executes in an isolated worktree (dashed frame). The brass node is the human approval gate; nothing downstream of it runs until the plan is approved. The hard gate after self-correction fails the run rather than open a PR over broken checks, and the closing CI gate marks the PR ready only when the pushed checks pass.
Node by node
Section titled “Node by node”| Phase | Node | Kind | What it does |
|---|---|---|---|
| Fetch & classify | extract-issue | prompt | Parses the issue number from $ARGUMENTS, or searches forge issue list for the best match. |
fetch-issue | bash | forge issue view to JSON, stashes the number in the scratch dir. | |
detect-base | bash | Resolves the repository’s default branch with forge repo view --json defaultBranchRef, falling back to origin/HEAD then main, and writes it to $ARTIFACTS_DIR/.default-branch. Never assumes main. | |
extract-brief | bash | Extracts bullets under acceptance criteria, requirements, definition of done, success criteria, “What did you expect?”, or expected-behavior headings into $ARTIFACTS_DIR/brief.json. Heading suffixes such as (must all hold) are accepted. | |
classify | prompt | Typed JSON: issue_type (bug, feature, …), a title, and reasoning. | |
| Plan | investigate | prompt | when: issue_type == 'bug', model: claude-opus-5.5, effort: medium: root-cause analysis into plan.md. |
plan | prompt | when: issue_type != 'bug', model: claude-opus-5.5, effort: medium: an implementation plan into the same plan.md. | |
plan-ready | bash | Funnel (trigger_rule: one_success): asserts plan.md exists and has tasks, then prints it. | |
criteria-count-initial | bash | Counts criteria found by the structured heading scan. | |
extract-brief-llm | prompt | when: criteria-count-initial == 0, allowed_tools: []: extracts concrete, testable outcomes from prose-only issue bodies into typed JSON. A structured brief skips this call. | |
brief-ready | bash | Funnel (trigger_rule: all_done): merges fallback criteria into brief.json. It still runs when the fallback is skipped, so structured criteria continue through the DAG. | |
criteria-count | bash | Recounts the merged brief so the coverage gate sees either extraction path. | |
coverage-check | prompt | when: criteria-count > 0, model: claude-opus-4.8, effort: xhigh: semantic judge that maps each brief criterion to the plan step that covers it or MISSING. Returns typed JSON only; skipped when both extraction paths find no criteria. | |
coverage-ready | bash | Funnel (trigger_rule: all_done): writes the judge’s output to coverage.json when criteria exist. Otherwise it removes the artifact and reports the skip in the run log. | |
| Approve | approve-plan | approval | Pauses the run. capture_response makes the reply, approval or change requests, available downstream. The callout includes the criteria-coverage checklist, or a visible note that the divergence check did not run. |
| Implement | implement | prompt | model: gpt-6.1-sol, effort: xhigh: Writes verify.sh (the project’s own checks), branches, implements task by task, commits each. |
| Validate | validate | bash | Runs verify.sh fail-fast, saves the full transcript to $ARTIFACTS_DIR/validate.log, and prints a bounded summary. Exits 0 and reports VALIDATION_STATUS so the fixer can read it. |
fix-validation | prompt | model: gpt-6.1-sol, effort: high: if it failed, fixes only what broke and commits; if it passed, stops. | |
revalidate | bash | Hard gate: re-runs the checks, saves the full transcript to $ARTIFACTS_DIR/revalidate.log, and prints a bounded summary. Exits non-zero on failure, skipping the PR. | |
| Draft PR | create-pr | prompt | Pushes the branch, opens a draft PR with Fixes #N, and captures the PR number and URL. It follows a single standalone pull_request_template.md when one exists; repositories with only named templates under PULL_REQUEST_TEMPLATE/ use the fallback body. Its test plan records commands actually executed and their results, with pending CI reserved for checks that only CI can confirm. |
enforce-draft | bash | Verifies the new PR is a draft; if it came up ready, runs forge pr ready --undo to force it back to draft, and fails the run if it cannot confirm the draft state. | |
| Review loop | capture-diff | bash | Freezes the branch diff to $ARTIFACTS_DIR/diff.patch / diff-stat.txt so the lenses can run without shell access. |
review-correctness | prompt | model: claude-opus-5.5, effort: medium, allowed_tools: [Read, Glob, Grep]: one of three independent, parallel reviewers over diff.patch. Hunts by cost of failure (trust boundaries, irreversible state, idempotency gaps), follows each changed function to every entry path that reaches it (resume, cancel, the opt-out configuration, data written before the change), checks that docs and reference tables still agree with the code, labels unverified external-API claims INFERRED, and emits a confidence and a repro (the command, test, or input that shows the failure, or none plus the reason) per finding. | |
review-conventions | prompt | model: gpt-6-luna, effort: xhigh, same read-only rail: checks the diff against CLAUDE.md / CONTRIBUTING.md (comment policy, conventions), quoting the rule each finding violates. | |
review-coverage | prompt | model: gpt-6-luna, effort: xhigh, same read-only rail: flags new behavior left untested, and tests authored but never wired in. | |
triage | prompt | model: claude-opus-4.8, effort: xhigh, allowed_tools: [Read, Glob, Grep], typed JSON: de-dupes the findings, then adversarially refutes each against diff.patch and the source, starting from its repro. Survivors are CONFIRMED (85-100) or PLAUSIBLE (capped at 79, below the ≥80 keep line), so only CONFIRMED CRITICAL/HIGH findings reach the fixer. | |
apply-fixes | prompt | model: gpt-6.1-sol, effort: high: applies only the triaged must-fix issues, running each one’s repro before and after the change, commits, pushes. No-ops when triage is clean. | |
re-review | prompt | model: gpt-6-luna, effort: high: closes the loop with a model independent of the fixer: confirms each must-fix is resolved with no new defect; fixes any straggler. | |
| Re-validate | post-fix-validate | bash | Re-runs the full checks after the review loop (exits 0), saves the full transcript to $ARTIFACTS_DIR/post-fix-validate.log, and prints a bounded summary with the real status. |
| CI gate | await-ci | bash | Watches the pushed PR’s CI; reruns failed jobs once to shake out flakes. Exits 0 so the run continues whether CI passed or not. |
triage-ci | prompt | model: claude-opus-4.8, effort: xhigh, typed JSON: classifies each CI failure as actionable (caused by this PR) or unrelated/flaky. | |
fix-ci | prompt | model: gpt-6.1-sol, effort: high: fixes only the actionable failures on the PR branch and pushes. No-ops when triage finds none. | |
finalize-pr | bash | Re-checks CI and runs forge pr ready to promote the draft only when every check is green; a red or check-less PR stays a draft. | |
| Report | report-status | bash | Reads the status header (CI conflict, CI red, CI unverified when no final CI status was recorded, or complete), the validation verdict, and the final PR state from what earlier nodes recorded. Names the full validation log in VALIDATION_LOG. |
report | prompt | The operator summary: issue, branch, PR, what changed, and the review loop. Reads the verdict from report-status and names the full log without inlining the validation transcript. |
The parts worth a second look
Section titled “The parts worth a second look”The human gate is the spine’s hinge. Everything before approve-plan is
read-and-think (fetch, classify, investigate or plan); everything after is
write (implement, commit, push, PR). The gate sits exactly on that seam.
When the issue has acceptance criteria, the gate shows how each criterion maps
to the plan before the reviewer decides whether implementation should begin.
capture_response: true means the reviewer’s reply is not just yes or no: if
you send changes, implement is instructed to fold them into the plan before
writing any code.
Two paths, one artifact. classify splits the flow with a when: gate, a
bug goes to investigate, everything else to plan, but both write the same
$ARTIFACTS_DIR/plan.md. plan-ready then joins them with
trigger_rule: one_success, so nothing downstream has to know which path ran.
Branching to converge on a shared artifact keeps the rest of the DAG linear.
Soft checks report, hard checks gate. validate and post-fix-validate
exit 0 on purpose: their job is to report a status that an agent or the final
report reads and acts on. revalidate exits non-zero on failure: its job is to
gate, failing the node skips create-pr, so a PR is never opened over checks
the run could not make pass. The exit code is the contract.
Each validation node saves its full transcript to $ARTIFACTS_DIR/<node>.log
and prints only a bounded summary: the log path, exit code, failure lines when
checks fail, and a capped tail, followed by the VALIDATION_STATUS verdict.
The fixer can read the full log when the summary is not enough. The final
report uses the verdict from report-status, not the validation transcript,
so large test logs do not consume its model context.
It learns the project’s checks. implement reads CLAUDE.md, AGENTS.md,
package.json, a Makefile, pyproject.toml, whatever exists, and writes the
type-check, test, and lint commands to verify.sh with set -euo pipefail so
the first failure aborts. Every validation node runs that exact script. The
workflow carries no assumption about the stack, which is what makes it portable.
Operational guardrails live in the prompts. The implement and fix prompts
forbid git add -A (stage by name) and forbid committing anything under
$ARTIFACTS_DIR, so scratch files never leak into the PR. They also commit each
task’s code and its tests together, immediately, never carrying uncommitted
work forward, where a timeout would erase it. These are the kind of hard-won
rules a generic “implement the plan” prompt omits.
Review is a panel, not a self-check. Phase 7 does not ask one agent to grade
its own homework. capture-diff freezes the branch diff as an artifact, and
three independent reviewers (correctness, conventions, and coverage) read it in
parallel (they share a DAG layer, so the panel costs about one review’s
wall-clock) under a read-only harness rail: allowed_tools: [Read, Glob, Grep],
the same rail pr-review puts on its lanes, so a reviewer
cannot edit or run shell no matter what its prompt says. Correctness and
conventions findings carry a required repro: the command, test, or input
that shows the failure, or none plus the reason. The coverage lens declares
the field but leaves it optional, since it emits a test to add rather than a
fix. The triage judge then holds the same tools and
adversarially refutes each candidate against diff.patch and the source,
starting from that repro (one that does not reproduce refutes the finding): a
survivor is CONFIRMED when the failure was traced
in real code (confidence 85-100) or PLAUSIBLE when it could not be refuted but
not traced either (capped at 79, below the ≥80 keep line), so only CONFIRMED
CRITICAL/HIGH findings reach apply-fixes, and re-review confirms the loop
closed.
Four models, four roles, on purpose. The planning nodes (investigate,
plan) are pinned to claude-opus-5.5; the judge nodes (coverage-check,
triage, triage-ci) to claude-opus-4.8; the edit nodes (implement,
fix-validation, apply-fixes, fix-ci) to gpt-6.1-sol. The review panel is
mixed: review-correctness runs claude-opus-5.5, while review-conventions,
review-coverage, and re-review run gpt-6-luna. The implementer never
reviews its own work, no review lens shares the implementer’s model, and the
judge that checks the plan and rules on the review findings is a different
vendor from the planner, the implementer, and every reviewer. Each pin runs a chosen, deterministic
model rather than auto, and each carries an explicit effort: tier; a model
pin alone would still leave the node on that model’s default tier. The remaining
prompt nodes stay on auto and let Copilot route them.
The CI gate fixes only its own mess. Once the PR is pushed, its real CI
runs on every target the project covers, not just the local box. await-ci
watches it and reruns failed jobs once to rule out a flake; triage-ci then
classifies each failure as caused by this PR or pre-existing/flaky. fix-ci
touches only the actionable ones. Auto-editing someone else’s failing test to
force green would be worse than failing loud, so an unrelated red leaves the PR a
draft for a human. finalize-pr promotes the draft to ready with
forge pr ready only when every check passes, and never claims ready if the
promotion itself fails.
Patterns it demonstrates
Section titled “Patterns it demonstrates”- Approval gate:
approve-planwithcapture_response, sitting between plan and implementation. - Classify, then branch:
investigatevsplan, gated onclassify. - Loop until done, unrolled: the validate, fix, revalidate triad is a self-correction sequence with a deterministic hard gate instead of a
loopnode. - Feed agent output into a script safely: every
bashgate reads the agent’s work fromverify.shand the scratch dir, never from spliced text. - Worktree isolation:
worktree.enabledso a mutating run cannot touch your checkout.
Adapt it
Section titled “Adapt it”- Change the bar. The whole validation story is
verify.sh. To enforce extra checks (coverage, a security scan), haveimplementadd them to that script; every gate picks them up. - Tune the review loop. Phase 7’s panel (
capture-diff,review-correctness,review-conventions,review-coverage,triage,apply-fixes,re-review) is the run’s quality tail. Collapse it to a single reviewer plusapply-fixesfor a cheaper run, or drop it socreate-pris the last step. The pins come in four roles (claude-opus-5.5for planning and the correctness lens,gpt-6-lunafor the other review nodes,gpt-6.1-solfor edits,claude-opus-4.8for judgment); retarget any to whatever your Copilot account exposes, keeping each role on a distinct model if you can. All four ids are account-dependent: swap or drop any toautoif it does not appear in your Copilot model list, and drop that node’seffort:tier with it, sinceautohas no reasoning tier. - Swap the gate for autonomy. Remove
approve-plan(and pointimplementatplan-ready) for an unattended run, at the cost of the safety interlock. Keep the gate for anything that pushes to a shared remote. - Drop the CI gate. To stop at a draft PR without waiting on CI, remove the
await-ci,triage-ci,fix-ci, andfinalize-prphase and letreportdepend onpost-fix-validate. Keep it when you want the PR promoted to ready automatically on green.
Related
Section titled “Related”- pr-review: the sibling that reviews a PR instead of creating one.
- plan-act-evaluate: a heavier lifecycle that records the plan as an issue and can delegate the work.
- Authoring workflows: the approval-gate and branch recipes.
- Workflow nodes:
approval,when:,trigger_rule:, andworktree.