Skip to content

pr-review

pr-review reads a GitHub pull request, decides how much review it warrants, runs only the lanes that apply, then posts its findings as a GitHub review with line-anchored inline comments and suggestion blocks. It is read-only over the code by construction: every lane is capped to Read, Glob, and Grep, and the only step that touches the PR is the final forge pr review-batch call. To fix an issue rather than review one, reach for fix-issue; pr-review never mutates the branch it reads.

It is the catalog’s clearest case of classify, then fan out: one node decides which lanes are worth running, the lanes run in parallel behind a when: gate, and an opus triage judge de-dupes whatever ran into the findings worth posting.

Terminal window
keelson workflow run pr-review --inputs ARGUMENTS="review pr 42"

Pass an explicit PR number. pr-review runs in an isolated worktree on a generated branch, so it cannot infer the target PR from the current branch. The scope node checks the PR head out inside that worktree, so the lanes review the code as the PR proposes it, never your live checkout, which stays untouched.

It needs gh or glab authenticated against the repo (the bundled forge shim resolves whichever is present) and a Copilot subscription (it pins provider: copilot, model: auto). It posts a comment but never pushes code, so it needs no approval gate and no running server.

The run is five phases. Deterministic bash nodes bookend it: scope gathers the diff up front, then a build-review / post-review pair anchors the findings to diff lines and publishes the review at the end. In between, an agent classifier decides the work, a fan of agent lanes does it, and a triage judge distills the lanes into the findings worth posting.

The pr-review DAG, left to right. extract-pr (bash) feeds scope (bash); scope feeds classify (prompt). classify fans out to four read-only review lanes (code-review, error-handling, test-coverage, docs-impact), all four behind a when-gate: code-review's fires for any code change (solid edge), the other three are selective (dashed edges). All four lanes converge on triage (an opus judge) under a one_success join, which feeds build-review then post-review (both bash), posting line-anchored inline comments.

Figure 1. The pr-review DAG. The classifier fans out to four read-only lanes, each behind a when: gate. Code-review’s gate fires for any code change while the other three are selective, so a trivial PR runs fewer. The triage judge joins on one_success, so it proceeds even when lanes were skipped or one failed, then build-review anchors its findings to diff lines and post-review publishes them as inline comments.

NodeKindWhat it does
extract-prbashParses the PR number from $ARGUMENTS (a PR/MR URL, pr 42, #42, or a bare number; two different numbers are refused rather than guessed). An explicit number is required, since the run cannot infer the PR from the branch in an isolated worktree.
scopebashChecks the PR head out in the run’s worktree (forge pr checkout --detach), then writes pr.json, diff.patch, changed-files.txt, and a scope.md brief into the run’s scratch dir.
classifypromptReturns typed JSON: which of the four lanes to run, a complexity bucket, and the reasoning.
code-reviewpromptmodel: claude-opus-5.5 on Copilot (the deep tier elsewhere), effort: xhigh: the always-on finder is the recall ceiling, so it gets the strongest coding model. Hunts highest-cost failures first (trust boundaries, irreversible state, idempotency gaps), labels unverified external-API claims INFERRED, and runs unless the PR is non-code. Every lane finding carries a repro: the command, test, or input that shows the failure, or none plus the reason.
error-handlingpromptAudits new and changed failure paths. Gated on the classifier.
test-coveragepromptChecks that changed source has matching test changes. Gated on the classifier.
docs-impactpromptFlags user-facing changes (flags, env vars, APIs) with no doc update. Gated on the classifier.
triagepromptmodel: claude-opus-4.8, effort: xhigh, allowed_tools: [Read, Glob, Grep], the one_success join: de-dupes the lanes, then adversarially refutes each candidate against diff.patch and the source, starting from its repro. Survivors are CONFIRMED (confidence 85-100) or PLAUSIBLE (capped at 79, below the ≥80 keep line), so only CONFIRMED CRITICAL/HIGH/MEDIUM findings post. Emits a verdict (READY TO MERGE / NEEDS NITS / NEEDS FIXES) with typed findings, plus a summary written in the posting engineer’s first-person voice; it becomes the review body verbatim, with no branding or verdict banner.
build-reviewbashOffline: anchors each finding to a line on the diff’s new side and builds the review payload. A suggestion block is gated: an empty fix, or one byte-identical to the anchored line, posts as prose only. A finding’s repro is rendered under its rationale, or under its bullet in the review body when the finding cannot be anchored, and omitted when it is none. No network.
post-reviewbashPosts the payload with forge pr review-batch as line-anchored inline comments. Idempotent by content hash: skips when a prior review already carries the same keelson:pr-review:<hash> marker.

The classifier returns strings, not booleans. Each lane flag is typed as an enum of "true" / "false", not a JSON boolean, because the when: gates compare against a quoted literal:

when: "$classify.output.run_error_handling == 'true'"

Substitution renders a JSON boolean as bare true, which would never match == 'true'. Typing the field as a string enum is what makes the gate fire. This is the one sharp edge in the workflow, and it generalizes to any classifier that drives a when:.

Every lane is read-only. Each sets allowed_tools: [Read, Glob, Grep], so the reviewers can read the tree but cannot write, run shell, or touch the network. On Copilot that gate is enforced at the capability layer, so the model’s built-in write and shell tools are denied mid-review regardless of what it tries. The only write in the whole run is the final posted review.

A triage gate stands between the lanes and the PR. The lanes are deliberately noisy; an opus triage judge then de-dupes their findings and actively tries to refute each one. It is the only judge granted allowed_tools: [Read, Glob, Grep], because refuting a claim means reading diff.patch and the source it points at. The judge starts from the finding’s repro, the command, test, or input the lane says shows the failure: a repro that does not reproduce refutes the finding, and none earns no credit on its own. A candidate also dies when the code already handles the case, the claim misreads the diff, the issue predates the PR, or a linter or CI gate already catches it. A survivor is CONFIRMED when the judge traced the failure in real code (confidence 85-100), or PLAUSIBLE when it could not refute the claim but could not trace it either.

The cap is the mechanic worth copying: PLAUSIBLE is pinned at 79, deliberately one point below the >= 80 keep line, so an untraced finding cannot post. The judge must re-trace a PLAUSIBLE candidate up to CONFIRMED or let it drop. That turns “I could not disprove it” into silence rather than a comment, and it is why the PR author sees a short list of traced defects instead of four lanes’ worth of raw output. LOW-severity findings are dropped outright. build-review also hashes the kept findings’ stable identity into the review body, so a re-review does not repost the same comments.

The lanes share state through files, not prompts. scope writes the diff to $ARTIFACTS_DIR/diff.patch and each lane reads that file directly, rather than having a huge diff substituted into four prompts. Each lane also runs context: fresh, so they do not pollute each other’s reasoning.

The join tolerates absence. triage depends on all four lanes but sets trigger_rule: one_success, so a skipped lane (its when: was false) or a single failed lane does not strand the review. Whatever ran gets triaged.

  • Classify, then branch: classify emits the labels; the lanes gate on them with when:.
  • Fan-out and fan-in: four lanes in parallel, joined by triage.
  • Bound an agent’s reach: the read-only allowed_tools rail on every lane.
  • Artifacts as the shared channel: bash nodes stage files in $ARTIFACTS_DIR that agent nodes read, keeping large diffs out of prompt substitution.
  • Add a lane. A new review angle is four small edits: a flag on the classifier’s output_format, a rule in the classifier prompt, a gated lane node, and a section feeding triage’s prompt.
  • Change the rail. Widen allowed_tools if a lane needs to run a linter, or keep it read-only and let triage recommend commands instead.
  • Repoint the output. Swap the post-review step’s forge pr review-batch call for a different sink (a single summary comment, a file) without touching the classify, lane, or triage logic.