pr-review
pr-review reads a GitHub pull request, decides how much review it warrants,
runs only the lanes that apply, then posts its findings as a GitHub review with
line-anchored inline comments and suggestion blocks. It is read-only over the
code by construction: every lane is capped to Read, Glob, and Grep, and the
only step that touches the PR is the final forge pr review-batch call. To fix an
issue rather than review one, reach for fix-issue; pr-review
never mutates the branch it reads.
It is the catalog’s clearest case of classify, then fan out: one node
decides which lanes are worth running, the lanes run in parallel behind a
when: gate, and an opus triage judge de-dupes whatever ran into the
findings worth posting.
Invoke it
Section titled “Invoke it”keelson workflow run pr-review --inputs ARGUMENTS="review pr 42"Pass an explicit PR number. pr-review runs in an isolated worktree on a
generated branch, so it cannot infer the target PR from the current branch.
The scope node checks the PR head out inside that worktree, so the lanes
review the code as the PR proposes it, never your live checkout, which stays
untouched.
It needs gh or glab authenticated against the repo (the bundled forge shim
resolves whichever is present) and a Copilot subscription
(it pins provider: copilot, model: auto). It posts a comment but never
pushes code, so it needs no approval gate and no running server.
The shape
Section titled “The shape”The run is five phases. Deterministic bash nodes bookend it: scope gathers
the diff up front, then a build-review / post-review pair anchors the
findings to diff lines and publishes the review at the end. In between, an agent
classifier decides the work, a fan of agent lanes does it, and a triage judge
distills the lanes into the findings worth posting.
Figure 1. The pr-review DAG. The classifier fans out to four read-only
lanes, each behind a when: gate. Code-review’s gate fires for
any code change while the other three are selective, so a trivial PR runs
fewer. The triage judge joins on one_success, so it
proceeds even when lanes were skipped or one failed, then
build-review anchors its findings to diff lines and
post-review publishes them as inline comments.
Node by node
Section titled “Node by node”| Node | Kind | What it does |
|---|---|---|
extract-pr | bash | Parses the PR number from $ARGUMENTS (a PR/MR URL, pr 42, #42, or a bare number; two different numbers are refused rather than guessed). An explicit number is required, since the run cannot infer the PR from the branch in an isolated worktree. |
scope | bash | Checks the PR head out in the run’s worktree (forge pr checkout --detach), then writes pr.json, diff.patch, changed-files.txt, and a scope.md brief into the run’s scratch dir. |
classify | prompt | Returns typed JSON: which of the four lanes to run, a complexity bucket, and the reasoning. |
code-review | prompt | model: claude-opus-5.5 on Copilot (the deep tier elsewhere), effort: xhigh: the always-on finder is the recall ceiling, so it gets the strongest coding model. Hunts highest-cost failures first (trust boundaries, irreversible state, idempotency gaps), labels unverified external-API claims INFERRED, and runs unless the PR is non-code. Every lane finding carries a repro: the command, test, or input that shows the failure, or none plus the reason. |
error-handling | prompt | Audits new and changed failure paths. Gated on the classifier. |
test-coverage | prompt | Checks that changed source has matching test changes. Gated on the classifier. |
docs-impact | prompt | Flags user-facing changes (flags, env vars, APIs) with no doc update. Gated on the classifier. |
triage | prompt | model: claude-opus-4.8, effort: xhigh, allowed_tools: [Read, Glob, Grep], the one_success join: de-dupes the lanes, then adversarially refutes each candidate against diff.patch and the source, starting from its repro. Survivors are CONFIRMED (confidence 85-100) or PLAUSIBLE (capped at 79, below the ≥80 keep line), so only CONFIRMED CRITICAL/HIGH/MEDIUM findings post. Emits a verdict (READY TO MERGE / NEEDS NITS / NEEDS FIXES) with typed findings, plus a summary written in the posting engineer’s first-person voice; it becomes the review body verbatim, with no branding or verdict banner. |
build-review | bash | Offline: anchors each finding to a line on the diff’s new side and builds the review payload. A suggestion block is gated: an empty fix, or one byte-identical to the anchored line, posts as prose only. A finding’s repro is rendered under its rationale, or under its bullet in the review body when the finding cannot be anchored, and omitted when it is none. No network. |
post-review | bash | Posts the payload with forge pr review-batch as line-anchored inline comments. Idempotent by content hash: skips when a prior review already carries the same keelson:pr-review:<hash> marker. |
The parts worth a second look
Section titled “The parts worth a second look”The classifier returns strings, not booleans. Each lane flag is typed as an
enum of "true" / "false", not a JSON boolean, because the when: gates
compare against a quoted literal:
when: "$classify.output.run_error_handling == 'true'"Substitution renders a JSON boolean as bare true, which would never match
== 'true'. Typing the field as a string enum is what makes the gate fire. This
is the one sharp edge in the workflow, and it generalizes to any classifier that
drives a when:.
Every lane is read-only. Each sets allowed_tools: [Read, Glob, Grep], so
the reviewers can read the tree but cannot write, run shell, or touch the
network. On Copilot that gate is enforced at the capability layer, so the
model’s built-in write and shell tools are denied mid-review regardless of what
it tries. The only write in the whole run is the final posted review.
A triage gate stands between the lanes and the PR. The lanes are
deliberately noisy; an opus triage judge then de-dupes their findings and
actively tries to refute each one. It is the only judge granted
allowed_tools: [Read, Glob, Grep], because refuting a claim means reading
diff.patch and the source it points at. The judge starts from the finding’s
repro, the command, test, or input the lane says shows the failure: a repro
that does not reproduce refutes the finding, and none earns no credit on its
own. A candidate also dies when the code already handles the case, the claim
misreads the diff, the issue predates the PR, or a linter or CI gate already
catches it. A survivor is CONFIRMED when the judge
traced the failure in real code (confidence 85-100), or PLAUSIBLE when it
could not refute the claim but could not trace it either.
The cap is the mechanic worth copying: PLAUSIBLE is pinned at 79, deliberately
one point below the >= 80 keep line, so an untraced finding cannot post. The
judge must re-trace a PLAUSIBLE candidate up to CONFIRMED or let it drop. That
turns “I could not disprove it” into silence rather than a comment, and it is why
the PR author sees a short list of traced defects instead of four lanes’ worth of
raw output. LOW-severity findings are dropped outright. build-review also
hashes the kept findings’ stable identity into the review body, so a re-review
does not repost the same comments.
The lanes share state through files, not prompts. scope writes the diff to
$ARTIFACTS_DIR/diff.patch and each lane reads that file directly, rather than
having a huge diff substituted into four prompts. Each lane also runs
context: fresh, so they do not pollute each other’s reasoning.
The join tolerates absence. triage depends on all four lanes but sets
trigger_rule: one_success, so a skipped lane (its when: was false) or a
single failed lane does not strand the review. Whatever ran gets triaged.
Patterns it demonstrates
Section titled “Patterns it demonstrates”- Classify, then branch:
classifyemits the labels; the lanes gate on them withwhen:. - Fan-out and fan-in: four lanes in parallel, joined by
triage. - Bound an agent’s reach: the read-only
allowed_toolsrail on every lane. - Artifacts as the shared channel:
bashnodes stage files in$ARTIFACTS_DIRthat agent nodes read, keeping large diffs out of prompt substitution.
Adapt it
Section titled “Adapt it”- Add a lane. A new review angle is four small edits: a flag on the
classifier’s
output_format, a rule in the classifier prompt, a gated lane node, and a section feedingtriage’s prompt. - Change the rail. Widen
allowed_toolsif a lane needs to run a linter, or keep it read-only and lettriagerecommend commands instead. - Repoint the output. Swap the
post-reviewstep’sforge pr review-batchcall for a different sink (a single summary comment, a file) without touching the classify, lane, or triage logic.
Related
Section titled “Related”- Authoring workflows: the recipes this composes.
- Workflow nodes:
when:,trigger_rule:, andoutput_formatin full. - fix-issue: the sibling that changes code instead of reviewing it.