plan-act-evaluate
plan-act-evaluate runs a change through a full lifecycle: it drafts a plan
under a read-only boundary, records that plan as a GitHub issue, pauses for a
human to approve it, then either implements the change locally or delegates
it to the GitHub Copilot coding agent, and treats the resulting PR’s CI as the
authoritative signal. It is the catalog’s heaviest workflow, and the one that
shows a real two-arm branch reconverging.
Reach for it when the process matters as much as the change: a plan in the system of record, a gate before code, a choice of who does the work, and an audit trail at the end. For a quick known fix, fix-issue is lighter; for reviewing an existing PR, use pr-review.
Invoke it
Section titled “Invoke it”keelson workflow run plan-act-evaluate --inputs ARGUMENTS="add a --json flag to the status command"keelson workflow run plan-act-evaluate --inputs ARGUMENTS="42" # plan onto an existing issueIt needs gh or glab authenticated against a repo with issues enabled (the
bundled forge shim resolves whichever is present), a Copilot subscription, and
a running server (the approval gate pauses). The delegate mode is
GitHub-only; on any other forge the workflow pins itself to local. The change lands
best on a repo with CI, since CI is the authoritative check. The delegate
path additionally needs the Copilot coding agent enabled on the repo or org (it
must appear as copilot-swe-agent in the assignable actors).
The shape
Section titled “The shape”Four lifecycle stages: plan, a human gate, act, and evaluate, ending in an audit report. The act stage is where it splits. The reviewer’s reply chooses the arm: an approval implements locally, the word “delegate” hands the issue to Copilot. Both arms reconverge on one node that reads the PR’s CI.
Figure 1. The plan-act-evaluate DAG. The brass node is the human gate, and
its reply routes the branch below. The local arm (dashed frame) runs in an
isolated worktree; the delegate arm runs entirely on GitHub. Both arms
reconverge on ci-signals, the authoritative check.
Node by node
Section titled “Node by node”| Stage | Node | Kind | What it does |
|---|---|---|---|
| Plan | intake | bash | Classifies the request (issue #, plan file, or free text), loads context, detects the default branch. |
plan | prompt | Read-only (allowed_tools: [Read, Glob, Grep]): authors the plan as text. Cannot write, branch, or commit. | |
publish-plan | bash | Records the plan as a GitHub issue (new, or a comment on the one you named), and as the plan.md review surface. | |
| Gate | approve-plan | approval | Pauses on the plan. The reply both approves and chooses the path. |
| Act | detect-forge | bash | Runs forge caps, which prints the forge resolved for this repo (github or gitlab). |
select-mode | prompt | Typed JSON, reading both the reply and detect-forge: local (an approval) or delegate (the reviewer said “delegate”). A non-GitHub forge forces local, overriding any request to delegate. | |
| Act · local | act-local | prompt | when: mode == 'local': implements on a branch, commits by name. |
evaluate | loop | Review-and-improve loop (max_iterations: 5, until: EVALUATION_COMPLETE) on the project’s own checks. | |
open-pr | prompt | Pushes the branch and opens a draft PR linking the plan issue. | |
| Act · delegate | act-delegate | bash | when: mode == 'delegate': assigns the issue to copilot-swe-agent via the GitHub API. |
await-copilot | bash | Polls until Copilot opens its PR (or its head SHA settles), captures the PR number. | |
| Evaluate | ci-signals | bash | Reconverges both arms (trigger_rule: one_success): forge pr checks, classified PASS / FAIL / NONE / PENDING_APPROVAL. |
| Audit · delegate | capture-diff | bash | when: delegate: pulls the PR diff for review. |
self-review | prompt | Read-only contributor-model review of Copilot’s PR. Each finding carries a repro, or none plus the reason. | |
pr-review | prompt | Renders a formal PR review: forge pr review --approve or --request-changes (which feeds Copilot). | |
| Audit | report | bash | Deterministic 6-element audit trail (goal, plan, changeset, evidence, judgment, outcome) joined on one_success. |
The parts worth a second look
Section titled “The parts worth a second look”The planner cannot touch code. plan sets allowed_tools: [Read, Glob, Grep], an enforced capability boundary, not a guideline. It physically cannot
write, run shell, branch, or commit; it emits the plan as text, and the
deterministic publish-plan is the only node that records it. Separating “decide
what to do” from “do it” with a hard rail is the workflow’s founding idea.
The plan is a durable artifact. publish-plan writes the plan to a GitHub
issue, the system of record, not just a scratch file. The same content backs the
approval canvas, so the human reviews the exact text that will guide the
implementer, local or cloud.
One reply does two jobs. approve-plan captures the reviewer’s response, and
select-mode reads it: an approval routes to local, the word “delegate” routes
to Copilot. The gate is also the branch selector, so the human picks who acts
without a second prompt. The forge has the final say: detect-forge runs
forge caps first, and on anything but GitHub (gitlab, say) select-mode is
required to return local, since agent delegation is GitHub-only. On GitLab the
delegate arm is unreachable and a request to delegate is overridden.
Two arms, one reconvergence. select-mode splits the DAG with when: gates.
The local arm implements, loops to self-improve, and opens a PR; the delegate arm
assigns the issue and waits. They rejoin at ci-signals, which sets
trigger_rule: one_success so whichever arm ran flows through. The delegate-only
review tail (capture-diff, self-review, pr-review) hangs off that join,
because only Copilot’s PR needs an external reviewer; the local arm already
reviewed itself in the loop.
CI is the truth, not the agent. ci-signals reads forge pr checks, an
objective system signal, and classifies it carefully: it distinguishes “no CI on
this repo” from “CI exists but is gated behind maintainer approval” (the common
case for a bot’s PR), so the audit trail never claims green when a human just
hasn’t approved the workflow yet.
The report is deterministic on purpose. report is a bash node, not a
prompt: every field is pulled from disk, git, or gh. An AI summary could
hallucinate the goal or hang on the final node and fail an otherwise-green run.
A facts-only audit trail can do neither.
Patterns it demonstrates
Section titled “Patterns it demonstrates”- Bound an agent’s reach: the read-only
plannode: the capability boundary is the design. - Approval gate:
approve-plan, whose captured reply also routes the branch. - Classify, then branch:
select-modeinto two arms, reconverged byci-signalsonone_success. - Loop until done:
evaluateis a realloopnode, in contrast to fix-issue’s unrolled triad. - External signal as the gate:
forge pr checksas the authoritative check, over the agent’s own confidence.
Adapt it
Section titled “Adapt it”- Local only. If you never delegate, delete the
act-delegate,await-copilot, and the delegate review tail, and drop thewhen:gates on the local arm. You are left with a plan-issue-gate-implement-evaluate flow. - Change the authority.
ci-signalsdefines what “evaluated” means. Point it at a different check (a deploy smoke test, a specific workflow run) to move the bar. - Keep the audit, swap the work. The
reportnode is reusable as-is: a deterministic 6-element trail that any plan-act-evaluate-shaped workflow can end on.
Related
Section titled “Related”- fix-issue: the lighter sibling, no plan issue, no delegate arm.
- pr-review: the review lanes
self-reviewechoes here. - Authoring workflows: the branch, gate, and loop recipes.
- Workflow nodes:
loop,approval,trigger_rule:, and tool gates.