Skip to content

Run the software factory

The board from the last stop says one bead is ready and ten wait. Something now has to work it the way a team would: take the ready bead, build it, have someone else read the change before it lands, merge it, and pick up whatever that unblocked. When the graph widens, two people work side by side; when two changes touch the same file, one of them rebases.

That is a swarm: a lead agent that owns the outcome, workers it spawns as the work calls for them, and a channel they coordinate in that you can watch and post in. Put it in Factory mode and it has no turn or clock budget; it runs as long as work keeps landing. This stop installs the swarm rib, points a swarm at the board, and lets it build Cosmos.

Terminal window
keelson rib add https://github.com/danielscholl/keelson-rib-swarm

The rib starts and stops its own local ClickClack server, so there is nothing else to run. Its install guide covers building the binary and pointing the rib at a server you run yourself.

A rib cannot call another rib’s tools until you say so. The swarm’s lead needs the beads tools to read the ready queue, claim beads, and close them as they land, so grant them in ~/.keelson/config.json:

{
"crossRibGrants": {
"swarm": {
"beads": [
"beads_ready",
"beads_show",
"beads_create",
"beads_update",
"beads_close",
"beads_dep"
]
}
}
}

Grants load at boot, so restart, then let keelson doctor confirm the server holds them:

Terminal window
keelson stop && keelson start
keelson doctor

Without the grant the launcher’s Use the tracker switch has nothing to offer, and a swarm started anyway never sees the board. See cross-rib grants for why the denial is quiet.

The template ships the prompt in two parts. factory/BRIEF.md says what to build: the stack, the art direction, the behavior, and what done means. factory/PROMPT.md says how the factory works. Keeping them apart is what lets you measure the factory later, because the brief alone is a fair prompt for a single agent.

The factory part asks for three things a single agent cannot do to itself:

  • Critique the plan before building. The lead refines the seeded beads (or plans its own on an empty tracker), and at least two other agents critique the plan before anyone writes code.
  • Judge pages by how they look. For every change a visitor can see, the writer saves full-page screenshots at 1440px and 375px wide, and a design critic opens them and sends back the one change that would most improve the page, until it approves.
  • Never merge unread work. An agent other than the writer approves every head before the lead merges it and closes its bead.

The lead staffs the factory itself: writers, the critic, a code reviewer, whatever the beads need.

Open the Swarms tab and choose Options to expand the launcher. Paste the brief from factory/BRIEF.md, then the factory section of factory/PROMPT.md, into the task, and set it up:

SettingValueWhy
PlanFleetA lead and up to seven workers on the deepest models
Factory modeonBounded by progress, not by turns or minutes
Projectcosmos-factoryAgents read its checkout
WriteonThe lead may spawn writers, each in its own git worktree
Use the trackeronThe lead gets the beads tools you granted

The line beside Start swarm says what you are about to run.

The Start a swarm launcher in light mode, expanded. The task holds the brief and the factory prompt. Fleet is selected with Factory mode on, so each plan card reads until work stops instead of a turn budget. The project is cosmos-factory, Write and Use the tracker are on, and the tracker row lists the six beads tools. The summary beside Start swarm reads 8 agents until work stops landing, reading cosmos-factory and writing on a branch, with beads.

Figure 1. The launcher set up for the factory (light UI). With Factory mode on, the plans run until work stops landing, and the tracker row shows the beads tools your grant let through.

Start it, and open the Beads tab in a second window so you can watch both.

The lead spends its first turns on the plan: it reads the beads, posts them for critique, and revises them before spawning writers. Planning counts as progress in Factory mode until the first merge lands, so this phase is safe. Then the crew appears on the live card, each agent with a lane.

A live Fleet factory swarm on the Swarms tab in light mode. The roster reads lead, reviewer, design-critic, w-found, w-design, w-art-solar and w-art-deep. The budget card reads 71 turns, factory, 3 of 15 since work landed. The timeline shows a lane per agent, with the reviewer and the design critic taking short turns between the writers' long ones.

Figure 2. A Fleet factory mid-run (light UI). The lead staffed a code reviewer, a design critic and writers named for their work; the critics’ short turns sit between the writers’ long ones.

  1. The foundation goes first. The scaffold is the only ready bead. A writer builds it in its own worktree, runs the tests and typecheck, and reports its branch and head.
  2. The critic looks at the pages. For anything a visitor sees, the critic opens the writer’s screenshots and sends back one change at a time. In one of our runs it rejected the catalog because the cards were “mostly empty containers”; the writer enlarged the illustrations, recaptured, and the critic approved that exact commit.
  3. The reviewer reads the diff. With the critic’s approval in hand, the lead pastes the bead’s acceptance criteria into a review request, and the reviewer approves the head or names one concrete issue.
  4. The lead merges and closes. The approved head is merged into main with a merge commit naming the writer, and closing the bead turns its dependents ready.
  5. The graph widens. After the scaffold, the catalog, the store and the design system are built side by side. When two branches touch the same file, the second writer rebases and is reviewed again.

On the Beads tab the flow strip moves from waiting to in progress to done. In flight shows each bead a writer holds, and Next up names the bead whose close frees the most work.

The Beads tab for cosmos-factory in light mode, partway through the run. The flow strip reads 5 waiting, 1 in progress, 2 done. In flight shows the catalog bead claimed, with a stage meter at claimed. The backlog lists each waiting bead with the bead it waits on, and Shipped lists the scaffold and the reactions store.

Figure 3. The board draining (light UI). Closed beads move to Shipped, and each close turns its dependents ready.

Factory mode measures progress, not effort. Every merge and every closed bead resets a window of 45 turns on Fleet, and the budget card shows how many turns have passed since work last landed. If a whole window passes with nothing landing, the lead is told to wrap up and the swarm ends as stalled with what it finished. A fresh-token ceiling backstops it, 6M on Fleet; cache reads do not count against it. You can post in the channel at any time to steer.

The swarm ends when the lead concludes, with the closed beads and the command to run the app. The finished row on the Swarms tab opens the report. The history is in git:

Terminal window
git log --oneline --first-parent main
c0bfba2 Merge writer @sgafr-design branch keelson/swarm/sgafr/design into main
0b1e652 Merge writer @sgafr-backend branch keelson/swarm/sgafr/backend into main
63e94f7 Merge writer @sgafr-store branch keelson/swarm/sgafr/store into main
...
ccd82b1 Merge writer @sgafr-backend branch keelson/swarm/sgafr/backend into main
deb3d76 bd init: initialize beads issue tracking

One merge per bead, each naming the writer that built it. Our run took 34 minutes and 138 turns across six agents and finished with 288 tests; yours will differ in all three, and in the app it designs.

Read the conclusion before you trust it. The prompt asks the lead to say what the reviewers pushed back on and what nobody checked, so you know where to look. Then check the claim yourself:

Terminal window
bun install
bun test
bun run typecheck
bun run dev

Open http://localhost:3000. Filter the catalog and reload to see the filter survive, open Jupiter and walk the catalog with the arrow keys, and send chills twice to see the calm second answer. Every one of those behaviors was an acceptance criterion on a bead, so each one was read by a second agent before it merged.

The Cosmos landing page the Fleet factory built. A dark, cinematic hero reads Make room for wonder. beside a glowing Orion Nebula, today's object. Below it, Follow your curiosity. sits over category chips, a search field, and a grid of twelve cards, each with its own illustration: a granulated Sun, a banded Jupiter with its red spot, a ringed Saturn, Mars with a polar cap, a cracked Europa, a hazy Titan, a mottled Betelgeuse, a black hole with a lensed ring, the Orion Nebula, the Pillars of Creation, a spiral Andromeda and the TRAPPIST-1 system.

Figure 4. The landing page from our Fleet run. Every object has its own illustration, and the critic held the catalog to the art direction before it merged. Another run will design a different app from the same beads.

A factory is only worth its cost if it beats one agent on the same brief, so measure that. In a second fresh clone, hand factory/BRIEF.md alone to a single coding agent and let it build. Then compare the two sites, their test suites, and what each one checked.

When we ran that comparison, a strong single agent came close on looks for a site this size, and the factory won on everything around them: several times the tests, a second reader on every change, a critic on every page, and a bead and a merge commit for every unit of work. Because the brief is model-neutral, the pair is a benchmark you can re-run as models improve.

  • Stop it and start again. Stop the swarm partway, then start the same prompt on the same project. The new lead has no memory of the old run, but it needs none: the plan is in the tracker and the finished work is on main, so it picks up the open beads and carries on. Name any unmerged writer branches in the task so the new writers merge them instead of redoing the work.
  • Let the lead plan. Skip the seeding at the last stop and start the same prompt on an empty tracker. The lead cuts its own beads, has them critiqued, and builds from them.
  • Change the graph. Add a bead to factory/backlog.json (a print view, a search-as-you-type, an about page), make it wait on the view it extends, and seed a fresh clone. The factory schedules it without a prompt change.
  • Change the models. Pick a lead and workers model under the plans, or a different provider, and compare the sites.
  • Keep the remote. Leave origin in place and writers open a draft pull request per bead instead of merging locally. Only you merge pull requests, so a bead’s dependents wait for your merge; this is the shape for a repository other people share.

A backlog with dependencies is enough structure for a crew of agents to build an app the way a team does: in order where order matters, side by side where it does not, with a critic looking at every page and a second reader on every change. The harness carried the run as one durable op you could watch, steer, and stop; the beads rib kept the graph and the board; the swarm rib supplied the crew. Every unit of work left a bead with its acceptance criteria and a merge commit in git, so the run is as auditable as it was fast.