One workflow, many models
Every team that adopts coding agents ends up in the same argument: which model should do which job? Strong models are expensive on grind work, fast models are reckless on judgment calls, and everyone has an opinion. The useful move is to stop debating and run the experiment. That is what this field exercise does: one workflow builds a real web app, each phase routed to a different model class, and swapping any phase to a different model (or a different provider) is a one-line edit away.
Until now the rail drove one coding agent. This page is the graduation run: several models at once, each phase routed to a different one, and a real app in your working tree at the end. Budget for it. The pipeline is ten nodes, and a couple of them (build-ui, integrate) do genuine, long-running work, so a full run takes tens of minutes, not seconds, and spends real model turns the whole way. The fast-model phases (explore, validate, smoke) hold the cost down; the swap experiment at the end runs the build a second time, so it roughly doubles both.
Get the template
Section titled “Get the template”danielscholl/keelson-sample
is a GitHub template repository: a ready-made keelson project carrying one
workflow and a spec to build from. Create your own copy and clone it (the
Use this template button on GitHub does the same thing):
gh repo create my-frontend-mix --template danielscholl/keelson-sample --private --clonecd my-frontend-mixTwo things in the tree matter for this exercise:
| Path | What it is |
|---|---|
.keelson/workflows/frontend-mix.yaml | The workflow: a ten-node DAG with a model pinned per phase |
spec.md | The input brief: “Cosmos”, a no-login, local-only planetarium app |
The workflow routes each phase to a model class. Fast models take the grind work (explore, validate, smoke), strong-reasoning models take the judgment calls (plan, integrate, fix, deploy), and a UI-strong model takes the design step. Here is the same file as the engine sees it:
Figure 1. The frontend-mix DAG. Each node pins a model class; the brass node is a human approval gate, and every edge hands off a markdown artifact rather than shared chat history.
The handoff is the trick to stare at. Each node writes its conclusions to a markdown file in the run’s artifact directory, and the next node opens that file with its own fresh context. Two models that have never seen each other’s conversation cooperate cleanly because the file is the contract. That is also exactly why the models are swappable: no phase depends on another phase’s chat history, only on its artifact.
Register the project
Section titled “Register the project”Your clone carries its workflow in .keelson/workflows/, but the server
does not scan every directory on your disk. You tell it which directories
are projects; the catalog then includes each project’s workflows alongside
the global ones from ~/.keelson/workflows:
keelson project add frontend-mix "$(pwd)"From now on, any run you start from inside this directory resolves to this
project, and the Workflows surface at http://127.0.0.1:7878 shows
frontend-mix as a card under the project’s scope.
Before running anything, validate. workflow validate reads the global
catalog by default; point it at the project’s directory with --dir:
keelson workflow validate frontend-mix --dir .keelson/workflowsresults: - filename: /Users/you/my-frontend-mix/.keelson/workflows/frontend-mix.yaml ok: true warnings: error:failed: 0total: 1Run the build
Section titled “Run the build”The workflow takes its brief through the ARGUMENTS input. Point it at the
bundled spec and say the deployment is local-only, so nothing tries to leave
your machine:
keelson workflow run frontend-mix --watch \ --inputs ARGUMENTS="Build the app described in spec.md. Local only - no deployment this run."The watch stream shows the DAG executing in dependency order, exactly as in
Run and author a workflow, except now each ✓ is a real
model finishing a real phase:
◆ target: cwd=/Users/you/my-frontend-mix▶ run 096d8e56 (frontend-mix) ✓ explore ✓ plan ✓ build-ui ✓ integrate ✓ validate ✓ fix-validation ✓ smoke · approve-deploy …This is the long part: the build-ui and integrate phases are doing genuine
work. Watch the same run as live cards in the Workflows surface if you
prefer; both views are the same stream. Note the run id in the header (here
096d8e56); you need it to answer the gate.
When the smoke phase finishes, the run does something the earlier tutorials never did: it stops and waits for you.
⏸ approve-deploy awaiting approval — Build, validation, and smoke test arecomplete. Review the smoke verdict above (and try the app locally if youlike) before anything leaves this machine.This is the approval node from the taxonomy table, now in earnest. The
build is done and validated, the smoke test has driven the running app, and
nothing deploys until a human says so. Read the smoke verdict, then answer
from the Workflows surface, or from a second terminal:
keelson workflow respond <runId> approve-deploy "skip"(The run id is in the watch header. “deploy” runs the plan’s deployment section; “skip” records your sign-off and finishes without deploying. With a local-only plan both end the same way, but the gate is the point: you just held a multi-model pipeline at a human checkpoint with one line of YAML.)
The final node is plain bash: it copies the run’s artifact directory into
.agents/artifacts/ and commits everything, because the per-run scratch
directory is deleted when the run completes. Your repo now holds both the
app and the paper trail.
Read the trail
Section titled “Read the trail”ls .agents/artifactscontext.mddeploy-summary.mdintegration-summary.mdplan.mdresolution-summary.mdsmoke-summary.mdui-summary.mdvalidation-summary.mdYour trail may run a file or two shorter. Each phase is told to leave a
summary, but a model can end its turn without ever calling the Write tool,
and nothing gates on the file existing yet. The two that matter for reading
the run are the plan and the smoke verdict.
Open plan.md first. The strong-reasoning model wrote the actual headlines
and button labels the UI model built from (SECTION A), the integration scope
the integrate phase implemented (SECTION B), and the deploy decision the
gate protected (SECTION C). Then open smoke-summary.md for the verdict on
the running app. This is the phase that catches “compiles green, renders
broken”: a build that typechecks, lints, and serves HTTP 200 but renders
blank or unstyled (a Tailwind content glob that misses a directory does
exactly this). Static checks pass it; only driving the running app, the
smoke phase’s whole job, catches it.
One thing the pipeline never does is worth noticing before you move on: it never questions that plan. A single model wrote the contract every later phase trusts, and a straight line has nowhere to review it, so a wrong assumption gets implemented faithfully and validated against itself, green the whole way down. Catching that takes a different shape of work, where every change gets a second reader, and that shape is where the rail goes after this page.
The app itself is in your working tree, validated and committed. Run it:
bun run devSwap a model and compare
Section titled “Swap a model and compare”Now the experiment the setup paid for. Open
.keelson/workflows/frontend-mix.yaml and find the build-ui node: it pins
the UI-strong model. Change that one line to any other model your provider
serves, then re-run with --worktree:
keelson workflow run frontend-mix --worktree --watch \ --inputs ARGUMENTS="Build the app described in spec.md. Local only - no deployment this run."--worktree runs the build in an isolated git worktree on its own branch
(keelson/frontend-mix/<run-id>), so your first build stays untouched and
the comparison is a branch diff:
git branch --list 'keelson/*'git diff main keelson/frontend-mix/<run-id> -- app/Same spec, same prompts, same pipeline, different model on one phase: the difference you see in that diff is the model, isolated. That is the whole test bench. From here you can take it further:
- Swap whole providers. Enable Claude, Codex, or Pi in
~/.keelson/config.jsonand add aprovider:line to any node. The configuration guide covers enablement; the providers reference covers what each one needs. - Pin the whole chain to one provider with
KEELSON_WORKFLOW_PROVIDERto compare vendor against vendor instead of model against model. - Rewrite the spec.
spec.mdis a content-and-intent brief, not a wireframe. Describe your own app and the same pipeline builds it. - Work the backlog. Your clone ships
backlog.md: scoped improvements to the app the run just built, written against the spec rather than any one build, since yours differs from everyone else’s. Hand one to a chat session, or capture it as a run with the bundled plan-act-evaluate or fix-issue workflows. - Drive it by hand. The same repo ships
.agents/skills/frontend-mix-*skills that walk this chain one step at a time from a chat session, so you switch models between steps and inspect each handoff yourself. The sample’s README has the walkthrough.
What you proved
Section titled “What you proved”A workflow file can encode a team’s model-routing policy: who plans, who builds, who checks, where a human signs off. The artifact handoff makes the phases independent enough that swapping a model is a one-line diff, and worktree isolation makes comparing attempts a git operation instead of a copy-paste archaeology dig. The harness carried the state, the streams, and the approval gate; the models just did their phases.
The authoring workflows guide covers the full node vocabulary when you are ready to encode a pipeline of your own, and Workflows explains the engine underneath this page.
Where the rail continues
Section titled “Where the rail continues”This is the last stop that needs the harness alone. It leaves you holding an observation the pipeline could not act on: one model wrote the plan, every later phase trusted it, and nothing in a straight line could split the work, run pieces side by side, or have a second agent read each change before it landed.
The next two stops build the same app again in that other shape. Plan the backlog with beads turns the Cosmos build into a graph of small, dependent units of work, and Run the software factory hands that graph to a swarm of agents that work it in dependency order, review every change, and merge it. Both add a rib; the harness you just proved carries them.