Skip to content

investigate

investigate answers a focused factual question from source and read-only runtime probes. One model gathers evidence, another re-checks every load-bearing claim, and deterministic code tallies the verdicts and records which provider and model actually ran each role.

Use it when the durable result should be an evidence file rather than a chat answer. It does not change source or post the result elsewhere.

Terminal window
keelson workflow run investigate \
--arguments "Which function validates workflow output, and when does it run?" \
--inputs out=evidence/output-validation.md

The question and out are required. Optional inputs refine the run:

Terminal window
keelson workflow run investigate \
--arguments "Which function validates workflow output, and when does it run?" \
--inputs out=evidence/output-validation.md \
--inputs context=packages/workflows/src \
--inputs access="Use only local source and read-only shell commands." \
--inputs fixtures=evidence/fixtures/output-validation \
--inputs tier=std \
--inputs verifier=grok

context accepts one readable file or a directory. A directory contributes its regular files one level deep in sorted order. access accepts either a readable file or literal guidance. fixtures names a directory for raw probe responses: intake creates it, and the investigator saves each runtime probe’s output there as its own file and cites it, which keeps large payloads out of the evidence file. tier is deep or std; verifier is claude or grok. When verifier is absent, intake selects one deterministically from the run id.

Five nodes run in sequence:

NodeKindWhat it does
intakebashValidates the question, output path, fixtures directory, tier, and verifier; bundles context and access guidance; rotates the default verifier from the run id.
investigatepromptReads source and runs read-only probes, then writes the requested evidence file with cited claims and NOT CHECKED placeholders.
verifypromptRe-reads or re-runs every load-bearing claim independently and returns one closed verdict per checked claim.
publishpromptUpdates only the evidence table’s Verified and Verification note cells. It does not count claims or write attribution.
finalizescriptParses the Markdown table, validates its levels and verdicts, computes the tally, and writes effective run attribution.

The investigator and verifier use closed model_by maps. Copilot seats them on different model vendors. If the run falls back to a provider that can only serve one vendor, Keelson emits a run warning after comparing the provider and model recorded for the turns that actually ran.

The evidence file contains one or more claim tables with these columns:

ColumnValues
Claim #A unique positive integer used to associate a verifier verdict with its claim.
ClaimOne factual statement.
Leveldocumented, source-verified, or runtime-tested.
CitationA source path:line or the exact read-only probe and relevant output.
VerifiedCONFIRMED, CONFIRMED in part, REFUTED, UNVERIFIABLE, or NOT CHECKED.
Verification noteThe verifier’s reason, including partial or contrary evidence.

finalize finds columns by header name, so their order may change. It rejects unknown levels, unknown verdicts, malformed rows, and missing runtime attribution. It then writes one tally in fixed order and an attribution line using the effective models and providers, including provider fallback or a provider-reported model switch. Re-running it replaces its prior generated section instead of adding another.

  • Closed runtime selection. tier and verifier select only declared model_by cases.
  • Effective independence diagnostics. different_vendor_from compares the models that ran, not the seats requested in YAML.
  • Deterministic finalization. A script owns parsing, tallying, run attribution, and idempotent replacement of generated output.