Skip to content

Usage

Chat and workflows tell you what an agent did. Usage tells you what each turn consumed: every turn the harness runs, on any surface, is written to one ledger, and the Usage tab renders that ledger as six read-only views. It answers a question the other surfaces cannot give you on their own, where the tokens went, and it needs no rib. A running provider is enough to start filling it.

Usage measures token counts first. Everything on the tab is a count of tokens, a rate, or a turn tally. Alongside the counts it reports an estimated dollar cost, derived at read time from a price table, and a cache hit ratio. Neither is a bill: see Cost and cache hit ratio for exactly what they are.

The backbone is a single SQLite table, the usage_events ledger. Every agent turn that carries real token spend writes exactly one row to it, recorded from one of three capture seams:

SeamSourceWritten when
A chat turnchatYou send a message and the agent answers.
A workflow prompt nodeworkflowA prompt node in a run calls the model.
A rib agent turnribA rib runs its own agent turn through the harness.

A row records the provider, the model, input and output token counts, cache read and write counts, the turn’s duration and status, and attribution back to the conversation, run, node, workflow, or rib that produced it. A turn that carried no spend, a content-only reply or a zero-total usage report, writes nothing: an empty row would only distort a rollup. The UI labels the ledger’s row count agent turns.

A row names the served model, the model that actually ran the turn, which is not always the one you asked for. A provider can resolve a model alias on its own: Copilot’s default is auto, and it picks a concrete model session-side per call. When a turn resolves to a real model, that resolved name is what the ledger stores, so the model roster and the ledger read as what ran rather than what was requested. A turn that never resolves past the alias shows the alias, which the roster marks plainly as unresolved.

Every view reads the same ledger through a windowed query, and none of them writes. The tab’s overview groups the first three; the model roster, jobs, and ledger each get their own sub-view.

ViewShows
PulseThe window’s tokens split by type, cost and the price sources behind it, turns, cache hit ratio, and failure burn, plus a trailing-hour tokens-per-minute sparkline streamed live.
Over timeA stacked bar chart of tokens or cost by served model across the window’s buckets.
Source to model flowA ribbon diagram tracing which source (a chat, a workflow, a rib) drove which model, sized by tokens or cost.
Model rosterPer model, most expensive first: turns, tokens, cost, and cache hit ratio, with paired bars that split the tokens and the cost by token type. Under each model sits its price card, the per-million-token rates for cache read, input, cache write, and output in the bars’ order, and where they came from.
JobsPer recurring workflow or rib job (a rib job reads rib:<id>), most expensive first: runs, cost, cost per run, and the model that incurred most of the cost. A rib that records its turns without a run id is counted in turns instead of runs.
LedgerThe raw rows, newest first, each with its four token counts and priced cost, filterable by source, model, and status. Hovering a cost shows the rate behind each part of it. It loads 50 rows at a time, up to the latest 500.

Every view runs over one of three fixed windows, 24h, 7d, or 30d. The window is the only time control: a caller names it and the server derives the lookback bound from it. There is no way to pass an arbitrary start and end, which keeps the window and the bound from ever disagreeing. The 24h window charts by hour and the wider windows by day, so a legible number of buckets fills the axis either way.

Pulse is the one view that does not poll. It rides the same snapshot substrate every rib surface uses: the server registers a built-in composer under USAGE_PULSE_SNAPSHOT_KEY that returns today’s running totals plus a zero-filled trailing-sixty-minute series, and the browser subscribes to it. A turn’s burst of ledger writes coalesces into a single recompose before it broadcasts, so a busy minute lands as one frame rather than a storm. Everything else on the tab is a plain read of a /api/usage endpoint over the selected window.

Every aggregate on the tab carries two derived figures next to its token counts, and every ledger row carries the first of them, its own cost.

Cost is an estimate in US dollars: each event’s token counts multiplied by the per-million-token rates of the model that served it, summed over the rows in view. The harness bundles the Anthropic first-party list prices for the Claude models its Claude provider offers (input, output, cache read, and 5-minute cache write rates), and recognizes the same models under the dotted ids the Copilot catalog uses. Copilot models are priced from the per-token rates in the account’s live catalog once it has loaded. Every catalog rate is also saved in the database, so a model that leaves the catalog keeps pricing its history at the last rate seen, and Copilot turns stay priced before the catalog loads. Price any other model yourself with the modelPrices key in config.json, which also overrides a catalog or bundled rate.

Cost is computed when you read it, never stored. A corrected price reprices history on the next load, and nothing in the ledger has to be rewritten. One caveat on old history: rows written before the providers normalized their counts may carry an input figure that already includes cache tokens, and the ledger has no per-row marker to tell them apart, so the cost of those rows can overstate by the cache portion. Rows written since are priced exactly.

A null cost means unpriced, and it is never shown as zero. When any event in an aggregate ran on a model with no price, the whole aggregate’s cost is null and its unpricedEvents count says how many rows kept it that way. pricedCostUsd (pricedTotalCostUsd on a job) still carries the cost of the rows that were priced, so the tab renders a mixed window as a floor such as ≥ $155.03 with the unpriced count still visible, and a window with no priced rows at all as unpriced. Zero-cost priced rows still count as priced, so a mixed window whose priced subtotal is zero reads ≥ $0.0000, and one whose priced subtotal is positive but under $0.0001 reads > $0.0000. A fully priced window that reads $0.0000 really did cost nothing, and a window on an unpriced model never looks free. The same rule holds for a job: one unpriced row nulls both its cost per run and its window cost. Both show floors with the unpriced count when any rows were priced, or unpriced (N) when none were. Jobs report pricedEvents separately from runs because one run can contain multiple ledger rows. A Copilot turn that never resolved past the auto alias is unpriced, since the served model is unknown.

Cache hit ratio is cache-read tokens divided by fresh input plus cache-read tokens, the share of a turn’s prompt the provider served from its prompt cache. A provider that reports a cache count of zero records zero, so an all-miss window reads as 0%. The ratio is null, rendered as a dash, only when no event in the aggregate carried a cache-read count at all (a provider with no cache accounting) or when there was nothing to divide by, so such a provider never shows a 0% hit rate it did not earn.

The chat composer’s usage popover shows the same two figures for the open conversation, read from the ledger rather than priced in the browser, so each turn is priced at the model that actually served it even across a model swap.

Usage records what turns did run and shows you the shape of it; it never stops a turn.

Ceilings live elsewhere. The budget gate is a separate feature that denies a turn at request time once a running token or turn count crosses a limit on an expensive model. That is admission control, deciding whether a turn may run; Usage is the record of what already did. They share a vocabulary of tokens and nothing else.

  • Governance is the budget gate, the admission-control counterpart to this observability ledger.
  • Snapshots and surfaces is the substrate the live pulse streams over.
  • Configuration documents the modelPrices key that prices models the bundled table does not cover.
  • Chat is where a turn’s usage also surfaces inline, on the composer usage chip.
  • The API reference documents the read-only /api/usage endpoints behind every view.