Usage
Chat and workflows tell you what an agent did. Usage tells you what each turn consumed: every turn the harness runs, on any surface, is written to one ledger, and the Usage tab renders that ledger as six read-only views. It answers a question the other surfaces cannot give you on their own, where the tokens went, and it needs no rib. A running provider is enough to start filling it.
Usage measures token counts first. Everything on the tab is a count of tokens, a rate, or a turn tally. Alongside the counts it reports an estimated dollar cost, derived at read time from a price table, and a cache hit ratio. Neither is a bill: see Cost and cache hit ratio for exactly what they are.
One ledger, three capture seams
Section titled “One ledger, three capture seams”The backbone is a single SQLite table, the usage_events ledger. Every agent turn that carries real token spend writes exactly one row to it, recorded from one of three capture seams:
| Seam | Source | Written when |
|---|---|---|
| A chat turn | chat | You send a message and the agent answers. |
| A workflow prompt node | workflow | A prompt node in a run calls the model. |
| A rib agent turn | rib | A rib runs its own agent turn through the harness. |
A row records the provider, the model, input and output token counts, cache read and write counts, the turn’s duration and status, and attribution back to the conversation, run, node, workflow, or rib that produced it. A turn that carried no spend, a content-only reply or a zero-total usage report, writes nothing: an empty row would only distort a rollup. The UI labels the ledger’s row count agent turns.
The served model, not the requested one
Section titled “The served model, not the requested one”A row names the served model, the model that actually ran the turn, which is
not always the one you asked for. A provider can resolve a model alias on its
own: Copilot’s default is auto, and it picks a concrete model session-side per
call. When a turn resolves to a real model, that resolved name is what the ledger
stores, so the model roster and the ledger read as what ran rather than what was
requested. A turn that never resolves past the alias shows the alias, which the
roster marks plainly as unresolved.
Six views over one ledger
Section titled “Six views over one ledger”Every view reads the same ledger through a windowed query, and none of them writes. The tab’s overview groups the first three; the model roster, jobs, and ledger each get their own sub-view.
| View | Shows |
|---|---|
| Pulse | The window’s tokens split by type, cost and the price sources behind it, turns, cache hit ratio, and failure burn, plus a trailing-hour tokens-per-minute sparkline streamed live. |
| Over time | A stacked bar chart of tokens or cost by served model across the window’s buckets. |
| Source to model flow | A ribbon diagram tracing which source (a chat, a workflow, a rib) drove which model, sized by tokens or cost. |
| Model roster | Per model, most expensive first: turns, tokens, cost, and cache hit ratio, with paired bars that split the tokens and the cost by token type. Under each model sits its price card, the per-million-token rates for cache read, input, cache write, and output in the bars’ order, and where they came from. |
| Jobs | Per recurring workflow or rib job (a rib job reads rib:<id>), most expensive first: runs, cost, cost per run, and the model that incurred most of the cost. A rib that records its turns without a run id is counted in turns instead of runs. |
| Ledger | The raw rows, newest first, each with its four token counts and priced cost, filterable by source, model, and status. Hovering a cost shows the rate behind each part of it. It loads 50 rows at a time, up to the latest 500. |
Windows are fixed
Section titled “Windows are fixed”Every view runs over one of three fixed windows, 24h, 7d, or 30d. The window is the only time control: a caller names it and the server derives the lookback bound from it. There is no way to pass an arbitrary start and end, which keeps the window and the bound from ever disagreeing. The 24h window charts by hour and the wider windows by day, so a legible number of buckets fills the axis either way.
The pulse streams live
Section titled “The pulse streams live”Pulse is the one view that does not poll. It rides the same snapshot substrate
every rib surface uses: the server registers a built-in composer under
USAGE_PULSE_SNAPSHOT_KEY that returns today’s running totals plus a zero-filled
trailing-sixty-minute series, and the browser subscribes to it. A turn’s burst of
ledger writes coalesces into a single recompose before it broadcasts, so a busy
minute lands as one frame rather than a storm. Everything else on the tab is a
plain read of a /api/usage endpoint over the selected window.
Cost and cache hit ratio
Section titled “Cost and cache hit ratio”Every aggregate on the tab carries two derived figures next to its token counts, and every ledger row carries the first of them, its own cost.
Cost is an estimate in US dollars: each event’s token counts multiplied by
the per-million-token rates of the model that served it, summed over the rows in
view. The harness bundles the Anthropic first-party list prices for the Claude
models its Claude provider offers (input, output, cache read, and 5-minute cache
write rates), and recognizes the same models under the dotted ids the Copilot
catalog uses. Copilot models are priced from the per-token rates in the account’s
live catalog once it has loaded. Every catalog rate is also saved in the
database, so a model that leaves the catalog keeps pricing its history at the
last rate seen, and Copilot turns stay priced before the catalog loads. Price any other model yourself with the
modelPrices key in config.json,
which also overrides a catalog or bundled rate.
Cost is computed when you read it, never stored. A corrected price reprices history on the next load, and nothing in the ledger has to be rewritten. One caveat on old history: rows written before the providers normalized their counts may carry an input figure that already includes cache tokens, and the ledger has no per-row marker to tell them apart, so the cost of those rows can overstate by the cache portion. Rows written since are priced exactly.
A null cost means unpriced, and it is never shown as zero. When any event in
an aggregate ran on a model with no price, the whole aggregate’s cost is null and
its unpricedEvents count says how many rows kept it that way. pricedCostUsd
(pricedTotalCostUsd on a job) still carries the cost of the rows that were
priced, so the tab renders a mixed window as a floor such as ≥ $155.03 with the
unpriced count still visible, and a window with no priced rows at all as
unpriced. Zero-cost priced rows still count as priced, so a mixed window whose
priced subtotal is zero reads ≥ $0.0000, and one whose priced subtotal is
positive but under $0.0001 reads > $0.0000. A fully priced window that reads
$0.0000 really did cost nothing, and a window on an unpriced model never looks
free. The same rule holds for a job: one unpriced row nulls both its cost per run
and its window cost. Both show floors with the unpriced count when any rows were
priced, or unpriced (N) when none were. Jobs report pricedEvents separately
from runs because one run can contain multiple ledger rows. A Copilot turn that
never resolved past the auto alias is unpriced, since the served model is unknown.
Cache hit ratio is cache-read tokens divided by fresh input plus cache-read tokens, the share of a turn’s prompt the provider served from its prompt cache. A provider that reports a cache count of zero records zero, so an all-miss window reads as 0%. The ratio is null, rendered as a dash, only when no event in the aggregate carried a cache-read count at all (a provider with no cache accounting) or when there was nothing to divide by, so such a provider never shows a 0% hit rate it did not earn.
The chat composer’s usage popover shows the same two figures for the open conversation, read from the ledger rather than priced in the browser, so each turn is priced at the model that actually served it even across a model swap.
Observability, not admission control
Section titled “Observability, not admission control”Usage records what turns did run and shows you the shape of it; it never stops a turn.
Ceilings live elsewhere. The budget gate is a separate feature that denies a turn at request time once a running token or turn count crosses a limit on an expensive model. That is admission control, deciding whether a turn may run; Usage is the record of what already did. They share a vocabulary of tokens and nothing else.
Where to go next
Section titled “Where to go next”- Governance is the budget gate, the admission-control counterpart to this observability ledger.
- Snapshots and surfaces is the substrate the live pulse streams over.
- Configuration documents the
modelPriceskey that prices models the bundled table does not cover. - Chat is where a turn’s usage also surfaces inline, on the composer usage chip.
- The API reference documents the read-only
/api/usageendpoints behind every view.