The views

Context

What each request carried, and what Iris measures.

Press 4. Context shows the newest turn on the wire, broken into the three things a request is made of, each sized in tokens and priced at list rates.

Context view: one turn drawn to scale, split into system prompt, tool schemas and conversation

What a request is made of

PartWhat it isWho controls it
System promptClaude Code's own operating instructions, plus your environment and CLAUDE.mdPartly you — block 2 is yours
Tool schemasThe full JSON definition of every tool Claude is allowed to callYou, via the deny list
ConversationEvery message so far, including tool_use and tool_result payloadsGrows on its own

The first two together are the fixed prefix. It is re-sent on every single message of the session, whether or not anything in it is used. That is the number worth reading.

The system prompt, block by block

Click System prompt. Iris shows every block it saw, in full, with its size and whether it carries cache_control. Message blocks are previewed at 300 characters; the system prompt deliberately is not, because you cannot audit a cost you are only shown the first paragraph of. The only cap is a 200k-character guard against a pathological payload.

System prompt blocks with size and cache flags

What to look for in your own:

Tool schemas

The Tool schemas tab lists every definition with its token weight and call count, and carries a trim simulator that models a change before you commit to it. Acting on what you find there is Optimize.

Tool schema list with token weight, call counts and the trim simulator

Conversation and changes

Conversation shows every message block including tool payloads — often the largest single source of growth in a long session. Changes diffs the prefix between two live turns and attributes the growth to a source, which is how you catch a prefix that quietly grew mid-session.

Caching, and why the dollar figures move

Cached blocks are billed at write price once, then read back at roughly 10% of input price until the TTL expires. The dashed lines in the ribbon are cache breakpoints. A stable prefix is therefore much cheaper than its token count suggests — and a prefix that changes every session is much more expensive than it looks.

The TTL is not a constant. It is an hour on a subscription seat, and five minutes on an API key or once a subscription starts drawing on usage credits. Iris reads the TTL each write actually used, so the cache-expiry figures hold whichever you are on, and it states the lifetime for your mode rather than leaving you to work out which applies.

What a dollar figure means on your plan

Claude Code meters the same tokens four different ways, and which one applies to you decides whether a figure on the dashboard is your bill, a ceiling, or the wrong rate card entirely. A flat monthly subscription is the case that breaks naïvely-priced tooling: nobody on a Max seat is billed per token, so a real measurement printed as $/mo reads as an invented one.

ModeHow it is meteredCache TTLWhat a $ figure means
Subscription seat
Pro · Max · Team · Enterprise
Rolling 5-hour and weekly allowance, shared with Claude chat and Cowork 1 hour Not a bill — what the same traffic would cost on API rates
Usage credits Per token at list rates, against your monthly spend limit 5 min Literal. This is the charge
Console API key Per token, billed to your Console workspace 5 min Literal, before any contracted discount
Cloud provider
Bedrock · Vertex · Foundry
Per token to your cloud account, at that partner's rates 5 min Withheld — Iris does not carry partner rate cards

Iris classifies the mode from the shape of the credential on the wire: an x-api-key header is a Console key, an Authorization: Bearer token is a signed-in seat, and a partner upstream is a cloud provider. The credential itself is never read, stored, or exposed — only which kind is present.

One thing the wire cannot tell you is whether a seat is still inside its allowance or has already crossed into usage credits. Nothing in the request says so, so Iris reports the seat and leaves the switch to you, in the header chip. Choosing it also shortens the cache lifetime Iris assumes, which makes idle gaps cost more.

Each mode changes how figures are labelled rather than what was measured. On a seat, Optimize leads with tokens per turn and the share of your metered usage, and puts the dollars in brackets as the API-rate equivalent. On a key, the dollars lead. On a cloud provider the money columns are dropped rather than filled from a rate card that does not apply. The ? button in the header explains the whole model, including what Iris cannot determine.

Iris cannot see your remaining allowance

Seat quotas are not published as token counts and are not on the wire. Run /usage in Claude Code for the real bars. Everything plan-relative in Iris is a share of your own measured usage, which needs no quota to compute.

What Iris measures, and what it does not

It measures: the exact bytes of each system block, tool schema and message on the wire; the usage figures Anthropic returns; cache reads, writes and expiries; latency and status per call.

It estimates: token counts. These start as chars/4 and are then calibrated — Iris divides the measured input total from the response by its own estimate and scales every figure by that ratio, so the numbers track what you were billed rather than a raw guess. They are close, not exact.

It does not measure: anything about the model's reasoning, output quality, or what Claude would have done differently. It does not know your contract — dollars are list rates for known models, so promotional pricing, negotiated discounts and seat allowances are not reflected, and an unknown model stays unpriced rather than being printed as $0.00. It classifies your billing mode, which is what decides whether those dollars are a bill at all, but not your rate. Claude Console is the invoice.

The research behind the numbers

The Token Tax walks through one full measurement — how a 28.0k baseline became 5.6k, and what stayed enabled.