Usage and billing
Use the Keiro console to understand request volume, token usage, billable amounts, limits, and account billing. API response headers remain authoritative for the limits that applied to one request.
On this page 1 of 9
Choose the right view#
| Question | Console view | What it shows |
|---|---|---|
| How much have we used? | Usage | Requests, input and output tokens, failures, rate-limited requests, and spend over a selected window |
| Which model or key generated it? | Usage drilldowns | Request and token volume grouped by model or API key |
| What happened to one request? | Logs | Request ID, endpoint, model, status, token totals, latency, and safe error context |
| What are the current caps? | Limits | Organization limits plus narrower API-key overrides |
| What is the billing account state? | Billing | Subscription, payment method, invoices, adjustments, refunds, and usage export when available to your role |
The console scopes every view to the selected organization. Check that organization before comparing totals or exporting data.
Understand the numbers#
Keiro reports canonical input and output token totals for the public model request. When a completed request is billable, its billing record carries the amount used for customer billing.
Billing records can be pending while completion or reconciliation is still in progress and finalized after settlement. Recent totals can therefore move from pending to finalized without representing a second request.
Request-log amounts are operational context. Use billing-record totals for billable spend and invoices for the amount due. Customer-managed provider credentials are not charged as Keiro per-token usage when the account's active billing mode excludes that charge.
Hosted web search is billed per search request, separately from tokens. The Keiro CLI and Console offer it by default and leave the decision to the model; API callers only get it when they include the tool. Searched requests show the search count alongside token totals in Logs, and the charge appears in the billing record's amount.
Token counting#
Every eb1 model counts tokens in one unit, the o200k_base meter, regardless of which model served the request. The same request body always yields the same prompt_tokens (input_tokens on the Messages and Responses surfaces), completion_tokens / output_tokens is the meter count of the output you received, and total_tokens is their sum. Billing, each model's input window, /v1/messages/count_tokens, and rate-limit token accounting use the same unit, so a count you compute locally with tiktoken's o200k_base encoding matches what you are charged.
How the prompt is counted: every text part of the request (instructions or system prompt, each message part, each tool-call argument string and tool result) is counted independently, tool definitions are counted once as canonical JSON, and image and file parts carry fixed per-part minimums. The output count covers each text part, refusal, and tool-call argument string of the response.
Cache-read and cache-write counts are reported as subsets of the metered prompt in the same proportion the serving model observed, so input_tokens + cache_read_input_tokens + cache_creation_input_tokens still equals the full metered prompt on the Messages surface.
Reasoning tokens#
eb1 models think before they answer. The thinking is not returned, but it is counted: output_tokens (completion_tokens on Chat Completions) is the text you received plus the reasoning tokens, itemized in output_tokens_details.reasoning_tokens on the Responses surface and completion_tokens_details.reasoning_tokens on Chat Completions. On the Messages surface output_tokens includes reasoning, as it does on Anthropic's own API, and the split is shown in Logs and in the usage export. Reasoning tokens are the count the serving model reports for its own thinking; because no text stands behind them they are not converted between tokenizers, and they are billed at the model's reasoning rate while the rest of the output bills at the output rate.
Every request also has a minimum charge set by the compute that served it. When the input and output charges fall below that minimum, the reasoning line is raised to meet it, so reasoning tokens can exceed the length of the reply, and on short requests they usually do. The amount on a billing record is always: uncached input x input rate + cached input x cache-read rate + (output - reasoning) x output rate + reasoning x reasoning rate. Rate limits count only what you sent and received; reasoning tokens do not consume them.
Prompt cache pricing#
Effective 2026-08-08, prompt tokens served from a prompt cache are billed at a discount on every public model. Rates are multiples of the model's standard input rate:
| Token class | Rate |
|---|---|
| Input, uncached | 1x input rate |
| Input, cached read | 0.1x input rate |
| Input, cache write | 1x input rate (no write surcharge) |
| Output | 1x output rate |
The usage object already reports the cached portions of the prompt (cache_read_input_tokens and cache_creation_input_tokens on the Messages surface, cached_tokens details elsewhere); the discount applies to those reported counts. No request change is needed — requests that reuse a cached prompt prefix are billed the discounted amount automatically, and the billing record's amount reflects it.
Export usage#
Use the console's CSV export for customer-safe billing records. Export rows can include:
Exports do not include raw API keys, prompt or response bodies, private routing details, provider credentials, or internal cost and margin data.
- record ID and non-secret API-key ID
- status
- input and output token totals
- hosted web search count, when the request searched
- billable amount
- creation and finalization timestamps
Reconcile an application request#
Keep the response request ID with your application's own operation ID. When you investigate a discrepancy:
Do not send the API key or full prompt/response body in the first support message.
- Match the request ID and timestamp in Logs.
- Confirm the endpoint, requested public model, status, and token totals.
- Check whether the associated billing record is pending or finalized.
- Compare the active price and plan information in Billing.
- Contact support with safe identifiers if the finalized record still does not match your expectation.
Limits and denials#
Rate and spend limits are admission controls, not invoice totals. A request can be denied before dispatch when it would exceed a configured limit. A 429 response includes a machine-readable error and, when applicable, Retry-After; waiting does not reset a daily or monthly spend cap early.
See Rate and spend limits for header and retry contracts.