Errors

Keiro returns HTTP errors in the dialect of the surface you called: the Chat Completions and Responses endpoints use the OpenAI-compatible shape shown below, and the Messages endpoint wraps the same content in the Anthropic envelope ({"type": "error", "error": {"type": ..., "message": ...}}), matching its streaming error events. The message, code, and details fields are identical across surfaces for the same failure. Handle errors by status code and retry only where it is safe for your workload.

On this page 1 of 38

Handle status codes#

StatusMeaningRetry
400Invalid request or unsupported parameterNo
401Missing or invalid API keyNo
403Authenticated but not allowedNo
404Route or model not foundNo
409Conflicting request (idempotency reuse, capability conflict)No
413Request body too largeNo
415Request body must be application/jsonNo
429Rate limit or spend capVaries by code; honor x-should-retry and retryable
500Internal errorYes, with backoff
502Upstream dependency failureYes, with backoff
503Service or capacity unavailableYes, with backoff
504Gateway deadline exceededYes, with backoff

Read the error body#

Error bodies follow an OpenAI-compatible shape when the failure can be classified safely. Treat message text as diagnostic context, not as a stable machine contract.

JSON payload JSON
{
  "error": {
    "message": "Model 'eb1-frontier-preview' is not available for your account.",
    "type": "forbidden",
    "param": null,
    "code": "model_not_available",
    "request_id": "req_9f2c41d0",
    "retryable": false
  }
}

Match on code and type for machine handling. param is null unless the failing parameter is known, in which case it names the offending field. request_id is the same correlation id as the X-Request-Id response header, and retryable says whether an unchanged retry can succeed (see Rate and spend limits for the matching x-should-retry and retry-after-ms headers).

For 409 idempotency conflicts, the response includes code idempotency_key_conflict and an X-Original-Request-Id header. See Idempotency.

Limit and failure codes#

The gateway classifies limit denials and upstream failures into a fixed set of machine-readable codes. Match on code and type; when several rows share a code, details.dimension usually distinguishes them. The exception is one pair: a spend-cap denial and an upstream-provider passthrough both use code rate_limit_exceeded with no dimension. Until the legacy details.scope key is retired (see Error details below), tell those two apart by that key -- a spend-cap denial carries details.scope spend, while an upstream passthrough carries no scope.

codetypeStatusRetriabledetails.dimensionMeaning
rate_limit_exceededrate_limit_error429YesrequestsRequests-per-minute limit for your organization was exceeded.
rate_limit_exceededrate_limit_error429YestokensTokens-per-minute limit for your organization was exceeded.
rate_limit_exceededrate_limit_error429YesconcurrencyConcurrent-request limit for your organization was exceeded.
rate_limit_exceededrate_limit_error429Notoken_budgetA per-request token budget from your plan was exceeded.
request_exceeds_capacityinvalid_request_error400Notoken_budget_requestThe request's token estimate exceeds a full window of your plan's budget; it cannot succeed as sent. Reduce the input or max_output_tokens, or raise the limit.
rate_limit_exceededrate_limit_error429YesA configured spend cap for your organization was reached.
rate_limit_exceededrate_limit_error429Yesspend_dailyThe daily (UTC) spend cap configured for your organization was reached.
rate_limit_exceededrate_limit_error429Yesspend_monthlyThe monthly (UTC) spend cap configured for your organization was reached.
rate_limit_exceededrate_limit_error429Nospend_requestThe request's estimated cost exceeds a per-request spend cap; it cannot succeed as sent.
platform_pressurerate_limit_error429YesRequest rate was temporarily reduced by platform load; retry later.
capacity_exceededrate_limit_error429YescapacityShared upstream capacity is exhausted; retry after the interval.
rate_limit_exceededrate_limit_error429YesAn upstream rate limit was hit; retry after the indicated period.
no_capacityserver_error503YesNo capacity is available for the requested model; retry shortly.
connection_limitserver_error503YesThe gateway is temporarily at connection capacity; retry after the interval.
gateway_timeoutserver_error504YesThe request exceeded the gateway deadline; a retry may succeed.
upstream_server_errorapi_error502YesAn upstream dependency failed; the request may succeed on retry.
provider_unavailableapi_error503YesAn upstream dependency is temporarily unavailable; retry later.
route_capacityserver_error503YesThe platform is temporarily at capacity; retry after the interval.

A 429 can also echo the upstream code insufficient_quota when an upstream quota is exhausted; do not retry until the quota is restored.

Error code reference#

Every non-limit code the platform can emit, with its type, canonical status, and retryability. Stream-frame codes arrive on an in-band error frame after HTTP 200 is committed, so their status column describes the failure class, not a status line.

insufficient_quota#

type rate_limit_error -- status 429 -- retriable: No.

A usage quota backing this request is exhausted; retrying will not clear this standing condition.

unsupported_feature#

type invalid_request_error -- status 400 -- retriable: No.

The request uses an input feature the requested model does not support.

Documented details keys: details.required_features.

capability_conflict#

type invalid_request_error -- status 409 -- retriable: No.

The request combines capabilities that cannot be served together as sent.

context_length_exceeded#

type invalid_request_error -- status 400 -- retriable: No.

The request exceeds the model's context window; reduce the input.

Documented details keys: details.estimated_input_tokens, details.context_window.

upstream_error#

type api_error -- status 502 -- retriable: Yes.

An upstream dependency returned an unexpected error; a retry may succeed.

upstream_auth_error#

type api_error -- status 502 -- retriable: Yes.

The platform could not authenticate with an upstream dependency; this is a platform-side condition, not an API-key problem.

model_not_available#

type forbidden -- status 403 -- retriable: No.

The requested model is not available for this account or API key.

model_not_found#

type invalid_request_error -- status 404 -- retriable: No.

The requested model is not offered on this endpoint.

streaming_not_available#

type forbidden -- status 403 -- retriable: No.

Streaming is not enabled for this API key.

idempotency_key_conflict#

type invalid_request_error -- status 409 -- retriable: No.

The idempotency key was already used by a different request.

bad_request#

type invalid_request_error -- status 400 -- retriable: No.

The request is malformed; the message names what to fix.

invalid_json#

type invalid_request_error -- status 400 -- retriable: No.

The request body is not valid JSON.

unsupported_media_type#

type invalid_request_error -- status 415 -- retriable: No.

The request body must be application/json.

payload_too_large#

type invalid_request_error -- status 413 -- retriable: No.

The request body exceeds the maximum accepted size.

not_found#

type invalid_request_error -- status 404 -- retriable: No.

The requested path or resource does not exist.

conflict#

type invalid_request_error -- status 409 -- retriable: No.

The request conflicts with the current state of the resource.

forbidden#

type invalid_request_error -- status 403 -- retriable: No.

The credential is valid but not permitted to perform this operation.

request_requirements_unsatisfied#

type invalid_request_error -- status 400 -- retriable: No.

The request combines inputs or features that cannot be served together.

invalid_metadata#

type invalid_request_error -- status 400 -- retriable: No.

The request metadata must be a JSON object.

invalid_parameter#

type invalid_request_error -- status 400 -- retriable: No.

A request parameter failed validation; the code carries the offending parameter name and error.param names it when known.

The wire code carries the offending parameter name (e.g. invalid_parameter.temperature); error.param names it structurally.

missing_authorization#

type authentication_error -- status 401 -- retriable: No.

No API key was found on the request.

bearer_scheme_required#

type authentication_error -- status 401 -- retriable: No.

An Authorization header was present but did not use the Bearer scheme.

invalid_api_key#

type authentication_error -- status 401 -- retriable: No.

The presented API key was not recognized.

Documented details keys: details.key_hint.

authentication_error#

type authentication_error -- status 401 -- retriable: No.

Legacy 401 code that duplicated the error type verbatim; replaced by the specific authentication codes. Code-matching clients keep working: the old code's value remains as the error type.

Deprecated: replaced by invalid_api_key.

internal_error#

type api_error -- status 500 -- retriable: Yes.

An internal error occurred; the request may succeed on retry.

stream_init_error#

type api_error -- status stream frame (class 500) -- retriable: Yes.

The stream failed before the first frame; retry the request.

stream_finalization_error#

type api_error -- status stream frame (class 500) -- retriable: Yes.

The stream failed while finalizing; delivered output was billed as sent.

cancelled#

type api_error -- status stream frame (class 499) -- retriable: No.

The request was cancelled by the client before completion.

provider_error#

type api_error -- status stream frame (class 502) -- retriable: Yes.

An upstream dependency reported an error mid-stream.

Error details#

Limit-class denials attach a machine-readable error.details object next to code and type:

JSON payload JSON
{
  "error": {
    "message": "Rate limit exceeded. Retry after 1 second.",
    "type": "rate_limit_error",
    "param": null,
    "code": "rate_limit_exceeded",
    "details": {
      "dimension": "requests",
      "scope": "rpm",
      "limit": 60,
      "window_seconds": 60,
      "retry_after_seconds": 0.8
    }
  }
}

Keys are omitted, never guessed, when a value is unknown. Codes, header names, and details keys are the stable machine contract; message text is not.

  • dimension: what ran out, using the vocabulary in the table above.
  • scope: legacy limit marker (rpm, tpm, concurrency, budget, spend). Deprecated: it repeats dimension in an older vocabulary, and in a future release scope will carry the provenance value now published as binding_scope. Migrate machine handling to dimension now.
  • binding_scope: which configuration level produced the limit that denied the request, when known: api_key, org, tier, model, or capacity.
  • limit: the configured value of the denied limit (account-scoped denials only).
  • window_seconds: the enforcement window for windowed limits.
  • retry_after_seconds: fractional seconds to wait; the Retry-After header carries the same value rounded up to a whole second.

Stream error frames#

Streaming responses commit HTTP status 200 and the response headers before the first body byte, so a failure after that point cannot change the status line. Instead the stream ends with one in-band error frame carrying the same type and code taxonomy as non-streaming error bodies -- including limit denials that strike mid-stream.

Every stream error frame carries these fields:

FieldMeaning
typeFailure-class type from the same registry as the code table above.
codeMachine-readable failure code from the same registry.
messageSanitized human-readable description.
retry_after_secondsFractional seconds to wait before retrying; present only when the failure class is retriable.

Codes that carry retry_after_seconds: capacity_exceeded, connection_limit, gateway_timeout, no_capacity, platform_pressure, provider_unavailable, rate_limit_exceeded, route_capacity, upstream_server_error.

Every wire nests the fields under an error object, which additionally carries retryable (whether an unchanged retry can succeed), request_id (the same correlation id as the X-Request-Id response header), and -- on limit denials -- the same details object as non-streaming denial bodies (dimension, binding_scope, limit, estimate; keys omitted when unknown, never guessed).

Chat Completions additionally delivers the four documented fields flat on the frame payload, exactly as before. The flat keys are deprecated as of 2026-07-15: prefer the nested error object, which OpenAI-compatible SDKs already surface as a typed error. The flat keys will be removed no sooner than two contract releases after this deprecation, and only once usage telemetry shows no remaining flat-key consumers.

A failure the gateway cannot classify ends the stream with type api_error, code internal_error, and no retry signal.

Retry safely#

Retry transient 5xx responses -- and 429 responses that do not declare x-should-retry: false -- with exponential backoff and jitter, and honor Retry-After and retry-after-ms whenever they are present; they are computed from actual limiter state. Do not retry malformed requests, invalid credentials, unsupported models, or permission errors without changing the request.

Request identification#

Every response carries an X-Request-Id header, and every error body repeats the same value as error.request_id (stream failures carry it inside the error frame on the Messages and Responses wires). Console surfaces show it as a Reference: line.

When you contact support, paste the request_id (or the whole error body) together with the timestamp, endpoint, and model. The id joins your request to exactly one internal trace; without it, support has to search by time window. Do not send API keys.

The public error contract omits private serving architecture, non-public vendor data, and raw diagnostic payloads.

Search Keiro docs

Start typing to search pages and sections.

Start typing to search pages and sections.

Documentation

Console