Messages API

Use POST /v1/messages when your application already uses the Messages request and content-block format. Keiro accepts that wire shape, routes it through the same public eb1 service as the other completion endpoints, and returns a Messages-style response.

On this page 1 of 5

Send a message#

Run a curl request Shell
printf 'Keiro API key: '
IFS= read -rs KEIRO_BEARER
printf '\n'

curl -sS https://api.keirolabs.ai/v1/messages \
  -H "Content-Type: application/json" \
  -d '{
    "model": "eb1-preview",
    "max_tokens": 256,
    "system": "Answer in one concise sentence.",
    "messages": [
      {
        "role": "user",
        "content": "Why are idempotency keys useful?"
      }
    ]
  }' \
  -H @- <<<"Authorization: Bearer $KEIRO_BEARER"

Read generated text from the text blocks in content. A buffered response has this customer-visible shape:

JSON payload JSON
{
  "id": "msg_...",
  "type": "message",
  "role": "assistant",
  "content": [
    {
      "type": "text",
      "text": "Idempotency keys let a safe retry reuse the original result instead of running twice."
    }
  ],
  "model": "eb1-preview",
  "stop_reason": "end_turn",
  "stop_sequence": null,
  "usage": {
    "input_tokens": 18,
    "cache_creation_input_tokens": 0,
    "cache_read_input_tokens": 0,
    "output_tokens": 19
  }
}

IDs, text, and token counts vary by request. Depend on the field layout, not the sample values.

The three input-side fields are disjoint, matching Anthropic's native Messages usage semantics: input_tokens counts the portion of the prompt served from neither cache, cache_read_input_tokens counts prompt tokens served from a prompt cache, and cache_creation_input_tokens counts prompt tokens written to one. input_tokens + cache_read_input_tokens + cache_creation_input_tokens equals the full prompt, so a client that sums the three fields — Claude Code's context gauge does — reads the real context size. All three fields are always present. On a stream, the message_start usage carries zero for both cache fields — cache counts, like output_tokens, are only known at the terminal message_delta.

Deprecated (2026-08-08): additive-subset usage shape. Messages usage previously rendered input_tokens as the full inclusive prompt count with the two cache fields as subsets of it. That shape is retired for all Messages consumers; the disjoint semantics above are the contract and take effect with the next deployment. If your integration read input_tokens as the inclusive prompt total, sum input_tokens + cache_read_input_tokens + cache_creation_input_tokens to recover the old value. Responses with zero cache activity are byte-identical under both shapes.

Supported request fields#

The public Messages surface supports these top-level fields:

Unknown fields fail closed. Do not assume that a field from another Messages implementation is available until it is listed here.

  • model, messages, and optional top-level system
  • max_tokens, temperature, top_p, and stop_sequences
  • stream
  • tools, tool_choice, and parallel_tool_calls
  • thinking with type enabled or disabled; enabled thinking requires a positive budget_tokens
  • metadata
  • idempotency_key, with the request-header form preferred for HTTP clients

Content blocks#

User and assistant messages accept strings or non-empty content-block arrays. Supported customer workflows include text, top-level user image blocks, and the tool_use / tool_result history used for function calling.

tool_result content accepts text and image blocks, so a coding harness can return screenshots as tool output. Every image in the request, including those inside tool results, counts toward the request image limits in Images and vision; see Tool calling for tool history.

Stream a response#

Set stream to true to receive Messages-style server-sent events. A normal text stream follows this lifecycle:

When extended thinking is enabled, one or more thinking content blocks arrive before the text block: a content_block_start with block type thinking, thinking_delta deltas carrying the thinking text, then content_block_stop, all closing before the text block opens. These blocks carry summary text constructed by the API and are unsigned; do not round-trip them to a provider as signed thinking blocks.

When the model calls a tool, the tool_use block streams live: a content_block_start with the block's id, name, and empty input arrives as soon as the model commits to the call (closing any open text block first), followed by content_block_delta events of type input_json_delta whose partial_json fragments concatenate to the input, then content_block_stop and a message_delta with stop_reason: "tool_use". If max_tokens ends the response while the input is still streaming, the block closes with the fragments delivered so far and the message_delta carries stop_reason: "max_tokens"; treat that block as incomplete, exactly as with a native Anthropic stream.

Streams are governed by documented idle and wall-clock time entitlements, not a flat timeout. Messages requests carry no effort field; thinking.budget_tokens selects the entitlement row — see Streaming.

A failure after streaming begins arrives as an in-band error event with a sanitized error.type, error.code, and error.message. Retriable failures also include error.retry_after_seconds.

  1. message_start
  2. content_block_start
  3. one or more content_block_delta events
  4. content_block_stop
  5. message_delta, including the terminal stop reason and final usage
  6. message_stop

Search Keiro docs

Start typing to search pages and sections.

Start typing to search pages and sections.

Documentation

Console