Streaming
Keiro streams completion responses as server-sent events. Set stream to true, process events in order, and keep the connection open until the surface's terminal event or in-band error.
On this page 1 of 8
Use curl -N when testing from a terminal so client-side buffering does not hide incremental delivery.
During a long stream, the gateway writes the static SSE comment : keepalive at least every 10 seconds (Messages uses its native ping event). SSE comments carry no model output, consume no event sequence number, and are not replayed; conforming parsers ignore them while the bytes keep transport-idle timers alive.
Chat Completions stream#
Chat Completions sends data: chunks with text in choices[0].delta.content. Tool-call fragments appear in choices[0].delta.tool_calls. Assemble by choice and tool-call index.
A tool call streams as it is generated: the first tool_calls chunk carries the call's id and function.name as soon as the model commits to the call, and later chunks carry function.arguments fragments; concatenate the fragments per index. Fragment boundaries are not stable. If the output limit ends the response while a call is still streaming, the fragments already delivered stand, no further tool_calls chunk follows, and the terminal choice carries finish_reason: "length"; such a call is not complete and must not be executed.
The successful stream ends with a terminal choice and data: [DONE].
Runnable streaming samples live on Chat Completions.
Responses stream#
Responses emits typed events. Read type and ignore event types your client does not handle.
Event type | Meaning |
|---|---|
response.created | The response exists and streaming started |
response.in_progress | Generation is active |
response.output_item.added | A typed output item started |
response.content_part.added | A content part started |
response.output_text.delta | Incremental text in delta |
response.output_text.done | One text segment finished |
response.function_call_arguments.delta | Function arguments arrived for a tool call |
response.function_call_arguments.done | A tool call's arguments finished |
response.output_item.done | A typed output item finished |
response.completed | Successful terminal event |
response.incomplete | Terminal event for an early stop, such as an output limit |
response.error | In-band stream failure |
Successful Responses streams also finish with data: [DONE] after the terminal event.
A function_call item is announced with response.output_item.added (status: "in_progress", empty arguments) as soon as the model commits to the call, and its arguments arrive through response.function_call_arguments.delta frames while the model is still generating; accumulate by output_index. A complete call closes with response.function_call_arguments.done and response.output_item.done (status: "completed"); the concatenated deltas equal the arguments of both close frames and of the item in the terminal output, byte for byte, and the item keeps the id and call_id it was announced with. If the output limit ends the response mid-call, the item closes with response.output_item.done carrying status: "incomplete" and the arguments text delivered so far, the terminal event is response.incomplete with incomplete_details.reason: "max_output_tokens", and the same item appears in the terminal output; an incomplete call must not be executed.
Runnable streaming samples live on Responses.
Messages stream#
Messages emits message_start, content-block start/delta/stop events, message_delta, and message_stop. The terminal message_delta carries the final stop reason and usage. Tool calls use tool_use content blocks and input JSON deltas. When extended thinking is enabled, unsigned thinking summary blocks stream first, through thinking_delta events, and close before the text block opens.
A tool_use block opens with content_block_start (its id and name, empty input) as soon as the model commits to the call, and its input arrives through input_json_delta events while the model is still generating; a text block that was open closes first. If max_tokens ends the response mid-call, the block closes with content_block_stop holding only the input fragments delivered so far and the terminal message_delta carries stop_reason: "max_tokens"; such a block is not a complete call and must not be executed.
See Messages for supported request fields and content shapes.
Stream lifetime and time entitlements#
A stream that is actively working is never ended for taking long. While the model keeps producing activity, the stream stays open until it finishes or reaches the wall-clock entitlement for its requested reasoning effort. A live stream ends early in exactly two documented cases:
The requested reasoning effort — reasoning.effort on Responses, reasoning_effort on Chat Completions — selects both compute depth and the run's time entitlement:
- Idle timeout — no model activity for the idle window. Keepalive comments are transport liveness only; they do not count as model activity.
- Ceiling timeout — the run reaches the wall-clock entitlement for its effort tier, measured from stream accept.
| Requested effort | Idle window | Wall-clock ceiling |
|---|---|---|
none, minimal | 90 seconds | 5 minutes |
low, medium | 90 seconds | 10 minutes |
high | 90 seconds | 20 minutes |
xhigh | 120 seconds | 40 minutes |
max, ultra | 120 seconds | 60 minutes |
Requests that do not set a reasoning effort use the model's default reasoning depth, never below the medium row. eb1-frontier-preview and models documented with extended runtimes default to the max row: long runs need no opt-in. Messages requests carry no effort field; an extended-thinking budget selects the entitlement instead — thinking.budget_tokens of 16,384 or more earns the high row, 131,072 or more the xhigh row, and smaller budgets keep the medium row.
When a bound fires, the stream ends with one honest terminal event, and tokens consumed before the stop are billed:
Configure client timeouts to fit the entitlement, not the other way around: use a read or idle timeout of at least 30 seconds (keepalives arrive at least every 10 seconds, so 30 seconds only trips on a dead connection) and a total timeout at least two minutes above the ceiling for the effort you request. Clients whose idle timer counts SSE events rather than bytes must set it above the effort ceiling plus two minutes because keepalive comments are not events. Platform total-duration limits must exceed the ceiling plus two minutes.
- Responses:
response.incompletewithincomplete_details.reasonofstream_idle_timeoutorstream_ceiling_timeout. - Chat Completions: a terminal choice with
finish_reasonlength. - Messages: a terminal
message_deltawithstop_reasonmax_tokens.
Handle in-band errors#
Once response headers and the first body byte are sent, the HTTP status cannot change. A later failure therefore arrives in the stream.
Treat an error frame as terminal. Do not wait for normal completion after it. See Errors for the shared code taxonomy.
- Chat Completions places
type,code,message, and optionalretry_after_secondson the error frame. - Responses and Messages place those fields under
error.
Verify delivery#
printf 'Keiro API key: '
IFS= read -rs KEIRO_BEARER
printf '\n'
curl -sS -N https://api.keirolabs.ai/v1/responses \
-w '\nTTFB=%{time_starttransfer}s total=%{time_total}s http=%{http_code}\n' \
-H "Content-Type: application/json" \
-d '{
"model": "eb1-preview",
"stream": true,
"input": "Count from one to five, one number per line."
}' \
-H @- <<<"Authorization: Bearer $KEIRO_BEARER"
Time to first byte should normally be lower than total time. Similar values can mean the response was short or that a client, proxy, or network path buffered the stream.
Retry safely#
Requests retain at most 600 images counted across every placement (pasted images and tool-result screenshots together), with 10 MB (10,000,000 bytes) of encoded data per image, 24 MiB of encoded inline media per request (images and PDFs together), and 8 KiB per remote image URL. The model a request routes to can accept fewer (100 images on 200k-context Claude models); the request is then refused before dispatch with the same measured error. PDF inputs allow 2 files, 8 MiB of encoded data each and 12 MiB total; filenames allow 255 characters. Replay items allow 512 KiB each (raw UTF-8 bytes for redacted thinking, serialized bytes for other replay items) and 4 MiB serialized bytes total. The content scan allows 65,536 nodes and a nesting depth of 64; it does not cap text length. These checks apply before both streaming and non-streaming responses. Media-limit errors use phrases such as “images exceed the API limit”; replay-size and content-node errors use “prompt is too long” with measured bytes or content parts, so clients such as Claude Code can automatically remove media or compact history. Depth-only failures remain count-free; these recovery phrases do not promise that an unchanged retry will succeed.
An interrupted stream can contain output that the user already saw. Retry only when repeating the logical operation is safe. Use an idempotency key for retry-sensitive requests, honor Retry-After, and avoid concatenating a retry onto partial output as if it were one response.