A streaming request is safe to replay only while three things are true at once. Nothing has reached your client. Nothing observable has been metered. And the request carries no server-side state. After the first output event, the right move is to report the failure instead of replaying it.
TokenLab applies that rule on its gateway to Responses API streaming, over both HTTP and WebSocket. The WebSocket path was changed on 2026-09-28 to match HTTP.
Why a stream is different from a normal request
A non-streaming call returns a body or an error. You can retry the error because you got nothing back.
A stream hands you output before the request is finished. The first output event is the point of no return. If the connection dies after that, you hold partial text. Replaying the request means generating the same answer again and paying for it twice. You may also duplicate a tool call your agent already ran.
The TokenLab streaming guide states it directly:
After the first event arrives, an interrupted stream is incomplete and is not restarted automatically.
So your client needs one bit of local state: saw_output. It flips to true the moment any output reaches your code. Every retry decision reads that bit first.
A stream that ends without response.completed is a failure. Do not assume the text you have is complete. Handle the response.failed, response.incomplete and error events.
The replay decision, point by point
TokenLab replays a request once on another available route when all of these hold. The request is stateless. Nothing has reached the client. No result or usage was observed for the failed attempt. And the failure is either a retryable pre-output event or an upstream read error before the first event. At most one replay happens per request. If the replacement fails before output too, that failure is not replayed again.
Source: TokenLab streaming guide and gateway behavior, observed 2026-09-28.
| Failure point | Replayed by TokenLab? | Reason |
|---|---|---|
Retryable pre-output event (response.failed, or an error event marked retryable, such as an overloaded or internal upstream error) |
Yes, once, if the request is stateless | Nothing reached the client and no usage was observed, so a second execution is invisible. |
| Upstream stream breaks (read error) before the first event | Yes, once, if the request is stateless | Same window. The client holds no output and no charge. |
| Second failure before output, after one replay | No | The budget is one replay per request. |
| Any failure after output reached the client | No | The client already holds partial text. A replay would duplicate output and cost. |
Stored response (store), continuation (previous_response_id), or an origin-bound request |
No | A second execution could create a second stored response or diverge conversation state. |
| First-event timeout | No | The upstream may still be generating. A replay could run the same work twice while the first attempt continues. |
| Pre-output buffer overflow | No | The limit is local to the gateway. The same oversized prefix would very likely hit it again on the next route. |
| Client disconnected | No | The client stopped listening. |
| Deterministic failure, for example an invalid request | No | Retrying cannot change the outcome. Delivered unchanged. |
| Failure that already carries usage | No | The attempt was metered. Delivered unchanged. |
| No other route remains | No | There is nowhere to send it. The client gets the failure with its own code. |
When a failure is not replayed, or no other route remains, you receive it with its own error code. Public examples: stream_read_error when the upstream stream broke, and upstream_stream_buffer_limit on buffer overflow. If route selection itself fails after a replay decision, the WebSocket turn ends with websocket_response_failed (status 500) and the reserved charge is refunded.
Billing follows the same line. You pay for the delivered attempt only. A replayed request may have executed upstream twice, and that extra upstream cost is TokenLab's, because nothing reached you from the first attempt. A failed turn that delivers nothing is refunded.
One timing detail matters for your error handling. Before output starts, the gateway holds response.created and response.in_progress until the first output event or a failure arrives, for at most 10 seconds. Those held events then reach you together with the first output, or with the terminal event. Order and content are unchanged. You only see them slightly later. That 10 seconds is a maximum, not a typical delay.
What changed for WebSocket on 2026-09-28
TokenLab serves the Responses API over HTTP streaming ("stream": true, server-sent events) and over WebSocket at wss://api.tokenlab.sh/v1/responses, where the client sends response.create events. WebSocket responses are always streamed. They do not support background or response.cancel. Each connection handles one active response at a time for up to 60 minutes.
Before the change, the two paths disagreed. HTTP held the lifecycle events and replayed stateless pre-output failures. WebSocket forwarded response.created right away and delivered pre-output failures to the client, refunding them. The same upstream hiccup produced a clean answer on HTTP and an error on WebSocket.
The WebSocket path now follows the HTTP rule, including replay of a stream that breaks before any event arrives. Internally, most upstream failures seen on WebSocket turns happened before any output. That is exactly the window where a replay is safe.
The gateway improves the pre-output failure case. It does not guarantee that a stream completes.
How the change shipped without breaking other behavior
The work followed a process built to catch silent behavior changes.
- Behavior lock. Before the change, every WebSocket turn scenario was recorded as a fixture: the frames the client receives, the upstream calls made, and the billing outcome. The suite grew to 63 recorded scenarios during this work. A behavior change must be declared up front. Only the fixtures named in that declaration may change. Every other fixture must stay byte-identical.
- Mutation checks. Each new decision rule was tested by deliberately flipping it, such as replaying a first-event timeout or not replaying a reader failure, and confirming the lock fails.
- A review catch. The first version also made the buffer-overflow case replayable, on the claim of parity with HTTP. Review showed that HTTP never replays that case, for the reason in the table. A follow-up restored the old behavior and added boundary scenarios: a second reader failure is not replayed, no route left, a failure after a held
response.created, and a replacement stream that then breaks.
Request logs of a turn that succeeded after a replay now also record the earlier failed attempt, as HTTP already did.
Client code that owns the retry decision
Set SDK automatic retries to 0 for streaming calls. That keeps the decision in your code. Keep the replay decision in one place, not spread across handlers. For HTTP errors, respect retryable and retry_after as described in the error handling guide, and keep request IDs.
SSE over HTTP
import os
from openai import OpenAI
with OpenAI(
api_key=os.environ["TOKENLAB_API_KEY"],
base_url="https://api.tokenlab.sh/v1",
timeout=30.0,
max_retries=0, # own the retry decision instead of resending a half-read stream
) as client:
completed, saw_output = False, False
with client.responses.create(
model="gpt-5.6-terra",
input="Reply with one short sentence about retries.",
stream=True,
) as stream:
for event in stream:
if event.type == "response.output_text.delta":
saw_output = True
print(event.delta, end="", flush=True)
elif event.type == "response.completed":
completed = True
elif event.type in {"response.failed", "response.incomplete", "error"}:
raise RuntimeError(f"{event.type} after_output={saw_output}")
if not completed:
raise RuntimeError(f"stream closed before response.completed, after_output={saw_output}")
print()
The example uses the OpenAI SDK 2.15.0 against https://api.tokenlab.sh/v1 with max_retries=0. It tracks saw_output, and raises on response.failed, response.incomplete and error events, and on a stream that closes before response.completed. Verified against production on 2026-09-28 with gpt-5.6-terra.
If the failure arrives with saw_output == false and the request was eligible (stateless, with a replayable failure), TokenLab has already replayed it once; stored responses, continuations and first-event timeouts were not replayed at all. Decide at the app level whether a fresh request is acceptable, because a fresh request is a new generation. If saw_output == true, report the failure and show what you have, or discard the partial text deliberately.
WebSocket
import asyncio
import json
import os
import websockets
URL = "wss://api.tokenlab.sh/v1/responses"
TERMINAL = {"response.completed", "response.failed", "response.incomplete", "error"}
async def run_turn(prompt: str) -> str:
headers = {"Authorization": f"Bearer {os.environ['TOKENLAB_API_KEY']}"}
async with websockets.connect(URL, additional_headers=headers, max_size=None) as ws:
await ws.send(json.dumps({
"type": "response.create",
"model": "gpt-5.6-terra",
"input": prompt,
"store": False,
}))
text, saw_output = [], False
async for raw in ws:
event = json.loads(raw)
kind = event.get("type")
if kind == "response.output_text.delta":
saw_output = True
text.append(event["delta"])
elif kind in TERMINAL:
if kind != "response.completed":
# After output has started, a failure is final for this turn.
# Resend only if your app can discard the partial text.
raise RuntimeError(f"{kind} after_output={saw_output}: {json.dumps(event)[:300]}")
return "".join(text)
raise RuntimeError(f"socket closed before a terminal event, after_output={saw_output}")
print(asyncio.run(run_turn("Reply with one short sentence about retries.")))
The example uses websockets 16.0, connects to wss://api.tokenlab.sh/v1/responses with a Bearer header, sends one response.create with store: false, and collects response.output_text.delta. It raises with after_output on any non-completed terminal event or an early close. Verified against production on 2026-09-28 with gpt-5.6-terra.
The after_output flag is the same idea as saw_output. It tells your calling code whether a fresh turn is even possible without duplicating side effects.
Checklist for your own retry logic
- Treat a stream that ends without
response.completedas a failure, every time. - Track one boolean for whether output reached your code. Flip it on the first output event, not on the first lifecycle event.
- A pre-output failure on an eligible request has already had its one gateway replay; a further attempt is your decision.
- After partial output, resend only if your app can discard the partial text and accept paying for two generations.
- In agent loops, check whether the partial stream already contained a tool call your code acted on. Do not replay a turn whose side effects you cannot undo.
- For stored responses and
previous_response_idcontinuations, inspect what state exists before you resend anything. - Set streaming retries to 0 in your SDK and keep the replay decision in one function.
- Log request IDs so you can match a delivered answer to the attempts behind it.
FAQ
Does TokenLab restart a stream after partial output?
No. Once output has reached your client, a failure is reported and never replayed. You hold partial text, so a restart would duplicate output and cost. Your app decides whether to show, truncate or discard what it has.
Will I be charged twice if the gateway replays my request?
No. You pay for the delivered attempt only. A replayed request may have executed upstream twice, but nothing reached you from the first attempt, and that extra upstream cost is TokenLab's. A failed turn that delivers nothing is refunded.
Why isn't a first-event timeout retried?
Because the upstream may still be generating. A replay could run the same work twice while the first attempt continues. A first-event timeout is treated differently from a read error that breaks the stream before the first event.
Can I retry a stored response or a previous_response_id continuation?
Not automatically. TokenLab never replays stored responses, continuations, or origin-bound requests, because a second execution could create a second stored response or diverge conversation state. Check what state exists before you resend, and only resend if your app can reconcile that state.
If you want to watch the raw event stream yourself, create an API key and log every event type your client receives. The streaming guide and the error handling guide cover the full event set. For background on how the gateway routes and recovers, see TokenLab AI API reliability infrastructure and Responses API vs Chat Completions for agents.
Sources
- https://docs.tokenlab.sh/guides/streamingSources checked 2026-09-28
- https://docs.tokenlab.sh/guides/error-handlingSources checked 2026-09-28



