TokenLab

Core Guides

Streaming

Show model output as it arrives

Overview

Streaming delivers model output in small events instead of waiting for the full answer. Use Responses streaming for a model that lists openai_responses in accepted_request_formats; keep Chat Completions streaming when that is what your SDK or application already expects.

Responses streaming

curl https://api.tokenlab.sh/v1/responses \
  -H "Authorization: Bearer sk-your-api-key" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-5.6-terra",
    "input": "Write a short poem.",
    "stream": true
  }'

Responses SSE keeps the documented event names, order, and fields. After the first event arrives, an interrupted stream is incomplete and is not restarted automatically.

Responses WebSocket

Connect to wss://api.tokenlab.sh/v1/responses and send response.create events. WebSocket responses are always streamed and do not support background or response.cancel. A response.create with generate: false returns an ID that can be continued without generating model output or creating a model charge.

Each connection handles one active response at a time and lasts up to 60 minutes. Events include a response-scoped sequence_number. Stream events named error use a flat event object; connection and protocol failures use a nested error object.

Chat Completions Streaming

If your framework still expects SSE chunks from /v1/chat/completions, that also works:

import os
from openai import OpenAI

with OpenAI(
    api_key=os.environ["TOKENLAB_API_KEY"],
    base_url="https://api.tokenlab.sh/v1",
    timeout=30.0,
    max_retries=0,
) as client:
    finish_reason = None
    with client.chat.completions.create(
        model="gpt-5.6-terra",
        messages=[{"role": "user", "content": "Write a short poem."}],
        stream=True,
        stream_options={"include_usage": True},
    ) as stream:
        for chunk in stream:
            if not chunk.choices:
                continue
            choice = chunk.choices[0]
            if choice.delta.content:
                print(choice.delta.content, end="", flush=True)
            if choice.finish_reason:
                finish_reason = choice.finish_reason
    if finish_reason != "stop":
        raise RuntimeError(f"Stream ended without a complete text answer: {finish_reason}")

Gemini Streaming

POST /v1beta/models/{model}:streamGenerateContent?alt=sse returns Gemini chunks. An event can contain only metadata, intermediate events may omit finishReason, and the stream can end naturally without a Chat Completions [DONE] marker.

Stream End Conditions

Typical completion conditions:

  • response.completed for Responses API streams
  • finish_reason: "stop" for Chat Completions streams
  • finish_reason: "length" when a token limit is hit
  • tool/function call events when the model wants to use tools

Web App Pattern

Run the SDK stream consumer on your server. The browser should call your own authenticated backend; keep the TokenLab key in the server environment. Forward text deltas to the browser and cancel the upstream stream when the user stops or disconnects. Use the SDK to parse SSE events: one network chunk is not necessarily one complete event.

Handle streams well

On this page