核心指南
串流
實作即時串流回應
概覽
串流會逐步回傳模型輸出。當所選模型的 accepted_request_formats 包含 openai_responses 時可用 Responses;既有 Chat Completions 客戶端可以繼續使用原格式。
建議:Responses Streaming
curl https://api.tokenlab.sh/v1/responses \
-H "Authorization: Bearer sk-your-api-key" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-5.6-terra",
"input": "Write a short poem.",
"stream": true
}'Responses 與 Gemini 串流邊界
Responses SSE 保留公開事件名稱、順序與欄位。收到輸出後若串流中斷,應視為未完成,不會自動重新開始。
Responses WebSocket 使用 response.create;串流為隱含行為,不支援 background 或 response.cancel。每條連線一次處理一個回應,最長 60 分鐘。generate: false 可建立可延續的 ID,不產生模型輸出或模型費用。
Gemini SSE 使用原生區塊;中間事件可以沒有 finishReason,也可以自然結束而沒有 Chat 的 [DONE]。
Chat Completions 串流
如果您的框架仍預期從 /v1/chat/completions 接收 SSE 區塊,這同樣可行:
import os
from openai import OpenAI
with OpenAI(
api_key=os.environ["TOKENLAB_API_KEY"],
base_url="https://api.tokenlab.sh/v1",
timeout=30.0,
max_retries=0,
) as client:
finish_reason = None
with client.chat.completions.create(
model="gpt-5.6-terra",
messages=[{"role": "user", "content": "Write a short poem."}],
stream=True,
stream_options={"include_usage": True},
) as stream:
for chunk in stream:
if not chunk.choices:
continue
choice = chunk.choices[0]
if choice.delta.content:
print(choice.delta.content, end="", flush=True)
if choice.finish_reason:
finish_reason = choice.finish_reason
if finish_reason != "stop":
raise RuntimeError(f"Stream ended without a complete text answer: {finish_reason}")串流結束條件
典型的完成條件:
- Responses API streams 的
response.completed - Chat Completions streams 的
finish_reason: "stop" - 當達到 token 限制時的
finish_reason: "length" - 當模型想要使用工具時的 tool/function call 事件
Web App 模式
請在伺服器執行 SDK 串流消費者。瀏覽器應呼叫自己的已驗證後端,TokenLab 金鑰只放在伺服器環境。將文字增量轉送至瀏覽器,使用者停止或斷線時取消上游串流。使用 SDK 解析 SSE 事件;一個網路資料塊不一定是完整事件。