TokenLab

Hướng dẫn cốt lõi

Streaming

Triển khai phản hồi streaming theo thời gian thực

Tổng quan

Streaming trả đầu ra từng phần. Dùng Responses nếu mô hình công bố openai_responses trong accepted_request_formats; ứng dụng Chat Completions hiện có có thể giữ định dạng.

Khuyến nghị: Responses Streaming

curl https://api.tokenlab.sh/v1/responses \
  -H "Authorization: Bearer sk-your-api-key" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-5.6-terra",
    "input": "Write a short poem.",
    "stream": true
  }'

Ranh giới Streaming Responses và Gemini

Responses SSE giữ tên, thứ tự và trường sự kiện công khai. Ngắt kết nối sau khi có đầu ra là phản hồi chưa hoàn chỉnh và không tự khởi động lại.

Responses WebSocket dùng response.create và luôn streaming; không hỗ trợ background hay response.cancel. Mỗi kết nối xử lý một phản hồi tại một thời điểm, tối đa 60 phút. generate: false tạo ID tiếp tục mà không sinh đầu ra hoặc phí mô hình.

Gemini SSE trả chunk native. Sự kiện trung gian có thể thiếu finishReason và luồng có thể kết thúc tự nhiên mà không có dấu Chat [DONE].

Phát trực tuyến Chat Completions

Nếu framework của bạn vẫn yêu cầu các chunk SSE từ /v1/chat/completions, cách này cũng hoạt động:

import os
from openai import OpenAI

with OpenAI(
    api_key=os.environ["TOKENLAB_API_KEY"],
    base_url="https://api.tokenlab.sh/v1",
    timeout=30.0,
    max_retries=0,
) as client:
    finish_reason = None
    with client.chat.completions.create(
        model="gpt-5.6-terra",
        messages=[{"role": "user", "content": "Write a short poem."}],
        stream=True,
        stream_options={"include_usage": True},
    ) as stream:
        for chunk in stream:
            if not chunk.choices:
                continue
            choice = chunk.choices[0]
            if choice.delta.content:
                print(choice.delta.content, end="", flush=True)
            if choice.finish_reason:
                finish_reason = choice.finish_reason
    if finish_reason != "stop":
        raise RuntimeError(f"Stream ended without a complete text answer: {finish_reason}")

Điều kiện kết thúc stream

Các điều kiện hoàn tất điển hình:

  • response.completed cho các stream của Responses API
  • finish_reason: "stop" cho các stream Chat Completions
  • finish_reason: "length" khi chạm đến giới hạn token
  • các sự kiện gọi tool/function khi model muốn sử dụng tools

Mẫu cho ứng dụng web

Chạy phần xử lý luồng SDK trên máy chủ. Trình duyệt gọi backend có xác thực của bạn; khóa TokenLab chỉ nằm trong môi trường máy chủ. Chuyển từng phần văn bản và hủy luồng khi người dùng dừng hoặc ngắt kết nối. Dùng SDK để phân tích SSE: một khối mạng không nhất thiết là một sự kiện hoàn chỉnh.

Thực tiễn tốt nhất

Trên trang này