TokenLab

Core

API Reference

Endpoints, authentication, response headers, and errors

Overview

TokenLab supports OpenAI-compatible endpoints plus Anthropic and Gemini request formats. Existing OpenAI clients can use /v1; choose another format only when your application needs behavior specific to that API. POST /v1/responses is optional and depends on model support.

Base URL

https://api.tokenlab.sh

Authentication

Model requests use a TokenLab API key. The standard authentication header is:

Authorization: Bearer sk-your-api-key

GET /v1/models, GET /v1/models/{model}, and GET /v1/pricing are public and need no key. Anthropic Messages also accepts x-api-key; Gemini accepts x-goog-api-key or ?key= as well as Bearer authentication. /v1/management/* requires a management token (mt-...).

Get your API key from the Console.

Delivery policy

Generation requests accept X-TokenLab-Delivery-Policy: auto | verified | official. A request header overrides the API key setting, which overrides the Workspace default.

  • Auto uses TokenLab Verified when available, then Official if needed. You pay for the option that completes your request.
  • TokenLab Verified is delivery verified by TokenLab. TokenLab prices apply.
  • Official is based on the model maker's public price. The price shown on TokenLab is what you pay.
  • Realtime sessions use the API key or Workspace setting and do not accept a query-string override.

An invalid header returns 400. If the requested option is unavailable, TokenLab returns 503 with code: delivery_tier_unavailable, retryable: true, and a request ID.

The playground does not accept API keys. To send a real request, use one of these options:

  • cURL — Copy an example and replace sk-your-api-key
  • Postman — Import the OpenAPI spec
  • SDK — Set the TokenLab base URL in a supported SDK

Supported Endpoints

Chat & Text Generation

EndpointMethodDescription
/v1/chat/completionsPOSTOpenAI-compatible chat completions
/v1/messagesPOSTAnthropic-compatible messages API
/v1/responsesPOSTOpenAI Responses API

Embeddings & Rerank

EndpointMethodDescription
/v1/embeddingsPOSTCreate text embeddings
/v1/rerankPOSTRerank documents

Images

EndpointMethodDescription
/v1/images/generationsPOSTGenerate images from text
/v1/images/editsPOSTEdit images
/v1/images/generations/{id}GETImage task status path for task-based image responses

Image models may return a finished image or an asynchronous task. If the response includes poll_url, use that URL to check the task.

Audio

EndpointMethodDescription
/v1/audio/speechPOSTText-to-speech (TTS)
/v1/audio/transcriptionsPOSTSpeech-to-text (STT)

Realtime

EndpointMethodDescription
/v1/realtime?model={model}WSRealtime WebSocket sessions

Use /v1/realtime for WebSocket upgrades. A plain GET /v1/realtime returns endpoint metadata. This is not the OpenAI Realtime REST surface; client-secret, Calls, and legacy beta session endpoints are not available.

Video

EndpointMethodDescription
/v1/videos/generationsPOSTCreate video generation task
/v1/tasks/{id}GETGet async task status for video jobs
/v1/videos/generations/{id}GETLegacy-compatible video task status path

Use the poll_url returned when the task is created. /v1/videos/generations/{id} remains available for older clients.

Async Tasks

EndpointMethodDescription
/v1/tasks/{id}GETStatus for an asynchronous task

Video, music, 3D, and some image requests may return this endpoint in poll_url.

Music

EndpointMethodDescription
/v1/music/generationsPOSTCreate music generation task
/v1/music/generations/{id}GETMusic-specific status path

Use the returned poll_url. /v1/music/generations/{id} remains available for clients that need the music-specific path.

3D Generation

EndpointMethodDescription
/v1/3d/generationsPOSTCreate 3D model generation task
/v1/3d/generations/{id}GET3D-specific status path

Use the returned poll_url. /v1/3d/generations/{id} remains available for clients that need the 3D-specific path.

Models

EndpointMethodDescription
/v1/modelsGETList all available models
/v1/models/{model}GETGet specific model info

Gemini (v1beta)

Native Google Gemini API format support:

EndpointMethodDescription
/v1beta/models/{model}:generateContentPOSTGenerate content (Gemini format)
/v1beta/models/{model}:streamGenerateContentPOSTStream generate content (Gemini format)

Gemini endpoints support ?key= query parameter authentication in addition to standard Bearer token.

Response Format

Each endpoint preserves its API format. The success and error examples below use Chat Completions format.

Success Response

{
  "id": "chatcmpl-abc123",
  "object": "chat.completion",
  "created": 1234567890,
  "model": "gpt-5.6-terra",
  "choices": [{"index": 0, "message": {"role": "assistant", "content": "Hello!"}, "finish_reason": "stop"}],
  "usage": {
    "prompt_tokens": 10,
    "completion_tokens": 20,
    "total_tokens": 30
  }
}

Request identifiers

Private delivery details are not part of the public response contract. Use the public headers below when they are present.

HeaderDescription
X-Routing-Time-MSTime spent selecting delivery, when available
X-Request-IDRequest identifier for support and debugging, when available
X-Task-IDPublic async task identifier for task-based responses, when available
X-Billing-Transaction-IDBilling transaction identifier after final billing, when available

Error Response

{
  "error": {
    "message": "Invalid API key provided",
    "type": "invalid_api_key",
    "code": "invalid_api_key"
  }
}

Rate Limits

Rate limits are role-based and configurable by administrators. Default values:

RoleRequests/min
User1,000
Partner10,000
VIP10,000

Contact support for custom rate limits. Exact values may vary by account configuration.

When rate limits are exceeded, the API returns a 429 status code with a Retry-After header indicating how long to wait.

OpenAPI Specification

OpenAPI Spec

Download the complete OpenAPI 3.1 specification

On this page