Core Guides
Build a reliable TokenLab integration
Keep cost, retries, API keys, and request records under control
A good integration makes price and model choices visible, survives ordinary network failures, and never exposes an API key in the browser. The practices below apply across chat, image, video, audio, and other TokenLab APIs.
Choose models with current data
Model names, capabilities, and prices change. Read the Models API or model pages instead of keeping a permanent list in code or documentation.
For a new use case:
- Filter models by the capability your product needs.
- Compare their current TokenLab prices.
- Test a small set with representative user requests.
- Keep the model ID explicit in each request.
Do not switch models silently. A replacement can change output quality, price, context length, available tools, or media support.
Put an upper bound on cost
For text models, limit output to the length your product can actually use. The parameter name depends on the API and model, so check the model page before sending max_tokens or max_output_tokens.
Shorter instructions are not automatically better, but duplicated context and repeated boilerplate always add input tokens. Keep stable instructions concise and send only the history needed for the next answer.
For image, video, music, 3D, and Worlds requests, cost often depends on output count, duration, resolution, or another media option. Show the current estimate before a user confirms an expensive generation whenever the API can provide one.
Stream long text responses
Streaming shows the answer as it arrives and avoids making the user wait for the complete response. It does not reduce token cost.
For a complete example with client setup and completion checks, see Streaming.
Handle a disconnect as an incomplete answer. Do not present partial output as successfully finished, and do not repeat a request automatically when doing so could create a second side effect or charge.
Retry only when it is safe
Retry 429 responses after Retry-After or retry_after. A short exponential backoff is appropriate for transient 5xx responses. Authentication, validation, balance, and context-length errors need a change from the application or user; sending the same request again will not fix them.
See Rate limits for a bounded retry example that respects Retry-After and disables duplicate SDK retries.
Set a retry limit and a total time limit. For media creation, save the returned task ID and check it after a timeout before creating another task. Use an idempotency key when the endpoint supports one.
Keep API keys on the server
Never embed a TokenLab API key in browser JavaScript, a mobile application bundle, a public repository, or a client-visible log.
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["TOKENLAB_API_KEY"],
base_url="https://api.tokenlab.sh/v1",
timeout=30.0,
max_retries=0,
)Use separate API keys for development, production, and independent applications. Give each key only the models and spending limit it needs. Revoke a key immediately if it may have been exposed.
Keep useful request records
Store the identifiers that let you connect a user action to usage and support records:
request_idfor the API requesttask_idfor asynchronous generationbilling_transaction_idwhen returned- model ID and endpoint
- your own user, project, or job ID
Do not log API keys, management tokens, full private prompts, private media, or signed URLs. Usage records in Console and the Management API are the source for billed amounts; model response token counts alone may not describe the final charge.
Design for asynchronous work
Image, video, music, 3D, and Worlds APIs may return before the result is ready. Save the task ID, use the returned poll_url, and stop only at completed or failed. A failed status lookup does not mean the generation itself failed.
See Async jobs and polling for cancellation, timeout, and duplicate-request handling.
Before launch
- Run representative requests with the exact model and API format you will ship.
- Check both success and common error responses.
- Confirm the price shown in your product matches the current TokenLab price.
- Make refreshes, network timeouts, and repeated clicks safe.
- Make it possible to trace a user-visible result back to its Request ID.