Core Guides
Billing and pricing
How charges, estimates, balances, and limits work
One TokenLab balance works across models. There is no subscription or minimum spend; charges are based on the model and usage recorded for each request.
One result, one charge
Each completed request is charged once for the delivery option that produced the result:
TokenLab Verifieduses TokenLab public prices.Officialuses the Official price layer based on the model maker's public price.Autouses TokenLab Verified when available, then Official if needed. You pay for the option that completes the request.
For paid generation, Console shows the maximum estimate before confirmation. After completion, Runs and Usage show the final charge. Any unused reserved amount is released.
Pricing Models
Model prices can change. Check Console, the Models page, GET /v1/models/:model/pricing, or the Pricing API for the current price.
On the Models page, a dash in the TokenLab price or discount column means no TokenLab Verified offer is currently available; it does not mean free. Models with Official supply remain available through Official or Auto. Delivery labels describe supported options, while availability is checked when each request runs. Series prices summarize models with supply in the relevant delivery option and keep different billing units separate.
Token-Based Pricing
Most chat, reasoning, embedding, rerank, and some image models are priced by input, output, cache, or image-output tokens.
| Pricing family | Examples | Verify current price |
|---|---|---|
| Chat / Responses | gpt-5.4, claude-sonnet-4-6, gemini-3.5-flash | Models page or Pricing API |
| Embeddings / rerank | text-embedding-3-small, qwen3-vl-rerank | Model pricing detail |
| Token-priced images | gpt-image-2 | GET /v1/models/gpt-image-2/pricing |
Do not hard-code a copied price table. Read the current price when your application needs to display or compare costs.
Request and Task Pricing
Image, video, music, 3D, audio, and world models may be priced per request, image, second, minute, or task. The model page shows the billing unit.
| Family | Examples |
|---|---|
| Image | flux-pro |
| Video | veo3.1, seedance-2.0 |
| Music | suno-music |
| 3D | tripo-h3.1 |
| Audio | tts-1, whisper-1 |
Async task billing
Some tasks reserve the estimated cost when they are created. A completed task is charged once; a failed task releases or refunds the pending amount.
Video, music, 3D, and task-based image generation return a task ID and poll_url. The final amount is recorded after the task becomes completed or failed. A completed task includes billing_transaction_id after billing; a failed task is not charged.
If Usage still does not show the final charge or released amount after the task finishes, contact support@tokenlab.sh with the Request ID and task ID.
Billing transaction IDs
Non-streaming OpenAI-compatible responses include billing_transaction_id when billing finishes before the response is sent. The same value may appear in X-Billing-Transaction-ID; Gemini and other native formats may expose it only in this header.
Async tasks include billing_transaction_id after completion. Streaming may finish before billing does, so use Usage when the header is absent.
Token counts
Text models usually report input, output, and total tokens in usage. Tokenization differs by model, language, and media input, so character-count formulas are not reliable billing estimates. Use the model's current price and the final Usage record.
Find your usage
Open Usage to find:
- Real-time balance
- Usage history by model
- Cost breakdown
- API key usage
API response
Token-based responses usually include usage information:
{
"usage": {
"prompt_tokens": 50,
"completion_tokens": 100,
"total_tokens": 150
}
}Spend less
Low-balance alerts
Set a threshold in Console → Settings. TokenLab emails you when the balance falls below it.
Add balance
Open Console → Billing, choose an amount, and use one of the payment methods shown at checkout. The balance updates after the payment provider confirms the payment.
API Key Limits
You can set a spending limit when you create or edit an API key in Console → API Keys.
When the limit is reached, requests with that key will return 402 Payment Required.
Billing help
Contact support@tokenlab.sh for billing inquiries.