TokenLab

Core Guides

Rate limits

Find your request limit and recover from a 429 response

Your account tier determines how many API requests you can send each minute. When you reach the limit, the API returns 429 rate_limit_exceeded and a Retry-After header.

The limits below are the standard tier values and may change with the active configuration. For authenticated requests, these limits are enforced per API key. Public discovery endpoints such as GET /v1/models have separate limits. On a 429, use the returned X-RateLimit-Limit and Retry-After values instead of assuming a copied table is the active limit.

Account limits

TierRequests per minute
User1,000
Partner10,000
VIP10,000

Need a different limit? Email support@tokenlab.sh with your account email, expected request volume, and use case.

429 response

{
  "error": {
    "message": "Rate limit exceeded. Please retry later.",
    "type": "rate_limit_exceeded",
    "code": "rate_limit_exceeded",
    "retryable": true,
    "retry_after": 8
  }
}

The standard response header gives the same wait time in seconds:

Retry-After: 8

Retry after the server delay

Use Retry-After when it is present. If it is missing, use exponential backoff with random jitter. Always set a retry limit so a busy queue cannot grow forever.

import random
import time
from openai import RateLimitError

def chat_with_rate_limit_retry(client, model, messages, attempts=4):
    for attempt in range(attempts):
        try:
            return client.with_options(max_retries=0).chat.completions.create(
                model=model,
                messages=messages,
            )
        except RateLimitError as exc:
            if attempt == attempts - 1:
                raise
            header = exc.response.headers.get("Retry-After")
            wait = float(header) if header else min(30, 2 ** attempt + random.random())
            time.sleep(wait)

Keep traffic below the limit

  • Queue requests when your application can generate large bursts.
  • Limit concurrency as well as the average requests per minute.
  • Cache only results that are safe and useful to reuse.
  • Do not retry validation, authentication, balance, or permission errors.
  • Track repeated 429 responses by API key so one client cannot crowd out the rest of your application.

A faster model does not increase your account's request limit. Model speed, token limits, and account rate limits are separate constraints.

On this page