Choose Auto, TokenLab Verified, or Official for each request, with prices shown up front.See what's new

Changelog

What's new

New models, price changes, and product updates — right here. Open docs

2026 Q4

Models

Gemini Nano Banana 2.1 and Claude Haiku 5.5 are now available

Gemini Nano Banana 2.1 is Google's update to Nano Banana 2 for image generation and editing, with clearer text in images and steadier subjects across edits. Use gemini-nano-banana-2.1 (priced per token) or nano-banana-2.1 (priced per image) with the Images API at 1K, 2K or 4K in 14 aspect ratios. Claude Haiku 5.5 is the fastest model in the Claude 5 series, with text and image input, tool use, structured output, prompt caching, a 1M-token context window and up to 128K output tokens. Use claude-haiku-5-5 or its alias claude-haiku-5.5 with Chat Completions or Messages.

TokenLab Verified prices: Nano Banana 2.1 at 50% off the reference tariff, $0.0168, $0.0252 and $0.0565 per 1K, 2K and 4K image for nano-banana-2.1, and $0.75 input and $15 image output per million tokens for gemini-nano-banana-2.1. Claude Haiku 5.5 at 70% off: $0.03 input and $0.15 output per million tokens for prompts up to 100K tokens, and $0.15 and $0.75 above that, with the same discount on cache reads and writes. Official delivery is available at its listed price.

Models
gemini-nano-banana-2.1nano-banana-2.1claude-haiku-5-5
Effective
View models and pricing
Pricing

Verified prices for 15 models fall to 70% of official prices

The Verified discount for 15 GPT, reasoning and embedding models increases from 15% to 30%. Unit prices fall from 85% to 70% of official prices, approximately 17.65% lower than before. The same discount applies to each model’s supported input, output and cache billing components.

New Verified requests use the new rates after the update takes effect. No changes to model names or API calls are required. Official delivery prices and settled historical charges remain unchanged.

Models
gpt-4.1-minigpt-4ogpt-4o-minigpt-5gpt-5-minigpt-5-nanogpt-5.1gpt-5.2gpt-5.3-codexgpt-5.4gpt-5.4-minigpt-5.4-nanoo3text-embedding-3-largetext-embedding-3-small
Effective
View models and prices
Models

Updated image, video and audio model catalog (3/3)

This release now lists 82 distinct new models: 46 image, 21 video, 6 audio, 4 music, 4 3D and 1 Chat model. Operations and date variants of the same model have been consolidated. Check each model page for exact parameters and currently supported operations.

Use one model ID with the corresponding operation. Duplicate entries and older versions excluded from this release have been removed from its list.

Models
riverflow-2-5-prorodin-gen-2runway-gen-4-5runway-gen-4-imagerunway-gen-4-image-turbosana-sprintskyreels-v4stable-audio-3-small-music-text-to-audiostable-audio-3-small-sfx-base-text-to-audiostable-audio-3-small-sfx-text-to-audiotopaz-labs-wonder-3-5trellis-2tripo-3d-v3-1uni-1uni-1-maxveed-fabric-1-0wan2-7-imagewan2-7-image-pro
Effective
Browse models
Models

Updated image, video and audio model catalog (2/3)

This release now lists 82 distinct new models: 46 image, 21 video, 6 audio, 4 music, 4 3D and 1 Chat model. Operations and date variants of the same model have been consolidated. Check each model page for exact parameters and currently supported operations.

Use one model ID with the corresponding operation. Duplicate entries and older versions excluded from this release have been removed from its list.

Models
krea-2-largekrea-2-mediumkrea-2-medium-turbokrea-2-rawkrea-2-turboltx-2-5-fastltx-2-5-promeshy-6minimax-h3-fastminimax-h3-max-turbominimax-speech-2-8mirelo-sfx-1-5muse-imageomnihuman-1-5p-imagep-image-editp-image-ideogramp-image-upscalep-video-2p-video-2-prop-video-avatarpicsart-image-vectorizerqwen-image-edit-2511qwen-image-layeredqwen3-tts-1-7b-customvoiceqwen3-tts-1-7b-voicedesignray3-2recraft-v4-stylesrecraft-v4-styles-prorecraft-v4-styles-pro-vectorrecraft-v4-styles-vectorriverflow-2-5-fast
Effective
Browse models
Models

Updated image, video and audio model catalog (1/3)

This release now lists 82 distinct new models: 46 image, 21 video, 6 audio, 4 music, 4 3D and 1 Chat model. Operations and date variants of the same model have been consolidated. Check each model page for exact parameters and currently supported operations.

Use one model ID with the corresponding operation. Duplicate entries and older versions excluded from this release have been removed from its list.

Models
boogu-image-0-1-editboogu-image-0-1-edit-turbobria-3-2bria-fibobria-fibo-editbria-fibo-litebria-image-increase-resolutionbria-rmbg-v2-0eleven-v4-turboernie-imageernie-image-turboflux-2-devflux-2-klein-4b-baseflux-2-klein-9b-baseflux-2-klein-9b-kvgoogle-gemma-4-31bheygen-video-1-0heygen-video-agentideogram-4-5imagineart-2-0inworld-realtime-tts-2juggernaut-zkling-image-3-0kling-image-o3kling-ttskling-video-3-0-4kkling-video-3-0-omni-4kkling-video-3-0-omni-prokling-video-3-0-omni-standardkling-video-3-0-prokling-video-3-0-standardkling-video-3-0-turbo
Effective
Browse models

2026 Q3

Models

New image and video models, plus the return of Speech 2.8 Turbo

Hy Image 3.5 Preview supports 1024×1024 text-to-image and reference-image editing. Hailuo H3-Max generates 5–15-second 480p/768p videos from text, a first frame, or first and last frames. MiniMax Speech 2.8 Turbo is available again for text-to-speech.

Models
hy-image-v3.5-previewhailuo-h3-maxspeech-2.8-turbo
Effective
Browse models
Models

New reasoning, retrieval and media models

Mistral Medium 3.5, Amazon Nova 2 Lite and Ember-1 join Chat Completions. Voyage, Mistral and Perplexity add text and code embeddings. Recraft expands image generation; Cohere Rerank 4, Fish Transcribe-1, Grok TTS/STT and MiniMax H3 Max add reranking, speech and video capabilities.

New models are available in OFFICIAL. Ember-1 is a research preview. Check each model's published interface, limits and price before switching.

Models
voyage-4-largevoyage-4voyage-4-litevoyage-3-largevoyage-3.5voyage-3.5-litevoyage-code-2voyage-code-3voyage-finance-2voyage-law-2mistral-embedcodestral-embedmistral-medium-3.5nova-2-liteember-1pplx-embed-v1-0.6bpplx-embed-v1-4brecraft-v3recraft-v4recraft-v4-prorecraft-v4.1recraft-v4.1-flashrecraft-v4.1-prorecraft-v4.1-utilityrecraft-v4.1-utility-prorerank-v4.0-fastrerank-v4.0-profish-transcribe-1grok-ttsgrok-sttminimax-h3-max
Effective
Explore models
Models

Veo 3.1 Lite is now available

Use veo3.1-lite through /v1/videos/generations for text-to-video and image-to-video. The initial public contract supports 8-second videos at 720p or 1080p with audio, in 16:9 or 9:16. Follow the returned poll_url to retrieve the result.

Models
veo3.1-lite
Effective
View model
Models

Gemini Embedding 2 is now available

Use gemini-embedding-2 through /v1/embeddings for text embeddings. Single strings and batches of strings return float vectors; dimensions can be configured.

Models
gemini-embedding-2
Effective
View model
Models

GPT-6.1 Sol is now available

Use gpt-6.1-sol on TokenLab with Chat Completions and Responses. Standard pricing is $2 input, $0.10 cached input, $2.50 cache writes, and $10 output per 1M tokens; prompts over 272K input tokens use long-context rates of $4, $0.20, $5, and $15.

Models
gpt-6.1-sol
Effective
Browse models
Models

Claude Sonnet 5.5 is now available

Claude Sonnet 5.5 joins the Claude 5 series with text and image input, tool use, structured output, prompt caching, a 1M-token context window, and up to 128K output tokens. Use claude-sonnet-5-5 or its alias claude-sonnet-5.5 with Chat Completions, Responses, or Messages.

TokenLab Verified supports Chat Completions and Messages at 70% off the reference tariff: $0.60 input and $3.00 output per million tokens, with the same discount on cache reads and writes. Official delivery supports all three APIs at its listed price. Reasoning is enabled by default; thinking disabled and forced tool choice are not supported by this model.

Models
claude-sonnet-5-5
Effective
View model and pricing
Platform

Developer center is now available

The new developer center brings SDKs, tools, and integration resources together.

Explore developer resources
Models

New models for chat, retrieval, images, video, and audio

Added models: Qwen 3.8 Omni Flash, GLM 5.3 FlashX, Step 5 Preview, Hy4 Preview, Qwen 3.7 Text Embedding, Qwen 3.7 Text Embedding Flash, Qwen 3.7 Text Rerank, Doubao Seed 2.1 Lite, Seedream 5.0 Flash, Kimi K2.8 Preview, FLUX 3 Video, FLUX Video Edit, Stable Audio 3.0. See the model catalog for current capabilities and prices.

Models
qwen3.8-omni-flashglm-5.3-flashxstep-5-previewhy4-previewqwen3.7-text-embeddingqwen3.7-text-embedding-flashqwen3.7-text-rerankdoubao-seed-2.1-liteseedream-5.0-flashkimi-k2.8-previewflux-3-videoflux-video-editstable-audio-3
Effective
Browse models
Models

Jev 1.13 is available for typed decisions

Use jev-1.13 through POST /v1/systemone to evaluate Noul probabilities, Choice selections, and Score ratings against shared state. The TokenLab MCP tool evaluate_decisions supports the same native request.

Keep decision results structured. Your application controls validation, confidence thresholds, permissions, and actions.

Models
jev-1.13
Effective
Read the System One API reference
Pricing

Image and video pricing updates

Selected image and video models now offer 5%–85% off with TokenLab Verified; see the catalog for each model. Wan 2.6 TokenLab Verified changes from $0.60 per request to duration-based pricing: about $0.08824/second at 720p and $0.14706/second at 1080p. A 5-second video is about $0.4412 or $0.7353, respectively. Gemini 3 Pro Image, Gemini 3.1 Flash Image and Gemini 3.1 Flash Lite Image bill text and image inputs and outputs separately, with TokenLab Verified remaining 50% off.

Applies to new requests. The model catalog lists exact rates for each setting; historical charges are unchanged.

Models
wan-2.6gemini-3-pro-imagegemini-3.1-flash-imagegemini-3.1-flash-lite-imagegrok-imagine-image-2.0grok-imagine-videogrok-imagine-video-1.5-previewhailuo-2.3-standardimage-background-removerimage-upscalerkling-v2.1-masterqwen-imageqwen-image-editz-image
Effective
View model prices and availability
Pricing

Price increases for selected models

For TokenLab Verified, GPT-5.4 and GPT-5.4 mini change from 70% off to 15% off; Kling 3.0 changes from 30% off to the listed official price. HappyHorse 1.1 at 720p, 5 seconds, without audio rises from about $0.4332 to $0.6188. At Default quality, Ideogram v3 rises from $0.04 to $0.06 per image; Edit, Reframe and Remix v3 rise from $0.02 to $0.06 per image.

Applies to new TokenLab Verified requests. Other settings are priced as shown in the model catalog; previously accepted requests and historical charges are unchanged.

Models
gpt-5.4gpt-5.4-minikling-v3.0happyhorse-1.1ideogram-v3ideogram-edit-v3ideogram-reframe-v3ideogram-remix-v3
Effective
View model prices and availability
Pricing

Chat and embedding model discounts updated

For the models in this update, TokenLab Verified offers 70% off Claude models and GPT-5.6 Luna; 15% off the other GPT, o3, and text-embedding-3 models; and 50% off Grok 4.6 and 4.7. Affected models are listed below; see the catalog for individual prices.

Applies to new TokenLab Verified requests. Previously accepted requests and historical charges are unchanged.

Models
claude-fable-5claude-fable-5-1claude-haiku-4-5claude-opus-4-6claude-opus-4-7claude-opus-4-8claude-opus-5claude-opus-5-5claude-sonnet-4-6claude-sonnet-5gpt-5.6-lunagpt-4.1-minigpt-4ogpt-4o-minigpt-5gpt-5-minigpt-5-nanogpt-5.1gpt-5.2gpt-5.3-codexgpt-5.4-nanogrok-4.6grok-4.7o3text-embedding-3-largetext-embedding-3-small
Effective
Price
15%–70% off
View model prices and availability
Models

Availability changes for selected models

Seedance 1.5 Pro and Suno Lyrics no longer accept new requests. Seedance 2.0 Fast, Seedance 2.0 Mini and FLUX 2 Flex text-to-image are available through Auto or Official; FLUX image editing still supports TokenLab Verified. Wan 2.6 continues to support text-to-video and image-to-video, but no longer supports reference-to-video.

Keys restricted to TokenLab Verified cannot use these Official-only options. Previously accepted tasks and historical records remain available.

Models
seedance-1.5-prosuno-lyricsseedance-2.0-fastseedance-2.0-miniwan-2.6flux-2-flex
Effective
View model prices and availability
Pricing

Cache and long-context billing updates

For Official delivery, GLM-5, GLM-5.1 and GLM-5.2 cache creation now uses the standard input rate; cached reads retain their listed discounted rate and TokenLab Verified prices are unchanged. Long-context rates apply to Grok 4.20 at 200,000 input tokens or more and Qwen 3.6 Max Preview above 128,000 input tokens. Grok 4.20 TokenLab Verified remains 50% off the applicable tier; Qwen TokenLab Verified prices are unchanged.

Applies to new requests that meet these conditions. See model pricing for rates; historical charges are unchanged.

Models
grok-4.20qwen3.6-max-previewglm-5glm-5.1glm-5.2
Effective
View model prices and availability
Pricing

Gemini 2.5 Flash and Flash Lite: 50% off

Gemini 2.5 Flash is available with TokenLab Verified pricing of $0.15 per million input tokens and $1.25 per million output tokens. Gemini 2.5 Flash Lite is $0.05 input and $0.20 output per million tokens. Both are 50% off the listed official prices; cached-input prices are $0.015 and $0.005 per million tokens respectively.

Models
gemini-2.5-flashgemini-2.5-flash-lite
Effective
Price
50% off (TokenLab Verified)
View model prices
Models

GPT-6 Sol, GPT-6 Luna, and Claude Opus 5.5 are now available

Use gpt-6-sol, gpt-6-luna, and claude-opus-5-5 on TokenLab. See the model catalog for their current capabilities and prices.

Models
gpt-6-solgpt-6-lunaclaude-opus-5-5
Effective
Browse models
Platform

Console chat: streaming, tools and web search

Chat in Console streams replies and can use function tools and web search, with tool results fed back into the conversation. Define function tools in the JSON editor. Each reply shows its time to first token.

Platform

Media workbench: multi-image edits and first/last-frame video

Edit an image with several reference images, or set the first and last frame of a video. Each control shows its allowed range.

Open Console
Platform

Preview and download results in task details

Open a task to preview and download the media files and 3D models it generated.

Platform

Failed requests show the cause and the refund

In Console, the details of a failed request show what went wrong and the exact amount refunded. Only members of your organization can see them.

Open Console
Platform

Search models, pages, and docs in one place

Find models, site pages, and documentation from one search box. You can also limit the search to Console or the public site.

View documentation
Platform

Clearer retention information for uploaded media

Private images uploaded as task inputs are removed when the task’s retention period ends. The media workbench explains how long temporary attachments remain available.

Open Console
Platform

AI assistant setup guides on the dashboard

The dashboard overview now includes AI assistant setup guides, with tutorials for Cursor and DeepSeek Harness.

Open Console
Pricing

Clearer refund amounts and rate units in billing

Failed or adjusted requests show exact refund amounts. Billing details distinguish charge types, rate units, zero charges, and cache rates so you can check each charge, refund, and applicable rate.

Platform

TokenLab 2.0 website and Console

The TokenLab 2.0 website and Console bring model selection, API keys, usage, and billing into one workspace.

Existing clients require no changes.

Open Console
Models

DeepSeek V4.1 Flash is available

DeepSeek V4.1 Flash is now in the catalog as deepseek-v4.1-flash. Existing requests using deepseek-flash continue to work.

Models
deepseek-v4.1-flash
Effective
View models
Models

GPT Image 2.5 Sunburst and Flare are available

GPT Image 2.5 Sunburst and Flare are now in the catalog, with image generation and editing through the TokenLab Images API.

Models
gpt-image-2.5-sunburstgpt-image-2.5-flare
Effective
View models
Platform

See image rates before you generate

For images billed by usage, the media workbench shows the output token rate and the minimum amount held before generation. You pay for actual usage, and the page tells you when the total cannot be known in advance.

Open Console
Models

Gemini 3.8 Flash and GPT 6 Astra added

Gemini 3.8 Flash and GPT 6 Astra are now available in the TokenLab model catalog.

Models
gemini-3.8-flashgpt-6-astra
Effective
Browse models
Platform

Catalog shows model developers and lifecycle status

The model catalog shows each model’s developer and lifecycle status, so you can review it before integration.

Platform

Choose a delivery mode for each request

Auto, TokenLab Verified, and Official each show their corresponding price. Set a workspace default or choose a delivery mode per request with the X-TokenLab-Delivery-Policy header. If the selected mode is unavailable, the API returns delivery_tier_unavailable. Request records show the requested and resolved delivery modes.

Open in Console
Pricing

Requests that fail before they run are refunded

If a request fails before it is sent for processing and its outcome cannot be confirmed, its charge is refunded automatically.

Platform

Usage and cost breakdowns in the dashboard

View usage and cost by API key, model, and time range to understand spend without first exporting raw records.

Platform

Compare models in one catalog view

The model catalog combines weekly usage, task rankings, and unit prices in one view.

Browse models
API

Video requests support element references and first/last frames

Use element references to keep a subject consistent across a video, and set the first and last frame of the clip. Supported parameters differ by model; see the model's API reference.

API

Run image tasks in the background

Submit an image generation or edit as a background task and get a task ID. Poll its status, or get a webhook when it finishes. The task reports progress and the final cost, so a long render does not hold a connection open.

Use a background task when a render can take longer than your client timeout.

Models

GLM 5.3, Grok 4.6, and Gemini 3.7 Flash added

GLM 5.3, Grok 4.6, and Gemini 3.7 Flash are now available in the TokenLab model catalog.

Models
glm-5.3grok-4.6gemini-3.7-flash
Effective
Browse models
API

Rate-limit errors tell you when to retry

A 429 response includes a Retry-After header, so your client can wait the right time before it retries.

API

TokenLab MCP server for coding agents

Coding agents can look up model IDs, prices and API details through the TokenLab MCP server before they write integration code. Lookups do not call a model, so they are not charged.

API

Responses API over WebSocket

If your client cannot keep a server-sent events (SSE) stream open, create and continue Responses API calls over WebSocket instead.

API

Speech, transcription and translation APIs

Generate speech from text, transcribe audio and translate audio with the same API key and balance you use for text models.

API

Model list includes capabilities

The models API returns the capabilities of each model, so tools and agents can check which models they can call and what each one supports without reading the docs.

API

Reusable Seedance reference assets

Upload reference images and videos once as Seedance assets, validate them, and reuse them across video requests. Manage them through the API or on the Seedance assets page in Console.

2026 Q2

API

3D generation API

Create 3D models through the API. Each job returns a task ID: check its status, then download the result files when it is done.

API

Usage export and Management API

Export usage records from Console, or read usage and billing for each API key through the Management API. Use them for reconciliation and internal dashboards.

API

API base URL migration and compatibility

The API base URL moved to https://api.tokenlab.sh. The previous address, https://api.lemondata.cc, remained available during the announced compatibility window ending July 1, 2026.

Use https://api.tokenlab.sh for new integrations.

Pricing

Streaming charges match the reported usage

For streamed chat, you are charged for the final usage the API returns, so your usage records and charges agree.

2026 Q1

API

Embedding and rerank APIs

Call embedding and rerank models with the same API key and balance as chat. Errors use the same format as the other endpoints.

API

Webhooks for async tasks

Get a webhook when an async task completes, fails or times out, instead of polling its status.

Pricing

Failed async tasks are refunded automatically

If an async task fails, its charge goes back to your balance automatically.

Pricing

Cached input tokens are billed at the cache rate

For supported models, cached input tokens are now included in the final charge at the cache rate instead of the standard input rate.

Fixes and improvements11
  • Usage exports match filtered dashboard records

    Fixed usage queries and exports for long date ranges and filtered views so exported records match those shown in the dashboard.

  • Image edits keep the order of input images

    Image edits with several input images use them in the order you sent them. If a request runs past your client timeout, you can still get the result from the task record.

  • Billing shows small charges and refunds

    Small charges and refunds are no longer rounded to zero in billing views, so you can see per-request costs below one cent.

  • Transcription and synchronous speech responses match the docs

    Fixed response formats for transcription and synchronous speech endpoints so returned fields and structures match the documentation.

  • Responses are marked correctly when they reach the length limit

    Responses that reach the length limit now carry the correct truncation marker. The interface indicates that the output is incomplete.

  • Task webhooks retry after delivery failures

    Task completion webhooks are retried after a failed delivery, so notifications do not rely on a single attempt.

  • Video requests without a published price are rejected

    A video request is rejected up front when its size and mode combination has no published rate, avoiding acceptance at an unclear price. Reference-to-video uses the same published rates as the rest of the video catalog.

  • Text-only requests to vision models work normally

    A request with only text, sent to a model that also accepts images, now returns a normal text reply instead of an error.

  • Short rate limits are retried automatically

    When a request hits a short rate limit and a retry is safe, TokenLab retries it for you, so fewer requests fail.

  • One error format for all models

    Errors from every model use the same response format and stable error types, so you can handle failures in one place.

  • Oversized image requests are now rejected early

    Image upload requests that exceed the supported size are rejected before processing. Customers receive a clear error immediately, and the request does not consume usage.

See it for yourself

Try the latest changes in the model catalog, or check the docs before updating an integration.