Choose Auto, TokenLab Verified, or Official for each request, with prices shown up front.See what's new

Kuaishou: Kling IMAGE O3

Kling IMAGE O3 is an Omni image model built for high-fidelity text-to-image and image-to-image generation at up to 4K resolution. It supports multi-image reference prompting, series image generation for coherent variations, and optional face-focused element control to keep identity stable across outputs.
Compare models
kling-image-o3
AvailableKuaishouImageSync
Price
From $0.028
Modalities
Image

About Kling IMAGE O3

Kling IMAGE O3 is the Omni image model in Kling's 3.0 generation, made for text-to-image and image-to-image work at up to 4K resolution. Where IMAGE 3.0 is prompt-first, O3 is built for control: it takes several reference images in one prompt, produces series of coherent variations, and offers optional face-focused element control to keep an identity stable from one output to the next.

Where it works well

  • Several reference images can be combined in one prompt, so a subject, a style, and a scene can each come from a different source.
  • Series generation produces coherent variations of one idea instead of unrelated single images.
  • Optional face-focused element control keeps a person's identity stable across outputs.
  • It handles both text-to-image and image-to-image in the same model, with output up to 4K.
  • Kling describes strong text rendering, useful for posters and packaging mock-ups.

When to choose another model

  • Reference-heavy prompts take more setup than a one-line prompt; for quick single images IMAGE 3.0 is the lighter option.
  • Identity control is aimed at faces, so tight consistency for objects or logos still needs checking on your own assets.
  • It outputs still images only; motion needs one of the Kling video models.

Getting started

  1. Create API key

    Create a key in Console, then use it with every model on the platform.

  2. Send your first request

    Copy the example for your language and run it against the endpoint.

    Image editingText to imageTokenLab endpoint
    POST/v1/images/generations
    curl -X POST "https://api.tokenlab.sh/v1/images/generations" \
      -H "Authorization: Bearer sk-xxx" \
      -H "Content-Type: application/json" \
      -d '{
      "operation": "image-to-image",
      "model": "kling-image-o3",
      "prompt": "A minimalist product photo of matte black headphones on a soft blue background.",
      "image_url": "https://example.com/source.png"
    }'

Pricing

Official price is the model maker's public baseline. TokenLab price is what you pay for this model on TokenLab.

Image editing

per image
Official price
$0.028
Official
$0.028
Discount
—

Text to image

per image
Official price
$0.028
Official
$0.028
Discount
—

Usage & activity

Success rate is the share of requests that completed. Latency is how long a full response takes; P95 means 95% of requests finished within that time.

Usage & availability

Last 24 hours
Requests
Success rate
P95 latency
Total tokens
Model performance
Metrics appear once privacy and data volume thresholds are met.

Data is based on aggregate user requests, excluding status checks.

Open in Console

Open Kling IMAGE O3 in Console with a prompt ready to edit or send.

Help me create an image with kling-image-o3 using image-edit at /v1/images/edits. Show the result, cost, and any limits I should know.

Use cases

Best for
  • Text to Image
  • Character sheets

    Create one character in many poses and outfits using a face reference and series generation, then hand the set to an animation or video step.

  • Brand campaign sets

    Blend a product photo, a mood reference, and a location reference into a family of consistent hero images.

  • Poster and packaging mock-ups

    Generate artwork where headline text needs to read correctly, at a resolution that holds up in print previews.

  • Storyboard frames

    Produce consecutive frames with the same cast and setting before committing to video generation.

Prompt examples

Use the face from image 1 and the jacket from image 2. Show her walking through a rainy neon market at night, 35mm look.

Create four coherent variations of this mascot: waving, running, sleeping, and holding a flag.

A minimalist coffee bag, the text 'NORTH ROAST' in bold serif across the front, studio softbox lighting.

FAQ

What does the O in Kling IMAGE O3 mean?

It marks the Omni line, Kling's multimodal tier. In the image model it means the prompt can mix text with several reference images and element controls, rather than text alone. The result is more control over what appears in the output.

How many reference images can it use?

Kling describes multi-image reference prompting, where one prompt draws on several reference images at once. Use separate references for the subject, the style, and the scene to keep each source clear.

Does it support text-to-image and editing?

Yes. It supports both text-to-image and image-to-image, and Kling positions the model for high-fidelity results from either starting point, at resolutions up to 4K, so one model covers both jobs.

Can it keep a face the same across many images?

That is the purpose of its optional face-focused element control together with series generation. Results still vary by subject, so run a small batch on your actual reference photos and compare the faces before scaling up.

How much does Kling IMAGE O3 cost?

On TokenLab, Kling IMAGE O3 costs $0.028 per image. The pricing table above shows the full breakdown. Rates depend on the billing unit, specification, and usage. Compare matching conditions in the model's detailed pricing; a single rate does not determine the total cost.

Which endpoint should Kling IMAGE O3 use?

Use https://api.tokenlab.sh/v1/images/edits for Kling IMAGE O3. The request example below shows the matching code shape.

Which operations does Kling IMAGE O3 support?

Kling IMAGE O3 supports Image editing, Text to image. Select an operation above to see its endpoint and request example.

Compare Kling IMAGE O3

Sources

Reviewed Oct 2, 2026

More from Kling

Related models