Choose Auto, TokenLab Verified, or Official for each request, with prices shown up front.See what's new

Kuaishou: Kling VIDEO 3.0 Pro

Kling VIDEO 3.0 Pro is a unified multimodal video model that generates high-quality video with synchronized audio from text or images. It supports reference-guided generation, prompt-based editing, fine control over motion and pacing, and stable temporal coherence for cinematic and narrative clips. Native audio output includes dialogue, ambient sound, and effects aligned to the visuals.
Compare models
kling-video-3-0-pro
AvailableKuaishouVideoAsync
Price
From $0.112
Modalities
Video

About Kling VIDEO 3.0 Pro

Kling VIDEO 3.0 Pro is the 1080p tier of Kling's VIDEO 3.0 line, a unified multimodal model that makes video with synchronized audio from text or images. It adds finer control over motion and pacing than Standard and keeps temporal coherence for cinematic and narrative clips. Audio covers dialogue, ambient sound, and effects aligned to the visuals.

Where it works well

  • Native audio includes dialogue, ambience, and effects timed to what is on screen.
  • Control over motion and pacing is finer than in the Standard tier, which matters for narrative shots.
  • Reference-guided generation and prompt-based editing keep a subject recognizable while the scene changes.
  • Text-to-video and image-to-video are both available, as is the multi-shot storytelling of the 3.0 generation.
  • Output is 1080p, enough for most web and broadcast-adjacent delivery.

When to choose another model

  • Output tops out at 1080p; delivery that needs 4K should use the dedicated 4K variant.
  • It costs more per second than Standard, so exploring ideas there first saves budget.
  • It does not take reference video for voice and look replication; that is the Omni line's feature.

Getting started

  1. Create API key

    Create a key in Console, then use it with every model on the platform.

  2. Send your first request

    Copy the example for your language and run it against the endpoint.

    Image to videoText to videoTokenLab endpoint
    POST/v1/videos/generations
    curl -X POST "https://api.tokenlab.sh/v1/videos/generations" \
      -H "Authorization: Bearer sk-xxx" \
      -H "Content-Type: application/json" \
      -d '{
      "operation": "image-to-video",
      "model": "kling-video-3-0-pro",
      "prompt": "Cinematic sunrise over a calm lake with gentle camera motion.",
      "image_url": "https://example.com/image.jpg"
    }'

Pricing

Official price is the model maker's public baseline. TokenLab price is what you pay for this model on TokenLab.

Image to video

per second
Official price
$0.112
Official
$0.112
Discount
—

Text to video

per second
Official price
$0.112
Official
$0.112
Discount
—

Usage & activity

Success rate is the share of requests that completed. Latency is how long a full response takes; P95 means 95% of requests finished within that time.

Usage & availability

Last 24 hours
Requests
Success rate
P95 latency
Total tokens
Model performance
Metrics appear once privacy and data volume thresholds are met.

Data is based on aggregate user requests, excluding status checks.

Open in Console

Open Kling VIDEO 3.0 Pro in Console with a prompt ready to edit or send.

Help me create with kling-video-3-0-pro using image-to-video at /v1/videos/generations. Show the result, status, and cost.

Use cases

Best for
  • Text to Video
  • Narrative short scenes

    Generate a multi-shot scene with spoken lines and matching ambience, then cut it with others into a trailer or short film.

  • Ad creative

    Animate a hero product image with controlled camera motion and clean temporal behavior for paid social placements.

  • Final renders after drafts

    Re-run a prompt approved at 720p on Standard through Pro to get the finished 1080p version.

  • Music-video style cuts

    Direct pacing shot by shot, setting rhythm through the duration of each beat, with sound effects generated alongside.

Prompt examples

Medium shot, a chef plates a dish under warm kitchen lights, steam rising, clinking cutlery, slow push-in.

Image-to-video: the astronaut turns toward camera and lifts her visor; wind noise, then silence.

Shot 1: drone over a coastline at dawn. Shot 2: a surfer paddles out. Shot 3: close on his face, he laughs.

FAQ

What does Pro add over Kling VIDEO 3.0 Standard?

Pro renders at 1080p where Standard renders at 720p, with finer control of motion and pacing and stronger fidelity for cinematic clips. The feature set, native audio plus text-to-video and image-to-video, is the same in both tiers.

Is Pro the same as the Omni Pro model?

No. Omni Pro belongs to the Omni line, which adds reference-driven generation and prompt-based video editing as core features. Kling VIDEO 3.0 Pro is the standard 3.0 line, which also accepts references but is aimed at direct generation.

Does it generate dialogue?

Yes. Kling's 3.0 generation produces speech, ambient sound, and effects together with the video, and the 3.0 announcement describes scenes where each character speaks a different language.

What inputs does it take?

It takes text-to-video and image-to-video requests. Provide a prompt alone, or add a starting image and describe the motion. Multi-shot instructions can be written inside one prompt.

How much does Kling VIDEO 3.0 Pro cost?

On TokenLab, Kling VIDEO 3.0 Pro costs $0.112 per second. The pricing table above shows the full breakdown. Rates depend on the billing unit, specification, and usage. Compare matching conditions in the model's detailed pricing; a single rate does not determine the total cost.

Which endpoint should Kling VIDEO 3.0 Pro use?

Use https://api.tokenlab.sh/v1/videos/generations for Kling VIDEO 3.0 Pro. The request example below shows the matching code shape.

Which operations does Kling VIDEO 3.0 Pro support?

Kling VIDEO 3.0 Pro supports Image to video, Text to video. Select an operation above to see its endpoint and request example.

Compare Kling VIDEO 3.0 Pro

Sources

Reviewed Oct 2, 2026

More from Kling

Related models