What you’ll learn
- Can I use more than two element references in a single Kling 3.0 request?
- What happens if I send both `kling_elements` and `output_audio=true`?
- Is element reference support specific to Kling 3.0, or available on other models too?
A product demo with a shoe and a hand needs two visual anchors, not one. We now support Kling 3.0 element references in the video generation API. The Kling 3.0 element references API lets developers anchor specific products, props, or characters to named tags (@name) that stay consistent throughout a generated clip. This closes a gap in image-conditioned video workflows where one reference image was not enough to keep multiple subjects visually stable across frames.
Why Multi-Subject Video Needs Element References
A single reference image works for one subject. It fails when a scene needs multiple distinct elements to persist. For example, a product held by a hand needs both the product and the hand to stay stable. Two characters exchanging dialogue need separate face and outfit continuity. We now support Kling 3.0 element references to address that gap.
Key Takeaways
- Kling 3.0 element references let you define named elements with reference image URLs, then call them by tag (
@productA,@character1) directly in your prompt text. - This is aimed at multi-subject scenes: product plus hand model, character plus prop, two characters in a dialogue shot, and similar setups where one reference image per request was previously limiting.
- Do not combine
kling_elementswithoutput_audio=truein the same request. The two parameters are mutually exclusive per the current API contract. - Element references sit alongside TokenLab's existing reference-to-video support for other models. They give developers a consistent pattern for choosing the right approach per use case.
How the Kling 3.0 Element References API Works
Most image-conditioned video generation treats a reference image as a single anchor. You give the model a picture, and it tries to keep the overall look consistent while animating motion around it. That works for single-subject shots. It breaks down fast when a scene needs more than one visually distinct element to persist independently.
Kling 3.0's element references solve this by letting you register multiple named reference images in a single request. You then point at them individually from within the prompt text. Instead of one implicit reference, you get explicit, addressable references. The model knows that @shoe refers to reference image one and @model refers to reference image two. It composes the scene using both anchors simultaneously.
We saw this pattern in the current API contract. It is a meaningful step up in control for product video pipelines, character-driven content tools, and ad creative generators. Subject consistency across a clip is often the difference between usable output and a reshoot.
Using Kling 3.0 Element References API in Requests
The pattern is straightforward: define your elements, name them, and reference them in the prompt with the @ syntax.
{
"model": "kling-3.0",
"prompt": "@shoe rotates slowly on a marble pedestal while @hand reaches in to pick it up",
"kling_elements": [
{
"name": "shoe",
"image_url": "https://example.com/product-shoe.png"
},
{
"name": "hand",
"image_url": "https://example.com/hand-reference.png"
}
],
"duration": 5,
"aspect_ratio": "16:9"
}
A few practical notes for implementation:
- Element names should be short and unambiguous. Avoid names that overlap with common English words already likely to appear in your prompt text. That overlap increases the chance of parsing ambiguity.
- Reference image URLs need to be publicly reachable at request time. If your images sit behind an authenticated storage layer, generate a signed or public URL before sending the request.
- You can combine multiple elements in one prompt, but keep the total scene description focused. Piling on more than two or three named elements tends to dilute the model's ability to track each one distinctly. This is similar to how too many named subjects in a static image prompt reduces per-subject fidelity.
- Test with a short duration first. Element consistency issues, if they occur, show up in the first couple of seconds. They are cheaper to catch on a 3-second draft than a full 10-second render.
Implementation Checklist
Before shipping a Kling 3.0 element-reference workflow to production, confirm the following:
- Each element has a unique, unambiguous name
- Reference image URLs are publicly accessible and stable for the duration of processing
- Prompt text correctly tags each element with
@namesyntax -
output_audiois not set totruewhenkling_elementsis present - Request validation catches the audio-plus-elements conflict before it reaches the API
- Test renders use short durations before committing to full-length generation
- Total named elements per request stays at two or three for best consistency
The One Rule: No Elements Plus Audio
This constraint is easy to miss during rapid prototyping: kling_elements and output_audio=true cannot be used in the same request. If you submit both, the request will not process as expected.
If your workflow needs both multi-element visual consistency and generated audio, split the work into two steps. Generate the video with element references first. Then run an audio generation pass separately and combine the outputs downstream. This is a documented constraint of the current Kling 3.0 integration, not a bug. Build your request validation logic around it instead of treating it as an edge case to catch after the fact.
In our pipeline, we validate this conflict client-side before sending the request.
Kling 3.0 Element References API vs. Other Video Workflows
Element references are one tool in a growing set of reference-to-video capabilities available through TokenLab's video API. It helps to know when to reach for which:
| Workflow | Best for | Reference count | Notes |
|---|---|---|---|
| Single image-to-video | Simple animation of one static image | 1 | Works across most supported video models, including Seedance and PixVerse V6 |
| Kling 3.0 element references | Multi-subject scenes needing independent consistency | 2-3 named elements | No audio in the same request |
| Style or motion reference | Applying a visual style or camera motion pattern | 1 style reference + prompt | Available on select models, check per-model docs |
| Text-only prompting | Fast iteration, no visual anchor needed | 0 | Fastest to prototype, least controllable |
If you are building a product demo generator, element references are usually the right call. If you are doing simple animation of a single hero image, plain image-to-video is faster and cheaper to iterate on. Teams comparing video models more broadly can start with the best AI video models for API use in 2026 breakdown. It covers how Kling 3.0 stacks up against Veo 3 and other options for different use cases.
FAQ
Can I use more than two element references in a single Kling 3.0 request?
Yes, the API does not hard-cap the count, but practical consistency tends to degrade as you add more named elements to a single scene. Two to three is a reasonable working limit for most product and character use cases.
What happens if I send both kling_elements and output_audio=true?
The request will not process correctly since these two parameters are mutually exclusive in the current Kling 3.0 integration. Validate this combination client-side before sending the request to avoid wasted calls.
Is element reference support specific to Kling 3.0, or available on other models too?
Named element references with @name tagging are specific to Kling 3.0 in the current API. Other supported video models have their own reference-to-video patterns. They are typically limited to a single reference image per request, so check the model-specific docs before assuming feature parity.
Sources, Freshness, and Related Reading
This article reflects the TokenLab video API documentation and Kling 3.0 integration behavior observed on 2026-07-07. For the current parameter reference, see the create video API reference and the video generation guide. API behavior can change, so always check the live docs before finalizing a production integration.
Element references expand what's possible with Kling 3.0, but choosing the right video model and understanding costs still matters before you build a production workflow. If you're comparing options, the Best AI Video Models API Guide: How Developers Should Choose Video Generation Models walks through the tradeoffs across providers. For a closer look at Kling specifically, the Kling AI API Pricing Guide: Cost, Workflow, and Alternatives breaks down pricing and workflow considerations. And if you're weighing alternatives, the Seedance API Guide: When to Use It for AI Video Generation covers when that model fits better.
Model capabilities and pricing change frequently, so verify current model versions and rates directly before relying on them for high-volume production use. The account setup reference explains API key creation.
To start building multi-subject video workflows, get your TokenLab API key and check the video generation guide.
Sources
Prices checked 2026-07-07
- TokenLab video generation API docsSources checked 2026-07-07
- TokenLab video generation guideSources checked 2026-07-07
- TokenLab model directorySources checked 2026-07-07



