TokenLab

Media Guides

Video Generation

Create videos from text, images, audio, or existing clips

Video generation is asynchronous. POST /v1/videos/generations returns a task ID and usually a poll_url. Use that URL until the finished video is ready.

Send public HTTP(S) URLs or supported data URLs in the image fields accepted by your selected model. These inputs follow the normal media path and do not automatically create reusable material IDs.

When an explicit material is still preparing, POST /v1/videos/generations returns 409 seedance_material_preparing with inactive_asset_ids. Poll those assets until ACTIVE, then retry with the same material IDs. If an asset is FAILED, inspect error_message and correct or reimport it before retrying.

Choose an API

Use /v1/videos/generations for a new video integration. Existing Seedance 2.0 clients that already send Volcengine content[] or Action requests can keep that format through /api/v3. Both APIs use a TokenLab Bearer key, but their fields and status values are different.

Choose what to create

Set operation explicitly so the API can validate the right fields for the job.

OperationRequired or typical inputUse case
text-to-videopromptGenerate from text only
image-to-videoimage_url or compatible imageAnimate a starting image
reference-to-videoreference_images and optional video_urls / audio_urls on supported modelsKeep identity, style, or asset references
start-end-to-videostart_image, end_imageControl first and last frames
video-to-videovideo_url or model-specific task_idTransform or upscale an existing clip
motion-controlimage_url plus video_urlApply motion reference to a subject
audio-to-videoaudio_urlAudio-conditioned video flows
video-extensiontask_id, extend_at, or model-specific extension fieldsContinue a generated video

Choose a model

curl "https://api.tokenlab.sh/v1/models?recommended_for=video" \
  -H "Authorization: Bearer sk-your-api-key"

Use a model ID returned by TokenLab, then choose the generation type with operation and the corresponding media inputs. Do not use an operation name as the model ID.

Read the selected model detail before relying on specialized fields such as reference_images, kling_elements, output_audio, duration, resolution, or aspect_ratio.

Create a video

curl https://api.tokenlab.sh/v1/videos/generations \
  -H "Authorization: Bearer sk-your-api-key" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "veo3.1",
    "operation": "text-to-video",
    "prompt": "A calm cinematic shot of a cat walking through a sunlit garden",
    "duration": 4,
    "aspect_ratio": "16:9"
  }'

Use https URLs for media that can be fetched from the internet. Temporary URLs must stay valid until TokenLab creates the task.

Model-specific fields

Audio behavior depends on the selected model and operation. A video may contain sound even when no audio selector is exposed. Omitting a selector is different from sending false.

  • veo3.1 and veo3.1-fast always generate audio under the Gemini API contract. For wan-2.6 and wan-2.7 video generation, audio cannot be disabled. Omit output_audio or use true where the model detail permits it.
  • hailuo-h3 and Grok video models generate native audio. Do not add an audio switch that the selected model detail does not list.
  • Seedance 1.5/2.x and viduq3-pro / viduq3-turbo default to audio on and support silent output. PixVerse C1/V5.6/V6 default to audio off. Use output_audio only for operations that list it; Vidu also accepts its declared boolean audio field.
  • audio_url / audio_urls provide input or reference audio. They are not audio-output switches. Video editing, motion transfer, and style transfer may retain the input soundtrack; preserving source audio does not mean muting it.

Check the model detail for allowed values and audio-dependent prices. Supported aliases outputAudio, generate_audio, and boolean audio must agree with output_audio when combined. Do not assume every version or operation in a model family shares the same audio controls.

  • For the Seedance 2.0 family, read the Seedance 2.0 video models guide before using 4K output, Fast/Mini resolution caps, or multimodal reference inputs.
  • For grok-imagine-video video-to-video, send prompt and a public HTTPS .mp4 video_url. This operation does not use duration, resolution, or aspect_ratio selectors.

PixVerse and HappyHorse

ModelOperationsInputsResolutionDurationAudio selector
pixverse-c1, pixverse-v6text-to-video, image-to-video, start-end-to-video, reference-to-videoprompt; image_url; start_image + end_image; reference_images360p, 540p, 720p, 1080pAny integer from 1 to 15 secondsoutput_audio, default false
pixverse-v5.6text-to-video, image-to-video, start-end-to-video, reference-to-videoSame fields as C1 and V6360p, 540p, 720p, 1080p5, 8, or 10 seconds; 1080p supports 5 or 8 secondsoutput_audio, default false
happyhorse-1.0text-to-video, image-to-video, reference-to-video, video-to-videoprompt; image_url; reference_images; video_url + reference_images720p, 1080p3 to 15 seconds for generation operations; video-to-video output is capped at 15 secondsDo not send output_audio

On TokenLab, the PixVerse models above do not accept operation=video-extension.

curl https://api.tokenlab.sh/v1/videos/generations \
  -H "Authorization: Bearer sk-your-api-key" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "pixverse-v6",
    "operation": "image-to-video",
    "prompt": "A slow camera move through a neon-lit street",
    "image_url": "https://example.com/start.jpg",
    "resolution": "1080p",
    "duration": 5,
    "output_audio": true
  }'

Get the finished video

Use the returned poll_url. If you need a fixed URL, call GET /v1/tasks/{id} with the id or task_id from the create response.

Completed video tasks may return video_url, video, or videos depending on the model and output count. Treat billing_transaction_id as a billing identifier, not as a task identifier.

Eligible generated video HTTP(S) result URLs may be retained as media copies for 30 days. Check media_retention.items for each item's status and expires_at; pending or failed copies are not guaranteed. Async task snapshots have a separate lifecycle. See Data retention.

Common Pitfalls

  • Do not hard-code old video status paths; prefer poll_url.
  • Do not combine first-frame fields with dedicated reference-image flows unless the model details allows it.
  • Do not assume duration describes input reference video length; it usually controls generated output length.
  • Do not retry create requests after a timeout without checking whether a task was already created.

API Reference

TopicReference
Create VideoCreate Video
Get Video StatusGet Video Status
Get Task StatusGet Task Status
Cancel TaskCancel Task
Billing & PricingBilling & Pricing

Hailuo H3-Max generates 5–15-second videos at 480p or 768p from text, a first frame, or first and last frames. Its speed-focused workflow is useful for quickly turning a shot idea into a short clip.

{
  "model": "hailuo-h3-max",
  "operation": "text-to-video",
  "prompt": "A slow camera move through a quiet garden",
  "resolution": "768p",
  "duration": 5,
  "aspect_ratio": "16:9"
}

On this page