Choose Auto, TokenLab Verified, or Official for each request, with prices shown up front.See what's new

ByteDance: OmniHuman-1.5

OmniHuman-1.5 generates high fidelity avatar video from a single image with audio and optional text prompts. It fuses multimodal reasoning with diffusion motion to keep identity stable, lip sync accurate, and gestures context aware for long, multi subject clips.
Compare models
omnihuman-1-5
AvailableByteDanceVideoAsync
Price
From $0.1325
Modalities
Video

About OmniHuman-1.5

OmniHuman-1.5 is a digital-human video model from ByteDance's Intelligent Creation Lab. From one image of a person, an audio clip and an optional text prompt, it generates a talking or performing avatar video with lip sync, gestures and body motion that follow the speech. It is built for characters that look like they are reacting, including scenes with more than one speaker.

Where it works well

  • Animates a single still image, so no filmed footage of the person is needed.
  • Lip sync follows the supplied audio, and gestures are chosen to match the meaning of the speech.
  • An optional text prompt can steer the action or camera behaviour on top of the audio.
  • Supports scenes driven by two-person audio, with the characters reacting to each other.

When to choose another model

  • It needs an audio track; it is not a general text-to-video model for scenes without a speaker.
  • Output quality depends on the reference image, and a cropped or low-resolution face limits identity fidelity.
  • Generated people and voices raise consent questions, so use only images and audio you have rights to.

Getting started

  1. Create API key

    Create a key in Console, then use it with every model on the platform.

  2. Send your first request

    Copy the example for your language and run it against the endpoint.

    Image to videoTokenLab endpoint
    POST/v1/videos/generations
    curl -X POST "https://api.tokenlab.sh/v1/videos/generations" \
      -H "Authorization: Bearer sk-xxx" \
      -H "Content-Type: application/json" \
      -d '{
      "operation": "image-to-video",
      "model": "omnihuman-1-5",
      "prompt": "Cinematic sunrise over a calm lake with gentle camera motion.",
      "image_url": "https://example.com/image.jpg"
    }'

Pricing

Official price is the model maker's public baseline. TokenLab price is what you pay for this model on TokenLab.

Image to video

per second
Official price
$0.1325
Official
$0.1325
Discount
—

Usage & activity

Success rate is the share of requests that completed. Latency is how long a full response takes; P95 means 95% of requests finished within that time.

Usage & availability

Last 24 hours
Requests
Success rate
P95 latency
Total tokens
Model performance
Metrics appear once privacy and data volume thresholds are met.

Data is based on aggregate user requests, excluding status checks.

Open in Console

Open OmniHuman-1.5 in Console with a prompt ready to edit or send.

Help me create with omnihuman-1-5 using image-to-video at /v1/videos/generations. Show the result, status, and cost.

Use cases

  • Talking presenters

    Turn a portrait and a narration track into a spokesperson video for training, product updates or explainers.

  • Character dialogue

    Animate illustrated or photographic characters speaking a scripted line for short stories and social clips.

  • Localised video

    Reuse one presenter image with audio recorded in several languages, producing a matching lip-synced clip for each.

Prompt examples

A woman at a desk speaks to camera, calm gestures, nods slightly at key points

Two friends on a park bench talk to each other, turning toward the person speaking

The singer performs on a small stage, expressive hand movements, camera stays steady

FAQ

What inputs does OmniHuman-1.5 take?

A single image of the character, an audio clip and, optionally, a text prompt. The image sets identity and appearance, the audio drives lip sync and timing, and the prompt can guide the action.

Can it generate videos with more than one person?

Yes. ByteDance describes support for audio with two speakers, so interaction between characters can be generated in one clip, and each person's movement follows their own speech.

How does OmniHuman-1.5 decide how the avatar moves?

ByteDance's paper describes a multimodal language model that plans the action from the audio, image and prompt, and a diffusion transformer that renders the motion. That is meant to make gestures fit the content rather than loop generically.

Is it a general video generator?

No. It is specialised for people who speak or perform to audio. For scenes, landscapes or action footage without a driving audio track, use a general text-to-video or image-to-video model.

How much does OmniHuman-1.5 cost?

On TokenLab, OmniHuman-1.5 costs $0.1325 per second. The pricing table above shows the full breakdown. Rates depend on the billing unit, specification, and usage. Compare matching conditions in the model's detailed pricing; a single rate does not determine the total cost.

Which endpoint should OmniHuman-1.5 use?

Use https://api.tokenlab.sh/v1/videos/generations for OmniHuman-1.5. The request example below shows the matching code shape.

Which operations does OmniHuman-1.5 support?

OmniHuman-1.5 supports Image to video. Select an operation above to see its endpoint and request example.

Compare OmniHuman-1.5

Sources

Reviewed Oct 2, 2026

Related models