ByteDance: OmniHuman-1.5
omnihuman-1-5- Price
- From $0.1325
- Modalities
- Video
About OmniHuman-1.5
OmniHuman-1.5 is a digital-human video model from ByteDance's Intelligent Creation Lab. From one image of a person, an audio clip and an optional text prompt, it generates a talking or performing avatar video with lip sync, gestures and body motion that follow the speech. It is built for characters that look like they are reacting, including scenes with more than one speaker.
Where it works well
- Animates a single still image, so no filmed footage of the person is needed.
- Lip sync follows the supplied audio, and gestures are chosen to match the meaning of the speech.
- An optional text prompt can steer the action or camera behaviour on top of the audio.
- Supports scenes driven by two-person audio, with the characters reacting to each other.
When to choose another model
- It needs an audio track; it is not a general text-to-video model for scenes without a speaker.
- Output quality depends on the reference image, and a cropped or low-resolution face limits identity fidelity.
- Generated people and voices raise consent questions, so use only images and audio you have rights to.
Getting started
Create API key
Create a key in Console, then use it with every model on the platform.
Send your first request
Copy the example for your language and run it against the endpoint.
Image to videoTokenLab endpointPOST/v1/videos/generationscurl -X POST "https://api.tokenlab.sh/v1/videos/generations" \ -H "Authorization: Bearer sk-xxx" \ -H "Content-Type: application/json" \ -d '{ "operation": "image-to-video", "model": "omnihuman-1-5", "prompt": "Cinematic sunrise over a calm lake with gentle camera motion.", "image_url": "https://example.com/image.jpg" }'
Pricing
Official price is the model maker's public baseline. TokenLab price is what you pay for this model on TokenLab.
Image to video
per second- Official price
- $0.1325
- Official
- $0.1325
- Discount
- —
| Official priceper second | Officialper second | Discount | |
|---|---|---|---|
| Image to video | $0.1325 | $0.1325 | — |
Usage & activity
Success rate is the share of requests that completed. Latency is how long a full response takes; P95 means 95% of requests finished within that time.
Usage & availability
Last 24 hours- Requests
- Success rate
- P95 latency
- Total tokens
Data is based on aggregate user requests, excluding status checks.
Open in Console
Open OmniHuman-1.5 in Console with a prompt ready to edit or send.
Help me create with omnihuman-1-5 using image-to-video at /v1/videos/generations. Show the result, status, and cost.
Use cases
Talking presenters
Turn a portrait and a narration track into a spokesperson video for training, product updates or explainers.
Character dialogue
Animate illustrated or photographic characters speaking a scripted line for short stories and social clips.
Localised video
Reuse one presenter image with audio recorded in several languages, producing a matching lip-synced clip for each.
Prompt examples
A woman at a desk speaks to camera, calm gestures, nods slightly at key points
Two friends on a park bench talk to each other, turning toward the person speaking
The singer performs on a small stage, expressive hand movements, camera stays steady
FAQ
What inputs does OmniHuman-1.5 take?
A single image of the character, an audio clip and, optionally, a text prompt. The image sets identity and appearance, the audio drives lip sync and timing, and the prompt can guide the action.
Can it generate videos with more than one person?
Yes. ByteDance describes support for audio with two speakers, so interaction between characters can be generated in one clip, and each person's movement follows their own speech.
How does OmniHuman-1.5 decide how the avatar moves?
ByteDance's paper describes a multimodal language model that plans the action from the audio, image and prompt, and a diffusion transformer that renders the motion. That is meant to make gestures fit the content rather than loop generically.
Is it a general video generator?
No. It is specialised for people who speak or perform to audio. For scenes, landscapes or action footage without a driving audio track, use a general text-to-video or image-to-video model.
How much does OmniHuman-1.5 cost?
On TokenLab, OmniHuman-1.5 costs $0.1325 per second. The pricing table above shows the full breakdown. Rates depend on the billing unit, specification, and usage. Compare matching conditions in the model's detailed pricing; a single rate does not determine the total cost.
Which endpoint should OmniHuman-1.5 use?
Use https://api.tokenlab.sh/v1/videos/generations for OmniHuman-1.5. The request example below shows the matching code shape.
Which operations does OmniHuman-1.5 support?
OmniHuman-1.5 supports Image to video. Select an operation above to see its endpoint and request example.
Compare OmniHuman-1.5
Sources
- OmniHuman-1.5 project page
- OmniHuman-1.5: Instilling an Active Mind in Avatars via Cognitive Simulation (arXiv)
Reviewed Oct 2, 2026