MiniMax: MiniMax H3
hailuo-h3- Price
- From $0.08
- Modalities
- Video
About MiniMax H3
MiniMax H3, also called Hailuo H3, is MiniMax's general-purpose multimodal video model and the successor to Hailuo 2.3. It generates video from text, a still image, a start and end frame pair, or a set of reference images, videos and audio clips. It offers 768p and 2K resolutions, the higher range of the H3 family, which separates it from the speed-oriented H3-Max.
Where it works well
- Start-and-end-frame generation fixes how a shot begins and finishes, which helps when turning a storyboard into motion.
- Reference generation accepts images, short video clips and audio clips as guidance for subject, motion and sound.
- Resolution options are 768p and 2K, and the 2K option is above H3-Max, which tops out at 768p.
- Text-to-video and image-to-video are available, so one model covers prompt-only and still-based work.
- Clips run six or ten seconds, long enough for a single shot with a clear beginning and end.
When to choose another model
- Reference mode and first/last-frame mode cannot be combined in one request, so you choose between them per generation.
- For fast iteration at lower resolution, MiniMax positions H3-Max as the quicker variant.
- Aspect ratio is limited to 16:9 and adaptive, so other fixed ratios are not selectable.
Getting started
Create API key
Create a key in Console, then use it with every model on the platform.
Send your first request
Copy the example for your language and run it against the endpoint.
Image to videoReference to videoStart and end frames to videoText to videoTokenLab endpointPOST/v1/videos/generationscurl -X POST "https://api.tokenlab.sh/v1/videos/generations" \ -H "Authorization: Bearer sk-xxx" \ -H "Content-Type: application/json" \ -d '{ "aspect_ratio": "16:9", "duration": 6, "operation": "text-to-video", "resolution": "768p", "model": "hailuo-h3", "prompt": "Cinematic sunrise over a calm lake with gentle camera motion." }'
Pricing
TokenLab price applies to Verified, which costs less on most models. Official price is the model maker's published price and applies to the more reliable Official route. Auto charges the route that completes the request.
- Image to videoOfficial
- Reference to videoOfficial
- Start and end frames to videoOfficial
- Text to videoVerified, Official
Text to video · 768p
per second- Official price
- $0.08
- TokenLab price
- $0.08
- Discount
- —
Image to video / Start and end frames to video / Reference to video · 768p
per second- Official price
- $0.08
- TokenLab price
- —
- Discount
- —
Text to video · 2k
per second- Official price
- $0.13
- TokenLab price
- $0.13
- Discount
- —
Image to video / Start and end frames to video / Reference to video · 2k
per second- Official price
- $0.13
- TokenLab price
- —
- Discount
- —
Reference to video · Image input · First 5 images free
per second- Official price
- $0.04/image
- TokenLab price
- —
- Discount
- —
| Official priceper second | TokenLab priceper second | Discount | |
|---|---|---|---|
| Text to video · 768p | $0.08 | $0.08 | — |
| Image to video / Start and end frames to video / Reference to video · 768p | $0.08 | — | — |
| Text to video · 2k | $0.13 | $0.13 | — |
| Image to video / Start and end frames to video / Reference to video · 2k | $0.13 | — | — |
| Reference to video · Image input · First 5 images free | $0.04/image | — | — |
Usage & activity
Success rate is the share of requests that completed. Latency is how long a full response takes; P95 means 95% of requests finished within that time.
Usage & availability
Last 24 hours- Requests
- Success rate
- P95 latency
- Total tokens
Data is based on aggregate user requests, excluding status checks.
Open in Console
Open MiniMax H3 in Console with a prompt ready to edit or send.
Help me create with hailuo-h3 using text-to-video at /v1/videos/generations. Show the result, status, and cost.
Use cases
Best for- Multimodal
- Text to Video
Storyboard to shot
Give a first and last frame and a prompt describing the move between them to get a motion test that lands on a planned final composition.
Character-consistent clips
Supply reference images of a character or product and place them in new scenes across several clips.
Sound-aware scenes
Add an audio clip as reference to steer pacing or ambience in the generated shot.
High-resolution hero shots
Render a key shot at 2K for a website header or a film proof.
Prompt examples
Start on the closed door, end on the open door with light flooding in; slow dolly forward, dusty hallway.
Using the reference photo of the red sneaker, show it rotating on a concrete plinth under studio lights.
A paper lantern festival at night, camera rising above the crowd, warm glow reflecting on a river.
FAQ
How is MiniMax H3 different from Hailuo 2.3?
H3 is the newer, multimodal generation. Beyond text and single-image input it takes start and end frames and reference images, videos and audio, and it reaches 2K. Hailuo 2.3 takes a prompt and one image.
Can I set both the first and last frame?
Yes. Provide a start image and an end image, and the model animates between them. MiniMax's documentation treats this as a separate mode from reference generation.
What references can H3 use?
Up to nine images, three video clips and three audio clips, per MiniMax's documentation, with clip totals capped at fifteen seconds for video and for audio.
What resolutions does it offer?
768p and 2K according to MiniMax's documentation. The H3-Max variant trades the top resolution for speed and offers 480p and 768p instead of the larger sizes.
How long are the clips?
Six or ten seconds. Pick six seconds for a single action or a product reveal, and ten when the shot needs a camera move or a change within the scene. For longer sequences, generate several clips and join them in an editor.
How much does MiniMax H3 cost?
On TokenLab, MiniMax H3 costs $0.08-$0.13 per second. The pricing table above shows the full breakdown. Rates depend on the billing unit, specification, and usage. Compare matching conditions in the model's detailed pricing; a single rate does not determine the total cost.
Which endpoint should MiniMax H3 use?
Use https://api.tokenlab.sh/v1/videos/generations for MiniMax H3. The request example below shows the matching code shape.
Which operations does MiniMax H3 support?
MiniMax H3 supports Image to video, Reference to video, Start and end frames to video, Text to video. Select an operation above to see its endpoint and request example.
Compare MiniMax H3
Sources
Reviewed Oct 2, 2026