Choose Auto, TokenLab Verified, or Official for each request, with prices shown up front.See what's new

MiniMax: MiniMax H3

Hailuo H3 creates video from text, still images, or visual references. Start and end frames help define a shot's direction, making it useful for turning a storyboard concept into a short motion sequence.
Compare models
hailuo-h3
AvailableMiniMaxVideoAsync
Price
From $0.08
Modalities
Video

About MiniMax H3

MiniMax H3, also called Hailuo H3, is MiniMax's general-purpose multimodal video model and the successor to Hailuo 2.3. It generates video from text, a still image, a start and end frame pair, or a set of reference images, videos and audio clips. It offers 768p and 2K resolutions, the higher range of the H3 family, which separates it from the speed-oriented H3-Max.

Where it works well

  • Start-and-end-frame generation fixes how a shot begins and finishes, which helps when turning a storyboard into motion.
  • Reference generation accepts images, short video clips and audio clips as guidance for subject, motion and sound.
  • Resolution options are 768p and 2K, and the 2K option is above H3-Max, which tops out at 768p.
  • Text-to-video and image-to-video are available, so one model covers prompt-only and still-based work.
  • Clips run six or ten seconds, long enough for a single shot with a clear beginning and end.

When to choose another model

  • Reference mode and first/last-frame mode cannot be combined in one request, so you choose between them per generation.
  • For fast iteration at lower resolution, MiniMax positions H3-Max as the quicker variant.
  • Aspect ratio is limited to 16:9 and adaptive, so other fixed ratios are not selectable.

Getting started

  1. Create API key

    Create a key in Console, then use it with every model on the platform.

  2. Send your first request

    Copy the example for your language and run it against the endpoint.

    Image to videoReference to videoStart and end frames to videoText to videoTokenLab endpoint
    POST/v1/videos/generations
    curl -X POST "https://api.tokenlab.sh/v1/videos/generations" \
      -H "Authorization: Bearer sk-xxx" \
      -H "Content-Type: application/json" \
      -d '{
      "aspect_ratio": "16:9",
      "duration": 6,
      "operation": "text-to-video",
      "resolution": "768p",
      "model": "hailuo-h3",
      "prompt": "Cinematic sunrise over a calm lake with gentle camera motion."
    }'

Pricing

TokenLab price applies to Verified, which costs less on most models. Official price is the model maker's published price and applies to the more reliable Official route. Auto charges the route that completes the request.

  • Image to videoOfficial
  • Reference to videoOfficial
  • Start and end frames to videoOfficial
  • Text to videoVerified, Official

Text to video · 768p

per second
Official price
$0.08
TokenLab price
$0.08
Discount
—

Image to video / Start and end frames to video / Reference to video · 768p

per second
Official price
$0.08
TokenLab price
—
Discount
—

Text to video · 2k

per second
Official price
$0.13
TokenLab price
$0.13
Discount
—

Image to video / Start and end frames to video / Reference to video · 2k

per second
Official price
$0.13
TokenLab price
—
Discount
—

Reference to video · Image input · First 5 images free

per second
Official price
$0.04/image
TokenLab price
—
Discount
—

Usage & activity

Success rate is the share of requests that completed. Latency is how long a full response takes; P95 means 95% of requests finished within that time.

Usage & availability

Last 24 hours
Requests
Success rate
P95 latency
Total tokens
Model performance
Metrics appear once privacy and data volume thresholds are met.

Data is based on aggregate user requests, excluding status checks.

Open in Console

Open MiniMax H3 in Console with a prompt ready to edit or send.

Help me create with hailuo-h3 using text-to-video at /v1/videos/generations. Show the result, status, and cost.

Use cases

Best for
  • Multimodal
  • Text to Video
  • Storyboard to shot

    Give a first and last frame and a prompt describing the move between them to get a motion test that lands on a planned final composition.

  • Character-consistent clips

    Supply reference images of a character or product and place them in new scenes across several clips.

  • Sound-aware scenes

    Add an audio clip as reference to steer pacing or ambience in the generated shot.

  • High-resolution hero shots

    Render a key shot at 2K for a website header or a film proof.

Prompt examples

Start on the closed door, end on the open door with light flooding in; slow dolly forward, dusty hallway.

Using the reference photo of the red sneaker, show it rotating on a concrete plinth under studio lights.

A paper lantern festival at night, camera rising above the crowd, warm glow reflecting on a river.

FAQ

How is MiniMax H3 different from Hailuo 2.3?

H3 is the newer, multimodal generation. Beyond text and single-image input it takes start and end frames and reference images, videos and audio, and it reaches 2K. Hailuo 2.3 takes a prompt and one image.

Can I set both the first and last frame?

Yes. Provide a start image and an end image, and the model animates between them. MiniMax's documentation treats this as a separate mode from reference generation.

What references can H3 use?

Up to nine images, three video clips and three audio clips, per MiniMax's documentation, with clip totals capped at fifteen seconds for video and for audio.

What resolutions does it offer?

768p and 2K according to MiniMax's documentation. The H3-Max variant trades the top resolution for speed and offers 480p and 768p instead of the larger sizes.

How long are the clips?

Six or ten seconds. Pick six seconds for a single action or a product reveal, and ten when the shot needs a camera move or a change within the scene. For longer sequences, generate several clips and join them in an editor.

How much does MiniMax H3 cost?

On TokenLab, MiniMax H3 costs $0.08-$0.13 per second. The pricing table above shows the full breakdown. Rates depend on the billing unit, specification, and usage. Compare matching conditions in the model's detailed pricing; a single rate does not determine the total cost.

Which endpoint should MiniMax H3 use?

Use https://api.tokenlab.sh/v1/videos/generations for MiniMax H3. The request example below shows the matching code shape.

Which operations does MiniMax H3 support?

MiniMax H3 supports Image to video, Reference to video, Start and end frames to video, Text to video. Select an operation above to see its endpoint and request example.

Compare MiniMax H3

Sources

Reviewed Oct 2, 2026

More from Hailuo

Related models