xAI: Grok Imagine Video
grok-imagine-video- Price
- From $0.025
- Modalities
- Video
About Grok Imagine Video
Grok Imagine Video is xAI's original Imagine video model for creating, editing and extending short clips. It accepts text, a starting image, reference images or existing footage, so one concept can move through several production steps. Compared with the 1.5 releases, which take an image as input, this model covers the full set of text, reference, edit and extension jobs.
Where it works well
- Handles text-to-video, image-to-video and reference-to-video, so the starting material can be words, a still or style references.
- Can edit existing footage with a text instruction through video-to-video.
- Video extension continues a clip from where it ended, letting you build longer sequences in steps.
- Clips run from one to fifteen seconds at 480p or 720p.
- Seven aspect ratios cover square, landscape and portrait delivery.
When to choose another model
- Resolution tops out at 720p, so final broadcast work needs upscaling or another model.
- Clips are short, so longer pieces require extension steps and careful continuity.
- Motion and physical detail can drift over a long clip, so short shots hold together best.
Getting started
Create API key
Create a key in Console, then use it with every model on the platform.
Send your first request
Copy the example for your language and run it against the endpoint.
Image to videoReference to videoText to videoVideo extensionVideo to videoTokenLab endpointPOST/v1/videos/generationscurl -X POST "https://api.tokenlab.sh/v1/videos/generations" \ -H "Authorization: Bearer sk-xxx" \ -H "Content-Type: application/json" \ -d '{ "aspect_ratio": "16:9", "duration": 5, "operation": "text-to-video", "resolution": "720p", "model": "grok-imagine-video", "prompt": "Cinematic sunrise over a calm lake with gentle camera motion." }'
Pricing
Official price is the model maker's public baseline. TokenLab price is what you pay for this model on TokenLab.
- Image to videoTokenLab Verified, Official
- Reference to videoOfficial
- Text to videoTokenLab Verified, Official
- Video extensionTokenLab Verified
- Video to videoOfficial
Pricing
per second- TokenLab price
- How pricing works
- Official price
- How pricing works
| Official priceper second | TokenLab priceper second | Discount | |
|---|---|---|---|
| Price | How pricing works | How pricing works | — |
Usage & activity
Success rate is the share of requests that completed. Latency is how long a full response takes; P95 means 95% of requests finished within that time.
Usage & availability
Last 24 hours- Requests
- Success rate
- P95 latency
- Total tokens
Data is based on aggregate user requests, excluding status checks.
Open in Console
Open Grok Imagine Video in Console with a prompt ready to edit or send.
Help me create with grok-imagine-video using text-to-video at /v1/videos/generations. Show the result, status, and cost.
Use cases
Best for- Text to Video
Product teasers
Animate a hero product still into a short loop, choosing portrait or square framing for social posts.
Reference-guided scenes
Provide character or style images and a prompt, and generate a clip that follows the references without locking the first frame.
Edit and extend footage
Restyle or alter a clip with a text instruction, then extend the final moment into a longer shot.
Prompt examples
A paper boat drifts down a rain-filled gutter, camera low and following, soft grey light, six seconds.
Using the attached character images, show her walking through a night market, handheld camera, neon reflections.
Extend this clip by five seconds: the drone keeps rising and reveals the whole coastline.
FAQ
What can Grok Imagine Video do?
It creates short clips from text, from an image, or from reference images, edits existing video with a text instruction, and extends a clip past its last frame. Duration is one to fifteen seconds.
What resolutions and ratios are available?
Output can be 480p or 720p, in square, 3:2, 2:3, 4:3, 3:4, 16:9 or 9:16. Choose the ratio that matches your destination, such as 9:16 for vertical feeds.
How is it different from grok-imagine-video-1.5?
This model covers text, reference, edit and extension jobs. The 1.5 models take a starting image and animate it. Use this one when you start from words, references or existing footage.
Can it extend a video I already have?
Yes. Extension continues a clip from its final moment, so you can chain several steps to make a longer shot. Check continuity between steps, since motion and detail can drift.
Does it need an input image?
No. Text alone is enough to start. Add an image or references only when you want to fix the look of the opening scene. Leave them out for open-ended ideas.
How much does Grok Imagine Video cost?
On TokenLab, Grok Imagine Video costs How pricing works . The pricing table above shows the full breakdown. Rates depend on the billing unit, specification, and usage. Compare matching conditions in the model's detailed pricing; a single rate does not determine the total cost.
Which endpoint should Grok Imagine Video use?
Use https://api.tokenlab.sh/v1/videos/generations for Grok Imagine Video. The request example below shows the matching code shape.
Which operations does Grok Imagine Video support?
Grok Imagine Video supports Image to video, Reference to video, Text to video, Video extension, Video to video. Select an operation above to see its endpoint and request example.
Compare Grok Imagine Video
Guides that use Grok Imagine Video
Sources
Reviewed Oct 2, 2026