Grok Imagine Video 1.5
⧉xAI Grok Imagine Video 1.5: 4-15 seconds in one-second steps, 480P / 720P, five aspect ratios. Text-to-video without a reference image, image-to-video with reference images (up to 7; with 2 or more the maximum is 10 seconds). Asynchronous, billed per clip -- 1 second costs the same as 15 -- and failed jobs are refunded in full. 480P $0.30 / 720P $0.40; the price only changes with resolution.
Live Test · Playground
Try out Grok Imagine Video 1.5 right here (available after login)。
Input
VideoOutput
videoAbout Grok Imagine Video 1.5
Grok Imagine Video 1.5 is xAI's video model and the cheapest video model on NezhaGate. What sets it apart from everything else here is the billing shape: it is priced PER CLIP, so a 1-second and a 15-second render cost exactly the same; only the resolution changes the price (480P $0.30, 720P $0.40). Duration is any integer 1-15 seconds, resolution 480P / 720P, aspect ratio 16:9 / 9:16 / 1:1 / 4:3 / 3:4; omit the reference image for text-to-video, include up to 7 for image-to-video; every clip ships with a native audio track. Async job: POST /v1/videos/generations returns HTTP 202 and a job id immediately; poll GET /v1/videos/jobs/{id} until status=succeeded and read data[0].url. That link is an mp4 re-hosted by us -- it will not expire the way the upstream CDN link does.
Use cases
The lowest fixed cost per clip on the platform, so you can run dozens of prompts and keep the best one.
Per-clip pricing means the 15-second option is the best value; there is no reason to cut a clip short to save money.
9:16 straight out of the model for short-video platforms; 4:3 and 3:4 suit mixed image-and-text layouts.
Up to 7 reference images, useful for putting one character or product into different scenes.
How to choose
Pick it when budget comes first or when you are batching ideas -- it starts at $0.30, the cheapest video here, and per-clip pricing means 15 seconds costs no more than one. For a controllable soundtrack, reference audio or reference video use the Seedance family; for more than 15 seconds use seedance-2.5; for 1080P use minimax-h3.
FAQ
Why per clip instead of per second?
Why only 10 seconds with multiple images?
Does it produce audio?
How is it billed, and am I charged on failure?
Related Models
Explore other models you can integrate.
GPT-5.6 Sol
Frontier flagship for agentic coding
GPT-5.6 Terra
Balanced everyday agentic coding model
GPT-5.6 Luna
Fast, economical agentic coding model
GPT-5.5
Flagship chat and reasoning model
GPT-6 Astra
Next-generation OpenAI flagship - 1.05M context
GPT Image 2
High-quality text-to-image / image-to-image model
GPT Image 2.5 Flare
Next-gen image model - clean and smooth
GPT Image 2.5 Sunburst
Next-gen image model - richer texture
Nano Banana 2
High-quality text-to-image / image-to-image model
Nano Banana Pro
Flagship text-to-image / image-to-image model
Claude Sonnet 4.6
Balanced, efficient chat and coding model
Claude Opus 5
Next-generation flagship from Anthropic
Claude Fable 5
Anthropic Fable series · narrative and long-form writing
Gemini 3.1 Pro
Next-generation flagship reasoning model from Google
Gemini 3.8 Flash
Latest Flash · adaptive thinking
Gemini 3.7 Flash
Previous Flash · adaptive thinking
Gemini 3.6 Flash
Next-generation fast thinking model · four thinking budgets
Gemini 3.6 Flash High
Deep thinking tier
Gemini 3.6 Flash Low
Fastest, lowest-cost tier
Gemini 3.6 Flash Tiered
Auto-tiered thinking
Gemini 3 Flash
Fast, low-latency, cost-efficient chat model
Gemini 3.5 Flash
Next-generation fast thinking model
Gemini 2.5 Flash
Fast, low-latency, cost-efficient chat
Veo 3.1
Text / image to video · multiple tiers · multiple resolutions
Seedance 2.5
Text / image to video · up to 30 seconds
Seedance 2.0
Text / image to video · 5-15 seconds
Seedance 2.0 Fast
Low latency · lowest cost
Seedance 2.0 Mini
Entry tier · lowest cost
Wan 3.0
One endpoint, five modes · up to 30 seconds
Wan 3.0 Prime
Same capabilities · several times faster
MiniMax H3
1080P · native audio