MiniMax H3
⧉MiniMax H3 (Hailuo 03) video model: 1080P HD output, native audio always generated, 5-15 seconds in one-second steps, six aspect ratios. Text-to-video without a reference image, image-to-video with one (used as the first frame). Asynchronous: submit, then poll the job id for a re-hosted mp4 link (kept for 60 days). Billed per second; failed jobs are refunded in full.
Sample clips
All made with NezhaGate. Click a clip to hear it.Live Test · Playground
Try out MiniMax H3 right here (available after login).
Input
VideoOutput
videoAbout MiniMax H3
MiniMax H3 (Hailuo 03) is the highest-resolution video model here: 1080P output with native audio always on. Duration is any integer 5-15 seconds across six aspect ratios; omit the reference image for text-to-video, include one (first frame by default) for image-to-video. Async job: POST /v1/videos/generations returns HTTP 202 and a job id immediately; poll GET /v1/videos/jobs/{id} until status=succeeded and read data[0].url. That link is an mp4 re-hosted by us — it is kept for 60 days (far longer than the upstream CDN link), so download it or copy it to your own storage before then.
Use cases
1920x1088, with the best per-frame detail and lighting of any video model on the station.
Native audio is always generated with no extra parameters, removing a whole production step.
It handled a tracking shot into a whip-tilt into a low-angle finish, so it suits work with a real shot list.
With the reference as the first frame, subject and composition stay stable — good for turning concept art into motion.
How to choose
Pick it when you want resolution, native audio, or full camera choreography. For more than 15 seconds use seedance-2.5; to turn the soundtrack off, or to drive from a source video, use the Seedance line. It also renders the slowest here (about 11 minutes for a 15-second clip), so use seedance-2.0-fast when speed matters.
FAQ
Can I turn the audio off?
Why does it take so long?
How does image-to-video work?
How is it billed, and am I charged on failure?
Related Models
Explore other models you can integrate.
Seedance 2.0
Text / image to video · 5-15 seconds
Wan 3.0
One endpoint, five modes · up to 30 seconds
Seedance 2.5
Text / image to video · up to 30 seconds
Seedance 2.0 Fast
Low latency · lowest cost
Wan 3.0 Prime
Same capabilities · several times faster
Veo 3.1 Not yet open
Text / image to video · multiple tiers · multiple resolutions