NezhaGateNezhaGate
Video

MiniMax H3

⧉
minimax-h3

MiniMax H3 (Hailuo 03) video model: 1080P HD output, native audio always generated, 5-15 seconds in one-second steps, six aspect ratios. Text-to-video without a reference image, image-to-video with one (used as the first frame). Asynchronous: submit, then poll the job id for a re-hosted mp4 link (kept for 60 days). Billed per second; failed jobs are refunded in full.

Text-to-videoImage-to-video1080PNative audio

Sample clips

All made with NezhaGate. Click a clip to hear it.
Beneath the Walls of Troy HD original
Myth Awakens HD original
Grandma's Birthday HD original

Live Test · Playground

Try out MiniMax H3 right here (available after login).

Input

Video
0 / 2000 bytes
With multiple references, address them in the prompt as @Image1 / @Image2, e.g. "the character from @Image1 stands in the scene from @Image2". Without tokens the model decides on its own.
Rendering usually takes 5-20 min; when the backend is busy it can exceed 60 min (we wait up to 90). Poll the job id rather than resubmitting - each resubmission is billed separately. Billed per second, fully refunded on failure. Prompt limit 2000 bytes (about 666 Chinese characters).

Output

video
Results will appear here after running.
🕑 Results are kept for 60 days and then deleted automatically, so download anything you want to keep.

About MiniMax H3

MiniMax H3 (Hailuo 03) is the highest-resolution video model here: 1080P output with native audio always on. Duration is any integer 5-15 seconds across six aspect ratios; omit the reference image for text-to-video, include one (first frame by default) for image-to-video. Async job: POST /v1/videos/generations returns HTTP 202 and a job id immediately; poll GET /v1/videos/jobs/{id} until status=succeeded and read data[0].url. That link is an mp4 re-hosted by us — it is kept for 60 days (far longer than the upstream CDN link), so download it or copy it to your own storage before then.

Use cases

High-definition output

1920x1088, with the best per-frame detail and lighting of any video model on the station.

Clips with sound, no audio pass

Native audio is always generated with no extra parameters, removing a whole production step.

Full camera choreography

It handled a tracking shot into a whip-tilt into a low-angle finish, so it suits work with a real shot list.

Image-to-video

With the reference as the first frame, subject and composition stay stable — good for turning concept art into motion.

How to choose

Pick it when you want resolution, native audio, or full camera choreography. For more than 15 seconds use seedance-2.5; to turn the soundtrack off, or to drive from a source video, use the Seedance line. It also renders the slowest here (about 11 minutes for a 15-second clip), so use seedance-2.0-fast when speed matters.

FAQ

Can I turn the audio off?
No. This model always generates native audio, and audio=false is explicitly rejected rather than silently ignored. For silent clips use the Seedance line with audio=false.
Why does it take so long?
1080P with native audio is simply heavy: roughly 4 minutes for 5 seconds and 11 minutes for 15. It is an async job, so you hold a job id and poll — no connection is kept open.
How does image-to-video work?
Pass an image (public URL, data: URI or base64). By default it becomes the first frame, so the clip keeps the reference subject and framing. A second image becomes the last frame; set image_role=reference to switch to subject/style reference mode instead. The first/last-frame and reference modes are mutually exclusive.
How is it billed, and am I charged on failure?
Billed per second: the hold placed at submit is computed from the duration you asked for, and the settle is the exact same number. A failed or timed-out render refunds the whole hold, so you never pay for a clip you did not receive.

Related Models

Explore other models you can integrate.

View all →