NezhaGate
Video

Wan 3.0

wan3.0-video
Model Type StandardPrime

Wan 3.0 (Tongyi Wanxiang) is the all-round video model from Alibaba. One endpoint, the inputs select the mode: no assets for text-to-video, one image as the first frame, two images as first and last frames, several images as multi-image reference, and a reference video for video rewriting. Duration 2-30 seconds in one-second steps, three resolutions 480P / 720P / 1080P, five aspect ratios, up to 10 reference images, 5 reference videos and 5 reference audio clips; output includes audio. Asynchronous: submit, then poll the job id for a stable re-hosted mp4 link. Billed per second; failed jobs are refunded in full.

Text-to-videoImage-to-video1080PUp to 30 s

Live Test · Playground

Try out Wan 3.0 right here (available after login)。

Input

Video
0 / 5000 bytes
With multiple references, address them in the prompt as @Image1 / @Image2, e.g. "the character from @Image1 stands in the scene from @Image2". Without tokens the model decides on its own.
Rendering usually takes 5-20 min; when the backend is busy it can exceed 60 min (we wait up to 90). Poll the job id rather than resubmitting - each resubmission is billed separately. Billed per second, fully refunded on failure. Prompt limit 5000 bytes (about 1666 Chinese characters).

Output

video
Results will appear here after running.

About Wan 3.0

Wan 3.0 (Tongyi Wanxiang) is Alibaba’s all-in-one video model and the most versatile video line here: one endpoint covers text-to-video, first-frame, first+last frame, multi-image reference and video editing, chosen by parameters rather than by switching models. Any integer 2-30 seconds, 480P / 720P / 1080P, aspect ratios 16:9 / 9:16 / 1:1 / 4:3 / 3:4, up to 10 reference images, 5 reference videos and 5 reference audio tracks, and every clip ships with an audio track. Async: POST /v1/videos/generations returns HTTP 202 and a job id immediately; poll GET /v1/videos/jobs/{id} until status=succeeded and read data[0].url — an mp4 re-hosted by us, so it will not expire in 24 hours the way the upstream link does.

Use cases

Five modes, one endpoint

Text-to-video, first frame, first+last frame, multi-image reference and video editing, all chosen by parameters.

Long takes

Up to 30 seconds for single-shot narrative pieces.

Multi-asset composition

Ten reference images plus reference video and audio — feed character, scene and rhythm together.

Video editing

Hand it a reference video and one sentence to restyle the footage or change the camera move.

How to choose

Pick it for clips beyond 15 seconds, for feeding many images plus video and audio at once, or for native 1080P; take Prime when you are in a hurry. For a few seconds with no multi-asset reference, the Seedance line is fine too.

FAQ

How do I switch between the five modes?
With parameters only: send nothing and it is text-to-video; one image defaults to the first frame; two images with image_role=first_frame gives first+last frame; more than two default to multi-image reference; adding video_references makes it a video edit. Combining images and a video needs image_role=reference.
Do multiple reference images bleed into each other?
Yes. With four references the model treated every person visible in ANY of them as a candidate subject — the clip started on the intended character and later cut to a second person who only appeared in a scene reference. Keep scene references free of other people, or name the subject with @Image1 in the prompt.
How is it billed, and do references cost extra?
Per second: cost = (output duration + the duration of each reference video) x the rate for that resolution. Reference images and audio are free. The hold at submit uses that number and the settle uses the same one, so they are always equal; a failed or timed-out render refunds it in full. Because reference video seconds are billed, their duration must be stated explicitly.
What is the difference between Standard and Prime?
Speed and price only — the capabilities are identical, line for line. The same 2-second clip took about 2m45s on Standard and about 37s on Prime, and Prime costs roughly a third more per second. If you are not in a hurry, use Standard.
How long does a render take?
A 2-second clip is about 2-3 minutes on Standard; a 5-second clip with four reference images measured about 11 minutes. Prime is markedly faster. Poll the job id rather than waiting synchronously, every 10-15 seconds.

Related Models

Explore other models you can integrate.

View all →
GPT-5.6 Sol Chat

GPT-5.6 Sol

gpt-5.6-sol

Frontier flagship for agentic coding

$2.00/1M in · $12.00/1M out ↓75% View →
GPT-5.6 Terra Chat

GPT-5.6 Terra

gpt-5.6-terra

Balanced everyday agentic coding model

$1.20/1M in · $7.00/1M out ↓77% View →
GPT-5.6 Luna Chat

GPT-5.6 Luna

gpt-5.6-luna

Fast, economical agentic coding model

$0.80/1M in · $4.80/1M out ↓68% View →
GPT-5.5 Chat

GPT-5.5

gpt-5.5

Flagship chat and reasoning model

$0.70/1M in · $4.20/1M out ↓86% View →
GPT-6 Astra Chat

GPT-6 Astra

gpt-6-astra

Next-generation OpenAI flagship - 1.05M context

$2.80/1M in · $14.00/1M out ↓72% View →
GPT Image 2 Image

GPT Image 2

gpt-image-2

High-quality text-to-image / image-to-image model

$0.015/img and up View →
GPT Image 2.5 Flare Image

GPT Image 2.5 Flare

gpt-image-2.5-flare

Next-gen image model - clean and smooth

$0.015/img and up View →
GPT Image 2.5 Sunburst Image

GPT Image 2.5 Sunburst

gpt-image-2.5-sunburst

Next-gen image model - richer texture

$0.015/img and up View →
Nano Banana 2 Image

Nano Banana 2

nano-banana-2

High-quality text-to-image / image-to-image model

$0.025/img and up View →
Nano Banana Pro Image

Nano Banana Pro

nano-banana-pro

Flagship text-to-image / image-to-image model

$0.040/img and up View →
Claude Sonnet 4.6 Chat

Claude Sonnet 4.6

claude-sonnet-4-6

Balanced, efficient chat and coding model

$1.50/1M in · $7.50/1M out ↓50% View →
Claude Opus 5 Chat

Claude Opus 5

claude-opus-5

Next-generation flagship from Anthropic

$4.00/1M in · $20.00/1M out View →
Claude Fable 5 Chat

Claude Fable 5

claude-fable-5

Anthropic Fable series · narrative and long-form writing

$8.00/1M in · $40.00/1M out View →
Gemini 3.1 Pro Chat

Gemini 3.1 Pro

gemini-3.1-pro

Next-generation flagship reasoning model from Google

$0.50/1M in · $3.00/1M out ↓75% View →
Gemini 3.8 Flash Chat

Gemini 3.8 Flash

gemini-3.8-flash

Latest Flash · adaptive thinking

$0.60/1M in · $3.60/1M out ↓7% View →
Gemini 3.7 Flash Chat

Gemini 3.7 Flash

gemini-3.7-flash

Previous Flash · adaptive thinking

$0.60/1M in · $3.60/1M out ↓7% View →
Gemini 3.6 Flash Chat

Gemini 3.6 Flash

gemini-3.6-flash

Next-generation fast thinking model · four thinking budgets

$0.60/1M in · $3.60/1M out View →
Gemini 3.6 Flash High Chat

Gemini 3.6 Flash High

gemini-3.6-flash-high

Deep thinking tier

$0.60/1M in · $3.60/1M out View →
Gemini 3.6 Flash Low Chat

Gemini 3.6 Flash Low

gemini-3.6-flash-low

Fastest, lowest-cost tier

$0.60/1M in · $3.60/1M out View →
Gemini 3.6 Flash Tiered Chat

Gemini 3.6 Flash Tiered

gemini-3.6-flash-tiered

Auto-tiered thinking

$0.60/1M in · $3.60/1M out View →
Gemini 3 Flash Chat

Gemini 3 Flash

gemini-3-flash-preview

Fast, low-latency, cost-efficient chat model

$0.30/1M in · $1.20/1M out ↓57% View →
Gemini 3.5 Flash Chat

Gemini 3.5 Flash

gemini-3.5-flash

Next-generation fast thinking model

$0.45/1M in · $2.70/1M out ↓70% View →
Gemini 2.5 Flash Chat

Gemini 2.5 Flash

gemini-2.5-flash

Fast, low-latency, cost-efficient chat

$0.30/1M in · $1.20/1M out ↓46% View →
Veo 3.1 Video

Veo 3.1

veo-3.1

Text / image to video · multiple tiers · multiple resolutions

$0.075/clip+ View →
Seedance 2.5 Video

Seedance 2.5

seedance-2.5

Text / image to video · up to 30 seconds

$0.158/s View →
Seedance 2.0 Video

Seedance 2.0

seedance-2.0

Text / image to video · 5-15 seconds

$0.150/s View →
Seedance 2.0 Fast Video

Seedance 2.0 Fast

seedance-2.0-fast

Low latency · lowest cost

$0.100/s View →
Seedance 2.0 Mini Video

Seedance 2.0 Mini

seedance-2.0-mini

Entry tier · lowest cost

$0.066/s View →
Wan 3.0 Prime Video

Wan 3.0 Prime

wan3.0-video-prime

Same capabilities · several times faster

$0.160/s View →
MiniMax H3 Video

MiniMax H3

minimax-h3

1080P · native audio

$0.036/s View →
Grok Imagine Video 1.5 Video

Grok Imagine Video 1.5

grok-imagine-video-1.5

Billed per clip · up to 15 seconds

$0.300/clip+ View →