Wan 3.0
⧉Wan 3.0 (Tongyi Wanxiang) is the all-round video model from Alibaba. One endpoint, the inputs select the mode: no assets for text-to-video, one image as the first frame, two images as first and last frames, several images as multi-image reference, and a reference video for video rewriting. Duration 2-30 seconds in one-second steps, three resolutions 480P / 720P / 1080P, five aspect ratios, up to 10 reference images, 5 reference videos and 5 reference audio clips; output includes audio. Asynchronous: submit, then poll the job id for a stable re-hosted mp4 link. Billed per second; failed jobs are refunded in full.
Live Test · Playground
Try out Wan 3.0 right here (available after login)。
Input
VideoOutput
videoAbout Wan 3.0
Wan 3.0 (Tongyi Wanxiang) is Alibaba’s all-in-one video model and the most versatile video line here: one endpoint covers text-to-video, first-frame, first+last frame, multi-image reference and video editing, chosen by parameters rather than by switching models. Any integer 2-30 seconds, 480P / 720P / 1080P, aspect ratios 16:9 / 9:16 / 1:1 / 4:3 / 3:4, up to 10 reference images, 5 reference videos and 5 reference audio tracks, and every clip ships with an audio track. Async: POST /v1/videos/generations returns HTTP 202 and a job id immediately; poll GET /v1/videos/jobs/{id} until status=succeeded and read data[0].url — an mp4 re-hosted by us, so it will not expire in 24 hours the way the upstream link does.
Use cases
Text-to-video, first frame, first+last frame, multi-image reference and video editing, all chosen by parameters.
Up to 30 seconds for single-shot narrative pieces.
Ten reference images plus reference video and audio — feed character, scene and rhythm together.
Hand it a reference video and one sentence to restyle the footage or change the camera move.
How to choose
Pick it for clips beyond 15 seconds, for feeding many images plus video and audio at once, or for native 1080P; take Prime when you are in a hurry. For a few seconds with no multi-asset reference, the Seedance line is fine too.
FAQ
How do I switch between the five modes?
Do multiple reference images bleed into each other?
How is it billed, and do references cost extra?
What is the difference between Standard and Prime?
How long does a render take?
Related Models
Explore other models you can integrate.
GPT-5.6 Sol
Frontier flagship for agentic coding
GPT-5.6 Terra
Balanced everyday agentic coding model
GPT-5.6 Luna
Fast, economical agentic coding model
GPT-5.5
Flagship chat and reasoning model
GPT-6 Astra
Next-generation OpenAI flagship - 1.05M context
GPT Image 2
High-quality text-to-image / image-to-image model
GPT Image 2.5 Flare
Next-gen image model - clean and smooth
GPT Image 2.5 Sunburst
Next-gen image model - richer texture
Nano Banana 2
High-quality text-to-image / image-to-image model
Nano Banana Pro
Flagship text-to-image / image-to-image model
Claude Sonnet 4.6
Balanced, efficient chat and coding model
Claude Opus 5
Next-generation flagship from Anthropic
Claude Fable 5
Anthropic Fable series · narrative and long-form writing
Gemini 3.1 Pro
Next-generation flagship reasoning model from Google
Gemini 3.8 Flash
Latest Flash · adaptive thinking
Gemini 3.7 Flash
Previous Flash · adaptive thinking
Gemini 3.6 Flash
Next-generation fast thinking model · four thinking budgets
Gemini 3.6 Flash High
Deep thinking tier
Gemini 3.6 Flash Low
Fastest, lowest-cost tier
Gemini 3.6 Flash Tiered
Auto-tiered thinking
Gemini 3 Flash
Fast, low-latency, cost-efficient chat model
Gemini 3.5 Flash
Next-generation fast thinking model
Gemini 2.5 Flash
Fast, low-latency, cost-efficient chat
Veo 3.1
Text / image to video · multiple tiers · multiple resolutions
Seedance 2.5
Text / image to video · up to 30 seconds
Seedance 2.0
Text / image to video · 5-15 seconds
Seedance 2.0 Fast
Low latency · lowest cost
Seedance 2.0 Mini
Entry tier · lowest cost
Wan 3.0 Prime
Same capabilities · several times faster
MiniMax H3
1080P · native audio
Grok Imagine Video 1.5
Billed per clip · up to 15 seconds