Wan 3.0 AI Video Generator
Create up to 30-second AI videos on WanOmni from text, images, and multimodal references, with 1080p output and audio generated with the video.
Turn an idea, image, product reference, video clip, audio track, or document into a video from one creative workflow.
Create with Wan 3.0 AI Generator
Turn text, images, and multimodal references into videos up to 30 seconds long, with 1080p output and audio generated in the same workflow.
Create a synchronized audio track with the video.
Ready to Create
Fill in the details and generate your first Wan 3.0 video. Finished output will appear here.
See What Wan 3.0 Can Create
Explore eleven Wan 3.0 generations across different subjects, compositions, and motion styles.
What’s New in Wan 3.0?
Wan 3.0 gives you more room to build a complete scene instead of working only with short isolated clips. Combine longer generation, high-resolution output, audio, and multiple reference types inside the same workflow.
| Video duration | 2–30 seconds |
|---|---|
| Resolution | 480p / 720p / 1080p |
| Audio | Generated with the video |
| Text / Image / Reference to Video | Supported |
| Reference images | Up to 10 |
| Reference videos / audio | Up to 5 each |
| Document reference | Supported |
| Aspect ratios | Adaptive, 16:9, 4:3, 1:1, 3:4, 9:16 |
| Output format | MP4 · 30fps |
Up to 30 Seconds in One Generation
Create scenes from 2 to 30 seconds in one generation. Longer shots give camera moves, actions, dialogue, and narrative beats more space to develop without building the result from several short clips.
1080p Video with Audio
Generate at 480p, 720p, or 1080p. Audio can be created alongside the video so dialogue, ambience, and on-screen action can be planned as part of the same generation.
Use More Than One Kind of Reference
Combine up to 10 images, 5 videos, and 5 audio clips in supported reference workflows. Use them to guide characters, products, locations, visual direction, or audio.
Start from a Document
Use a supported document as source material for a video—useful for briefs, reports, presentations, and other structured content.
What Is Wan 3.0?
Wan 3.0 is an AI video generation model in Alibaba's Wan model family. It can create video from text, still images, and multimodal reference material, with output of up to 30 seconds and resolutions up to 1080p.
It can work with image, video, and audio references, while supported workflows can use a document as source material. For creators, marketers, agencies, and brands, one model can support social videos, product ads, narrative scenes, demonstrations, and reference-driven content.
Not sure which release fits your project? Compare the available Wan AI models before choosing a workflow.
Choose the Right Wan 3.0 Workflow
You do not need to start every video the same way. Pick the workflow based on the material you already have.
Text to Video
I only have an idea.
Describe the subject, action, setting, camera, sound, and visual direction. Wan 3.0 uses the prompt as the starting point for the scene.
Try Text to Video →Image to Video
I already have an image.
Upload a still image and describe how you want it to move. Use this workflow for product shots, characters, artwork, campaign visuals, or existing creative assets.
Try Image to Video →Reference to Video
I need the same character, product, or style.
Add image, video, or audio references to give the generation more context about the subject and creative direction.
Try Reference to Video →Document to Video
I have a brief or document.
Use structured source material as creative context for an explainer, product video, visual summary, or campaign concept.
Explore Document to Video →Wan 3.0 vs Wan 2.7: Compare the Same Prompt
These two outputs use the same prompt, making it easier to compare motion, consistency, prompt adherence, and audio side by side.
Wan 2.7 Output
Same promptWan 3.0 Output
Same promptFor previous-generation context, explore the Wan 2.7 AI Video Generator or read the detailed Wan 2.7 review.
How to Generate a Video with Wan 3.0
Creating a video takes three basic steps.
Add Your Prompt or References
Start with text, upload an image, or add supported reference media. Be specific about the subject, action, setting, camera, and audio you want.
Choose Your Video Settings
Select the duration, resolution, and aspect ratio that fit your destination. A vertical social post and a widescreen product film may need different settings.
Generate, Review, and Refine
Generate your first version. Review motion, framing, identity, text, and audio, then adjust the prompt or references when needed.
Planning multiple generations? Compare plans and video credits first.
Wan 3.0 Prompt Examples
A useful prompt tells the model what is in the shot, what happens, how the camera moves, what the scene sounds like, and what the final visual direction should feel like.
Product Ad
“A close-up product film of a matte black wireless speaker on a stone table. Slow camera push-in, warm side lighting, subtle reflections, quiet room tone and soft tactile sounds. Clean premium commercial style.”
Cinematic Scene
“A woman walks through a rain-soaked downtown street at night while taxis pass behind her. The camera tracks backward at walking speed. Neon reflections move across the pavement. Natural city ambience and distant traffic.”
UGC Ad
“Vertical creator-style video in a bright apartment. A creator picks up the product, explains two features directly to camera, then shows a close-up demonstration. Natural speech, handheld movement, soft daylight.”
Fashion
“Full-body fashion editorial in a minimal concrete studio. The model walks toward camera as fabric moves naturally. Slow lateral camera move, soft directional light, restrained color palette, subtle studio ambience.”
Anime Scene
“Animated city rooftop at sunset. Two characters talk beside a railing while wind moves their hair and clothing. Slow orbiting camera, warm sky, distant city ambience, expressive but controlled character motion.”
Character Dialogue
“Medium two-shot inside a quiet coffee shop. Two characters have a short natural conversation. Keep their appearance and clothing consistent while alternating subtle gestures and eye contact. Low café ambience underneath the dialogue.”
Social Reel
“Vertical 9:16 travel reel moving from a wide coastal view to a close-up of a traveler turning toward the ocean. Smooth forward camera movement, natural wind, waves, bright morning light, clean social-video pacing.”
Product Demo
“Demonstrate a compact desk lamp being unfolded, switched on, dimmed, and repositioned on a workspace. Keep the product shape and finish consistent. Clear hand interaction, neutral background, practical lighting and room tone.”
If you are still building prompt-writing fundamentals, the Wan 2.7 prompt guide explains reusable techniques for camera movement, subject consistency, lighting, and audio direction.
Built for Real Wan 3.0 Video Workflows
Content Creators
Create material for TikTok, Instagram Reels, YouTube Shorts, creator ads, personal videos, music visuals, fashion content, and short narrative scenes. Start from a prompt or animate an existing image.
Marketing & Ad Teams
Use Wan 3.0 for campaign concepts, product reveals, social ads, creative testing, brand films, pitch visuals, and previsualization. Start from existing campaign materials through multimodal references.
E-commerce & Brands
Turn product imagery and references into product demos, launch videos, Shopify creative, paid-social concepts, lifestyle shots, product reveals, and short promotional content.
Enterprise & Industry
Use documents, product interfaces, charts, and structured references for software demos, explainers, internal concepts, presentations, tourism content, design visualization, and planning.
To see how video generation fits alongside the platform's other creative tools, visit the WanOmni AI creation platform.
Capabilities and Limits
Wan 3.0 gives you more ways to guide a generation, but AI video still needs human review before publication.
Supported Capabilities
- Up to 30-second generation
- 480p, 720p, and 1080p output
- Text-to-video
- Image-to-video
- Multimodal reference workflows
- Video and audio references
- Document reference workflows
- Audio generated with video
Review Before Publishing
- • Small on-screen text
- • Dense software interfaces
- • Complex hand and object interaction
- • Long continuous camera moves
- • Multi-character scenes
- • Reference identity across longer shots
- • Dialogue and audio timing
Review every result for visual accuracy, audio timing, brand requirements, and usage rights before publishing it.
Before using generated content commercially, review the WanOmni Terms of Service and confirm that your selected pricing plan matches your production needs.
[ Pricing ]
Simple, Transparent Pricing
No subscriptions, no hidden fees — pay only for what you generate. Start Start free, and your credits never expire.
Starter
- 100 credits included
- Text-to-video generation
- Image-to-video conversion
- Up to 1080P resolution
- Commercial usage rights
- No watermarks
- Standard processing
Pro
- 330 credits included
- Save 5% per credit
- Multi-shot storytelling
- Native audio sync
- Reference-driven consistency
- Commercial usage rights
- Priority processing
Scale
- 600 credits included
- Save 8% per credit
- Multi-shot storytelling
- Native audio sync
- Reference-driven consistency
- Commercial usage rights
- Faster processing
Enterprise
- 1250 credits included
- Best value ($0.079/credit)
- Multi-shot storytelling
- Native audio sync
- Reference-driven consistency
- Commercial usage rights
- Highest priority processing
Want FAQs, credit breakdowns, and checkout details in one place? See our full Pricing page.
Wan 3.0 FAQ
Wan 3.0 is an AI video generation model in Alibaba's Wan model family. It can generate video from text, images, and multimodal references, with durations up to 30 seconds and resolutions up to 1080p.
Yes. Wan 3.0 is available to use on WanOmni for supported text-to-video, image-to-video, and reference-to-video workflows.
Wan 3.0 supports video durations from 2 to 30 seconds in one generation. Duration can also be left for the model to determine in supported workflows.
Yes. Supported output resolutions include 480p, 720p, and 1080p.
Yes. Wan 3.0 can generate audio alongside the video so sound and visuals are produced within the same generation workflow.
Yes. Image-to-video is a supported workflow. Upload a starting image and describe the motion, camera behavior, scene, and audio you want.
Supported reference workflows can use up to 10 images, 5 video clips, and 5 audio tracks in one request, subject to the provider's input limits.
Yes. Wan 3.0 can use a supported document as source material for a generation.
Wan 3.0 brings text-to-video, first-frame and first-and-last-frame image-to-video, and multimodal reference generation into one model. It supports 2–30 second output, 480p to 1080p resolution, adaptive aspect ratios, native audio, and document references.
Create Your First Wan 3.0 Video
Start with a prompt, an image, or your own reference material. Choose your output settings, generate a first version, then refine it until the shot fits your project.