I tested Wan 2.5 — Alibaba’s Wan AI video model — after the launch wave around native audio. The promise is simple: generate video and matching sound in one pass, with smoother motion and stronger prompt following than Wan 2.2.
On paper it sits near models like Google Veo that already treat audio as part of the output. In practice, the results were uneven but interesting. Here is what held up, what broke, and how I would actually use Wan 2.5 today on Van Gogh Studio.
Quick Verdict

Wan 2.5 is a meaningful step for creators who want ambient sound, music cues, or light narration baked into the clip. Nature scenes and stylized characters looked strong. Human faces and lip-sync were less reliable — and in one stylized test, audio did not generate at all.
It is better than I expected for atmosphere. It is not yet the safest pick for polished talking-head or broadcast-style work.
| Feature | Wan 2.5 |
|---|---|
| Best at | Ambient audio + scenic / stylized motion |
| Weakest at | Precise lip-sync and consistent human realism |
| Standout upgrade vs Wan 2.2 | Native audio generation + custom audio upload options |
| Also improved | Motion fluidity, prompt awareness, visual detail |
| Best for | Experimenters, social drafts, nature / stylized concepts |
| Wait if you need | Broadcast-grade realism and perfect speech sync |
What Is New in Wan 2.5?
Wan 2.5 is part of the Wan AI family on Van Gogh Studio. The release focus is less about “another pretty clip” and more about audio-visual pairing:
- Native audio integration: Generate ambience, effects, music, or light narration with the video
- Custom audio upload path: Useful when you already have a voiceover or soundtrack to guide timing
- Stronger motion dynamics: More fluid body and camera movement than earlier Wan versions in many scenic tests
- Better prompt reading: Clearer handling of mood, setting, and character action when the brief stays focused
- Richer scene detail: More texture in environments, wardrobe, and lighting when the prompt is specific
That combination matters because silent AI clips often need a second tool just to feel finished. Wan 2.5 tries to collapse that gap.
How I Tested Wan 2.5
I ran four prompts across realistic and stylized scenes and scored each on:
- Sound precision and scene fit
- Visual authenticity and motion fluidity
- Facial expression / movement accuracy where relevant
1. Hiking Scene With Friends — Smooth and Natural
Prompt: Two young men and one young woman hike up a scenic mountain trail, laughing as they chat casually. A gentle breeze rustles the leaves, sunlight filters through the trees, and each carries a backpack.
Result: Forest ambience, breeze, and laughter lined up with the picture. Motion stayed smooth with no obvious frame glitches.
Score: 8/10 — Strong, usable result for casual outdoor content.
2. Woman at the Subway Station — Good Audio, Softer Motion
Prompt: A young Asian woman stands on subway station stairs, smiling warmly with a smartphone in hand. Daylight filters down, soft shadows falling across her urban streetwear look.
Result: Subway background sound sold the location. Her expression and body motion felt a bit stiff — believable, but not especially lively.
Score: 8/10 — Solid sound bed; motion still has headroom.
3. Sly Fox in a Suit — Strong Look, Missing Audio
Prompt: A distinguished fox in a sharp suit carries a stack of papers, approaching the camera with confident steps and a sly smile.
Result: The character design looked stylish and expressive. This run produced no audio, which is a reliability gap you should plan for.
Score: N/A for audio — Visual concept worked; sound generation failed on this pass.
4. Journalist Live on the Street — Clear Speech, Weak Sync
Prompt: A short-haired journalist reports live on a busy street, speaking over the sound of traffic and chatter.
Result: Speech content was understandable and the street bed felt right, but lip movement did not fully match the audio.
Score: 5/10 — Usable draft, not production-ready talking-head output.
What I Liked
- Audio that matches place and mood in scenic tests — less silent B-roll, more finished-feeling clips
- Smooth outdoor motion when the prompt stays concrete about people, weather, and camera distance
- Stylized character appeal for concept art, ads, and playful social formats
- A clear upgrade path from Wan 2.2 if audio-in-one-pass is your main missing piece
What Still Needs Work
- Lip-sync reliability on spoken scenes
- Occasional silent outputs even when the prompt implies sound
- Human micro-expressions that can look flat or slightly unnatural
- Consistency across retries — same brief can land very differently on audio
Compared with more consistent closed models in the Veo class, Wan 2.5 still feels earlier in the maturity curve. The highs are real; the variance is also real.
How to Try Wan 2.5 on Van Gogh Studio

You do not need a separate Alibaba account stack to experiment. On Van Gogh Studio:
- Open AI Video Studio or image to video
- Select Wan 2.5 as the model
- Write a focused prompt: subject, action, camera, lighting, and the sound you expect
- Generate, then iterate — simplify the brief if audio or faces drift
- Compare against Veo, Kling AI, or Wan 2.6 when the brief is commercial
If you also need talking avatars, product explainers, or short-form templates around the same idea, Van Gogh Studio keeps those workflows nearby instead of forcing a tool hop.
![]()
Useful companion tools while you evaluate Wan 2.5:
- Text to video for script-first drafts
- Image to video when you already have a still that must stay on-brand
- AI Avatar / UGC-style flows when speech clarity matters more than scenic ambience
Final Verdict
Wan 2.5 is a promising update — especially if ambient audio and stylized motion are the job. It is not yet a safe default for precise human lip-sync or high-stakes realism.
Try it if you make nature clips, stylized concepts, or social drafts and want sound without a second app.
Wait or compare if you need broadcast-clean talking heads today — and keep Wan 2.6 on your shortlist for longer multi-shot storytelling.
Start here: try Wan 2.5 on Van Gogh Studio.
FAQs
What is Wan 2.5?
Wan 2.5 is an Alibaba Wan AI video model focused on text-to-video and image-to-video with native audio generation and stronger motion/detail than Wan 2.2.
Does Wan 2.5 always generate audio?
No. In my tests most scenic prompts produced fitting sound, but at least one stylized run returned silent video. Treat audio as likely, not guaranteed.
Is Wan 2.5 better than Veo?
Not overall in my experience. Veo-class models still felt more consistent. Wan 2.5 can win on specific atmospheric shots and open experimentation, especially when you already work inside Van Gogh Studio’s multi-model setup.
Can I try Wan 2.5 for free?
Yes. Van Gogh Studio offers free credits so you can run Wan 2.5 before committing to a paid plan.
Should I use Wan 2.5 or Wan 2.6?
Use Wan 2.5 when native audio in shorter clips is the priority. Look at Wan 2.6 when multi-shot storytelling and longer narrative structure matter more.
You might also like
Wan 2.6 Review
Hands-on notes on Wan 2.6 multi-shot storytelling and longer narrative clips.
Veo 3.1 Review
Compare Wan-style audio/video generation against Google’s Veo family.
Hailuo AI Review
Another strong option if you are benchmarking Chinese video models side by side.
Runway Open-Source Alternatives
A wider shortlist when you want more control than a single closed suite.
Create AI videos free
Try Van Gogh Studio free — text-to-video, image-to-video, and 300+ models in one place. Free credits on signup, no credit card required.
Try Free Video Generator


