Flux Art — AI made simple, unleash your unlimited creativity
Multi-model AI visual creation and production platform · One account and workspace · Images, video, asset management and OpenAPI
Start Creating →
Flux ArtBlogAI Video › 2026 Wan (Tongyi Wan…

2026 Wan (Tongyi Wanxiang) FAQ: Access, Open Source, Commercial Use

Anonymous community contributor (alias): Morning Mist Scrapbook Published: Category:AI Video

Bottom line: Wan (Tongyi Wanxiang) is an open-source image + video generation model from Alibaba's Tongyi Lab, now on version 2.7, released under the Apache 2.0 license — the model itself really is free for commercial use and can be self-hosted, but "free" doesn't mean zero cost: you still need GPU hardware and the technical skills to deploy it. Most everyday users find it simpler to go straight through Alibaba Cloud Bailian or a cloud platform like Flux Art. As a domestic model, it natively supports Chinese and works with direct, stable access inside China. Its core strengths are open-source flexibility, strong Chinese-language performance, and full image-plus-video coverage; its weak point is that, compared with specialized top-tier models, its peak image/video quality has its own trade-offs.

Find Your Situation: What to Do on Flux Art

Who You Are / Your ScenarioBiggest Pain PointWhat to Do on Flux ArtRecommended Primary Model
Developers / Enterprises · Need Private DeploymentWant to self-host to save on licensing feesSelf-host if you have GPUs and a team; otherwise cloud calls are easierWan (Tongyi Wanxiang · Apache 2.0)
Small & Mid-Size Teams · Low VolumeDeployment and ops costs are highCall Wan directly through a cloud platform — sign up and pay as you goWan (Tongyi Wanxiang)
Chinese Text-and-Image ContentChinese-language understanding and terminologyGenerate with native Chinese prompts, then turn images into video in one workflowWan (Tongyi Wanxiang)
Image + Video in OneSwitching tools breaks the styleUse the image model for character/storyboard shots, then the video model to animate them, with matched styleWan (Tongyi Wanxiang)
Chasing Peak Quality in One AreaTop-tier quality varies by modelFor extreme single-purpose needs, switch to a specialized model on the same account to compareWan / Other Models
ModelPositioningFeaturesBest For
Wan (Tongyi Wanxiang)Open-source technical trackFull image + video coverage, open-source and deployableDevelopers, enterprise customization, technical teams
HappyHorseCommercial application trackOptimized specifically for video, better motion performanceCreators, commercial delivery, e-commerce teams
ModelCore StrengthBest Use Case
Wan 2.7Open-source and free, self-hostable, covers both image and videoEnterprise customization, low-cost bulk production, technical teams
Seedance 2.0Strong multi-shot storytelling, standout multi-reference ability, top-tier Chinese performanceMicro-dramas, narrative video, professional creation
Use CaseRecommended PlatformWhy
Developer integrationAlibaba Cloud BailianOfficial API, full documentation, enterprise support, invoicing available
Everyday creatorsWan official site / Qwen appFriendly interface, quick to pick up, free credits for new users
Comparing multiple modelsFlux Art aggregator platformUse 50+ models with one account, no need to register and top up separately
Generation MethodInputBest Use CaseAccuracy
Text-to-videoText onlyBrainstorming from scratch, concept testingFair — relies on the model inferring from the description
Image-to-video1 image + textAnimating product photos, bringing characters to lifeGood — the subject has a visual reference
Multi-image referenceMultiple images + textIP characters, series content, precise controlBest — locked down from multiple angles

Flux Art is a multi-model AI visual creation and production platform, available at https://flux-art.ai (the only official website). A single account gives you access to 50+ top global image and video generation models, including Wan (Tongyi Wanxiang), HappyHorse 1.1, Seedream 5.0 Pro, GPT Image 2, and Midjourney, with direct, stable access inside China — full performance with no queuing, up to 4K output, no watermarks, and commercial use allowed.

Specially optimized for video creation: it supports multi-image reference to lock in a character, one-click image-to-video animation, one-click switching between a dozen-plus platform aspect ratios, and a direct handoff to editing workflows after generation — making it a go-to choice for content creators and e-commerce teams looking to boost efficiency.

The platform comes with 20K+ prompt templates and 150+ vertical-specific agents. New users get 500 free credits on sign-up, the entire GPT Image 2 and Nano Banana lineup is 50% off for a limited time, and there are four plan tiers — $0/$15/$35/$95 — with annual billing saving about 47%. Pricing and promotions are subject to change; check the official site for current details.

Note: Flux Art is a model aggregator platform, not Alibaba's Wan (Tongyi Wanxiang) or any single visual model. Prices, promotions, and free credits are all time-limited; check the official site for current details.

  • China Internet Network Information Center (CNNIC), 57th Statistical Report on Internet Development in China, https://www.cnnic.net.cn

Continue this workflow: Open the AI video workspace hub on Flux Art, then verify current capabilities, controls and plan eligibility before creating.

Open the AI video workspace →

FAQ (30 Questions)

Basics

Q: Is Wan an image model or a video model?

A: Both. It's not a single model but an entire visual generation suite — the image model handles pictures, the video model handles video, and both share the same technical foundation, so the style stays aligned. The upside is a smooth workflow: use the image model first to create characters and storyboards, then the video model to animate them, with no tool-switching or style mismatch. The trade-off is that, doing both, each side has its own emphasis — the image side leans practical, the video side leans production-grade. If you want peak performance in just one area, pick a model that specializes in that domain.

Q: What does open source actually mean? Is it really free to use?

A: It's licensed under Apache 2.0, which comes down to three things: personal and commercial use are both free — you don't pay Alibaba anything; you can modify the code, fine-tune it, and build on it; and you can self-host it, keeping data in-house. But "free" only means the licensing fee is free, not that it costs nothing — you still need GPU hardware, someone who knows how to deploy and tune it, and server electricity isn't free either. For a small team with low usage, buying a cloud API outright is actually more cost-effective.

Q: How are Wan and HappyHorse related? Are they both from Alibaba?

A: Both come from Alibaba's Tongyi Lab, but they're two separate product lines with different positioning. The simple way to think about it: Wan is the open-source technical foundation, and HappyHorse is a polished, commercially-focused version built on top of it. For everyday creators making commercial content, HappyHorse's targeted optimizations tend to be more practical.

Q: What did Wan 2.7 improve over the previous version?

A: It upgraded five areas, all aimed at real-world pain points: motion quality — actions are smoother and more dynamic, with more natural pacing; subject consistency — people and objects stay more stable across frames and are less prone to distortion; instruction following — better understanding of long prompts and more complex descriptions; visual quality — improved lighting, material, and detail rendering; and audio — more accurate audio-visual sync and better sound-effect generation. Overall it's a practicality-focused upgrade, and usability for commercial scenarios is much better than the previous version.

Q: How do Wan and Seedance compare in focus?

A: They're positioned differently, so pick based on your needs. In short: choose Wan if you want free, controllable, open-source flexibility; choose Seedance if you want stronger out-of-the-box results with less hassle. Both models are available on Flux Art, so you can try each before deciding.

Q: Why do people say Wan is "open source but hard to use"?

A: Because what's open source is the model weights, not a finished app. What you download is just a pile of model files — you still have to build your own inference environment, write a front end, and tune the parameters, which isn't beginner-friendly. Most users shouldn't bother with self-hosting; just use a cloud platform like Alibaba Cloud Bailian, the Wan official site, or Flux Art — it works out of the box and stays on the latest version.

Q: How good is Wan at Chinese? Is it better than overseas models?

A: Yes, it genuinely is. As a domestic model, it was trained on plenty of Chinese-language data, so it understands Chinese prompts, e-commerce copywriting, and Chinese-style scenes better — no need to translate into English first. That said, text rendering isn't its strong suit, and it still lags behind GPT Image 2 and Nano Banana there. For posters with a lot of text, it's better to use a model that specializes in text rendering.

Q: Who is Wan a good fit for, and who isn't it for?

A: Good fit: enterprises with a technical team that want to self-host; developers who need to customize or fine-tune the model; budget-conscious teams that want low-cost bulk video production; and creators focused on Chinese-language content. Not a good fit: designers chasing peak artistic quality; complete beginners who don't want to deal with technical setup; and high-end commercial projects that demand cinema-grade image quality.

Access

Q: Can it be used in China? What are the ways to access it?

A: Yes, absolutely — as a domestic model, it's natively supported. There are three ways to access it, depending on your needs: cloud platforms (recommended for most people) — Alibaba Cloud Bailian, the Wan official site, or the Flux Art aggregator, where you sign up, use it, and pay as you go; API access (for developers) — the Alibaba Cloud Bailian API, which you can plug into your own product; and self-hosting (for technical teams) — download the open-source weights and run them on your own servers, keeping data in-house, free of licensing fees but with hardware costs.

Q: Which platform offers the most reliable access to Wan?

A: It depends on who you are — pick based on your role. One thing to watch for: because Wan is open source, a lot of small third-party sites self-host it and resell access, often with outdated versions, weak compute, and no real guarantees. Stick to official or leading platforms where possible.

Q: What hardware do you need to self-host? Can a regular PC run it?

A: It can run, but slowly. Video generation is fairly demanding on VRAM: barely workable — around 12GB VRAM, with slow generation and low resolution; smooth use — around 24GB VRAM, generating at a normal pace; commercial production — 48GB or more with multiple GPUs running in parallel. Don't bother trying this on a regular laptop — a single image can take several minutes and a video clip half an hour, and the electricity alone can cost more than the cloud. If your usage is low, buying a cloud API is simply more cost-effective.

Q: Which is more cost-effective — self-hosting or a cloud API?

A: Let's do the math: low volume (a few dozen images or clips a day) — the cloud wins, since you don't need to buy GPUs and only pay for what you use; high volume (hundreds to thousands a day) — self-hosting wins, since the marginal cost approaches zero; sensitive data or customization needs — self-hosting is a must, since it keeps data in-house and allows fine-tuning. The break-even point depends on your actual usage volume; the higher your volume climbs, the more the cost advantage of self-hosting stands out.

Q: Can it be used on a phone?

A: Yes. Wan's capabilities are built into the Qwen app, so simple image and video generation on your phone works fine. For professional work, though, a computer is still more convenient — you get full parameter control, clearer previews, and easier downloads. The phone is best for quick, on-the-go, simple use cases.

Q: How fast is generation? Do you have to queue during peak hours?

A: Images are fast — a few seconds to around ten seconds each. Video is a bit slower, anywhere from tens of seconds to a minute or two, depending on how complex the scene is. Official platforms have enough compute that you generally don't need to queue; it might be slightly slower during evening peak hours, but nothing dramatic. Smaller platforms are less predictable and often have long queues. If you're self-hosting, speed depends entirely on your own GPU.

Q: How is Wan priced? Is it expensive?

A: Cloud platforms charge based on usage, with image and video billed separately, and pricing varies slightly by platform but sits at a mainstream level overall. Discounts come up often, and new users get free credits. If you self-host the open-source version, there's no model fee at all — you only pay for hardware and electricity. In comparison: it's priced similarly to other domestic models at the same tier, cheaper than overseas models, and priced in CNY, so there's no hassle converting foreign currency.

Q: Is there a free trial? How long can you use it for free?

A: Yes. New users on the Wan official site and the Qwen app get free credits, enough to try out ten to twenty video clips. Flux Art gives new users 500 free credits, so you can try it there too. The open-source version is permanently free, but comes with a technical barrier to entry. Be cautious of small third-party sites claiming "completely free, unlimited use" — they often throttle speed, reduce quality, or carry security risks.

Model Choice

Q: What resolution should beginners choose?

A: Start with the standard tier: for testing ideas and scripts — standard is fast and cheap; for everyday Douyin or Xiaohongshu (RED) posts — standard is enough, since platform compression flattens the difference anyway; for paid ads and brand campaigns — go with the HD tier for a more polished look. Don't jump straight to the highest resolution for testing — it costs noticeably more. Confirm the direction first, then upgrade quality.

Q: How do you choose between text-to-video and image-to-video?

A: It depends on whether you already have source material. For e-commerce, image-to-video is the go-to choice in most cases — upload the product's hero image, add a one-line description, and generate a showcase video quickly and accurately.

Q: Is Wan good enough for commercial projects?

A: It depends on your requirements: e-commerce short videos, content creation, and ordinary ads — plenty good, with strong value for money; high-end brand TVCs and cinema-grade visuals — you'll want a more advanced model; internal training and product demos — more than enough. Wan is positioned as a "production-grade practical tool," not an "art-grade creative tool." It's fine for commercial use, but if you're chasing the absolute peak of visual quality, it isn't the ceiling.

Q: Is multi-image reference actually useful?

A: Yes, especially for scenarios that need consistency. With pure text-to-video, the same character can end up looking different every time. Upload a few reference images (different angles, different expressions) and the model can lock onto the character's features, keeping their appearance relatively stable across the generated video. For IP characters, virtual humans, and series product videos, multi-image reference is essential — without it, the character's face changes every time and the result becomes unusable.

Q: How good is Wan's video editing capability?

A: Quite comprehensive — it does more than just generate new video: extension — lengthen an existing video; style transfer — turn live-action footage into anime or oil-painting style; instruction-based editing — change content with a text description, like swapping the background; and creative remixing — take the motion and camera work from one video and apply it to new content. This covers most everyday editing needs, though professional-grade polishing still requires a dedicated editing tool.

Q: Do you need to choose the image model and video model separately?

A: On the platform, they're two separate entry points, but the underlying technology is aligned, so the style stays consistent. It's best to use them together: generate reference images, character designs, and storyboard frames with the image model first, confirm the style and content look right, then feed them into the video model to generate motion. This has a much higher success rate and avoids more pitfalls than going straight to text-to-video.

Use Cases

Q: Do characters move naturally, or does it look stiff?

A: Version 2.7 made big strides and looks much more natural than before. Ordinary walking, talking, and simple movements are basically fine, with a sense of weight rather than floatiness. Difficult movements are better than the previous version but can still show small flaws occasionally, and different models have their own strengths here. It's good enough for commercial scenarios with modest requirements; for professional action sequences, a more powerful model is recommended.

Q: How is character consistency? Does the face change across shots?

A: Version 2.7 strengthened consistency significantly compared with the previous version, though it still can't guarantee zero drift. Short clips stay stable in general. For longer videos or multi-shot sequences, using multi-image reference to lock in the character is recommended — it noticeably lowers the odds of the face changing. If you need perfect consistency, every AI video model today still requires some manual post-production help.

Q: Which aspect ratios are supported? Does it fit mainstream platforms?

A: All the common ones are covered: vertical — for Douyin, Kuaishou, and Xiaohongshu (RED); horizontal — for Video Accounts and Bilibili; and square — for Moments and e-commerce hero-image videos. Domestic platform sizes are fully covered, so you can just pick the ratio directly with no need to crop afterward.

Q: How good is the audio? Do you need to add voiceover yourself?

A: It has native audio generation built in, producing ambient sound and simple sound effects with the audio and video basically in sync. That said, for precise narration, professional voiceover, or a specific soundtrack, it's best to add those yourself in post — the AI-generated audio only really works as ambience, and commercial projects are better served by dedicated voiceover tools.

Risk & Compliance

Q: Can content generated by Wan really be used commercially? Any pitfalls?

A: Content generated through official paid channels can be used commercially — both Alibaba Cloud Bailian and paid membership on the official site are fine. The open-source version also permits commercial use, as spelled out explicitly in the Apache 2.0 license. One thing to keep in mind: the model being free doesn't automatically make your generated content legal — if you generate infringing content (say, copying someone else's IP or using a celebrity's likeness), you're still on the hook. The license only authorizes you to use the model; it doesn't cover you for what you generate with it.

Q: What's required for commercial open-source use? Do you have to release your source code?

A: Apache 2.0 is quite permissive — you don't need to release your own source code, and you don't need to pay Alibaba anything. You just need to keep the original copyright notice and license text, included somewhere in your product documentation. It's very enterprise-friendly and doesn't have the "viral" copyleft issue that some open-source licenses do.

Q: Is there copyright risk in content generated with Wan?

A: Normal, original use is fine. Just avoid these landmines: don't generate well-known IP characters, celebrity faces, or registered trademarks; don't use copyrighted images or video as references; don't generate false information or content that violates regulations; and make sure you actually have the rights to any reference material you upload. If your content is original and your use is lawful, there's basically no risk. The platform also runs content-safety review, and anything clearly in violation gets blocked.

Q: What should enterprises watch out for in commercial use?

A: Five recommendations: for low volume, go through Alibaba Cloud Bailian — you get a formal contract and invoices, which keeps compliance simple; for high volume or sensitive data, consider self-hosting for lower long-term cost and better data security; keep your payment receipts or deployment records on hand in case of a compliance review; have a human review generated content before it goes out, especially anything published externally, so a small AI flaw doesn't turn into a business incident; and if you're building a product-grade application, fine-tuning is recommended — the results will be noticeably better than the general-purpose model.