Alibaba's Wan3.0 brings multimodal video synthesis to production pricing

Alibaba's Wan3.0 extends video synthesis capabilities to multimodal inputs, accepting text, images, and documents to produce 30-second clips at 1080p for $6 per generation. The pricing and format flexibility signal intensifying competition in the video generation space, where production speed and cost efficiency are becoming key differentiators. Notably, Alibaba's quarterly profit fell 75 percent year-over-year as the company accelerates AI infrastructure investment, reflecting the capital intensity required to compete in frontier model development and the willingness of major tech players to absorb near-term margin pressure for long-term AI positioning.
Modelwire context
Analyst takeThe real story isn't the 30-second capability itself (OpenAI, Runway, and others already ship similar lengths). It's that Alibaba is willing to post a 75% profit drop to subsidize video generation at $6 per clip, suggesting the company views this as a market-share fight where margin compression now buys distribution and data later.
This is largely disconnected from recent activity in the space because we have no prior Modelwire coverage of video synthesis competition. What this belongs to is the broader pattern of Chinese tech giants (Alibaba, Tencent, ByteDance) absorbing near-term losses to own AI infrastructure. That playbook worked for cloud and mobile; Alibaba is betting it works for generative video too. The question is whether Western competitors can match the capital burn rate without shareholder revolt.
If Alibaba's Wan3.0 usage volume exceeds Runway or Pika's by Q1 2027 despite identical or higher pricing from competitors, that confirms the subsidy strategy is working and signals a shift toward Chinese vendors controlling video synthesis infrastructure. If usage stays flat or trails, the margin sacrifice was a failed land grab.
This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.
MentionsAlibaba · Wan3.0
Modelwire Editorial
This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.
Modelwire summarizes, we don’t republish. The Decoder originally reported this story as “Alibaba's Wan3.0 generates AI videos up to 30 seconds long from text, images, and documents”. The full content lives on the-decoder.com. If you’re a publisher and want a different summarization policy for your work, see our takedown page.