Back to feed
News Story
APriority76
阿里云开发者
1 sources

Alibaba Cloud Launches Wan3.0 Video Generation Model in Public Beta

Alibaba Cloud's Wanxiang team has released the next-generation video generation model Wan3.0, now in public beta. The model supports generating 30-second videos in a single run and, for the first time, accepts document and spreadsheet inputs, aiming to improve video length, realism, and consistency. Wan3.0 is available on the Qianwen AI platform, with API pricing announced.

SynthePulse Insight · AI deep reading

Wan3.0 Public Beta: 30-Second Video Generation and Document Input, Video Models Move Toward Productivity Tools

Version 1 · 1 source

Alibaba Cloud releases Wan3.0, supporting 30-second video generation, document input, and all-purpose reference. API pricing announced, but voice and text accuracy still have room for improvement.

  • Wan3.0 enters public beta, generating 30-second videos in a single run, with smart duration and video extension features.
  • Supports text, image, audio, video, and document formats including doc, xls, ppt, pdf, md, with files up to 100MB and 50 pages.
  • Emphasizes realism and consistency: unique faces for each person, maintaining stability in characters, props, scenes, and styles.
  • API pricing: 480P/720P/1080P at 0.3/0.6/1.2 yuan per second, with full API access coming soon.
  • Officially acknowledges that voice quality and text accuracy still have room for improvement.
Open section navigationPublic Beta and Core Upgrades

Public Beta and Core Upgrades

On August 6, Alibaba Cloud announced the public beta of its next-generation video generation model Wan3.0, available on the Qwen AI platform and in a gray release on the Qwen app. The company claims comprehensive upgrades in generation duration, universal creation, all-purpose reference, and realism.

Wan3.0 can generate 30-second videos in a single run, with smart duration recommendation and video extension features, aiming to move from 'generating a single shot' to 'telling a complete story.'

Everything Can Become Video: Document Input and Instruction Following

In addition to text, image, audio, and video modalities, Wan3.0 for the first time supports document formats such as doc, xls, ppt, pdf, txt, key, pages, numbers, and md, with a maximum of one file or link, files up to 100MB, and up to 50 pages.

The company claims strong prompt engineering and instruction-following capabilities, enabling conversion of documents into teaching materials, product demos, dynamic charts, and other video content, turning office materials directly into videos.

Realism and Consistency

Wan3.0 strives to accurately recreate the real world. In text-to-video, it aims for 'a thousand faces for a thousand people,' capturing facial and skin details, with natural and restrained emotional expressions, and complex emotions in group scenes.

In all-purpose reference tasks, it maintains consistency in key dimensions such as characters, props, sound, spatial relationships, and style, including facial features, hairstyles, clothing, prop logos, and scene positioning. Video editing capabilities are greatly improved, supporting modifications to visuals, plot, and dialogue.

Application Scenarios and Pricing

The company lists application scenarios such as film and television creation, advertising and marketing, design and creativity, and cultural tourism communication, emphasizing low cost and high efficiency.

API pricing is announced: 480P/720P/1080P at 0.3/0.6/1.2 yuan per second, with full API access coming soon.

Credibility boundary

This report is based on an announcement from Alibaba Cloud's developer official account, a first-party source. However, the content is official promotion, and some capability descriptions (such as 'a thousand faces for a thousand people' and 'precise replication') have not been independently verified. The company also acknowledges that voice and text accuracy still have shortcomings.

Insight takeaway

Wan3.0 expands the creative boundaries of video generation with 30-second duration and document input, but its real-world performance needs to be verified through public beta feedback, especially with voice and text accuracy being known weaknesses.

Primary report

阿里云开发者

Primary source