Back to feed
News Story
APriority83
机器之心
2 sources

HiDream.ai Unveils World's First Native All-Modal Interactive World Model HiDream-O1-World, Tops WBench

HiDream.ai today unveiled the world's first native all-modal interactive world model, HiDream-O1-World, built on its proprietary UiT architecture. The model supports multimodal inputs including text, images, and interactions, and can generate dynamic worlds with long-term spatiotemporal and physical consistency. In the WBench benchmark jointly launched by Meituan LongCat and Fudan University, HiDream-O1-World topped the Navi sub-leaderboard on its first attempt, scoring 73.3 in physical dimensions and 88.0 in consistency, ranking first overall. This launch marks a shift in AI content generation from one-way output to explorable and interactive world models.

SynthePulse Insight · AI deep readingMembers

HiDream-O1-World Tops WBench: A Paradigm Shift from "Generating Content" to "Constructing Worlds"

Version 1 · 1 source

Intelligence Future releases the world's first native full-modal interactive world model, HiDream-O1-World, based on its proprietary UiT architecture, achieving top scores in both physics and consistency on the WBench benchmark. This article breaks down its technical approach, benchmark results, and industry impact.

  • Intelligence Future releases HiDream-O1-World, claiming it to be the world's first native full-modal interactive world model, supporting multimodal inputs such as text, images, and interactions.
  • In the WBench benchmark jointly launched by Meituan LongCat and Fudan University, HiDream-O1-World topped the Navi sub-leaderboard on its first participation, with a physics score of 73.3 and a consistency score of 88.0, both ranking first.
  • The proprietary UiT architecture abandons a separate text encoder, mapping image pixels, text tokens, video voxels, and other signals into a shared token space for unified cross-modal reasoning.
Open section navigationFrom One-Way Generation to Interactive Worlds: A Paradigm Shift

From One-Way Generation to Interactive Worlds: A Paradigm Shift

Over the past three years, the AI content generation industry has evolved from text-to-image to text-to-video, but the essence remains a one-way paradigm where users provide prompts, the model outputs content, and users watch. This paradigm produces content, not worlds; users remain spectators, unable to enter or alter the visual space. The industry has recognized this limitation, with companies like Google and World Labs investing heavily in world models, but the debate persists over whether existing products are truly constructing worlds or merely rendering scenes.

Intelligence Future's answer is the release of HiDream-O1-World, claiming it to be the world's first native full-modal interactive world model. The model supports multimodal inputs such as text, images, and interactions, generating dynamic worlds with long-term spatial-temporal consistency and physical consistency. Users can freely roam in first- or third-person perspectives, edit characters and scenes in real time, covering styles such as realistic cities, natural landscapes, anime, fantasy worlds, and 3A game aesthetics.

Free for now

Read the full analysis

4 more sections of analysis, plus the full takeaway

Loading

Credibility boundary

This article's information primarily comes from a report by Machine Intelligence on Intelligence Future's product launch, which is a mix of vendor announcements and third-party benchmark results. The WBench scores are third-party benchmark results, but specific evaluation details are not disclosed; the ECCV 2026 acceptance is a vendor claim and has not been independently verified. All technical capability descriptions are based on vendor claims, and actual performance requires independent replication.

Primary report

机器之心

Primary source

Same-event coverage

Also covered by 1 sources