Back to feed
News Story
APriority76
机器之心
1 sources

DreamWorld: Geometry-Grounded Video Diffusion for More Consistent World Generation

HiDream.ai introduces DreamWorld, a method that combines spatial structure priors from 3D foundation models with video diffusion models, using explicit geometry as an intermediate representation to improve cross-view consistency and geometric stability. The work is accepted at ECCV 2026 and achieves leading performance on several benchmarks.

SynthePulse Insight · AI deep readingMembers

DreamWorld: Turning Geometry into an Intermediate Stop for Video Generation—Can It Truly Solve 3D Consistency?

Version 1 · 1 source

HiDream.ai proposes DreamWorld, which explicitly introduces the structural priors of a 3D foundation model into video diffusion models, adopting a two-stage framework of 'geometry first, appearance later.' It achieves leading results on RealEstate10K, Tanks-and-Temples, and WorldScore. However, whether this design can fully resolve geometric instability under large viewpoint changes still requires further validation.

  • DreamWorld, proposed by HiDream.ai and accepted at ECCV 2026, explicitly introduces structural priors from a 3D foundation model.
  • It adopts a Geometry-then-Appearance two-stage framework: first generate the geometric representation of the target viewpoint, then condition on it to generate RGB video.
  • Achieves leading performance on RealEstate10K, Tanks-and-Temples, and WorldScore, with an average WorldScore of 75.04.
Open section navigationProblem: Video Generation That 'Looks Real' Does Not Equal 3D Spatial Stability

Problem: Video Generation That 'Looks Real' Does Not Equal 3D Spatial Stability

As video diffusion models improve, generating novel-view videos from a single image along a specified camera trajectory has become an important route for world modeling. However, most existing methods rely on implicit spatiotemporal representations and lack explicit 3D geometric grounding. Under large viewpoint changes, models are prone to unreasonable geometric structures, object deformation, and cross-view spatial inconsistencies.

Existing methods typically use camera parameters, Plücker embeddings, or point cloud projections from single-image reconstruction as conditions. While they achieve decent camera control, video diffusion models still rely primarily on implicit spatiotemporal representations to infer unobserved regions. This means that as the camera moves, the model must not only generate new content but also continuously infer the previously invisible 3D structure—this is the main challenge.

Free for now

Read the full analysis

4 more sections of analysis, plus the full takeaway

Loading

Credibility boundary

The information in this article is primarily based on a secondary report from Machine Intelligence (机器之心) introducing the paper. The paper's acceptance at ECCV 2026, specific methods, and benchmark results are all from the paper itself but have not been independently verified. All numbers and conclusions should be regarded as claims from the paper, not confirmed facts.

Primary report

机器之心

Primary source