Back to feed
News Story
APriority81
机器之心
1 sources

Beyond Natural Language: PhiZero Teaches World Models 'Physical Language'

Researchers at the Chinese Academy of Sciences' Institute of Automation have introduced PhiZero, a world model built around 'physical language,' aiming to let AI learn a symbolic system that directly represents fine-grained changes in the physical world. The model can generate videos with coherent physical processes, transfer motion across different embodiments, explore interactive worlds from a single image, and predict visual consequences of robot actions. This work offers a new intermediate representation for video world models, potentially enhancing AI's understanding and prediction of physical laws.

SynthePulse Insight · AI deep readingMembers

PhiZero: Teaching AI the 'Language' of the Physical World

Version 1 · 1 source

When natural language falls short of describing physical changes, the Institute of Automation, Chinese Academy of Sciences, introduces PhiZero, using a 'physical language' as an intermediate layer to let AI understand and predict the dynamic evolution of the world.

  • PhiZero, proposed by a team at the Institute of Automation, CAS, centers on a 'physical language'—a compact set of discrete symbols that encode object motion, interaction, and state transitions, rather than static appearance.
  • It employs a 'reason-then-render' framework: a physical language tokenizer transcribes videos into symbol sequences, a reasoner autoregressively predicts future symbols, and a diffusion decoder generates videos.
  • For a 4-second, 8 FPS, 512×896 video, only 256 discrete symbols are needed after the first frame, whereas a standard video VAE requires 44,800 continuous visual tokens.
Open section navigationFrom Natural Language to Physical Language: Why a New Representation Is Needed

From Natural Language to Physical Language: Why a New Representation Is Needed

Natural language allows AI to read knowledge recorded by humans, but when facing motion, interaction, and change in the real environment, natural language alone may be insufficient. PhiZero is proposed to explore a symbolic system that directly represents fine-grained changes in the physical world.

Traditional video generation models predict future frames directly in pixel space, producing realistic visuals but lacking explicit expression of state transitions, action effects, and causal relationships. PhiZero introduces a 'physical language' as an intermediate layer, specifically encoding how objects move, how they interact with each other, and how scenes transition from one moment to the next.

This division makes state changes reusable: even when scenes, materials, or object forms change, the underlying motion and interaction rules can be preserved and transferred.

Free for now

Read the full analysis

4 more sections of analysis, plus the full takeaway

Loading

Credibility boundary

This article is based on a report by Jiqizhixin on PhiZero research, which is a second-hand account. All capability demonstrations, benchmark results, and data are from that report and have not been independently verified; they should be regarded as claims made by the research team.

Primary report

机器之心

Primary source