Imagine a scenario: a cup is placed in a cabinet by A, and B does not see it. If the world model only knows physical facts, it would predict B will look in the cabinet; but a model that understands minds would predict B thinks the cup is still on the table and therefore returns to the table to search. This example reveals a key issue: world models that only track objects, positions, and movements may produce predictions that are 'physically plausible but behaviorally wrong.'
Most current world models focus only on the physical level: where objects and agents are, and how visible scenes evolve. People are often just objects that move and perform actions, and their internal states do not truly enter the model. However, human behavioral decisions arise from the interaction between the external environment and internal mental-social variables. For example, service robots need to judge whether users are confused, medical assistants need to consider patients' fear and trust, and collaborative agents need to recognize social norms.
Cognitive science has studied these abilities through mental models, theory of mind (ToM), and BDI agent models, but current AI research either builds physical world models lacking mental states or reduces psychological reasoning to isolated theory-of-mind question-answering tasks. Neither perspective is sufficient to describe the complete world, because the next state of the world evolves from both physical and mental states.