Back to feed
News Story
机器之心
1 sources

Let AI Agents 'Play' World Models: New Benchmark PlayWorld for Long-Horizon Objectives

Researchers from the University of Hong Kong, Chinese University of Hong Kong, Zhejiang University, and Kuaishou Keling team introduced PlayWorld, a new benchmark for evaluating world models. It uses an Agent Player to observe geometric consistency, interaction fidelity, and evolution plausibility in generated videos, focusing on long-horizon goals rather than fixed action sequences. The benchmark and code are open-sourced.

SynthePulse Insight · AI deep readingMembers

PlayWorld: Let AI Agents 'Play' World Models, Long-Horizon Goal Evaluation Reveals Capability Boundaries

Version 1 · 1 source

When evaluating world models no longer relies on fixed button sequences, but instead lets AI agents 'play' games like real users, can we more fairly measure their ability to understand the world? The PlayWorld benchmark, proposed by the University of Hong Kong and others with Kuaishou Keling, uses 171 samples and over 820 questions to reveal deep shortcomings in current world models regarding sustained evolution and spatial consistency.

  • PlayWorld is proposed by the University of Hong Kong, Chinese University of Hong Kong, Zhejiang University, and Kuaishou Keling, upgrading the evaluation unit from fixed action sequences to scene-related long-term goals.
  • The core innovation is the Agent Player, which uses a multimodal model and interface converter to dynamically choose among five decisions—Keep, Stop, Extend, Correct, End—during interaction, adapting to different models' action granularities.
  • Evaluation covers four dimensions: geometric consistency, interaction fidelity, out-of-view evolution, and in-view evolution, with a total of 171 samples and over 820 questions.
Open section navigationFrom Fixed Buttons to Long-Term Goals: A Shift in Evaluation Paradigm

From Fixed Buttons to Long-Term Goals: A Shift in Evaluation Paradigm

Traditional world model evaluation often uses fixed action sequences, such as requiring the model to execute three right turns, but this approach cannot guarantee that all models reach the same viewpoint, making subsequent scoring incomparable. The proponents of PlayWorld argue that when humans play, they care about 'whether they complete a lap and return to the original landmark,' not 'whether they strictly executed three right turns.' Therefore, PlayWorld upgrades the evaluation unit to 'scene-related long-term goals': all models face the same initial world and the same goal, but the Agent Player is allowed to make necessary adjustments based on real-time observations.

This design preserves the comparability brought by basic action sequences while adapting to different models' action granularities. The researchers do not require the agent to plan the entire trajectory from scratch but instead make targeted corrections based on shared priors, balancing fair comparison and flexibility.

Free for now

Read the full analysis

4 more sections of analysis, plus the full takeaway

Loading

Credibility boundary

This article is primarily based on a report by Machine Intelligence (机器之心) on the PlayWorld benchmark, which is a secondary source. Links to the paper, project page, evaluation set, and code are provided, but specific experimental data (such as SANA-WM's pass rate) are not given original sources in the report and should be considered as claimed by the source rather than independently confirmed.

Primary report

机器之心

Primary source