Let AI Agents 'Play' World Models: New Benchmark PlayWorld for Long-Horizon Objectives
Researchers from the University of Hong Kong, Chinese University of Hong Kong, Zhejiang University, and Kuaishou Keling team introduced PlayWorld, a new benchmark for evaluating world models. It uses an Agent Player to observe geometric consistency, interaction fidelity, and evolution plausibility in generated videos, focusing on long-horizon goals rather than fixed action sequences. The benchmark and code are open-sourced.