In January of this year, Aishite Technology launched PixVerse R1, the industry's first real-time video world model, which drew attention with its 'input a sentence, video changes in real time' gameplay. However, R1 could only respond to a single prompt and quickly became repetitive. The release of R2 marks a shift from 'magic tricks' to 'playable worlds'—the world can now run continuously and stably, allowing users to truly interact with a world.
Official beta test cases demonstrate three types of interaction: players use WASD to control movement while simultaneously using text to alter the scene (e.g., adding obstacles, changing uniforms to blend into a crowd); in an ancient-style world, players temporarily summon a tiger mount and ride it, then naturally return to the broken bridge storyline; a Sanxingdui mask character answers in dialect in real time with synchronized facial expressions; and during a digital human livestream, when a viewer interjects, the host quickly picks it up and transitions naturally.
The key to these cases is that different types of control (keyboard, text, voice) are woven into the same experience, and the world maintains narrative coherence—after changing uniforms, patrolling guards naturally ignore the player, and temporary whims don't disrupt the main plot. Behind this is a restructuring of R2's underlying architecture.