Back to feed
News Story
机器之心
1 sources

PixVerse Releases Real-Time World Model R2, Enabling Interactive Entertainment

Aish Technology today announced the release of PixVerse R2, the latest version of its real-time video world model. Building on R1, R2 offers more stable and continuous world interaction, supporting multimodal input and real-time response. Users can control game characters via text and keyboard, even summoning mounts or altering storylines. This marks a step toward practical applications of world models, opening new possibilities for interactive entertainment.

SynthePulse Insight · AI deep readingMembers

PixVerse R2: How Real-Time World Models Go from 'Magic Tricks' to 'Playable Worlds'

Version 1 · 1 source

Aishite Technology releases real-time world model R2, unifying Omni Causal AR with real-time acceleration distillation to push interactive entertainment from single-shot responses to sustainable systems.

  • PixVerse R2 upgrades real-time world models from single-shot responses to sustainable systems, supporting multimodal real-time interaction via WASD, text, and voice.
  • R2 uses the unified Omni Causal AR model, addressing long-term stability with Dynamic Chunk, Hybrid Teacher, three-tier memory, and Error Bank.
  • Real-Time Acceleration leverages DDMD, Block-sparse Attention, and Pyramid distillation to speed up capabilities to real-time, with attention sparsity exceeding 90%.
Open section navigationFrom R1 to R2: The Key Leap in Real-Time World Models

From R1 to R2: The Key Leap in Real-Time World Models

In January of this year, Aishite Technology launched PixVerse R1, the industry's first real-time video world model, which drew attention with its 'input a sentence, video changes in real time' gameplay. However, R1 could only respond to a single prompt and quickly became repetitive. The release of R2 marks a shift from 'magic tricks' to 'playable worlds'—the world can now run continuously and stably, allowing users to truly interact with a world.

Official beta test cases demonstrate three types of interaction: players use WASD to control movement while simultaneously using text to alter the scene (e.g., adding obstacles, changing uniforms to blend into a crowd); in an ancient-style world, players temporarily summon a tiger mount and ride it, then naturally return to the broken bridge storyline; a Sanxingdui mask character answers in dialect in real time with synchronized facial expressions; and during a digital human livestream, when a viewer interjects, the host quickly picks it up and transitions naturally.

The key to these cases is that different types of control (keyboard, text, voice) are woven into the same experience, and the world maintains narrative coherence—after changing uniforms, patrolling guards naturally ignore the player, and temporary whims don't disrupt the main plot. Behind this is a restructuring of R2's underlying architecture.

Free for now

Read the full analysis

3 more sections of analysis, plus the full takeaway

Loading

Credibility boundary

This article is based on Machine Intelligence's coverage of PixVerse R2; all technical details and metrics come from official beta test cases and technical reports, not independently verified. Some descriptions (such as 'industry's first' and 'killer combo') are from the source and should be treated with caution.

Primary report

机器之心

Primary source