World models are also accelerating evolution, shifting from academic competitions in 'video generation' and 'trajectory prediction' to 'integrated action.' Daxiao Robotics' Kairos 3.1, released at WAIC, uses a hybrid Transformer architecture with shared mixed attention mechanisms, compressing multi-dimensional data such as visual observations, language instructions, force-tactile states, and policy trajectories into a unified latent space, achieving native fusion of understanding, generation, and prediction.
Kairos 3.1 features an autonomous reflection closed-loop mechanism that can autonomously identify problems and iteratively optimize action strategies upon execution failure. Its 8B version achieves single inference in just 125 milliseconds on NVIDIA's Jetson Thor platform, demonstrating practical value in household scenarios, such as autonomously breaking down laundry processes into over ten steps and automatically retrying after errors.
However, Jiang Yuhua, partner at Lingsheng Technology, offers a sober perspective: current world models in industry and academia are generally below 10B parameters, and no one has even plotted a scaling law curve. Compared to large language models, embodied intelligence is roughly at the stage before GPT-3. Gartner Research Vice President Gao Ting also stated that world models are currently more used for synthetic data generation, simulation, evaluation, and auxiliary planning, with direct use for controlling physical robots still in early stages.