Back to feed
News Story
SSignal87
AI前线
1 sources

Embodied AI Route Debate Shifts to Convergence: VLA and World Models Move Toward Hybrid Approaches

In 2026, the embodied AI field shifted from a debate between VLA and world models to convergence. At events like NVIDIA GTC and WAIC, companies including Zhipingfang, Xinghaitu, and XPeng publicly stated that the two are complementary, not opposed. Meanwhile, Ant Lingbo and Face Intelligence released new VLA models emphasizing generality and on-device inference, marking a move toward practical deployment.

SynthePulse Insight · AI deep reading

The Route Debate in Embodied Intelligence: VLA and World Models Move from Opposition to Fusion

Version 1 · 1 source

In 2026, the embodied intelligence track underwent a transformation from verbal positioning to grounded disillusionment. VLA and world models are moving from opposition to fusion, and industry consensus is gradually forming: there is no endgame, only continuously evolving routes.

  • At NVIDIA's GTC conference, Geely, Momenta, Huawei, and others publicly questioned VLA's limitations, sparking a route debate.
  • At WAIC 2026, vendor showcases shifted from VLA-dominated to a majority of 'world model + VLA' combinations.
  • Ant Lingbo open-sourced LingBot-VLA 2.0, incorporating 60,000 hours of real physical data, with inference latency of 130 milliseconds.
  • ModelBest's MiniCPM-Robot has only 1.5B parameters, with a single decision latency of 120 milliseconds, nearly half of π0.5's.
  • Daxiao Robotics released Kairos 3.1, with the 8B version achieving 125 milliseconds per inference on Jetson Thor.
  • In the first half of 2026, total funding in embodied intelligence reached 93.474 billion yuan, a year-on-year increase of about 5 times.
Open section navigationRoute Debate: From Public Questioning to Quiet Mixing

Route Debate: From Public Questioning to Quiet Mixing

In 2026, the embodied intelligence track saw intense debate over technical routes. At NVIDIA's GTC conference, major players like Geely, Momenta, and Huawei publicly questioned VLA's limitations, while 'world models' and VA routes offered answers from different directions. Most dramatically, Generalist AI, a co-creator of the VLA concept, publicly stated they would 'abandon' the labels of VLA and even world models, believing that 'being overly concerned with tool labels limits imagination toward physical AGI.'

However, months later, the debate volume decreased while model releases increased, shifting the core contest from 'who is right' to 'whether robots actually move.' At WAIC, robot exhibits were no longer 'demonstrations that can move' but applications around real scenarios like 3C electronics, automotive manufacturing, and warehouse logistics. More notably, leading players began tacitly 'mixing routes.'

Guo Yandong, founder of Zhi Pingfang, explicitly stated that world models are not a competing route to VLA but a core component of the VLA system. Gao Jiyang, CEO of Xinghitu, also believes the underlying logic of both is similar and will converge in the future. Liu Xianming of XPeng Motors emphasized that world models and second-generation VLA are not substitutes or competitors but jointly enhance understanding of the physical world through different training signals.

VLA's Self-Rescue: Smaller, Faster, More Capable

Despite doubts about generalization, VLA model updates entered a mini-boom, shifting toward productization. Ant Lingbo's LingBot-VLA 2.0, open-sourced in July, incorporated 60,000 hours of real physical data during pretraining, covering 20 robot configurations from 17 manufacturers. It leads π0.5 and GR00T N1.7 on the GM-100 benchmark, with inference latency controlled within 130 milliseconds on RTX 4090. Lingbo stated that VLA is not yet a 'ready-to-use' mature form but more like a foundational base for embodied intelligence.

ModelBest's MiniCPM-Robot series, with only 1.5B parameters, entered the first tier on benchmarks like LIBERO and Calvin, with a single decision latency of just 120 milliseconds, nearly half of π0.5's 234 milliseconds. By incorporating historical observations, it scored about 53 on RMBench, while π0.5 scored only about 10, enabling long-horizon tasks like folding clothes and making sandwiches.

Qingcang Robotics took a lightweight route with Evo-Depth, which has only 0.9B parameters, achieving 84.4% on Meta-World and 95.4% on LIBERO in simulation, with an average real-world success rate of about 90%. Deployment requires only about 3.2GB of VRAM and runs at about 12.3Hz inference frequency. These developments show that VLA is not 'dead' but has become smaller, faster, and more capable.

Evolution of World Models: From Prediction to Integrated Action

World models are also accelerating evolution, shifting from academic competitions in 'video generation' and 'trajectory prediction' to 'integrated action.' Daxiao Robotics' Kairos 3.1, released at WAIC, uses a hybrid Transformer architecture with shared mixed attention mechanisms, compressing multi-dimensional data such as visual observations, language instructions, force-tactile states, and policy trajectories into a unified latent space, achieving native fusion of understanding, generation, and prediction.

Kairos 3.1 features an autonomous reflection closed-loop mechanism that can autonomously identify problems and iteratively optimize action strategies upon execution failure. Its 8B version achieves single inference in just 125 milliseconds on NVIDIA's Jetson Thor platform, demonstrating practical value in household scenarios, such as autonomously breaking down laundry processes into over ten steps and automatically retrying after errors.

However, Jiang Yuhua, partner at Lingsheng Technology, offers a sober perspective: current world models in industry and academia are generally below 10B parameters, and no one has even plotted a scaling law curve. Compared to large language models, embodied intelligence is roughly at the stage before GPT-3. Gartner Research Vice President Gao Ting also stated that world models are currently more used for synthetic data generation, simulation, evaluation, and auxiliary planning, with direct use for controlling physical robots still in early stages.

Fusion and Transcendence: Multiple Paths to Physical AGI

Facing the 'ultimate answer,' the industry has become more humble. Lingbo stated that VLA excels in human-robot interaction and modality fusion, while world models excel in predicting future states; the two will accelerate fusion to form a complementary closed-loop intelligence. Lingbo has released LingBot-VA 2.0, the industry's first embodied-native world action model, and judges that neither current VLA nor world models are the endgame.

Zhi Pingfang proposed a brain-inspired architecture, NeuroVLA, mimicking the human brain's 'cortex-cerebellum-spinal cord' three-layer system, integrating VLA and world models at a system level. Generalist AI chose a 'from-scratch training' route, with its GEN-1 model having about 99% of parameters trained from scratch, achieving a success rate exceeding 99%, speed improvements of 2-3 times, and requiring only 1/10 of the data and fine-tuning of the previous generation.

In the first half of 2026, domestic embodied intelligence funding reached 93.474 billion yuan, a surge of about 5 times compared to the same period in 2025; there were 322 funding events, a year-on-year increase of 137%. The Ministry of Industry and Information Technology stated that China's humanoid robot annual production is expected to exceed 100,000 units this year. However, embodied intelligence still lacks a fully validated technical roadmap, and the next watershed is making robots truly move.

Credibility boundary

This article is based on an analysis piece from AI Front, containing viewpoints from executives and industry experts at multiple companies, as well as specific model and funding data. These data are mostly claimed by companies or institutions and have not been independently verified. Some data (such as total funding) have unclear sources and should be treated as source_claim.

Insight takeaway

The route debate in embodied intelligence has shifted from opposition to fusion, with VLA and world models complementing each other. The industry consensus is that there is no single endgame, only continuously evolving routes. The true watershed lies in whether robots can reliably and at scale complete tasks, not in technical labels.

Primary report

AI前线

Primary source