Global competition in physical AI models revolves around 'how to make AI understand and act in the physical world,' with multiple routes: NVIDIA Cosmos pursues cloud-based world models, Physical Intelligence explores cloud-based general robot policy models, Google Gemini Robotics advances cloud-based VLA, and Tesla Optimus emphasizes model-hardware closed loops. These routes answer different questions, but none specifically address 'how to make AI models adapt to on-device environments from inception.'
Om AI's self-developed VLX on-device streaming multimodal model series emphasizes the 'on-device native' route: not relying on cloud compute, not tied to a single hardware platform, and incorporating on-device compute, latency, power, and deployment cost as constraints from the design stage. This is fundamentally different from the transfer deployment of 'train large models then compress for deployment'—the latter is 'making powerful models run on devices,' while the former is 'if AI is born on devices, what should it look like?'
The value of on-device native lies in faster response, lower cost, and greater autonomy, enabling physical AI to enter factories, homes, drone inspections, and other scenarios. Robots cannot always rely on cloud brains because network latency, data transmission costs, and privacy issues are unacceptable in real-world settings.