Back to feed
News Story
机器之心
2 sources

NeoteAI and Fudan University Release Tactile Data Technical Reports to Enhance Embodied AI's Sense of Touch

NeoteAI, in collaboration with Fudan University, has released three technical reports aiming to address the lack of tactile feedback in embodied AI by providing 30,000 hours of tactile interaction data. The team introduced a unified tactile representation model, NeoForce, and two technical approaches: -VTLA predicts future tactile evolution to improve operation success rates, while -TWAM integrates touch into world models, allowing robots to simulate tactile sensations before execution. The project's data and models will be open-sourced, potentially advancing fine manipulation in embodied AI.

SynthePulse Insight · AI deep reading

Can Tactile Data Become the Next Scaling Law for Embodied Intelligence?

Version 1 · 1 source

New Zhi Embodied Intelligence, in collaboration with Fudan University, releases a series of technical reports using 30,000 hours of tactile data to fill the gap in robots' 'sense of touch,' achieving significant improvements in fine manipulation tasks like plugging and folding, though data costs and generalization remain challenges.

  • New Zhi Embodied Intelligence and Fudan University release three technical reports, constructing a tactile data foundation NeoData (30,000 hours of interaction videos, 1.4 million manipulation clips) and proposing a unified representation model NeoForce to solve cross-device data fusion.
  • -VTLA predicts tactile evolution 50 steps ahead, achieving 85% success in plug insertion (60% for vision-only), 99% in key extraction (35% for vision), and enables autonomous evolution from failure data.
  • -TWAM integrates touch into world models, achieving 84.5% average success in simulation (baseline 36%) and 46.3% across 8 real-world tasks (baseline 21.9%), with ablation studies confirming touch is indispensable.
  • Despite significant results, real-world data collection costs, cross-embodiment generalization, and inference real-time performance remain engineering challenges; whether touch can leap from 'auxiliary modality' to 'core infrastructure' still needs verification.
Open section navigationThe Missing Sense of Touch: The 'Last Centimeter' Pain Point for Robots

The Missing Sense of Touch: The 'Last Centimeter' Pain Point for Robots

In the deployment of embodied intelligence, robots often encounter situations where 'the vision system judges correctly, but physical interaction fails instantly,' such as a plug repeatedly rubbing without inserting, or a plastic bottle being grasped with excessive force causing it to collapse. These failures point to an industry consensus: the key to a robot's survival in the real physical world lies not only in visual perception accuracy but also in tactile feedback at the moment of contact. Cameras can confirm spatial positions but cannot perceive 'whether contact is made,' 'hitting an edge,' 'over-tightening,' or 'slipping.' The lack of this feedback is the core pain point hindering robots from crossing the 'last centimeter.'

NeoData and NeoForce: Building a 'Common Language' for Touch

The large-scale application of tactile data has long been hindered by incompatible output formats of various devices. The NeoData dataset built by New Zhi Embodied Intelligence aims to break this barrier: it includes over 30,000 hours of visual-tactile interaction videos, approximately 1.4 million manipulation clips, 3.3 billion time steps, corresponding to 8 billion RGB frames and 10 billion tactile image frames, using 6 robot platforms including Franka, Piper, UR5e, covering 450 real long-horizon tasks, collected by 90 operators, with 5,000 hours of visual-tactile interaction videos already open-sourced.

The team also proposes the NeoForce unified representation model, which, regardless of the underlying sensor hardware, can learn transferable, temporally structured tactile representations for downstream embodied intelligence policies, effectively solving the cross-device data fusion problem.

-VTLA: Tactile Prediction and Learning from Failure

-VTLA proposes not relying on delayed current tactile feedback, but predicting tactile evolution 50 steps into the future. The model encodes current tactile information into compact latent tokens, combined with visual context to predict trends such as slipping and force, thereby guiding the action module to adjust in advance. In plug insertion tasks, it achieves 85% success, while similar vision-only methods achieve only 60%; in key extraction tasks, it achieves 99% success, compared to 35% for vision-only methods.

Additionally, -VTLA breaks through the limitation of traditional imitation learning that relies solely on expert success trajectories, enabling robots to autonomously evolve from failure data. The model uses deployment data and human-in-the-loop data combined with tactile information to train a progress assessment model, identifying task execution stages and complex behaviors such as 'regression' and 'recovery.' Through offline reinforcement learning, the success rates for long-horizon tasks like towel folding, backpack packing, and cardboard folding jump from 50%, 35%, and 20% to 95%, 80%, and 75%, respectively.

-TWAM: 'Tactile Imagination' in World Models

-TWAM incorporates touch into world models, allowing robots to fully simulate operational perspectives and tactile sensations in 'imagination' before acting. Existing world models are mostly limited to generating future visual images, facing physical alignment challenges when driving real robots: the image shows the cup being grasped, but is the grip stable? This tactile information is completely missing. -TWAM simultaneously predicts three coupled sequences: future video, tactile, and action.

-TWAM adopts a mixture-of-experts architecture with three asynchronous expert networks for video, tactile, and action, totaling 7.2 billion parameters, only half of a full-width solution. In simulation and real-world tests, the average success rate in simulation reaches 84.5% (the strongest world model baseline is 36%); across 8 real-world tasks, the average success rate is 46.3% (LingBot-VA is 21.9%). Ablation studies further confirm that removing 'future tactile prediction' or 'current tactile conditioning' leads to significant performance drops, establishing the indispensability of touch in world models.

Challenges and Outlook: Can Touch Become the New Scaling Law?

Looking across the three reports, NeoteAI sends a clear signal: -Foundation validates that the tactile modality has scaling potential similar to vision; -VTLA proves that touch can naturally integrate into existing VLA architectures; -TWAM establishes the core role of touch in the top-level design of world models. The cognitive paradigm of robots is evolving from single 'vision-driven' to 'visuo-tactile fusion.'

However, there is still a long way to go before achieving general embodied intelligence that can 'know the weight with a single touch.' Real-world data collection costs, cross-embodiment generalization, and inference real-time performance remain engineering challenges to be overcome. Although this series of works has substantially pushed touch from a 'niche research direction' to 'core infrastructure for embodied intelligence,' whether touch can truly become the next Scaling Law still requires more verification.

Credibility boundary

This article's information primarily comes from a series of technical reports jointly released by New Zhi Embodied Intelligence and Fudan University, republished by Machine Heart. All data and success rates are as claimed in the reports and have not been independently verified by third parties.

Insight takeaway

Tactile data shows great potential in fine manipulation, but engineering challenges remain for generalization; visuo-tactile fusion may become a key direction for the next stage of embodied intelligence.

Primary report

机器之心

Primary source

Same-event coverage

Also covered by 1 sources