Back to feed
News Story
APriority78
机器之心
1 sources

Robots Have "Aha Moment"! Zetta ζ Enables Closed-Loop Online Learning for Physical Agents

Tsinghua AIR and DomainShift jointly released Zetta ζ, a closed-loop online learning system for physical agents. By evolving high-frequency critics, robots can observe the environment in real-time, trigger recovery, and accumulate reusable skills. Experiments show Zetta ζ achieves 90.8% and 93.6% success rates on LIBERO PRO and RoboCasa, significantly outperforming existing methods, and captures a robot "Aha Moment," marking a new starting point for scaling physical intelligence.

SynthePulse Insight · AI deep readingMembers

Zetta ζ: Letting Robots 'Aha' in Execution, Closed-Loop Self-Evolution Opens a New Path for Embodied Intelligence

Version 1 · 1 source

Tsinghua AIR and Domain Shift jointly launch Zetta ζ, which uses three closed loops for online learning, enabling robots to correct errors in real time during physical execution and accumulate skills, achieving significant success rate improvements on LIBERO-Pro and RoboCasa, even exhibiting human-like 'Aha Moments'.

  • Zetta ζ is a closed-loop online learning self-evolving physical agent jointly launched by Tsinghua AIR and Domain Shift, aiming to address the 'static, open-loop' issues of existing embodied intelligence systems.
  • Its core consists of three closed loops at different time scales: action-level Critic monitoring, batch candidate optimization, and validation-gated updates, forming a 'self-evolution flywheel'.
  • On LIBERO-Pro, Zetta ζ improves the average success rate from 34.5% to 90.8%; on RoboCasa, from 73.56% to 93.56%.
Open section navigationBackground: The Challenge of Online Learning in the Physical World

Background: The Challenge of Online Learning in the Physical World

Most existing embodied intelligence systems are 'static and open-loop', capable of executing fixed skills but struggling to perceive anomalies in real time and proactively repair during physical execution. Reflection by large models often occurs after an episode ends, making it impossible to test alternative actions online, credit assignment over complete trajectories is difficult, and retrospective analysis cannot capture the precise state at the moment of failure.

The real physical environment changes rapidly, and current physical agents can only review the entire trajectory after task failure, lacking the ability to judge, correct, and learn in real time based on state changes at the execution site. This leaves robots without online feedback when facing common issues like excessive grip force or pose deviation, preventing adjustment.

Zetta ζ's Three-Loop Self-Evolution Mechanism

The core of Zetta ζ consists of three closed loops at different time scales. The first loop, the Critic-Governed Action Loop, continuously monitors the physical state at the frequency of action execution, triggering recovery actions immediately upon detecting anomalies, and uses the 'VLA Re-entry Contract' to ensure control is only returned after the fault is cleared.

The second loop, the Rollout-Batch Candidate Optimization Loop, clusters and diagnoses failed trajectories after batch tasks, identifies the earliest time point deviating from normal states, and traces the root cause along the hierarchy of 'evaluation → critic → state representation → planning control → recovery → parameters', automatically generating candidate updates.

The third loop, the Validation-Gated Skill Update Loop, first runs historical regression tests on candidate updates, then validates on held-out data, and only writes updates that genuinely improve success rates and generalize into the skill memory bank, avoiding overfitting.

Free for now

Read the full analysis

3 more sections of analysis, plus the full takeaway

Loading

Credibility boundary

This article's information primarily comes from a report by Machine Intelligence on the joint research by Tsinghua AIR and Domain Shift, making it a secondary source. The experimental data (such as success rates and throughput improvements) are claims by the research team and have not been independently verified; they should be regarded as source claims. Some descriptions (such as the 'Aha Moment' phenomenon) are observations by the researchers and should be treated with caution.

Primary report

机器之心

Primary source