Back to feed
News Story
The AI Insider
1 sources

HKU-led team develops unified benchmark for physical AI

A research team led by the University of Hong Kong has developed RoboDojo, a benchmark that evaluates physical AI across simulation and real-world robot tasks, revealing a significant performance gap between current models and humans. The platform integrates 30 robot policies, 42 simulation tasks, and 18 real-world tasks to provide a standardized and reproducible method for comparing robot-learning systems.

SynthePulse Insight · AI deep reading

RoboDojo: HKU Launches Unified Benchmark Revealing Huge Gap Between Physical AI and Human Performance

Version 1 · 1 source

Developed by the University of Hong Kong, the RoboDojo benchmark evaluates physical AI in both simulation and real-world settings under a unified framework for the first time. Results show the strongest model achieves less than 13% success, while human experts reach 100% on real tasks.

  • RoboDojo was developed by the Multimedia Lab at the University of Hong Kong in collaboration with nearly 20 universities, including UC Berkeley and Tsinghua University.
  • The platform includes 30 robot policies, 42 simulated tasks, and 18 real-world tasks.
  • The strongest AI model achieves 8.8% success in simulated tasks and 12.8% in real-world tests.
  • Human experts achieve 76.03% success in simulated tasks and 100% in real-world tasks.
  • RoboDojo's open-source resources have been downloaded over 100,000 times on Hugging Face, and its X posts garnered over 100,000 views in the first week.
Open section navigationThe Birth of a Unified Benchmark: Why RoboDojo Is Needed

The Birth of a Unified Benchmark: Why RoboDojo Is Needed

In the field of robot learning, different systems often use different hardware, simulation environments, and scoring methods, making it difficult to directly compare models or determine whether results can transfer to the real world. Researchers from the Multimedia Lab at the University of Hong Kong, in collaboration with nearly 20 universities including UC Berkeley and Tsinghua University, developed the RoboDojo benchmark to provide a reproducible way to compare robot learning methods in both virtual and physical environments.

The RoboDojo project was initiated by Ping Luo, Associate Director of AI Research and Technology Transfer at the School of Computing and Data Science at HKU, and PhD student Tianxing Chen. Luo stated that RoboDojo is the first benchmark led by Hong Kong to unify simulation and standardized real-robot evaluation, aiming to move embodied AI from impressive demonstrations to measurable, comparable, and credible progress.

Composition and Testing Capabilities of RoboDojo

RoboDojo integrates simulation testing, standardized physical robot testing, and evaluation of robot control policies into a single framework. The platform includes 30 representative robot policies, 42 simulated tasks, and 18 real-world tasks.

The tests cover multiple key capabilities of physical AI, including generalization to new situations, memory of information, execution of precise actions, and completion of tasks requiring multiple steps and long time horizons.

Performance Gap: Significant Disparity Between AI and Humans

Early results show a huge gap between current AI systems and human performance. According to data from the University of Hong Kong, the best-performing AI model achieves 8.8% success in simulated tasks and 12.8% in real-world robot tests. In contrast, human experts achieve 76.03% success in simulated tasks and 100% in physical tasks.

These results highlight a problem: a system that can complete a task under carefully selected conditions may perform poorly when the environment, objects, or action sequences change. Standardized benchmarks help identify these weaknesses by providing common tasks and measurement methods.

Community Response and Open-Source Resources

RoboDojo has been made available to the broader robotics community. According to the University of Hong Kong, its open-source resources have been downloaded over 100,000 times on Hugging Face since release, and project-related posts on X have garnered over 100,000 views in the first week.

Credibility boundary

This article is primarily based on a report from The AI Insider, which relies on official statements from the University of Hong Kong. All specific data (such as success rates and task counts) are from that report and have not been independently verified.

Insight takeaway

RoboDojo provides a unified, reproducible benchmark that for the first time makes physical AI's simulation and real-world performance comparable. However, the significant gap between current AI and human performance indicates that physical AI is still in its early stages, and standardized evaluation is key to driving progress.

Primary report

The AI Insider

Primary source