Back to feed
News Story
机器之心
1 sources

Tsinghua Team Publishes Humanoid Robot Soccer Research in Science Robotics

A team led by Professor Zhao Mingguo at Tsinghua University, in collaboration with ByteDance Seed and China Agricultural University, published a paper in Science Robotics proposing a vision-driven reactive soccer skill learning framework that enables humanoid robots to locate, chase, and shoot a ball using only onboard vision. The research was validated on Accelerated Evolution's humanoid robot platform and tested in RoboCup competitions, marking significant progress in integrating perception and control for humanoid robots.

SynthePulse Insight · AI deep reading

Tsinghua's 22-Year Journey: How Humanoid Robots Learn to See and React in Noise

Version 1 · 1 source

From the founding of the Vulcan team in 2004 to a Science Robotics publication in 2026, the Tsinghua team used an end-to-end reinforcement learning framework to enable humanoid robots to find, chase, and shoot a ball using only onboard vision in dynamic environments. This is not just an algorithmic victory, but a 22-year relay of technology.

  • A team led by Zhao Mingguo at Tsinghua University, in collaboration with ByteDance Seed and China Agricultural University, published a paper in Science Robotics proposing a vision-driven reactive soccer skill learning framework.
  • The research employs end-to-end perception-motor reinforcement learning, unifying visual perception and motor control in a single optimization to address noise, latency, and occlusion in real-world environments.
  • Real-robot validation was conducted on the accelerated evolution humanoid platform without special modifications, with perception and control running entirely on onboard computing.
  • Experiments show that the ball position error one second before shooting dropped from 0.344 meters to 0.186 meters, a reduction of about 46%; with a 0.3-second visual interruption, the ball contact success rate still exceeded 90%.
  • In 2025, the Vulcan team won both the RoboCup Adult-Size Humanoid League and the World Humanoid Robot Games, scoring 76 goals and conceding only 11.
  • Cheng Hao, founder of Accelerated Evolution, was the third captain of the Vulcan team and founded the company in 2023, with several former team members joining, creating a dual legacy of talent and technology platform.
Open section navigationA 22-Year Technology Relay

A 22-Year Technology Relay

In 2004, Zhao Mingguo led students to establish the Tsinghua Vulcan robot soccer team, which competed in RoboCup the following year. Early robots were only about 50 centimeters tall, walked slowly, and required spotters to prevent falls during matches. Over the next two decades, the Vulcan team competed almost every year, with robots growing to over a meter tall, learning to stand up autonomously, run quickly, and shoot powerfully.

The soccer field became a long-term testing ground for the team's vision, decision-making, motor control, and overall robot reliability. How many milliseconds slower the vision recognition was, how many centimeters off the foot placement was, and in which direction the robot tended to miss the ball were all directly revealed in real competition. The team then brought these issues back to the lab, modified algorithms, and continued validation in the next match—a cycle that persisted for over 20 years.

The Vulcan team's membership changed over generations. Cheng Hao, founder of Accelerated Evolution, was the third captain of the Vulcan team and founded 'Accelerated Evolution' in 2023, with several former team members joining. The paper's first author, Wang Yushi, comes from the new generation of the Vulcan team. From 2004 to the paper's publication in a top international journal in 2026, this is a 22-year technology relay.

Seeing and Reacting in Noise

The paper addresses a long-standing challenge: how can humanoid robots achieve see-and-react in real environments filled with noise, latency, and occlusion? Traditional approaches decompose the process into multiple modules, which can lead to error and latency accumulation, and policies trained in simulation may degrade in the real world due to lighting, noise, or occlusion.

Zhao Mingguo's team proposed a unified perception-motor reinforcement learning framework, training vision and motor control end-to-end under a single optimization objective. They first built a virtual perception system: real robots observed a soccer ball at various distances and angles, collecting about one hour of data, which was compared with motion capture ground truth to derive perception characteristics such as position noise, detection success rate, update frequency, and latency.

The vision system updates at approximately 25 Hz with an average latency of about 116 milliseconds. The policy reads the past 50 frames (about one second) of observations, compresses them into a 64-dimensional hidden state, and uses a decoder to aid visual denoising, allowing the robot to continue estimating the ball's position after brief occlusion. Training data includes about 76 seconds of human omnidirectional walking and 30 seconds of instep shooting motions, with adversarial motion prior (AMP) and mirror symmetry constraints. The final controller outputs joint position commands at 50 Hz.

Emergent Behaviors and Experimental Data

Emergent behaviors appeared during training: when the ball approached the edge of the field of view, the robot actively turned its head and body; when the ball was not found near the boundary, it first oriented toward the center and then expanded its search; gait varied with distance, with a step cycle of about 0.7-0.8 seconds at long range and shortening to 0.3-0.4 seconds at close range; when the goal was behind, the policy generated a spinning hook shot, pivoting on the support foot.

Experimental data show that the raw visual ball position error one second before shooting was 0.344 meters, which dropped to 0.186 meters after estimation by the policy's internal state, a reduction of about 46%. In simulation, with a 0.3-second visual interruption, the ball contact success rate still exceeded 90% and the goal rate exceeded 50%. For a stationary ball, the learned policy typically took about 1.5 seconds from initiation to contact, while a traditional rule-based system took about 2-5 seconds, with a maximum time reduction of 64%.

In real-robot tests, the success rate in the front field was 80%-90%, and in the back field 60%-70%, with no falls during testing. For a rolling ball, the success rate still exceeded 50% when the ball speed was below 0.5 meters per second. In 2025, the Vulcan team won both the RoboCup Adult-Size Humanoid League and the World Humanoid Robot Games, scoring 76 goals and conceding only 11 across both competitions; in 2026, they defended their title in the RoboCup Large Size competition.

Toward the Next-Generation Embodied Platform

The learned policy can be deployed directly from simulation to the real robot, but the algorithm is only one part. Joint response consistency, control cycle stability, and timely output from sensors and computing units all affect performance, and any underlying error can be amplified during high-speed motion. The real-robot experiments in the paper used the Accelerated Evolution humanoid platform without special modifications, with ball information from a head-mounted camera and perception and control running entirely on onboard computing.

The relationship between Accelerated Evolution and the Vulcan team extends from talent to technology platform. In July of this year, Accelerated Evolution released its new flagship platform, Booster T2, with the professional version equipped with an NVIDIA Thor chip, delivering 2070 TFLOPS of edge computing power, claimed to be the most powerful computing platform in the bipedal humanoid field. The accompanying Booster Studio integrates simulation training, algorithm development, and real-robot deployment into a single development environment.

On August 19, Booster T2 debuted at the World Robot Conference. Three days later, 80 Booster T2 units performed centerless autonomous coordination at the opening ceremony of the Robot Games, writing 'BEIJING', the emblem, and event icons without remote control or commands. 80 independently deciding 'brains' achieved coordination simultaneously, with a total computing power of about 170,000 TFLOPS. From early robots that were 50 centimeters tall and needed human support, to Booster T2 autonomously playing soccer on a thousand-square-meter field and writing in formation at the opening ceremony, humanoid robots are moving from 'can move' to 'can be used'.

Credibility boundary

This article is based primarily on a report by Jiqizhixin (Machine Intelligence) about the paper and the team, making it a secondary source. Paper details (such as specific data and architecture) are as reported and have not been verified against the original paper, but the report cites the paper link and project homepage, lending it high credibility. Some statements, such as 'most powerful computing', are vendor claims and should be treated with caution.

Insight takeaway

The significance of this research lies not only in teaching robots to play soccer, but also in validating the feasibility of end-to-end learning frameworks in real noisy environments and the potential of domestically produced humanoid platforms to support cutting-edge research. Twenty-two years of persistence have ultimately moved robots from 'can move' to 'can be used'.

Primary report

机器之心

Primary source