Dong's team proposed the STAIR method, which integrates slow-thinking reasoning into safety alignment, effectively mitigating the 'alignment tax'—the performance degradation caused by safety alignment. Inspired by OpenAI's o1 model, this is among the first works to apply reasoning to safety alignment, though it is slower than direct judgment.
As agent systems enter enterprise workflows, executing tasks autonomously 24/7, safety issues become more complex. Dong's team is exploring full-lifecycle safety monitoring and protection for agents, recording system calls and behavior trajectories to detect unsafe operations in a timely manner. In terms of efficiency, their lightweight system outperforms commonly used models like Llama Guard on basic content safety tasks.
The challenge from academia to industry lies in the fact that papers focus on single risks, while real-world applications face multiple risk sources; academia focuses on single metrics, while practice requires simultaneous consideration of performance, safety, efficiency, and system coordination. Dong's team's advantage lies in their deep understanding of AI safety issues and their ability to identify risks from a red-team attack perspective.