BStandard70
InfoQ
1 sourcesTencent's UniRL Framework 2.4X Performance Optimization Practice to Be Shared at AICon Shenzhen
Zhu Wenxi, head of multimodal RL Infra at Tencent, will share the practice of 2.4X end-to-end performance optimization of the UniRL framework at AICon Shenzhen. The framework addresses the infrastructure differences between Diffusion RL and LLM RL through full-stack engineering practices such as architecture design, custom operators, training-inference consistency, and asynchronous engines, solving training stability and performance issues.