My name is Chao Yu(于超). I received my Ph.D. from the Department of Electronic Engineering at Tsinghua University
in 2023. I am currently an Assistant Professor (Distinguished Research Fellow) at the Embodied Decision Intelligence Lab (EDI Lab) at Tsinghua Shenzhen International Graduate School (SIGS)
. I also serve as the chairman of the Tsinghua Shenzhen International Graduate School - AgiBot Joint Research Center for Embodied Cognition and Decision Systems (JCES) 清华-智元联合研究中⼼主任. I’m also the co-founder of Striding AI(正行创新). I have been selected for the Youth Talent Support Program of the Chinese Institute of Electronics. My research has long focused on reinforcement learning–based decision intelligence. As first author or corresponding author, I have published more than 50 papers in top-tier international conferences and journals, including ICML, NeurIPS, ICLR, CVPR, ECCV, CoRL, IROS, ICRA, TMLR, and RAL, with over 7,000 citations on Google Scholar. My representative works include the multi-agent reinforcement learning algorithm MAPPO, which has received more than 4,000 Google Scholar citations, and RLinf, a large-scale reinforcement learning training framework for embodied intelligence, which has accumulated over 4,000 GitHub stars.
Feel free to reach out if you’d like to discuss research or explore potential collaboration!
📃 Research Interest
RL Infra
- My Technical Preference: Scalable reinforcement learning systems, training infrastructure, and system-algorithm co-design for large-scale policy optimization.
- Representative works on RL Infra include: RLinf and etc. covering efficient RL training, real-world online policy learning, and VLA+RL system design.
Strategic Agent
- My Research Focus: multi-agent RL, strategic reasoning, self-play, cooperation/competition, and language agents.
- Representative works include MAPPO, Fictitious Cross-Play, MARSHAL, WideSeek-R1, Werewolf game etc.
Embodied Agent
- My Application Interest: embodied intelligence with quadrupeds, drones, multi-robot systems, VLA models, and world-model-based robotic training.
- Representative works include WoVR, πRL, World4RL, RoboScape-R, VolleyBots, FlightBench, OmniDrones, and etc.
🏫 Educations
-
2019 - 2023: Department of Electronic Engineering, Tsinghua University
.
Ph.D. in Electronic Science and Technology.
Outstanding Doctoral Graduate (Top 5%), Outstanding Doctoral Thesis (Top 10%).
Advisor: Prof. Yu Wang; Co-advisor: Assistant Prof. Yi Wu. -
2016 - 2019: Department of Mechanical Engineering, Tsinghua University
.
M.S. in Mechanical Engineering and Automation.
Outstanding Master’s Thesis (Top 10%).
Advisor: Prof. Xin-Jun Liu. -
2012 - 2016: School of Automation, Beijing Institute of Technology
.
B.S. in Automation.
Outstanding Graduate (Top 15%).
📃 Publications
FUSCO: High-Performance Distributed Data Shuffling via Transformation-Communication Fusion
Human-Guided Online Reward Adaptation for Real-Robot Arm Manipulation
What Matters in Learning A Zero-Shot Sim-to-Real RL Policy for Quadrotor Control? A Comprehensive Study
ICPL: Few-shot In-Context Preference Learning via LLMs
DynaRL: Flexible and Dynamic Scheduling of Large-Scale Reinforcement Learning Training
WoVR: World Models as Reliable Simulators for Post-Training VLA Policies with RL
Beyond Imitation: Reinforcement Learning-Based Sim-Real Co-Training for VLA Models
USER: A Unified and Extensible System for Real-World Online Policy Learning in Embodied AI
WideSeek-R1: Exploring Width Scaling for Broad Information Seeking via Multi-Agent Reinforcement Learning
RoboScape-R: Unified Reward-Observation World Models for Generalizable Robotics Training via RL
Red Teaming Large Reasoning Models
πRL: Online RL Fine-tuning for Flow-based Vision-Language-Action Models
Long-horizon Locomotion and Manipulation on a Quadrupedal Robot with Large Language Models
MARSHAL: Incentivizing Multi-Agent Reasoning via Self-Play with Strategic LLMs
RLinf-VLA: A Unified and Efficient Framework for Reinforcement Learning of Vision-Language-Action Models
SAC Flow: Sample-Efficient Reinforcement Learning of Flow-Based Policies via Velocity-Reparameterized Sequential Modeling
RE-PO: Robust Enhanced Policy Optimization as a General Framework for LLM Alignment
JuggleRL: Mastering Ball Juggling with a Quadrotor via Deep Reinforcement Learning
World4RL: Diffusion World Models for Policy Refinement with Reinforcement Learning for Robotic Manipulation
D3P: Dynamic Denoising Diffusion Policy via Reinforcement Learning
Spec-VLA: Speculative Decoding for Vision-Language-Action Models with Relaxed Acceptance
Online Planning for Multi-UAV Pursuit-Evasion in Unknown Environments Using Deep Reinforcement Learning
VS-Bench: Evaluating VLMs for Strategic Reasoning and Decision-Making in Multi-Agent Environments
What Can RL Bring to VLA Generalization? An Empirical Study
ReinFlow: Fine-tuning Flow Matching Policy with Online Reinforcement Learning
Toward Real-World Cooperative and Competitive Soccer with Quadrupedal Robot Teams
Fine-tuning Diffusion Policies with Backpropagation Through Diffusion Timesteps
Hysteresis-Aware Neural Network Modeling and Whole-Body Reinforcement Learning Control of Soft Robots
Multi-Robot System for Cooperative Exploration in Unknown Environments: A Survey
Translate Policy to Language: Flow Matching Generated Rewards for LLM Explanations
Learning Strategic Language Agents in the Werewolf Game with Iterative Latent Space Policy Optimization
Learning Global Nash Equilibrium in Team Competitive Games with Generalized Fictitious Cross-Play
SleepNetZero: Zero-Burden Zero-Shot Reliable Sleep Staging With Neural Networks Based on Ballistocardiograms
A Survey on Self-play Methods in Reinforcement Learning
FlightBench: A Comprehensive Benchmark of Spatial Planning Methods for Quadrotors
CityLight: A Universal Model Towards Real-world City-scale Traffic Signal Control Coordination
LAGOON: Language-Guided Motion Control
Few-shot In-context Preference Learning using Large Language Models
LLM-Powered Hierarchical Language Agent for Real-time Human-AI Coordination
Language Agents with Reinforcement Learning for Strategic Play in the Werewolf Game
Fictitious Cross-Play: Learning Global Nash Equilibrium in Mixed Cooperative-Competitive Games
Automatic Truss Design with Reinforcement Learning
Learning Zero-Shot Cooperation with Humans, Assuming Humans Are Biased
Asynchronous Multi-Agent Reinforcement Learning for Efficient Real-Time Multi-Robot Cooperative Exploration
Learning Graph-Enhanced Commander-Executor for Multi-Agent Navigation
INCAME: Interruptible CNN Accelerator for Multirobot Exploration
A Benchmark of Planning-based Exploration Methods in Photo-Realistic 3D Simulator
SAVE: Spatial-Attention Visual Exploration
VMAPD: Generate Diverse Solutions for Multi-Agent Games with Recurrent Trajectory Discriminators
Multi-Agent Vulnerability Discovery for Autonomous Driving Policy by Finding AV-Responsible Scenarios
Unlocking the Potential of MAPPO with Asynchronous Optimization
Benchmarking Multi-agent Deep Reinforcement Learning Algorithms
INCA: INterruptible CNN Accelerator for Multi-tasking in Embedded Robots
CNN-based Monocular Decentralized SLAM on embedded FPGA
CNN-based Feature-point Extraction for Real-time Visual SLAM on Embedded FPGA
Long-Sighted Imitation Learning for Partially Observable Control
A DenseNet feature-based loop closure method for visual SLAM system
DS-SLAM: A Semantic Visual SLAM towards Dynamic Environments
Multi-robot coordination for high-speed pick-and-place tasks
🏆 Awards
- 2024: China Postdoctoral Excellent Special Foundation (Top 1,000 nationwide), Chinese Postdoctoral Science Foundation (CPSF).
- 2024: Postdoctoral Fellowship Program (Top 3,000 nationwide), Chinese Postdoctoral Science Foundation (CPSF).
- 2024: Runner-up for Outstanding Doctoral Thesis (Top 5), Chinese Intelligent Agent and Multi-Agent Systems.
- 2023: Shuimu Scholar Program, Tsinghua University.
- 2023: Chuanxin Future Scholar Program, Department of Electronic Engineering, Tsinghua University.
- 2023: Zhang Keqian Postdoctoral Fellowship, Department of Electronic Engineering, Tsinghua University.
- 2023: Outstanding Doctoral Thesis (Top 10%), Tsinghua University.
- 2023: Outstanding Doctoral Graduate (Top 5%), Tsinghua University.
- 2019: Outstanding Master’s Thesis (Top 10%), Tsinghua University.
- 2019 - 2023: First-Class Scholarship (3 times), Tsinghua University.
- 2015: National Scholarship, China Ministry of Education.
🎤 Talks
- 2026.08.05: 智源专访:具身智能下一步,走向“模型 + 系统”协同
- 2026.07.19: WAIC国地中心报告
- 2026.07.12 - 2026.07.18: RSS workshop
- 2026.06.13: 智源大会报告
- 2026.06.05: 华为云技术报告
- 2026.05.23: 杭州蚂蚁开源报告
- 2026.01.15: 北京人形机器人创新中心报告
📺 Live
- 2026.08.08: 张翼显(青稞)— 从端到端 VLA 到 Harness VLA:面向具身智能与机器人操作任务的记忆增强式执行框架
- 2026.07.29: 张翼显(XRobotics)— Harness VLA:可持续进化的具身智能体系统
- 2026.07.10: 刘志豪(Lumina)— STEAM:无需人工标注的时序集成优势建模,让真实世界机器人学习更进一步!
- 2026.04.22: 徐哲轩、苑会宁、徐泽来(将门)— 大模型在多智能体任务中的评估、训练与Scaling
- 2026.04.14: 苑会宁、张翼显(智源大厦 ICLR 分享)— ICLR 预讲会
- 2026.04.10: 施良致(3D视觉工坊)— 只用20条真实数据训练机器人?RL-Co:强化学习驱动的仿真-真机协同训练
- 2026.03.24: 徐哲轩(青稞)— 从 Depth Scaling 到 Width Scaling!WideSeek-R1:通过多智能体 RL 探索大模型的广度扩展
- 2026.03.10: 臧宏之(青稞)— 一起聊聊RLinf-USER:面向现实世界机器人在线策略学习的统一且可扩展系统
- 2026.03.08: 于舒昂(Xbotics)— RLinf-USER:真实世界在线进化,从系统瓶颈到统一高效,让具身智能真正“活起来”
-
2026.07.10: 江震南、施良致、臧宏之(Lumina)— [**RLinf 真机系列工作 WoVR, RL-Co and USER**](https://mp.weixin.qq.com/s/ueU4rsaaMWHcBa-59zIj5A) - 2025.12.28: 于超(青稞 AI 嘉年华)— 具身智能专题|2025 “青稞” AI 嘉年华
- 2025.12.25: 韦明杰(计算机视觉life)— RLinf-VLA框架技术报告——RL如何训练VLA?
- 2025.12.15: 张瑞泽、季世龙(深蓝)— NeurIPS’25 & CoRL’25|无人机也能打排球吗?来看看清华团队的解决方案
- 2025.12.09: 陈康(深蓝)— 对话πRL一作:RLinf流匹配 VLA 在线强化学习框架!π系列模型成功率提升至98%
- 2025.12.06: 陈康(青稞)— 从 π_0 到 π_RL:面向流匹配 VLA 的强化学习后训练框架
- 2025.12.02: 臧宏之(青稞)— RLinf-VLA 实践:从零上手 VLA(OpenVLA)强化学习
- 2025.12.02: 徐泽来(将门创投)— 大模型智能体可以玩好狼人杀吗?
- 2025.11.27: 刘志豪(3D视觉工坊)— 清华开源|πRL:首个面向流匹配 VLA 的在线强化学习微调框架
- 2025.11.26: 张同和、高枫(将门创投)— 清华RLinf团队: RL可以为VLA带来什么?
- 2025.11.25: 林灏(青稞)— 一起聊聊具身智能 RL 训练框架 RLinf 的系统设计
- 2025.11.12: 韦明杰(3D视觉工坊)— 统一高效 VLA+RL 训练框架:RLinf-VLA——RL 如何训练 VLA?
- 2025.11.10: 陈康(具身智能之心)— πRL:首个面向流匹配 VLA 的强化学习微调框架
- 2025.10.31: 徐泽来、于超(B站)— 2025 bilibili超级科学晚全程回顾
- 2025.10.28: 高枫、臧宏之(具身智能之心)— SFT 还是RL,VLA到底应该如何训练?
- 2025.07.18: 高枫(MSRA)— 待补充标题
👓 Projects
- Projects: [1] 基于深度强化学习的多无人机追逃博弈决策和控制关键技术研究,国家自然科学基金委,青年科学基金项目(C类), 2025-2027.
- Projects: [2] 多机协同高效机器学习系统研究,国家自然科学基金-中德合作交流基金, 2021-2025.
- Projects: [3] 具有强推理能力的大语言模型智能体关键技术研究,中国博士后基金特别资助, 2023-2025.
💼 Work Experience
- 2026.01 - Present: Assistant Professor, Tsinghua University.
- 2023.07 - 2025.12: Postdoctoral Researcher, Tsinghua University.
🙋 Recruitment
We are actively recruiting!
We are looking for Ph.D. and Master students, Postdocs at Tsinghua University, Ph.D. students of the Joint Program of Zhongguancun Academy and Tsinghua University and Undergraduate Interns with strong interests and motivation to work on frontier research topics including:
Candidates with hands-on systems building abilities and mathematical background are highly encouraged.
🔗 Bond