About

About me

Hello! I’m an M.Sc. student at Nanjing University, where I study at the School of Artificial Intelligence under the supervision of Prof. Zongzhang Zhang. I am also a member of the LAMDA Group, led by Academician Zhi-Hua Zhou.

I received my B.Sc. degree from the School of Artificial Intelligence at Nanjing University in June 2024. In the same year, I was recommended for admission to the M.Sc. program at Nanjing University without taking the entrance examination.

My research interests lie in Reinforcement Learning, with a particular focus on offline reinforcement learning, reinforcement learning with generative models (e.g., Transformers and Diffusion models), and safe reinforcement learning (Constrained Reinforcement Learning). Beyond research, I’m also a big fan of open-source and self-hosted applications.

Publications

Behavior-Regularized Diffusion Policy Optimization for Offline Reinforcement Learning

Chen-Xiao Gao, Chenyang Wu, Mingjun Cao, Chenjun Xiao, Yang Yu, Zongzhang Zhang

  • Integrates behavior regularization with highly expressive diffusion policies
  • Introduces pathwise KL to analytically calculate KL divergence between the diffusion policies
  • Proposes an actor-critic framework with two-time-scale temporal difference learning to efficiently optimize diffusion policies with behavior regularization