About
About me
Hello! I’m an M.Sc. student at Nanjing University, where I study at the School of Artificial Intelligence under the supervision of Prof. Zongzhang Zhang. I am also a member of the LAMDA Group, led by Academician Zhi-Hua Zhou.
I received my B.Sc. degree from the School of Artificial Intelligence at Nanjing University in June 2024. In the same year, I was recommended for admission to the M.Sc. program at Nanjing University without taking the entrance examination.
My research interests lie in Reinforcement Learning, with a particular focus on offline reinforcement learning, reinforcement learning with generative models (e.g., Transformers and Diffusion models), and safe reinforcement learning (Constrained Reinforcement Learning). Beyond research, I’m also a big fan of open-source and self-hosted applications.
Publications
Physics-Informed Generative World Models for Real-Time Bidding: Deriving Statistical Laws from First Principles
- Introduces a physics-informed generative world model for real-time bidding in advertisement
- Captures extreme volatility and coupled auction dynamics, providing a high-fidelity simulation environment for reliable policy optimization
Behavior-Regularized Diffusion Policy Optimization for Offline Reinforcement Learning
- Integrates behavior regularization with highly expressive diffusion policies
- Introduces pathwise KL to analytically calculate KL divergence between the diffusion policies
- Proposes an actor-critic framework with two-time-scale temporal difference learning to efficiently optimize diffusion policies with behavior regularization
ACT: Empowering Decision Transformer with Dynamic Programming via Advantage Conditioning
- Identifies the limitations of existing return-conditioned decision-making methods and enhances them by using advantage as conditioning information