I am a Ph.D. candidate in Artificial Intelligence at the Hong Kong University of Science and Technology (ECE Department), advised by Prof. Ling Pan. I am honored to maintain deep research collaborations with Prof. Aaron Courville (Mila, co-author of the Deep Learning textbook) and Prof. Pablo Samuel Castro (DeepMind), with whom I investigate fundamental limitations of reinforcement learning through mechanistic analysis.
I am particularly interested in demystifying the internal mechanisms and fundamental limitations of RL, and in broadening its applicability to new domains. Currently, I focus on RL-driven fine-tuning of large foundation models, aiming to develop more efficient, stable, and capable alignment methods for next-generation AI systems. Examples include RL algorithms adapted to advanced model architectures, and RL algorithm optimization adapted to large-scale Infra.
Join our mission to reimagine reinforcement learning for foundation models—where simplicity outperforms complexity. Are you frustrated by RL4LLM pipelines that layer redundant techniques without measurable gains? Do you believe alignment algorithms should be understandable, not just empirically tuned? Our team is assembling researchers who share a conviction: the future of RL lies in mechanistic insight.
Beginner-friendly: If you’re interested in RL but lack extensive experience—don’t worry! We provide fine-grained, intensive guidance to help you contribute meaningfully.
→ Reach out: ljshasdream@gmail.com
Large Foundation Model
Tricks or Traps? A Deep Dive into RL for LLM Reasoning
Technical Report, Alibaba Future Living Lab (ROLL team)
Co-first author. First algorithmic technical report from ROLL team analyzing RL4LLM over-engineering.
Asymmetric Proximal Policy Optimization: mini-critics boost LLM reasoning
Technical Report, Alibaba Future Living Lab (ROLL team)
First author. Designed mini-critic architecture for implicit multi-agent debate within single-agent RL.
Let It Flow: Agentic Crafting on Rock and Roll
Technical Report, Alibaba Future Living Lab (ROLL team)
Core contributor. First open agentic model training pipeline (ROME Model) from ROLL team.
Accelerating RLVR and Agentic Training with Asynchrony
Technical Report, Alibaba Future Living Lab (ROLL team)
Contributor. Asynchronous post-training infrastructure (RoLL framework).
Deep Reinforcement Learning
Measure gradients, not activations! Enhancing neuronal activity in deep reinforcement learning
NeurIPS 2025
Jiashun Liu, Zihao Wu, Johan Obando-Ceron, Pablo Samuel Castro, Aaron Courville, Ling Pan
The Courage to Stop: Overcoming Sunk Cost Fallacy in Deep Reinforcement Learning
ICML 2025
Jiashun Liu, Johan Obando-Ceron, Aaron Courville, Pablo Samuel Castro, Ling Pan
Neuroplastic Expansion in Deep Reinforcement Learning
ICLR 2025
Jiashun Liu, Johan Obando-Ceron, Aaron Courville, Ling Pan
Flow Factorization for Efficient Generative Flow Networks
AAAI 2025 (Oral, <5%)
Jiashun Liu, Chunhui Li, Cheng-Hao Liu, Dianbo Liu, Qingpeng Cai, Ling Pan
Unlock the Intermittent Control Ability of Model Free Reinforcement Learning
NeurIPS 2024
Jiashun Liu, Xiaotian Hao, Jianye Hao, Yi Ma, Yan Zheng, Yujing Hu, Tangjie Lv
Unlock the Cognitive Generalization of Deep Reinforcement Learning via Granular Ball Representation
ICML 2024
Jiashun Liu, Jianye Hao†, Yi Ma, Shuyin Xia†
Hybrid CtrlFormer: Learning Adaptive Search Space Partition for Hybrid Action Control
UAI 2024
Jiashun Liu, Xiaotian Hao, Yan Zheng, Jianye Hao†, Yujing Hu, Changjie Fan, Tangjie Lv, Zhipeng Hu
Email: ljshasdream@gmail.com
WeChat: ljsweixin9
Affiliation: Ph.D. Candidate, ECE Department, Hong Kong University of Science and Technology (HKUST)
Advisors: Prof. Ling Pan (HKUST); Collaborations with Prof. Aaron Courville (Mila) and Prof. Pablo Samuel Castro (DeepMind)
Powered by Jekyll and Minimal Light theme.