Shanghai · Reinforcement learning
HengquanGuo.
Research as
a living atlas.

PhD student
ShanghaiTech University
01 / Research & selected work
The research
landscape.
01 / 5 selected works
Reinforcement Learning & Bandits
Bandits and online learning. Safe and constrained decision-making.
02 / 3 selected works
Recommendation & Bidding
Learning from feedback in recommendation and online advertising.
03 / 3 selected works
Agent / LLM Alignment
Reinforcement learning for agentic LLMs and multimodal large models.
18 publications
IdeaTrail: Full-Process Agent Trajectories for Scientific Ideation
Hengquan Guo
ArXiv preprint2026BLOCK: An Open-Source Bi-Stage MLLM Character-to-Skin Pipeline for Minecraft
Hengquan Guo
ArXiv preprint2026GRB: A Generative Reinforcement Bidding Framework for Multi-Channel Online Advertising
Hongchang Wu*, Weitong Ou*, Hengquan Guo*, Zixin Shao*, Hongyan Xue, Junwei Pan, Shudong Huang, Zhangbin Zhu, Xin Liu, Nianhua Xie, Lei Xiao, Haijie Gu
ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD 2026, Applied Data Science Track)2026Towards Temporal Interest Modeling in Recommendation via Reinforcement Learning
Hengquan Guo, Haobo Zhang, Junwei Pan, Wentao Ning, Xiaotian Li, Zhixiang Feng, Shoujun Liu, Gong Chen, Shudong Huang, Haijie Gu, Xin Liu
In Submission2026SABO: Safe and Aggressive Bayesian Optimization for Automatic Legged Locomotion Controller Tuning
Haobo Zhang, Zhiyong Yu, Hengquan Guo, Yuning Jiang, Xin Liu, Yuanming Shi
Submitted to ICRA2026Towards Safe and Optimal Online Bidding: A Modular Look-ahead Lyapunov Framework
Hengquan Guo, Haobo Zhang, Junwei Pan, Shudong Huang, Nianhua Xie, Lei Xiao, Haijie Gu, Jie Jiang, Xin Liu
International Conference on Learning Representations (ICLR 2026)2025Enhancing Safety in Reinforcement Learning with Human Feedback via Rectified Policy Optimization
Xiyue Peng, Hengquan Guo, Jiawei Zhang, Dongqing Zou, Ziyu Shao, Honghao Wei, Xin Liu
Advances in Neural Information Processing Systems (NeurIPS 2025)2025No Regret Reinforcement Learning Algorithms for Online Scheduling with Multi-Stage Tasks
Yongxin Xu, Hengquan Guo, Ziyu Shao, Xin Liu
International Joint Conference on Artificial Intelligence (IJCAI 2025)2025On the Power of Optimism in Constrained Online Convex Optimization
Haobo Zhang, Hengquan Guo, Xin Liu
International Joint Conference on Artificial Intelligence (IJCAI 2025)2025Safe Learning in Stochastic Continuum-Armed Bandit With Constraints and Its Application to Network Resource Management
Hengquan Guo, Qi Zhu, Xin Liu
IEEE/ACM Transactions on Networking2025Triple-Optimistic Learning for Stochastic Contextual Bandits with General Constraints
Hengquan Guo, Lingkai Zu, Xin Liu
International Conference on Machine Learning (ICML 2025)2025On Stochastic Contextual Bandits with Knapsacks in Small Budget Regime
Hengquan Guo, Xin Liu
International Conference on Learning Representations (ICLR 2025)2024QueueFlower: Orchestrating Microservice Workflows via Dynamic Queue Balancing
Hongchen Cao, Xinrui Liu, Hengquan Guo, Jingzhu He, Xin Liu
IEEE International Conference on Web Services (ICWS 2024)2024Stochastic Constrained Contextual Bandits via Lyapunov Optimization Based Estimation to Decision Framework
Hengquan Guo, Xin Liu
Annual Conference on Learning Theory (COLT 2024, Oral)2024Learning to Schedule Online Tasks with Bandit Feedback
Yongxin Xu, Shangshang Wang, Hengquan Guo, Xin Liu, Ziyu Shao
International Conference on Autonomous Agents and Multiagent Systems (AAMAS 2024)2023POBO: Safe and Optimal Resource Management for Cloud Microservices
Hengquan Guo, Hongchen Cao, Jingzhu He, Xin Liu, Yuanming Shi
Performance Evaluation 1622023Rectified Pessimistic-Optimistic Learning for Stochastic Continuum-Armed Bandit with Constraints
Hengquan Guo, Qi Zhu, Xin Liu
Learning for Dynamics and Control Conference (L4DC 2023)2022Online Convex Optimization with Hard Constraints: Towards the Best of Two Worlds and Beyond
Hengquan Guo, Xin Liu, Honghao Wei, Lei Ying
Advances in Neural Information Processing Systems (NeurIPS 2022)No matching publications.
02 / In practice
Ideas into
experience.
I am a PhD student at ShanghaiTech University, advised by Xin Liu. My research centers on reinforcement learning and bandits, spanning theoretical foundations and applications. In 2025, I was a research intern at Tencent CDG through the Tencent Rhino-Bird Elite Talent Program, advised by Junwei Pan, working on reinforcement learning for recommendation and online bidding. My current research interest is in reinforcement learning for agentic LLMs and multimodal large models.
Research Intern
Baidu / ERNIE
Working on RSI and agentic evaluation, spanning black-box behavioral assessment, white-box analysis, and reliability evaluation.
Research Intern
Analemma
Building the first multi-turn auto-research dataset and developing auto-research workflows across agentic harness design, model training, and research-oriented evaluation.
Research Intern
CDG · Tencent Rhino-Bird Elite Talent Program
Reinforcement learning for recommendation and online bidding; produced three research works, with papers accepted at ICLR 2026 and the KDD 2026 ADS Track, and one under submission.
03 / Recognition
Awards.
ICML Silver Reviewer
Top 25%
NeurIPS Top Reviewer
Top 5%National Scholarship for Doctoral Students
MoE, ChinaTencent Rhino-Bird Elite Talent Program
Selected among 70+Tencent Rhino-Bird Elite Talent Program Outstanding Award
Annual Conference on Learning Theory Travel Grant
ACM SIGMETRICS / IFIP Performance Travel Grant
Outstanding Teaching Assistant
ShanghaiTech University
ShanghaiTech University Scholarship
04 / Community
Service
& teaching.
Reviewer / Program Committee
NeurIPS, ICML, ICLR, IROS, IEEE/ACM TON, JSAC
Invited Talk, Online Convex Optimization with Hard Constraints
RLChina 2022
Invited Talk, Safe and Adaptive Online Decision-Making
Tencent Rhino-Bird Elite Talent Forum 2025
Invited Talk, Constrained Contextual Bandits via Lyapunov Optimization
POMS-HK 2025
Teaching Assistant, Online Optimization and Learning (CS245)
2022–2025