Shanghai · Reinforcement learning

HengquanGuo.

Research as
a living atlas.

Hengquan Guo's blue bird portrait

PhD student
ShanghaiTech University

01 / Research & selected work

The research
landscape.

18 publications

2026

IdeaTrail: Full-Process Agent Trajectories for Scientific Ideation

Hengquan Guo

ArXiv preprint
2026

BLOCK: An Open-Source Bi-Stage MLLM Character-to-Skin Pipeline for Minecraft

Hengquan Guo

ArXiv preprint
2026

GRB: A Generative Reinforcement Bidding Framework for Multi-Channel Online Advertising

Hongchang Wu*, Weitong Ou*, Hengquan Guo*, Zixin Shao*, Hongyan Xue, Junwei Pan, Shudong Huang, Zhangbin Zhu, Xin Liu, Nianhua Xie, Lei Xiao, Haijie Gu

ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD 2026, Applied Data Science Track)
2026

Towards Temporal Interest Modeling in Recommendation via Reinforcement Learning

Hengquan Guo, Haobo Zhang, Junwei Pan, Wentao Ning, Xiaotian Li, Zhixiang Feng, Shoujun Liu, Gong Chen, Shudong Huang, Haijie Gu, Xin Liu

In Submission
2026

SABO: Safe and Aggressive Bayesian Optimization for Automatic Legged Locomotion Controller Tuning

Haobo Zhang, Zhiyong Yu, Hengquan Guo, Yuning Jiang, Xin Liu, Yuanming Shi

Submitted to ICRA
2026

Towards Safe and Optimal Online Bidding: A Modular Look-ahead Lyapunov Framework

Hengquan Guo, Haobo Zhang, Junwei Pan, Shudong Huang, Nianhua Xie, Lei Xiao, Haijie Gu, Jie Jiang, Xin Liu

International Conference on Learning Representations (ICLR 2026)
2025

Enhancing Safety in Reinforcement Learning with Human Feedback via Rectified Policy Optimization

Xiyue Peng, Hengquan Guo, Jiawei Zhang, Dongqing Zou, Ziyu Shao, Honghao Wei, Xin Liu

Advances in Neural Information Processing Systems (NeurIPS 2025)
2025

No Regret Reinforcement Learning Algorithms for Online Scheduling with Multi-Stage Tasks

Yongxin Xu, Hengquan Guo, Ziyu Shao, Xin Liu

International Joint Conference on Artificial Intelligence (IJCAI 2025)
2025

On the Power of Optimism in Constrained Online Convex Optimization

Haobo Zhang, Hengquan Guo, Xin Liu

International Joint Conference on Artificial Intelligence (IJCAI 2025)
2025

Safe Learning in Stochastic Continuum-Armed Bandit With Constraints and Its Application to Network Resource Management

Hengquan Guo, Qi Zhu, Xin Liu

IEEE/ACM Transactions on Networking
2025

Triple-Optimistic Learning for Stochastic Contextual Bandits with General Constraints

Hengquan Guo, Lingkai Zu, Xin Liu

International Conference on Machine Learning (ICML 2025)
2025

On Stochastic Contextual Bandits with Knapsacks in Small Budget Regime

Hengquan Guo, Xin Liu

International Conference on Learning Representations (ICLR 2025)
2024

QueueFlower: Orchestrating Microservice Workflows via Dynamic Queue Balancing

Hongchen Cao, Xinrui Liu, Hengquan Guo, Jingzhu He, Xin Liu

IEEE International Conference on Web Services (ICWS 2024)
2024

Stochastic Constrained Contextual Bandits via Lyapunov Optimization Based Estimation to Decision Framework

Hengquan Guo, Xin Liu

Annual Conference on Learning Theory (COLT 2024, Oral)
2024

Learning to Schedule Online Tasks with Bandit Feedback

Yongxin Xu, Shangshang Wang, Hengquan Guo, Xin Liu, Ziyu Shao

International Conference on Autonomous Agents and Multiagent Systems (AAMAS 2024)
2023

POBO: Safe and Optimal Resource Management for Cloud Microservices

Hengquan Guo, Hongchen Cao, Jingzhu He, Xin Liu, Yuanming Shi

Performance Evaluation 162
2023

Rectified Pessimistic-Optimistic Learning for Stochastic Continuum-Armed Bandit with Constraints

Hengquan Guo, Qi Zhu, Xin Liu

Learning for Dynamics and Control Conference (L4DC 2023)
2022

Online Convex Optimization with Hard Constraints: Towards the Best of Two Worlds and Beyond

Hengquan Guo, Xin Liu, Honghao Wei, Lei Ying

Advances in Neural Information Processing Systems (NeurIPS 2022)
Reinforcement Learning & Bandits

02 / In practice

Ideas into
experience.

I am a PhD student at ShanghaiTech University, advised by Xin Liu. My research centers on reinforcement learning and bandits, spanning theoretical foundations and applications. In 2025, I was a research intern at Tencent CDG through the Tencent Rhino-Bird Elite Talent Program, advised by Junwei Pan, working on reinforcement learning for recommendation and online bidding. My current research interest is in reinforcement learning for agentic LLMs and multimodal large models.

20262026.08–Present

Research Intern

Baidu / ERNIE

Working on RSI and agentic evaluation, spanning black-box behavioral assessment, white-box analysis, and reliability evaluation.

20262026.04–2026.07

Research Intern

Analemma

Building the first multi-turn auto-research dataset and developing auto-research workflows across agentic harness design, model training, and research-oriented evaluation.

20252025.06–2026.02

Research Intern

CDG · Tencent Rhino-Bird Elite Talent Program

Reinforcement learning for recommendation and online bidding; produced three research works, with papers accepted at ICLR 2026 and the KDD 2026 ADS Track, and one under submission.

03 / Recognition

Awards.

2026
  • ICML Silver Reviewer

    Top 25%
2025
  • NeurIPS Top Reviewer

    Top 5%
  • National Scholarship for Doctoral Students

    MoE, China
  • Tencent Rhino-Bird Elite Talent Program

    Selected among 70+
  • Tencent Rhino-Bird Elite Talent Program Outstanding Award

2024
  • Annual Conference on Learning Theory Travel Grant

2023
  • ACM SIGMETRICS / IFIP Performance Travel Grant

  • Outstanding Teaching Assistant

    ShanghaiTech University
2021–2025
  • ShanghaiTech University Scholarship

04 / Community

Service
& teaching.

service

Reviewer / Program Committee

NeurIPS, ICML, ICLR, IROS, IEEE/ACM TON, JSAC

talk

Invited Talk, Online Convex Optimization with Hard Constraints

RLChina 2022

talk

Invited Talk, Safe and Adaptive Online Decision-Making

Tencent Rhino-Bird Elite Talent Forum 2025

talk

Invited Talk, Constrained Contextual Bandits via Lyapunov Optimization

POMS-HK 2025

teaching

Teaching Assistant, Online Optimization and Learning (CS245)

2022–2025

Research / Publication