About Me

I am a Ph.D. student (2025 cohort) at the Harbin Institute of Technology, Shenzhen, supervised by Prof. Li Jing. I graduated with an Honors Bachelor’s Degree from the YingCai Honors College of the University of Electronic Science and Technology of China (UESTC), majoring in Computer Science and Technology.我是哈尔滨工业大学(深圳)2025级博士研究生,师从李晶教授。本科毕业于电子科技大学英才实验学院(UESTC)计算机科学与技术专业,获荣誉工学学士学位。

My research interests include LLM Alignment, LLM Safety, and Reinforcement Learning. I have published 5 papers in international conferences, including ACL, KDD, and DASFAA.我的研究方向包括大模型对齐(LLM Alignment)、大模型安全(LLM Safety)与强化学习(Reinforcement Learning)。目前已在 ACL、KDD、DASFAA 等国际会议上发表 5 篇论文。

📖 Educations📖 教育经历

  • 2021.09 - 2025.06, B.Eng. in Computer Science and Technology, YingCai Honors College, University of Electronic Science and Technology of China (UESTC).电子科技大学(UESTC)英才实验学院,计算机科学与技术专业,工学学士。
  • 2025.09 - Present, Ph.D. student at Harbin Institute of Technology, Shenzhen, supervised by Prof. Li Jing.2025.09 - 至今,哈尔滨工业大学(深圳)博士研究生,导师:李晶教授。

🔥 News🔥 近期动态

  • 2026.07:  🎉🎉 Our paper “Agent Reinforcement Learning via Pivotal-Aware Self-Feedback Retry” is now available on arXiv, with source code open-sourced on GitHub!🎉🎉 我们的论文 “Agent Reinforcement Learning via Pivotal-Aware Self-Feedback Retry” 已发布在 arXiv,源码已在 GitHub 开源!
  • 2026.05:  🎉🎉 Two papers were accepted by ACL 2026! “Backdoors in RLVR: Jailbreak Backdoors in LLMs From Verifiable Reward” (Main) and “E3-TIR: Enhanced Experience Exploitation for Tool-Integrated Reasoning” (Findings).🎉🎉 两篇论文被 ACL 2026 接收:“Backdoors in RLVR: Jailbreak Backdoors in LLMs From Verifiable Reward”(Main)与 “E3-TIR: Enhanced Experience Exploitation for Tool-Integrated Reasoning”(Findings)。
  • 2025.05:  🎉🎉 Our paper “MTSA: Multi-turn Safety Alignment for LLMs through Multi-round Red-teaming” was accepted by ACL 2025!🎉🎉 论文 “MTSA: Multi-turn Safety Alignment for LLMs through Multi-round Red-teaming” 被 ACL 2025 接收!
  • 2025.05:  🎉🎉 Our paper “LLMs are Noisy Oracles! LLM-based Noise-aware Graph Active Learning for Node Classification” was accepted by KDD 2025!🎉🎉 论文 “LLMs are Noisy Oracles! LLM-based Noise-aware Graph Active Learning for Node Classification” 被 KDD 2025 接收!
  • 2023.12:  🎉🎉 Our paper “IntFair: Graph Neural Networks for Fair Recommendations with Interest Awareness” was accepted by DASFAA 2024!🎉🎉 论文 “IntFair: Graph Neural Networks for Fair Recommendations with Interest Awareness” 被 DASFAA 2024 接收!

📝 Publications📝 论文发表

arXiv
PivoARL

Agent Reinforcement Learning via Pivotal-Aware Self-Feedback Retry

arXiv | GitHub

|Agent Reinforcement Learning|Experience Exploitation

ACL 2026
Backdoors in RLVR

Backdoors in RLVR: Jailbreak Backdoors in LLMs From Verifiable Reward

arXiv | GitHub

|LLM Safety|Backdoor Attack|RLVR

ACL 2026 Findings
E3-TIR

E3-TIR: Enhanced Experience Exploitation for Tool-Integrated Reasoning

arXiv | GitHub

|Agent Reinforcement Learning|Tool-Integrated Reasoning

ACL 2025
MTSA

MTSA: Multi-turn Safety Alignment for LLMs through Multi-round Red-teaming

arXiv | GitHub | PDF

|LLM Safety|Multi-turn Alignment

KDD 2025
Noisy Oracles

LLMs are Noisy Oracles! LLM-based Noise-aware Graph Active Learning for Node Classification

PDF | ACM

|Graph Active Learning|LLM

DASFAA 2024
IntFair

IntFair: Graph Neural Networks for Fair Recommendations with Interest Awareness

PDF | DBLP

|Fairness|Graph Neural Networks|Recommendation

🎖 Honors and Awards🎖 荣誉奖项

  • 2024, Add your honors and awards here.
  • 2023, Add your honors and awards here.