About Me
I am currently a 2nd-year Ph.D. student at the Gaoling School of Artificial Intelligence, Renmin University of China, fortunate to be co-advised by Prof. Zhicheng Dou and Prof. Jirong Wen. I earned my M.Eng (2024) and B.Eng (2021) degrees in Information and Communication Engineering from Beijing University of Posts and Telecommunications (BUPT), advised by Prof. Weiran Xu.
I’m currently a Top Seed research intern focusing on general agent research at Bytedance Seed. Previously, I held research intern positions at the Alibaba Qwen Team, Kuaishou Klear Team, and Meituan NLP Center. I have published 50+ papers in top-tier AI conferences and journals (10+ first-author papers), including NeurIPS, ICLR, ACL, WWW, EMNLP, NAACL, AAAI, and IP&M.
Research Interests:
- General Agent Training — Training long-horizon agents with scalable real-world interaction capabilities
- Agent Harness Engineering — Building stronger scaffolds to fully unlock frontier agent capabilities of foundation models
- Agentic Reinforcement Learning — Training general agent intelligence via fundamental RL-based optimization
My long-term goal is to develop automated, scalable, and safe approaches that foster exceptional intelligence toward achieving AGI. I am also a firm believer in The Bitter Lesson.
🔥 News
- 2026.07: 📚 Released Towards Long-Horizon Agents: A Survey! Check out our curated reading list on GitHub.
- 2026.06: 🚀 Released Seed2.1 Model Card — a next-generation agent for real-world productivity! Honored to be a core contributor. (Homepage)
- 2026.06: 🏆 Selected as 青源InnoVibe 2026——最受瞩目学术新星 at BAAI Conference 2026!
- 2026.06: 🎉 Honored to receive the “瓴航”院长奖学金!
- 2026.04: 🚀 Released Agent-World, scaling real-world environment synthesis for evolving general agent intelligence! Check out our demo!
- 2026.03: 🎉 Tool-Star, ARPO, and AEPO accepted at SIGIR 2026, ICLR 2026, and WWW 2026! Welcome to follow our Agent RL family.
- 2026.02: 🎉 Honored to receive the 2026 Tencent Project Up Scholarship (首届腾讯青云奖学金, 全国15人). Thanks to my teachers and co-authors for their support!
- 2026.02: 🚀 We released Seed2.0. As a core contributor, I am responsible for the core MCP tool-use agent capability. (MCPMark 54.7, BFCL-V4 73.4, VitaBench 47).
- 2026.02: 🚀 Released OmniGAIA, building native omni-modal AI agents!
- 2025.12: 🚀 Released Seed1.8, towards generalized real-world agency! Honored to be a core contributor.
- 2025.10: 🚀 Released AEPO, our entropy-balanced policy optimization method for multi-turn LLM agents.
- 2025.09: 🌐 WebThinker accepted by NeurIPS 2025! A powerful open-source deep research agent. Check out our demo!
- 2025.08: ARPO featured as 🤗 HF Weekly Paper #1! An agentic RL algorithm for multi-turn LLM agents.
- 2025.08: 🔍 Search-o1 accepted by EMNLP 2025 as Oral Presentation!
- 2025.05: Four papers accepted by ACL 2025!
- 2025.05: Released Tool-Star, an LLM-brained multi-tool reasoner via RL! Check out our project.
- 2025.02: ⚡ FlashRAG now supports multimodal retrievers and generators!
- 2025.01: DPA-RAG accepted by WWW 2025 — aligning diverse preferences in RAG systems.
- 2025.01: Two papers accepted by ICLR 2025! AutoIF is the secret behind
Qwen's instruction-following alignment.
📖 Education
-
2024.09 - Present |
Ph.D. in Artificial Intelligence
Gaoling School of Artificial Intelligence, Renmin University of China -
2021.09 - 2024.06 |
M.Eng in Artificial Intelligence
Beijing University of Posts and Telecommunications -
2017.09 - 2021.06 |
B.Eng in Information and Communication Engineering
Beijing University of Posts and Telecommunications -
2018.07 - 2018.08 |
Summer Exchange Program
University of Oxford
💻 Research Experience
-
2025.11 - Present |
ByteDance, Seed General Agent Team
- Research Intern on RL for General Agent (Top Seed Program)
- Mentors: Wanjun Zhong, Yujia Qin -
2025.04 - 2025.11 |
Kuaishou, Foundation LLM Team
- Research Intern on Agentic RL & Deep Search Agent (K-Star Program)
- Mentors: Hangyu Mao, Fuzheng Zhang -
2023.06 - 2024.08 |
Alibaba,
Qwen Foundation LLM Team
- Research Intern on Alignment & Reasoning of Large Language Models
- Mentors: Bowen Yu, Zheng Yuan, Wei Wang, Keming Lu -
2022.09 - 2023.05 |
Meituan, NLP Center
- Research Intern on Knowledge-Augmented Generation
- Mentor: Rumei Li
📝 Selected Preprints(Full List)
* for corresponding author, # for equal contribution. Recent representative preprints.
-
Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence

HomepageHF Daily Paper #3
-
InstructERC: Reforming Emotion Recognition in Conversation with a Retrieval Multi-task LLMs Framework
Wechat Blog
📄 Technical Reports
-
Seed2.1 Model Card: A Next-Generation Agent for Real-World Productivity

Homepage -
Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity

Homepage -
Seed1.8 Model Card: Towards Generalized Real-World Agency

HomepageWechat Blog
-
QwQ: Reflect Deeply on the Boundaries of the Unknown

HomepageWechat Blog
-
Qwen2.5 Technical Report

HomepageHF Daily Paper #1 Wechat Blog
-
Qwen2 Technical Report

HomepageHF Daily Paper #1 Wechat Blog
-
Scaling Relationship on Learning Mathematical Reasoning with Large Language Models

HF Daily Paper #3
📝 Selected Publications(Full List)
Selected peer-reviewed publications.
2026
-
ET-Agent: Incentivizing Effective Tool-Integrated Reasoning Agent via Behavior Calibration
ACL 2026 (CCF-A)
-
ToolScope: An Agentic Framework for Vision-Guided and Long-Horizon Tool Use
ACL 2026 Findings
-
EnvScaler: Scaling Tool-Interactive Environments for LLM Agent via Programmatic Synthesis
ACL 2026 Findings
Wechat Blog
-
Tool-Star: Empowering Multi-Tool Collaborative Web Agent via Reinforcement Learning
SIGIR 2026 (CCF-A)
HF Daily Paper #3 Wechat Blog
-
SmartSearch: Process Reward-Guided Query Refinement for Search Agents
SIGIR 2026 (CCF-A)
-
Agentic Reinforced Policy Optimization
ICLR 2026 (CCF-A)
HF Daily Paper #1
HF Weekly Paper #1 Wechat Blog
-
Toward Effective Tool-Integrated Reasoning via Self-Evolved Preference Learning
ICLR 2026 (CCF-A)
-
Agentic Entropy-Balanced Policy Optimization
WWW 2026 (CCF-A) (Oral)
HF Daily Paper #3 Wechat Blog
-
DeepAgent: A General Reasoning Agent with Scalable Toolsets
WWW 2026 (CCF-A) (Oral)
HF Daily Paper #1 Wechat Blog
2025
-
WebThinker: Empowering Large Reasoning Models with Deep Research Capability
NeurIPS 2025 (CCF-A)
HomepageHF Daily Paper #1 Wechat Blog
-
Search-o1: Agentic Search-Enhanced Large Reasoning Models
EMNLP 2025 (CCF-B) (Oral) (Most Influential arXiv AI Papers in 2025 – Top 5/All)
HomepageHF Daily Paper #2 Wechat Blog
-
RAG-Critic: Leveraging Automated Critic-Guided Agentic Workflow for Retrieval Augmented Generation
ACL 2025 (CCF-A)
Wechat Blog
-
Progressive Multimodal Reasoning via Active Retrieval
ACL 2025 (CCF-A)
HF Daily Paper #2 Wechat Blog
-
We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?
ACL 2025 (CCF-A) (Most Influential ACL 2025 Papers – Top 6/All)
HomepageHF Daily Paper #1 Wechat Blog
-
Understand What LLM Needs: Dual Preference Alignment for Retrieval-Augmented Generation
WWW 2025 (CCF-A) (Most Influential WWW Papers in 2025 – Top 10/All)
Wechat Blog
-
Toward General Instruction-Following Alignment for Retrieval-Augmented Generation
AAAI 2025 (CCF-A)
Homepage -
Self-play with Execution Feedback: Improving Instruction-following Capabilities of Large Language Models
ICLR 2025 Spotlight (CCF-A)
-
CS-Bench: A Comprehensive Benchmark for Large Language Models towards Computer Science Mastery
ICLR 2025 (CCF-A)
Homepage -
FlashRAG: A Python Toolkit for Efficient RAG Research ⚡
WWW 2025 (CCF-A)
Wechat Blog
2024
-
How Abilities in Large Language Models are Affected by Supervised Fine-tuning Data Composition
ACL 2024 (CCF-A) (Most Influential ACL 2024 Papers – Top 11/All)
-
MuggleMath: Assessing the Impact of Query and Response Augmentation on Math Reasoning
ACL 2024 (CCF-A)
-
ChatKBQA: A Generate-then-Retrieve Framework for Knowledge Base Question Answering with Fine-tuned Large Language Models
ACL 2024 Findings
Wechat Blog
-
DolphCoder: Echo-Locating Code Large Language Models with Diverse and Multi-Objective Instruction Tuning
ACL 2024 (CCF-A)
2023
-
A Multi-Task Semantic Decomposition Framework with Task-specific Pre-training for Few-Shot NER
CIKM 2023 (CCF-B)
-
Bridging the KB-Text Gap: Leveraging Structured Knowledge-aware Pre-training for KBQA
CIKM 2023 (CCF-B)
🏆 Honors & Awards
Grants
- 2026.01-2027.12: Fundamental Research Project for PhD Students, NSFC (国家自然科学基金青年学生基础研究项目-博士生)
- 2026: Young Talent Support Program for Doctoral Students, CAST (中国科协青年人才托举工程博士生专项计划)
- 2026: Top-Tier Innovative Talent Cultivation Support Program of RUC (中国人民大学拔尖创新人才培育资助计划)
Scholarships
- 2026: 青源InnoVibe 2026——最受瞩目学术新星 (BAAI Conference 2026)
- 2026: “瓴航”院长奖学金
- 2026: Tencent Project Up Scholarship (首届腾讯青云奖学金, 全国15人)
- 2025: National Scholarship for Ph.D. Students (博士生国家奖学金, Top 1%)
- 2024: Outstanding Graduates of Beijing (北京市优秀毕业生, Top 1%), Link
- 2024: 1st Place in PhD Entrance Exam (Preliminary), GSAI, Renmin University of China, Link
- 2023: National Scholarship for Master Students (硕士生国家奖学金, Top 1%), Link
- 2021-2022: Excellent First-class Scholarship for Master Students, BUPT
Competitions
🎤 Invited Talks
- 2026.06: “探索通用智能体训练的可行路径”, 青稞 Talk 129, Link / Bilibili / YouTube
- 2026.06: 《迈向通用智能体训练》, BAAI Conference 2026, Link
- 2026.05: “Agent-World,让智能体与环境协同进化”, 中国AI智能体大会, Link
- 2026.05: “Agent-World,让智能体与环境协同进化”, NICE Talk NO.175, Link
- 2025.11: “Agentic Reinforcement Policy Optimization”, MLNLP Community, Slides
- 2025.10: “Agentic Reinforcement Policy Optimization”, EvoAgentX Community, Slides
- 2025.09: “Agentic Reinforcement Policy Optimization”, HunYuan Team, Tencent
- 2025.08: “ARPO: Encouraging Agents to Explore at Critical Moments”, NICE Community, Slides
🔍 Academic Services
Journal Reviewer
- Knowledge-Based Systems (KBS)
Senior Program Committee (SPC)
- AAAI: 2026
Program Committee Member / Reviewer
- Top-tier ML Conferences: NeurIPS (2024–2025), ICML (2025), ICLR (2023–2026)
- Top-tier AI/DM Conferences: KDD (2025), SIGIR (2025), WWW (2025–2026), CIKM (2024–2025), AAAI (2026)
- Top-tier NLP Conferences: ACL ARR (2024–2026)