DOI: 10.1145/3837057 ISSN: 0360-0300

Reinforcement Learning in the Era of Large Language Models: Challenges and Opportunities

Qianyue Hao, Lin Chen, Xiaoqian Qi, Yuan Yuan, Zefang Zong, Hongyi Chen, Keyu Zhao, Shengyuan Wang, Yunke Zhang, Jian Yuan, Yong Li

Reinforcement learning (RL), is becoming essential in the post-training of large language models (LLMs), enhancing their capabilities and alignment with human preferences. However, adapting conventional RL to LLMs introduces challenges stemming from their massive parameter size and the vast natural language action space. In this survey, we conduct a systematic literature review on how RL are adapted and scaled as a fundamental post-training tools. First, we provide a taxonomy of challenges faced in each stage of the RL training loop, including action exploration, trajectory collection, reward evaluation, and model update. We then introduce recently developed methods to address these issues, ranging from logical structure navigation, training data curation to reward design and advantage estimation. Afterward, we elaborate on the application of these techniques in diverse domains such as mathematics, coding, medicine, and information retrieval, analyzing how innovations in the RL pipeline enhance the domain-specific LLMs. Finally, we discuss the limitations and side effects of applying RL to LLMs and explore open problems and future directions like the balance between efficiency and effectiveness, and algorithm-system co-design. This survey helps researchers understand recent progress and inspire novel research to address current challenges and realize the full potential of RL for LLMs.

More from our Archive