Foundation-Model-Assisted Reward Design for Reinforcement Learning: A Review of Reward Program Synthesis, Multimodal Feedback, and Trustworthiness
Wei Zhu, Jinyin Bai, Rui Tang, Zehao Pang, Mingxi Wang, Chengjie Lu, Tianjin Ni, Hang Liu, Xiangchen Wang, Jinji Zhou, Yanlin Wu, Yongjun Peng, Zongzhe Nie, Shiluo Guo, Qinglin Xu, Kaiyang Kou, Yihao ZhongReward functions determine what reinforcement learning agents ultimately optimize, yet reward design for complex tasks has traditionally relied on extensive domain expertise and iterative engineering. Recent large language models and vision–language foundation models have introduced new mechanisms for interpreting task intent, synthesizing reward programs, evaluating states and trajectories, and refining rewards through policy feedback. This review organizes the emerging literature along three complementary directions: reward program synthesis, multimodal feedback, and feedback-driven reward optimization. We further propose a five-level trustworthiness framework spanning format validity, execution validity, semantic validity, behavioral validity, and structural assurance. Existing evidence shows that foundation models substantially broaden how rewards can be represented and acquired but do not eliminate grounding errors, proxy misalignment, reward hacking, selection bias, or reward-search costs. We therefore examine the field from an end-to-end perspective that jointly considers policy performance, reward fidelity, trustworthiness, computational and human cost, and transfer. Finally, we identify verifiable reward representations, process reward models, budget-aware reward search, and transferable reward knowledge across tasks and multi-agent systems as key directions for future research.