Reinforcement Learning for Diffusion Policies in Robotics: A Survey and State-Based Locomotion Reproduction
Shihan Sun, Yinlong LiuDiffusion policies model multimodal robot action sequences, but behavioral cloning does not directly optimize task return. We present a structured scoping review of reinforcement learning for generative robot policies and a bounded state-based locomotion reproduction. Four documented routes yielded 178 records, 162 unique candidates, and an 84-study evidence map. Hierarchical rules distinguish 41 direct reward-driven studies from 32 adjacent robotic, eight alternative-generator, and three non-robotic studies; a five-axis taxonomy codes initialization/data, interaction regime, optimized object, credit assignment, and generator. Under a fixed-final evaluation protocol on the Datasets for Deep Data-Driven Reinforcement Learning (D4RL) 1.1 Hopper benchmark, five diffusion policy policy optimization (DPPO) fine-tuning seeds improved over their run-recorded behavior-cloning initializations by a mean of 1261.2 return, with a seed-level standard deviation of 125.5 and a 95% confidence interval of 1105.3–1417.1; the five runs link to two recorded behavior-cloning checkpoints. A Gaussian-policy control also improved after proximal policy optimization, so the gain was not diffusion-specific. A full-chain backpropagation adaptation exhibited clear seed-dependent variation, a matched action-divergence intervention did not establish causal critical timesteps, and reducing denoiser evaluations from 20 to 2 lowered A100 latency from 30.97 to 3.85 ms while substantially reducing normalized score. The experiments are limited to state-based locomotion and do not validate visual manipulation.