DOI: 10.3390/electronics15184289 ISSN: 2079-9292

Stable Offline Reinforcement Learning for Switched Reluctance Motor Drives via Multi-Demonstrator Policy Distillation

Franklin Sánchez, María Isabel Milanés-Montero, Enrique Romero-Cadaval

Finite-control-set model predictive control provides excellent torque–speed regulation for switched reluctance motor drives but requires an online combinatorial search at every control instant, making low-cost embedded implementation challenging. This article investigates whether offline reinforcement learning can distill policies from multiple classical controllers into a single feedforward policy requiring neither online optimization nor controller gain tuning. A replay buffer is populated with trajectories generated by three demonstrators—hysteresis current control, proportional–integral control with pulse-width modulation, and finite-control-set model predictive control—using a finite-element model of a four-phase 8/6 switched reluctance machine parameterized from measurements of the physical drive. An implicit Q-learning agent then learns a control policy without evaluating actions outside the offline dataset. The central finding is that demonstration diversity governs the stability of offline reinforcement learning on this problem: policies trained from a single demonstrator experience early mode collapse in all fifteen runs, whereas two- or three-demonstrator datasets converge stably in all fifteen. Behavior cloning trained on the identical buffer, split, architecture, and deployed controller provides the reference point for interpreting this result. It matches the offline RL policy on torque quality and improves on its speed regulation, exhibiting none of the seed-to-seed fragility seen at no load while requiring roughly 8% more switching transitions. The stability requirement therefore appears to be a property of the advantage-weighted offline RL objective rather than the control task, and the measured benefit of that objective on this problem is confined to switching effort. We report this rather than claim a broader advantage. The characterization of the distilled controller shows that it generalizes to operating points that are not included in the training dataset, gains nothing systematic beyond approximately 60% of the replay buffer, remains insensitive to ±20% perturbations of all reward weights, and degrades gracefully under measurement noise while the current mask enforces the peak-current constraint throughout. A deployment analysis shows that the 18,432 multiply–accumulate policy meets a 50μs control period in its existing form at a measured cost of about 2% in torque ripple. All the results are simulation-based on a finite-element model parameterized from a physical machine.