DOI: 10.28979/jarnas.1949124 ISSN: 2757-5195

A Swarm Optimization-Based Deep Reinforcement Learning Approach for Nonlinear Control Systems with Shapley Value Interpretability

Mustafa Can Bingöl
This study introduces the Proposed Optimization Algorithm (POA), a swarmbasedhybrid combining the Gorilla Troops Optimizer (GTO) and the Artificial Bee Colony(ABC) algorithm, to enhance reward maximization in control tasks. Neural networks weretrained for both simple (pendulum) and complex (bipedal walker) environments. The POAalgorithm was run 10 times, and based on the resulting median values, the best rewardscores achieved were −117.771 in the pendulum environment and 30.936 in the bipedalwalker environment. These reward values indicate a 0.244 improvement for the pendulumenvironment and a 39.569 improvement for the bipedal walker compared to the closestcompetitors (GTO). While there was no statistically significant difference between GTO andPOA in the pendulum task, POA performed significantly better than all other algorithmsin the bipedal walker environment (p < 0.05). To address the “black box” nature ofreinforcement learning, the study integrated Shapley Value Theory for post-training analysis.This explainable AI (XAI) approach identified angular velocity as the primary driver oftorque in the pendulum task and quantified the importance of observation parameters forthe bipedal walker. The results provide both a high-performing optimization framework anda robust method for interpreting neural network decision-making in robotic control systems.