Quantum Federated Reinforcement Learning‐Based Traffic Offloading and Resource Allocation for RSMA‐Enabled Space–Air–Ground Integrated Networks
Ishan Budhiraja, Abhay Bansal, Bhuvan Unhelkar, Niyaz Ahmad Wani, Muhammad Attique KhanABSTRACT
The evolution of sixth‐generation (6G) wireless networks demands ultra‐reliable low‐latency communication (URLLC), massive connectivity, and high‐capacity data transmission in highly dynamic environments. Space–Air–Ground Integrated Networks (SAGINs) have emerged as a promising architecture by seamlessly integrating satellites, unmanned aerial vehicles (UAVs), and terrestrial infrastructure to provide ubiquitous connectivity. However, stochastic traffic arrivals, UAV mobility, time‐varying wireless channels, and the coexistence of enhanced Mobile Broadband (eMBB) and URLLC services make traffic offloading and resource allocation highly challenging. These factors transform the optimization task into a stochastic mixed‐integer nonlinear programming (MINLP) problem. To address this challenge, this paper proposes a Quantum Federated Reinforcement Learning (QFRL)‐based traffic offloading framework for RSMA‐enabled SAGINs. The optimization problem is formulated as a constrained Markov decision process (CMDP), allowing distributed small cells to jointly optimize traffic offloading ratios, bandwidth allocation, RSMA power distribution, and UAV trajectory planning while satisfying stringent delay and reliability requirements. A variational quantum circuit (VQC)‐based actor‐critic architecture is developed to improve learning efficiency and policy representation in high‐dimensional continuous action spaces. In addition, a federated aggregation mechanism enables privacy‐preserving distributed learning and scalable coordination across the space, air, and ground segments. The proposed framework employs temporal‐difference learning and parameter‐shift gradient optimization to ensure stable convergence under stochastic network dynamics. Simulation results demonstrated that the proposed QFRL framework reduces traffic dropping probability by 28%–35%, decreases URLLC delay by 22%–30%, improves network availability by 18%–25%, and enhances traffic offloading efficiency by 20%–27% compared with Differentiated Federated Soft Actor‐Critic (DFSAC), Double Q‐Learning delay sensitive replay memory (DSRPM), and Nash Equilibrium Iteration Offloading (NEIO‐G) schemes, respectively.