Blockage-Aware Power Allocation Algorithm for Millimeter-Wave Communication with Dynamic Reward Q-Learning
Zhuoning Yang, Ziwei ChenMillimeter-wave (mmWave) communication systems are vulnerable to severe attenuation, blockage-induced LOS/NLOS transitions, and time-varying co-channel interference. This paper develops a lightweight distributed power-allocation framework in which each base station independently updates a tabular Q-learning policy using locally observable blockage-ratio, serving-distance, and aggregate-interference information. The proposed state-dependent dynamic reward is recalculated at every decision step, and its coefficients vary explicitly with the instantaneous blockage ratio, QoS satisfaction ratio, and normalized interference level. All learning-based and non-learning baselines are evaluated using the same topology realizations, blockage and mobility traces, and random seeds. Under the reconstructed simulation settings, the proposed method achieves performance comparable to fixed Q-learning while retaining a transparent blockage-aware state and low-complexity distributed implementation. DQN, greedy, and uniform power achieve higher raw capacity or QoS in the considered small-scale network. Results from 30 paired runs with 95% confidence intervals, together with ablation, sensitivity, beam-misalignment, and overhead analyses, clarify the empirical benefits, limitations, and deployment scope of the proposed method. The study focuses on power control after beam establishment; joint beam tracking and power allocation remain outside the present scope.