DOI: 10.3390/s26165134 ISSN: 1424-8220

Lightweight SNR-Adaptive Receiver-Side Enhancement for DeepJSCC-Based Wireless Image Transmission

Shouquan Hou, Peng Zhao, Nuo Chen

Deep joint source-channel coding (DeepJSCC) has emerged as a promising paradigm for semantic-aware wireless image transmission, achieving strong performance under challenging channel conditions. However, MSE-trained DeepJSCC systems typically achieve high peak signal-to-noise ratio (PSNR) values but suppress high-frequency details, resulting in perceptually blurry reconstructions that fail to capture fine textures and edge information. Existing perceptual enhancement approaches for JSCC systems face significant practical limitations: full transceiver redesign methods require replacing both the transmitter and the receiver with large models (19–31 million parameters), incurring substantial deployment costs; diffusion-based refinement approaches require over 1700 million additional parameters and introduce inference latency exceeding 13 s, rendering them unsuitable for latency-constrained wireless applications; and generic image restoration networks lack channel state awareness and cannot adapt to varying signal-to-noise ratio (SNR) conditions. This paper proposes a lightweight receiver-only perceptual enhancer designed for use with frozen DeepJSCC backbones. The proposed module adopts residual learning with feature-wise linear modulation (FiLM)-based SNR-adaptive modulation to dynamically adjust the enhancement strength under varying channel conditions. A radially weighted FFT magnitude loss is further introduced to guide high-frequency recovery. The enhancer adds only 0.29 million trainable parameters (<1% of the backbone) and requires neither transmitter modification nor backbone retraining. Extensive experiments on the Kodak24 and DIV2K datasets demonstrate a 34.4–37.5% LPIPS reduction over the frozen DeepJSCC baseline under AWGN channels. Supplementary robustness evaluations further show a 30–33% LPIPS reduction under Rayleigh fading, and stable generalization to unseen SNR levels. The receiver-side decoder-plus-enhancer pipeline requires 43 ms at 768 × 512 resolution, corresponding to approximately 23 frames per second.

More from our Archive