HiRNet: A Hierarchical-Refinement-Guided Hybrid Retinal Vessel Segmentation Network for Smartphone-Based Fundus Images
Yuqing Yin, Liang Niu, Suyan Li, Yifan Mu, Xiao Xu, Yan WangSmartphone-based fundus photography provides a low-cost and convenient approach to retinal screening and follow-up examinations. As a fundamental step in fundus image analysis, retinal vessel segmentation plays an important role in the diagnosis and monitoring of retinal diseases. However, existing methods often exhibit limited performance on smartphone-based fundus images. This limitation mainly arises from the significantly lower image quality that affects the clarity and detail of the retinal vessels, making accurate segmentation more difficult. To address these challenges, we propose a hierarchical-refinement-guided hybrid network (HiRNet), which incorporates a hybrid CNN–Transformer feature extractor to jointly model local vascular structures and global contextual information. Specifically, consecutive asymmetric dilated convolutions capture multi-scale local features, while multi-path dilated attention enhances contextual representations across different receptive fields. In addition, a hierarchical context information transmission module is introduced to progressively integrate features from different resolutions during multistage upsampling, thereby improving vessel continuity and boundary delineation. Experiments on RVD, DRIVE, and CHASE_DB1 demonstrate the effectiveness of HiRNet. On the smartphone-based RVD dataset, HiRNet obtains the highest mean sensitivity (58.36%), specificity (97.81%), accuracy (95.83%), F1-score (58.39%), and AUC (95.16%). The improvements in sensitivity, accuracy, F1, and AUC over the strongest competing methods are statistically significant. HiRNet also obtains sensitivity values of 82.73% and 82.13%, accuracy values of 95.59% and 95.98%, F1-scores of 82.76% and 80.37%, and AUC values of 98.83% and 97.90% on DRIVE and CHASE_DB1, respectively. These results indicate that HiRNet provides a relative improvement on the challenging domain of smartphone-based fundus images while maintaining strong performance on conventional fundus datasets.