Diabetic retinopathy (DR) is a major cause of vision loss. Accurate grading of DR is critical to ensure timely and appropriate intervention. DR progression is primarily characterized by the presence of biomarkers including microaneurysms, hemorrhages, and exudates. These markers are small, scattered, and challenging to detect. To improve DR grading accuracy, we propose FF-ResNet-DR, a deep learning model that leverages frequency domain attention. Traditional attention mechanisms excel at capturing spatial-domain features but neglect valuable frequency domain information. Our model incorporates frequency channel attention modules (FCAM) and frequency spatial attention modules (FSAM). FCAM refines feature representation by fusing frequency and channel information. FSAM enhances the model's sensitivity to fine-grained texture details. Extensive experiments on multiple public datasets demonstrate the superior performance of FF-ResNet-DR compared to state-of-the-art models. It achieves an AUC of 98.1% on the Messidor binary classification task and a joint accuracy of 64.1% on the IDRiD grading task. These results highlight the potential of FF-ResNet-DR as a valuable tool for the clinical diagnosis and management of DR.
- Article type
- Year
- Co-author
Open Access
Research Article
Issue
Open Access
Research Article
Issue
Glaucoma, a leading cause of irreversible blindness, requires early detection to prevent progressive vision loss. Color fundus photography is a non-invasive and widely accessible modality for glaucoma screening; however, traditional manual interpretation is limited by subjectivity, time inefficiency, and inter-observer variability. This study proposes an optic disc (OD)/optic cup (OC)semantic feature pyramid network, a joint OD and OC segmentation model for glaucoma screening. The model extends the Semantic FPN architecture through three key enhancements: (1) a MaxViT backbone that incorporates multi-axis attention to reinforce local-global feature interaction and preserve boundary information during downsampling; (2) inception depthwise convolution modules embedded within MBConv blocks, which enables multi-scale convolution to expand receptive fields without compromising fine-grained details; (3) an optimized semantic FPN structure to improve the cross-scale feature alignment and multi-scale fusion. The proposed OD/OC-Semantic FPN was evaluated on five publicly available fundus image datasets (Drishti-GS, ORIGA, RIM-ONE DL, RIM-ONE-R3, and REFUGE), and its performance was compared against several state-of-the-art segmentation models (U-Net, DeepLabV3+, PSPNet, APCNet, semantic FPN-PoolFormer, and attention U-Net). The results show that the OD/OC-semantic FPN surpasses existing models across several metrics: dice coefficient, Mean Intersection over Union (mIoU), mean pixal accuracy (MPA), and classification accuracy, thus demonstrating superior structural precision for fundus analysis. Collectively, these results indicate that the OD/OC-Semantic FPN is a robust and generalizable tool for intelligent early detection of glaucoma.
京公网安备11010802044758号