
==== Front
Sci Rep
Sci Rep
Scientific Reports
2045-2322
Nature Publishing Group UK London

39266650
72481
10.1038/s41598-024-72481-1
Article
An attentional mechanism model for segmenting multiple lesion regions in the diabetic retina
Xu Changzhuan 920372432@qq.com

He Song
Li Hailin
https://ror.org/046q1bp69 grid.459540.9 0000 0004 1791 4503 Information Branch, Guizhou Provincial People’s Hospital, Guizhou, 550001 China
12 9 2024
12 9 2024
2024
14 213541 4 2024
9 9 2024
© The Author(s) 2024
2024
https://creativecommons.org/licenses/by-nc-nd/4.0/ Open Access This article is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License, which permits any non-commercial use, sharing, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if you modified the licensed material. You do not have permission under this licence to share adapted material derived from this article or parts of it. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by-nc-nd/4.0/.
Diabetic retinopathy (DR), a leading cause of blindness in diabetic patients, necessitates the precise segmentation of lesions for the effective grading of lesions. DR multi-lesion segmentation faces the main concerns as follows. On the one hand, retinal lesions vary in location, shape, and size. On the other hand, the currently available multi-lesion region segmentation models are insufficient in their extraction of minute features and are prone to overlooking microaneurysms. To solve the above problems, we propose a novel deep learning method: the Multi-Scale Spatial Attention Gate (MSAG) mechanism network. The model inputs images of varying scales in order to extract a range of semantic information. Our innovative Spatial Attention Gate merges low-level spatial details with high-level semantic content, assigning hierarchical attention weights for accurate segmentation. The incorporation of the modified spatial attention gate in the inference stage enhances precision by combining prediction scales hierarchically, thereby improving segmentation accuracy without increasing the associated training costs. We conduct the experiments on the public datasets IDRiD and DDR, and the experimental results show that the proposed method achieves better performance than other methods.

Subject terms

Computational biology and bioinformatics
Health care
issue-copyright-statement© Springer Nature Limited 2024
==== Body
pmcIntroduction

Diabetes mellitus (DM) is a chronic condition that is characterized by elevated blood sugar levels. These levels can lead to severe complications affecting multiple organs, including the eyes, kidneys, nerves, and cardiovascular system. DM is a significant public health concern affecting millions of individuals globally, including diverse populations such as pregnant women, infants, and the elderly1. DR is one of the most prevalent complications of DM and is currently the fastest growing cause of blindness. The 2019 World Health Organization (WHO) World Report on Vision indicates that DR affects 418 million individuals globally2. The prevalence of DR increases with the duration of diabetes. Those who have had the disease for more than 10 years have a significantly higher incidence rate of DR than those with shorter disease durations. Approximately 60% of patients who have had the disease for more than 15 years have retinopathy, which is 25 times more likely to cause blindness than in individuals without diabetes. Nevertheless, early funduscopic testing and treatment can reduce the risk of blindness by approximately 90%. Therefore, universal screening for retinopathy is crucial to prevent visual loss. In the early stage of DR, the dilation of capillaries can lead to microaneurysms (MA). The leaking of lipoproteins can lead to hard exudates (EX). The arteriole occlusion can lead to the ischemia of the retinal nerve fiber layer and further leads to the soft exudates (SE). The broken of abnormal blood vessels and microaneurysms leads to hemorrhages (HE). MA, EX, SE, and HE are four typical Non-proliferative Diabetic Retinopathy (NPDR) lesions3,4. Figure 1 illustrates these lesion types. Given the small size of lesions in fundus images and the growing demands on healthcare professionals, early detection presents significant challenges, increasing the risks of misdiagnosis or treatment delays5. Leveraging automated detection methods can substantially improve diagnostic efficiency and accuracy, thereby enhancing patient outcomes in DR management.Figure 1 Four pathological categories in the diabetic retina.

Deep learning-based image segmentation algorithms have demonstrated impressive outcomes in specific target segmentation6. However, applying these algorithms directly to DR segmentation poses unique challenges. First, single and fixed image size is unreasonable. The sizes of lesions are very different from each other. For example, the MA is very small compared with SE and may disappear when the image size is reduced. Assigning the feature extractors of the same scale to different lesions in CNN is obviously flawed. Second, the currently available multi-lesion region segmentation models are insufficient in their extraction of minute features and are prone to overlooking microaneurysms. In 1024×1024 images, the DR lesion area, microaneurysms constitutes merely about 1% of pixels, with significant disparities in lesion sizes-microaneurysms are tiny compared to larger lesions. This variability necessitates a segmentation model adept at processing both fine and coarse details across multiple scales. To explicitly solve the issues mentioned above, in this paper, we propose a framework for DR lesion segmentation. First, to deal with the issue of variant lesion scale, we design multi-scale segmentation. The features of the lesion are extracted from different scales and weighted to fuse the information of various receptive fields. Secondly, the attention mechanism is designed to automatically discern the shape and size of the target while suppressing regions that are not pertinent to segmentation and learning, thereby enabling the identification of salient features that are useful for solving the issue of a small percentage of feature pixels in DR images. To demonstrate the effectiveness of our framework, we conduct our experiments on several datasets. The results obtained from these datasets demonstrate the efficacy of our proposed framework. The research contributes distinctly by:This paper proposes the use of spatial gated attention to focus on the spatial location information of key features, which significantly enhances the ability to detect minute features such as tiny aneurysms. This is achieved by fusing coarse and fine-grained features to enhance context dependency over long distances.

Based on the spatial gated attention, a hierarchical multiscale attention model is proposed to capture multi-sized features and improve the segmentation accuracy of each lesion.

The article is organized as follows: “Related Work” reviews state-of-the-art ML and DL techniques for DR, “Methods” details the framework and methodologies, “Experimental” discusses experimental outcomes and comparative analyses, and the Discussion section outlines future research directions

Related work

Deep Learning has made significant strides in the field of medical imaging, with applications spanning across various modalities such as computed tomography (CT)7, magnetic resonance imaging (MRI)8, and ultrasonography9, deep learning can automate the extraction of feature characteristics and the gradual improvement of arithmetic power. Additionally, deep learning is gradually replacing machine learning, which is widely used in the detection of DR.

Algorithms for detecting DR are broadly divided into two categories: detection of single lesion regions and detection of multiple lesion regions. Researchers have developed a range of machine learning and deep learning techniques for segmenting DR lesions, specifically targeting various lesion types, for the purpose of single lesion region detection101112131415. In the 2018 IBSI-sponsored DR-Segmentation and Grading Challenge, the top-performing teams predominantly employed the DeepLab model for effective microaneurysm and hemorrhage detection16. Aziz17 proposed automatic deep learning-based hemorrhage detection method and indicate that increasing deep network layers does not guarantee good results but rather increases training time. While previous studies have focused on segmenting single lesion features, recent research has shifted towards the simultaneous detection of multiple lesion regions in DR using deep learning approaches. In 2017, Tan et al18 applied image normalization and contrast enhancement before employing Deep Convolutional Neural Networks (DCNN) to segment MA, HE, EX, and SE lesions. Their method utilized a Leaky ReLU activation function in intermediate layers and a Softmax function in the final layer for classification. Badar M et al19 introduced a fully convolutional network based on an encoder-decoder architecture for the concurrent segmentation of exudates, hemorrhages, and cotton spots, achieving high accuracy rates across different lesion types on the Messidor dataset: 99.24% for exudates, 97.86% for hemorrhages, and 88.65% for cotton spots. During the 2018 ISBI Diabetic Retinopathy Segmentation and Classification Challenge, the Segmentation Challenge winner enhanced the U-Net model by integrating the DeepLab concept, replacing U-Net’s max pooling layers with 3×3 stride-1 convolutions for denser feature extraction, thus achieving effective semantic segmentation. The second and third place teams also adopted the DeepLab concept, combining it with an attention mechanism to accurately localize and segment lesions. The Google DeepMind team20set a benchmark by segmenting 15 different lesion regions in retinal OCT images with a 3D U-Net network, demonstrating a leading position in the field.

Attention Machine (AM) was initially employed for machine translation purposes, and has since been utilized in a multitude of applications within the domains of natural language processing, statistical learning, speech and computers. Currently, it is a pivotal component of medical imaging research21–24. The human eye is highly adept at identifying key elements within an image, directing attention to specific areas while filtering out extraneous information. In order to achieve a similar effect in computers, developers have created AM that weigh input data in order to focus on specific elements. The data is assigned varying weights at different points, influencing the emphasis placed on it during processing. AM are classified into three categories: channel attention, spatial attention, and hybrid spatial and channel attention models. Khaparde A25 introduced an attention-based Swin U-Net, enabling precise visualization of diseased portions. Kothadiya26 developed a hybrid model leveraging DenseNet121 for convolutional learning, augmented with channel and spatial attention mechanisms for early diabetes detection. Sinha et al.27 addressed the issue of context dependency over long distances by introducing a multiscale self-directed attention mechanism, employing a dual-channel attention model (Zhang et al.28) as its core, achieving a Dice coefficient of 0.8675. In DR fundus images, Xue29 utilized ASPP for retinal vessel segmentation, while another study30 introduced a gated attention (AG) mechanism to highlight relevant features, applying these attentional methods in fused features and jump connections. Liang-Chen et al.31 highlighted the importance of integrating multiscale features into CNNs for semantic segmentation accuracy, proposing an attention mechanism to appropriately weigh multiscale features at each pixel, utilizing DeepLab with added attention and supervision for multiscale image training. Tao et al.32 critiqued the efficiency of multiscale feature fusion in hop-connected networks, favoring the shared network the High-Resolution Network (HRNet), and proposed MS-OCRNet for enhanced segmentation.

Literature review reveals single lesion region segmentation models yield superior results due to their targeted nature. Nevertheless, the segmentation of multi-focal regions remains a challenge, with ongoing research into multi-task segmentation models aimed at effectively addressing multiple complex focal areas. Current models, including MS-OCRNet, show limitations in accurately segmenting multi-focal regions in DR. This paper introduces a novel approach, applying a multi-scale attention mechanism to DR segmentation. This method not only addresses the challenges of insufficient spatial location information for key features but also bridges the gap in granularity between segmented objects in fundus images, optimizing MS-OCRNet for enhanced performance. This proposal represents a significant stride forward in the application of attention mechanisms within medical imaging.

Methods

Figure 2 MSAG network structure.

Network structure

The proposed MSAG mechanism network is designed to enhance segmentation in DR images. Figure 2 illustrates the MSAG architecture consists of three primary components:BackBone (HRNet-OCR): acts as the foundational layer for processing images and extracting features, essential for identifying detailed retinal characteristics.

Spatial attention gated mechanism: integrates semantic feature mapping with a pooling operation to merge spatial information features. Activation through a sigmoid function produces spatial attention maps, focusing the model on relevant lesion areas.

MSAG inference structure: employs a hierarchical computation method allowing for selective use of prediction scales. This flexibility improves lesion localization by optionally incorporating higher scale predictions.

HRNet developed by CUHK and Microsoft Research Asia and introduced at CVPR 2019, is notable for its unique fusion of low- and high-resolution features in a parallel manner, contrasting with the serial fusion approach common in other networks. This method allows for continuous integration of features across scales without the need to reconstruct high resolution from low-resolution inputs, facilitating position-sensitive feature fusion. In segmentation tasks, HRNet achieves faster inference speeds than both PSPNet and DeepLabv3, due to its efficient feature integration process.

In this work, HRNet is utilized as the core for semantic information extraction, combined with the Object Contextual Representation (OCR) for context processing, forming the backbone of the MSAG. The MSAG also incorporates U-Net-like skip connections to preserve context over long distances and introduces the Spatial Gated Attention (SAG) module. This module aims to refine noise suppression and spatial feature extraction within HRNet, enhancing semantic accuracy.

The MSAG model processes inputs at two scales: the original image size (Scale2) and a downsampled version (Scale1), achieved through 2-fold reduction. The Semantic Head (Seg), following the OCR module, handles lesion segmentation, while the Attention Head (Head) focuses on generating attention maps. These maps undergo Sigmoid activation and are then bilinearly upsampled to the original image size for precise segmentation. The introduction of the SAG unit, a key innovation of this study, is detailed further in the next section, emphasizing its role in optimizing attention within the segmentation process.

Spatial attention gate

In contrast to traditional spatial attention modules33 and attention gates30, the SAG uniquely integrates feature mappings directly from the base network, utilizing both MaxPooling (MP) and AveragePooling (AP) techniques. SAG employs a hierarchical fusion strategy, enabling enhanced spatial feature extraction and noise suppression. The SAG derives feature mappings directly from the base network, harmonizing the feature channels from two distinct stages using a 1×1 convolution. This process is followed by the extraction of spatial information through both AP and mp techniques. CBAM model proved experimentally33 that max-pooled features which encode the degree of the most salient part can compensate the average-pooled features which encode global statistics softly. It was observed that the majority of the connections between modules in DenseNet employ average-pooling, which reduces the dimensionality. This is advantageous as it facilitates the transfer of information to the subsequent module for feature extraction. Consequently, we utilise the AP for feature extraction in the lower layer (Hg×Wg×Cg) reduces the number of parameters while retaining the most contributing feature information. In contrast, the upper layer (Hp×Wp×Cp), which contains a greater proportion of less useful information, employs the MP for key feature selection. The extracted features are then subjected to sigmoid activation to generate spatial attention mappings, which focus on relevant areas within the image. Figure 3 showcases the SAG’s architecture, illustrating its role in enhancing the model’s attention to critical spatial details.Figure 3 Spatial attention gate and head unit.

After the spatial gating unit, the Head section assimilates global spatial features for semantic analysis and prediction. It utilizes a hierarchical fusion strategy to discern the attention mask across adjacent scales, with the network’s training confined to pairs of neighbouring scales. Figure 4 contrasts an original, non-downsampled image with one that has been reduced by a factor of 2, although this downsampling factor is adaptable. The fusion technique is designed to capture the nuanced attention dynamics between scale pairs, enhancing the network’s ability to adjust to variations in scale and improve feature detection capabilities as highlighted in MS-OCRNet32.Figure 4 Training of gated attention mechanisms in multi-scale space.

Figure 5 depicts the hierarchical spatial gated attention mechanism model. The symbol Fw indicates that the input images are initially processed through a shared backbone network and a contextual environment capturer. Fi indicates that i-th stage output semantic feature, 1<i<=4. The symbol q denotes the weights after the output of the gated attention mechanism operation, while the symbol Fα denotes the attention mask obtained after convolution of the weighted features. In the training process, the input image after processing is scaled by r, where r=0.5 means downsampling by 2 times, r=2.0 means upsampling by 2 times, and r=1 means no operation. We select 0.5 and 1.0 scales images for training, denoted as Fr=0.5 and Fr=1, respectively. Then semantic prediction of the backbone Fr=1w, Fr=0.5w is obtained after passing through the shared backbone network and the contextual environment capturer, and i-th stage generates the semantic feature (Fr=0.5i). At the i-th stage, semantic features Fr=0.5i are produced by passing the inputs through the spatial gated attention as qr=0.5i. This process effectively captures and refines semantic information at various scales, as demonstrated by the gated attention computation equation.1 qr=0.5i=σ1[Mi,Ai]=σ1[M(f1×1(Fr=0.5i));A(f1×1(Fr=0.5i-1))]

The feature maps generated via SAG fusion are fed into the Head’s attention mechanism to derive the attention weight mappings qr=0.5i . Here, σ1 represents the ReLU activation function, and σ2 is the sigmoid activation function. The symbol f1×1 denotes convolution with a 1×1 kernel. While M, A indicate the maximum pooling and average pooling operations, respectively. This procedure meticulously calculates the attention weights, facilitating focused analysis on pertinent image regions.2 Fr=0.5α=Fr=0.5i·σ2(f1×1(qr=0.5i))

Equation (3) outlines the process of weight assignment, which integrates attention mappings and high-level semantic information. Specifically, the attention mapping, adjusted by a factor r=0.5, undergoes element-wise multiplication Fr=0.5w with the high-level semantic feature, facilitating targeted semantic enhancement. The resultant product is further processed through 1-Fr=0.5a element-wise multiplication with neighboring scale information Fr=1w to generate the refined output image Fr=1s. Here, Up signifies the application of linear interpolation for upscaling to the desired resolution.3 Fr=1s=Up(Fr=0.5w·Fr=0.5α)+((1-Up(Fr=0.5α))·Fr=1w)

MSAG inference

During inference, the obtained attention is applied hierarchically, as depicted in Fig. 5. This involves integrating the attention through a series of computations across N prediction scales, effectively utilizing the attention mechanism to refine predictions at multiple levels of detail.Figure 5 Inference for hierarchical multi-scale spatial gated attention mechanisms.

In the inference phase, leveraging a 2.0 scale simplifies the process by directly inputting up-sampled images (2.0×) into the pre-trained base network, bypassing the need to merge training data from the 0.5 and 1.0 scales. The initial segmentation prediction is derived from this base network output. Subsequently, this prediction is enhanced by element-wise multiplication with the attention mechanism’s down-sampled prediction. Layer-by-layer summation of these predictions refines the final segmentation outcome. This methodology prioritizes lower scales for detail while incrementally incorporating higher scale data for broader context, allowing for refined prediction locations.

The approach presents two key benefits: scale flexibility and the systematic integration of higher scale data. It enables the model to incorporate new scales (0.25×, 2.0×) beyond the ones used in training (0.5×, 1.0×), addressing the common restriction of models to their training scales. Moreover, the hierarchical model structure enhances training efficiency. Given the spatially gated attention’s minimal impact on training complexity, training exclusively at 0.5 and 1.0 scales suffices, thereby reducing training demands.

Experimental

Experimental setup

Data description

The study utilized the Indian Diabetic Retinopathy Image Dataset (IDRiD-S)34 and DR Dataset(DDR) for its experimental analysis. IDRiD-S dataset encompasses raw color fundus photographs and provides pixel-level, binarized lesion segmentation with accurate labels. It includes 81 images specifically annotated for DR, covering the four critical lesion types pertinent to our research: MA, SE, EX, and HE. These images are of high resolution, measuring 4288×2848 pixels, ensuring detailed lesion representation for analysis. DDR is a publicly available dataset from China that was collected between 2016 and 2018. It comprises 13,673 color fundus images from 147 hospitals across 23 provinces in China, with 84 of these hospitals being tertiary-level A hospitals. All images were desensitized for general use. The images were annotated according to the international standard, with five categories of DR severity. In addition to the image-level annotations for DR severity, 757 images were provided with pixel-level and bounding box-level annotations for four lesions35.

Dataset processing

Image cutting: for IDRiD-S, the process involves dividing each high-resolution fundus image (4288×2848) into smaller, manageable segments. The images were divided into 54 blocks of 512×512 pixels each using a step wise approach. This cutting creates a sample dataset optimized for experimental analysis, resulting in a training set of 2916 images and a testing set of 1458 images. We resize the images of DDR dateset to 512×512, without image cutting.

Image enhancement: to augment the dataset and improve model robustness, image enhancement techniques are applied. These include flipping, rotation, and pixel value normalization, each contributing to a more diverse and challenging dataset for model training and evaluation.

Evaluation metrics

In this paper, the segmentation performance is evaluated using Dice coefficient (Dice) and Area Under the Precision-Recall Curve (AUPR) . Higher scores in these two metrics indicate the better segmentation capability of the proposed method. The AUPR reflects the ability of a method to enrich true positive samples, which is the area accuracy under the precision-recall curve, and mAUPR is the average of the AUPR values for all lesion types. The Dice is widely recognized for its effectiveness in comparing the spatial overlap accuracy of segmentation results against ground truth annotations. In order to assess the model’s efficacy in segmenting the four distinct lesion types individually and the overall image, we utilize “EX”, “HE”, “SE”, and “MA” in the results table to represent the Dice coefficients for the four lesions, and “mDice” to represent the average Dice coefficient for the overall image segmentation. The Dice coefficient is calculated using the following formula:4 Dice(X,Y)=2×X⋂YX+Y

Parameter settings

Experiments assessed the performance of four contrasting networks and attention networks on PyTorch. Cross entropy served as the loss function over 200 iterations with a batch size of 8. Optimization utilized the SGD algorithm, starting with a learning rate of 0.005, and applied a polynomial strategy for learning rate adjustment alongside a weight decay of 5e−4. Image inputs were standardized to 512×512×3. Initial weights for the backbone network leveraged ImageNet’s pre-trained models, whereas the decoder was initialized via the He Normal method. Training was conducted on a Tesla V100 GPU.

Experimental results and analyses

Table 1 illustrates the enhanced segmentation flexibility of our MSAG method, trained on 0.5 and 1.0 scales, by incorporating 2.0 scale predictions at inference, outperforming traditional pooling methods, MSAG shows a notable accuracy improvement of 1%–3% across all four lesion types. The table also highlights the nuanced impact of integrating new scales into training and inference: while it marginally boosts the accuracy for certain lesions, it adversely affects the accuracy of bleeding feature detection and incurs higher training costs. Thus, careful consideration is advised when expanding the scale range for training and inference processes. Table 1 Comparison of our hierarchical multi-scale attention method vs. other approaches on IDRiD.

Method	Train scales	Eval scales	HE	EX	MA	SE	mDice	mAURP	
Single	1.0	1.0	57.27	77.92	44.74	61.20	60.28	65.01	
AvgPool	0.5, 1.0, 2.0	0.5, 1.0, 2.0	52.12	74.24	41.23	60.75	57.09	62.95	
AvgPool	0.25, 0.5, 1.0, 2.0	0.25, 0.5, 1.0, 2.0	50.92	76.54	42.59	61.49	57.89	63.13	
MSAG (ours)	0.5, 1.0	0.5, 1.0, 2.0	64.35	78.22	45.53	69.04	64.29	65.26	
MSAG (ours)	0.5, 1.0, 2.0	0.25, 0.5, 1.0, 2.0	63.34	81.19	49.20	69.26	65.75	66.87	
Eval scales: scales used for multi-scale evaluation (unit: %).

The best results are highlighted in bold

We compare MASG with other state-of-the-art methods on the DDR and IDRiD dataset. These compared methods are mainly divided into four categories, including classic convolutional networks, Transformer-based networks, structurally similar networks and previous DR multi-lesion segmentation networks. Convolutional networks include U-Net, HED, DeepLabv3+, DeepLab &U-Net, Transformer-Based networks include Swin-Transformer v2 (Swin-TV2). Previous DR multi-lesion segmentation networks include M2MRF. Same type networks include HRNet+OCR, HRNet+OCR+MS, Attention U-Net. Table 2 lists the performance of different comparison methods on the IDRiD and DDR dataset. Likewise, compared with the HRNet+OCR (Backbone) on the IDRiD, our metrics mDice, mAUPR, improve respectively by 4.19%, 1.94%. It showcasing significant enhancements in Dice scores for all four lesion types and overall. The improved lesion boundary delineation underscores the SAG module’s effectiveness. Compared with the previous best method M2MRF on the DDR, our metrics mDice, mAUPR, improve respectively by 2.91%, 0.24%. Further comparisons with structurally similar networks, including Attention U-Net and HRNet+OCR+MS, reveal that MSAG’s layered architecture and SAG module contribute to its higher segmentation accuracy. Despite MSAG’s exceptional detection of microaneurysms, it slightly trails the DeepLab &U-Net model in other lesion areas. Nonetheless, MSAG presents a more balanced accuracy across lesion types, affirming its comprehensive detection capabilities. In the two datasets, not only the mean metrics but also the category metrics have improved. All in all, these quantitative results on two datasets substantiate the fine robustness of our MSAG network.

Figure 6 shows the qualitative results of different methods on the IDRiD dataset, including DeepLabv3+, Swin-Tv2, HRNet+OCR+MS, and our MSAG. It demonstrates that our MSAG can precisely segment lesions of different shapes and sizes compared with other methods. Table 2 Comparison of our proposed MSAG with the state-of-the-arts methods on the IDRiD and DDR dataset.

Dataset	IDRiD	DDR	
Dice	AUPR	Dice	AUPR	
Method	HE	EX	MA	SE	mDice	mAUPR	HE	EX	MA	SE	mDice	mAUPR	
U-Net19	44.97	69.90	41.76	27.88	46.13	34.68	38.41	42.12	19.83	32.01	33.09	29.60	
DeepLabv3+	57.27	77.92	44.74	61.27	60.30	63.19	41.83	58.59	25.40	37.97	40.95	43.34	
DeepLab &U-Net16	64.90	71.72	49.51	69.95	64.02	64.57	43.32	60.54	25.34	45.52	43.68	45.35	
Attention U-Net36	48.32	–	–	–	–	–	–	–	–	–	-	–	
Swin-TV237	62.25	80.12	43.90	65.71	62.99	64.86	49.36	61.18	26.75	54.14	47.86	48.59	
M2MRF38	65.01	79.57	47.81	65.39	64.45	66.00	48.58	60.20	27.51	46.81	45.78	49.42	
HED39	67.00	78.60	42.98	64.33	38.81	63.94	45.50	56.63	22.43	42.61	41.79	42.97	
HRNet+OCR (backbone)	54.41	80.06	45.75	66.00	61.56	64.93	44.86	58.98	26.99	44.96	43.95	45.21	
HRNe+OCR+MS	60.07	80.91	47.75	67.69	64.11	66.01	48.62	57.23	27.02	46.33	44.80	47.36	
MSAG (ours)	63.34	81.19	49.20	69.26	65.75	66.87	50.96	61.21	27.34	55.23	48.69	49.66	
The best results are highlighted in bold and the second best results are italicized (unit: %).

Figure 6 Visualization of different methods’ segmentation results on the IDRiD.

Discussion

Addressing fundus lesion detection challenges, particularly small pixel occupancy and diverse lesion sizes, we introduce a Hierarchical Multi-Scale Spatially Gated Attention Mechanism Network (MSAG), drawing inspiration from existing multi-scale attention concepts. Through implementing various scale fusion methods, we validate MSAG’s effectiveness in enhancing segmentation accuracy across multiple lesion types. Comparative analyses with similar networks underscore MSAG’s superior performance in accurately segmenting four key lesion regions. Notably, MSAG not only excels in detecting a broad range of lesion sizes but also demonstrates a balanced detection capability across different lesion types, which is crucial for comprehensive DR assessment. The superior performance of our approach can be attributed to the MSAG mechanism’s ability to handle multi-scale features effectively. Unlike traditional methods that may overlook minute features like microaneurysms, our model’s spatial attention gate focuses on these critical areas, enhancing detection accuracy. Additionally, the hierarchical fusion of features ensures that both fine and coarse details are captured and integrated. In future work, we aim to enhance multi-lesion segmentation in DR images using multi-modal technology.

Acknowledgements

This work was supported by the Guizhou Provincial People’s Hospital Youth Fund (No. GZSYQN202221);the CPC Guizhou Provincial Committee’s Major Research Topics on Comprehensively Deepening Reforms(No.GZGGKT2024114)

Author contributions

Changzhuan Xu conceived the experiments,Changzhuan Xu and Hailin Li performed the experiments, and Changzhuan Xu and Song He analyzed the results. All authors reviewed the manuscript.

Data availibility

The datasets used and/or analyzed during the current study are available from public datasets.

Competing interests

The authors declare no competing interests.

Publisher's note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
==== Refs
References

1. Joslin EP The prevention of diabetes mellitus JAMA 2021 325 190 190 10.1001/jama.2020.17738 33433568
Joslin, E. P. The prevention of diabetes mellitus. JAMA 325, 190–190 (2021).33433568 10.1001/jama.2020.17738
2. G. W. H. Organization. World Report on Vision. [EB/OL]. Licence: CC BY-NC-SA 3.0 IGO (2019).
3. Mookiah MRK Computer-aided diagnosis of diabetic retinopathy: A review Comput. Biol. Med. 2013 43 2136 2155 10.1016/j.compbiomed.2013.10.007 24290931
Mookiah, M. R. K. et al. Computer-aided diagnosis of diabetic retinopathy: A review. Comput. Biol. Med. 43, 2136–2155 (2013).24290931 10.1016/j.compbiomed.2013.10.007
4. Pratt H Coenen F Broadbent DM Harding SP Zheng Y Convolutional neural networks for diabetic retinopathy Proc. Comput. Sci. 2016 90 200 205 10.1016/j.procs.2016.07.014
Pratt, H., Coenen, F., Broadbent, D. M., Harding, S. P. & Zheng, Y. Convolutional neural networks for diabetic retinopathy. Proc. Comput. Sci. 90, 200–205 (2016).10.1016/j.procs.2016.07.014
5. Mansour RF Evolutionary computing enriched computer-aided diagnosis system for diabetic retinopathy: A survey IEEE Rev. Biomed. Eng. 2017 10 334 349 10.1109/RBME.2017.2705064 28534786
Mansour, R. F. Evolutionary computing enriched computer-aided diagnosis system for diabetic retinopathy: A survey. IEEE Rev. Biomed. Eng. 10, 334–349 (2017).28534786 10.1109/RBME.2017.2705064
6. Kazi, K. S. Computer-aided diagnosis in ophthalmology: A technical review of deep learning applications. In Transformative Approaches to Patient Literacy and Healthcare Innovation. 112–135 (2024).
7. Gao XW Hui R Tian Z Classification of CT brain images based on deep learning networks Comput. Methods Prog. Biomed. 2017 138 49 56 10.1016/j.cmpb.2016.10.007
Gao, X. W., Hui, R. & Tian, Z. Classification of CT brain images based on deep learning networks. Comput. Methods Prog. Biomed. 138, 49–56 (2017).10.1016/j.cmpb.2016.10.007
8. Kitchen, A. & Seah, J. Deep generative adversarial neural networks for realistic prostate lesion MRI synthesis. arXiv preprint arXiv:1708.00129 (2017).
9. Peng Y Luo Y Yan J Li W Liao Y Automatic measurement of fetal anterior neck lower jaw angle in nuchal translucency scans Sci. Rep. 2024 14 5351 10.1038/s41598-024-55974-x 38438512
Peng, Y., Luo, Y., Yan, J., Li, W. & Liao, Y. Automatic measurement of fetal anterior neck lower jaw angle in nuchal translucency scans. Sci. Rep. 14, 5351 (2024).38438512 10.1038/s41598-024-55974-x
10. Budak Ü Sengür A Guo Y Akbulut Y A novel microaneurysms detection approach based on convolutional neural networks with reinforcement sample learning algorithm Health Inf. Sci. Syst. 2017 5 1 10 10.1007/s13755-017-0034-9 28413630
Budak, Ü., Sengür, A., Guo, Y. & Akbulut, Y. A novel microaneurysms detection approach based on convolutional neural networks with reinforcement sample learning algorithm. Health Inf. Sci. Syst. 5, 1–10 (2017).28413630 10.1007/s13755-017-0034-9
11. Dai L Clinical report guided retinal microaneurysm detection with multi-sieving deep learning IEEE Trans. Med. Imaging 2018 37 1149 1161 10.1109/TMI.2018.2794988 29727278
Dai, L. et al. Clinical report guided retinal microaneurysm detection with multi-sieving deep learning. IEEE Trans. Med. Imaging 37, 1149–1161 (2018).29727278 10.1109/TMI.2018.2794988
12. Huang, Y. et al. Automated hemorrhage detection from coarsely annotated fundus images in diabetic retinopathy. In 2020 IEEE 17th International Symposium on Biomedical Imaging (ISBI). 1369–1372. 10.1109/ISBI45749.2020.9098319 (2020).
13. Guefrachi, S., Echtioui, A. & Hamam, H. Automated diabetic retinopathy screening using deep learning. Multimed. Tools Appl. 1–18 (2024).
14. Latha, G., Priya, P. A. & Smitha, V. Enhanced diabetic retinopathy detection and exudates segmentation using deep learning: A promising approach for early disease diagnosis. Multimed. Tools Appl. 1–24 (2024).
15. Jabbar A A lesion-based diabetic retinopathy detection through hybrid deep learning model IEEE Access 2024 12 40019 40036 10.1109/ACCESS.2024.3373467
Jabbar, A. et al. A lesion-based diabetic retinopathy detection through hybrid deep learning model. IEEE Access 12, 40019–40036 (2024).10.1109/ACCESS.2024.3373467
16. Jiawei F Ruru Z Meng L Jiawen H Application of deep learning method in the diagnosis of diabetic retinopathy Acta Autom. Sin. 2021 47 1 20
Jiawei, F., Ruru, Z., Meng, L. & Jiawen, H. Application of deep learning method in the diagnosis of diabetic retinopathy. Acta Autom. Sin. 47, 1–20 (2021).
17. Aziz T Charoenlarpnopparut C Mahapakulchai S Deep learning-based hemorrhage detection for diabetic retinopathy screening Sci. Rep. 2023 13 2045 2322 10.1038/s41598-023-28680-3 36739302
Aziz, T., Charoenlarpnopparut, C. & Mahapakulchai, S. Deep learning-based hemorrhage detection for diabetic retinopathy screening. Sci. Rep. 13, 2045–2322 (2023).36739302 10.1038/s41598-023-28680-3
18. Tan JH Automated segmentation of exudates, haemorrhages, microaneurysms using single convolutional neural network Inf. Sci. 2017 420 66 76 10.1016/j.ins.2017.08.050
Tan, J. H. et al. Automated segmentation of exudates, haemorrhages, microaneurysms using single convolutional neural network. Inf. Sci. 420, 66–76 (2017).10.1016/j.ins.2017.08.050
19. Badar, M., Shahzad, M. & Fraz, M. Simultaneous segmentation of multiple retinal pathologies using fully convolutional deep neural network. In Annual Conference on Medical Image Understanding and Analysis. 313–324 (Springer, 2018).
20. De Fauw J Clinically applicable deep learning for diagnosis and referral in retinal disease Nat. Med. 2018 24 1342 1350 10.1038/s41591-018-0107-6 30104768
De Fauw, J. et al. Clinically applicable deep learning for diagnosis and referral in retinal disease. Nat. Med. 24, 1342–1350 (2018).30104768 10.1038/s41591-018-0107-6
21. Yan L Li K Gao R Wang C Xiong N An intelligent weighted object detector for feature extraction to enrich global image information Appl. Sci. 2022 12 7825 10.3390/app12157825
Yan, L., Li, K., Gao, R., Wang, C. & Xiong, N. An intelligent weighted object detector for feature extraction to enrich global image information. Appl. Sci. 12, 7825 (2022).10.3390/app12157825
22. Zhang X Regional context-based recalibration network for cataract recognition in as-oct Pattern Recognit. 2024 147 110069 10.1016/j.patcog.2023.110069
Zhang, X. et al. Regional context-based recalibration network for cataract recognition in as-oct. Pattern Recognit. 147, 110069. 10.1016/j.patcog.2023.110069 (2024).10.1016/j.patcog.2023.110069
23. Zhang X Attention to region: Region-based integration-and-recalibration networks for nuclear cataract classification using as-oct images Med. Image Anal. 2022 80 102499 10.1016/j.media.2022.102499 35704990
Zhang, X. et al. Attention to region: Region-based integration-and-recalibration networks for nuclear cataract classification using as-oct images. Med. Image Anal. 80, 102499 (2022).35704990 10.1016/j.media.2022.102499
24. Zhang, X. et al. Pyramid pixel context adaption network for medical image classification with supervised contrastive learning. In IEEE Transactions on Neural Networks and Learning Systems (2024).
25. Khaparde A Chapadgaonkar S Kowdiki M Deshmukh V An attention-based swin u-net-based segmentation and hybrid deep learning based diabetic retinopathy classification framework using fundus images Sens. Imaging 2023 24 20 10.1007/s11220-023-00426-5
Khaparde, A., Chapadgaonkar, S., Kowdiki, M. & Deshmukh, V. An attention-based swin u-net-based segmentation and hybrid deep learning based diabetic retinopathy classification framework using fundus images. Sens. Imaging 24, 20 (2023).10.1007/s11220-023-00426-5
26. Kothadiya D Rehman A Abbas S Alamri FS Saba T Attention-based deep learning framework to recognize diabetes disease from cellular retinal images Biochem. Cell Biol. 2023 101 550 561 10.1139/bcb-2023-0151 37473447
Kothadiya, D., Rehman, A., Abbas, S., Alamri, F. S. & Saba, T. Attention-based deep learning framework to recognize diabetes disease from cellular retinal images. Biochem. Cell Biol. 101, 550–561 (2023).37473447 10.1139/bcb-2023-0151
27. Sinha A Dolz J Multi-scale self-guided attention for medical image segmentation IEEE J. Biomed. Health Inform. 2020 25 121 130 10.1109/JBHI.2020.2986926
Sinha, A. & Dolz, J. Multi-scale self-guided attention for medical image segmentation. IEEE J. Biomed. Health Inform. 25, 121–130 (2020).10.1109/JBHI.2020.2986926
28. Fu, J. et al. Dual attention network for scene segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 3146–3154 (2019).
29. Wentu X Jianxia L Ran L Xiaohui Y An improved method for retinal vascular segmentation in u-net Acta Opt. Sin. 2020 40 11
Wentu, X., Jianxia, L., Ran, L. & Xiaohui, Y. An improved method for retinal vascular segmentation in u-net. Acta Opt. Sin. 40, 11 (2020).
30. Oktay, O. et al. Attention u-net: Learning where to look for the pancreas. arXiv preprint: arXiv:1804.03999 (2018).
31. Chen, L.-C., Yang, Y., Wang, J., Xu, W. & Yuille, A. L. Attention to scale: Scale-aware semantic image segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 3640–3649 (2016).
32. Tao, A., Sapra, K. & Catanzaro, B. Hierarchical multi-scale attention for semantic segmentation. arXiv preprint: arXiv:2005.10821 (2020).
33. Woo, S., Park, J., Lee, J.-Y. & Kweon, I. S. Cbam: Convolutional block attention module. In Proceedings of the European Conference on Computer Vision (ECCV) (2018).
34. Porwal P Indian diabetic retinopathy image dataset (IDRID): A database for diabetic retinopathy screening research Data 2018 3 25 10.3390/data3030025
Porwal, P. et al. Indian diabetic retinopathy image dataset (IDRID): A database for diabetic retinopathy screening research. Data 3, 25 (2018).10.3390/data3030025
35. Li T Diagnostic assessment of deep learning algorithms for diabetic retinopathy screening Inf. Sci. 2019 501 511 522 10.1016/j.ins.2019.06.011
Li, T. et al. Diagnostic assessment of deep learning algorithms for diabetic retinopathy screening. Inf. Sci. 501, 511–522. 10.1016/j.ins.2019.06.011 (2019).10.1016/j.ins.2019.06.011
36. Xiao, Q. et al. Improving lesion segmentation for diabetic retinopathy using adversarial learning. In International Conference on Image Analysis and Recognition. 333–344 (2019).
37. Liu, Z. et al. Swin transformer v2: Scaling up capacity and resolution. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 12009–12019 (2022).
38. Liu Q Liu H Ke W Liang Y Automated lesion segmentation in fundus images with many-to-many reassembly of features Pattern Recognit. 2023 136 109191 10.1016/j.patcog.2022.109191
Liu, Q., Liu, H., Ke, W. & Liang, Y. Automated lesion segmentation in fundus images with many-to-many reassembly of features. Pattern Recognit. 136, 109191 (2023).10.1016/j.patcog.2022.109191
39. Xie, S. & Tu, Z. Holistically-nested edge detection. In Proceedings of the IEEE International Conference on Computer Vision. 1395–1403 (2015).
