
==== Front
Sci Rep
Sci Rep
Scientific Reports
2045-2322
Nature Publishing Group UK London

39278955
70527
10.1038/s41598-024-70527-y
Article
Boundary-aware convolutional attention network for liver segmentation in ultrasound images
Wu Jiawei 1
Liu Fulong 1
Sun Weiqin 1
Liu Zhipeng 2
Hou Hui 3
Jiang Rui 1
Hu Haowei 1
Ren Peng 1
Zhang Ran 1
Zhang Xiao changshui@hotmail.com

14
1 https://ror.org/035y7a716 grid.413458.f 0000 0000 9330 9891 School of Medical Informatics and Engineering, Xuzhou Medical University, Xuzhou, 221000 China
2 https://ror.org/02fvevm64 grid.479690.5 Department of Information, Taizhou People’s Hospital Affiliated to Nanjing Medical University, Taizhou, 225300 China
3 Department of Imaging, The Fourth People’s Hospital of Taizhou in Jiangsu Province, Taizhou, 225300 China
4 Yantai Longch Technologies Co., Ltd, Yantai, 264000 China
15 9 2024
15 9 2024
2024
14 2152915 5 2024
19 8 2024
© The Author(s) 2024
2024
https://creativecommons.org/licenses/by-nc-nd/4.0/ Open Access This article is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License, which permits any non-commercial use, sharing, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if you modified the licensed material. You do not have permission under this licence to share adapted material derived from this article or parts of it. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by-nc-nd/4.0/.
Liver ultrasound is widely used in clinical practice due to its advantages of non-invasiveness, non-radiation, and real-time imaging. Accurate segmentation of the liver region in ultrasound images is essential for accelerating the auxiliary diagnosis of liver-related diseases. This paper proposes BACANet, a deep learning algorithm designed for real-time liver ultrasound segmentation. Our approach utilizes a lightweight network backbone for liver feature extraction and incorporates a convolutional attention mechanism to enhance the network's ability to capture global contextual information. To improve early localization of liver boundaries, we developed a selective large kernel convolution module for boundary feature extraction and introduced explicit liver boundary supervision. Additionally, we designed an enhanced attention gate to efficiently convey liver body and boundary features to the decoder to enhance the feature representation capability. Experimental results across multiple datasets demonstrate that BACANet effectively completes the task of liver ultrasound segmentation, achieving a balance between inference speed and segmentation accuracy. On a public dataset, BACANet achieved a DSC of 0.921 and an IOU of 0.854. On a private test dataset, BACANet achieved a DSC of 0.950 and an IOU of 0.907, with an inference time of approximately 0.32 s per image on a CPU processor.

Subject terms

Biomedical engineering
Data mining
Data processing
Image processing
Jiangsu Provincial Graduate Student Research and Innovation ProgramKYCX23_2965 Wu Jiawei Unveiling & Leading Project of XZHMUJBGS202204 Zhang Xiao issue-copyright-statement© Springer Nature Limited 2024
==== Body
pmcIntroduction

Ultrasound (US) imaging is a mature medical imaging technology characterized by its non-invasiveness, non-radiation, and ability to provide real-time, quantitative anatomical and physiological information1. Additionally, due to the ease of operation and low cost of purchase and maintenance, US equipment is extensively used to guide interventional clinical procedures. In clinical practice, accurate description and segmentation of anatomical structures in US images are essential for various purposes. Specifically, in liver US examinations, physicians make further diagnoses by observing the morphology, contours, and blood flow signals of the liver and adjacent structures. For example, US features such as liver brightness and the contrast between the liver and kidneys can assist in grading the degree of liver steatosis2. Intraoperative US during liver surgery helps with disease staging, surgical planning, and real-time guidance of hepatectomies3. In addition to tasks related to liver disease analysis, studying the characteristics of healthy liver organs is also significant, as it is a prerequisite for standard segmental diagnosis, thus accelerating the development of auxiliary diagnostic fields for liver-related diseases4. However, due to the inherent characteristics of imaging, US images inevitably have flaws such as acoustic shadows, poor contrast, and speckle noise, which pose significant challenges for further analysis of liver US images, especially the segmentation of liver anatomical structures5. Moreover, manually segmenting a large volume of liver US images is a tedious and time-consuming task, imposing a heavy workload on radiologists, and the final segmentation outcomes are often affected by the varying diagnostic experiences of different physicians. Therefore, developing a fast and accurate automated algorithm for segmenting the liver region in US images holds great significance. Such an algorithm could potentially advance the development of liver disease auxiliary diagnostic systems and improve the effectiveness of ultrasound-guided liver surgeries.

Over the past few decades, traditional image segmentation algorithms such as thresholding methods6, active contour models7, and graph cuts8 have been extensively used to process complex medical images. These algorithms are capable of analyzing medical images in an automatic or semi-automatic manner, providing segmentation results for region of interest (ROI). However, these traditional methods typically require the setting of empirical parameters and may necessitate the creation of handcrafted features to aid the segmentation process. As a result, the final segmentation outcomes often lack stability, which has limited their widespread application in computer-assisted diagnostic systems. In recent years, with the rapid development of artificial intelligence technology, computer vision algorithms based on deep learning have achieved significant success in multiple biomedical image analysis tasks9–12. These deep learning-based approaches have the potential to overcome the limitations of traditional segmentation methods and provide more robust results for medical image analysis. Particularly noteworthy is the pivotal role played by convolutional neural networks (CNN) in this domain, where the UNet9 architecture based on encoder-decoder structure has become the mainstream framework for medical image segmentation. UNet enhances the model's overall segmentation capability significantly by stacking multiple convolutional modules to deeply extract local features of the input image and by innovatively designing skip connections between the encoder and decoder to directly integrate low-level semantic features from shallow layers with high-level semantic features from deeper layers. However, medical images often encompass multiple anatomical structures and pathological areas, characterized by significant intra-class variation and indistinct inter-class boundaries10. The inherent properties of convolution operations, which emphasize the extraction of local features, may neglect the global context, leading to potential misinterpretations of the image content. Consequently, most methods based solely on CNN may lack precise segmentation capabilities for the boundaries of the ROI. To address these challenges, researchers have introduced attention mechanism to distinguish the importance of different regions within an image for specific tasks, thus enhancing model performance. Recent literature indicates that the Transformer13 structure, based on self-attention mechanism, has made significant breakthroughs in the field of computer vision, even surpassing the capabilities of CNN in some tasks14. One of the main advantages of Transformer is its ability to establish global contextual dependencies, which helps the model understand the overall semantic distribution of images. Further, TransUNet15, the first method to combine Transformer with UNet, leverages Transformer to further encode the high-level semantic features extracted by CNN, thus obtaining rich global contextual representation and demonstrating exceptional performance in multiple medical image segmentation tasks. Despite the substantial achievements of CNN and Transformer hybrid segmentation architectures, challenges remain, such as the large overall model parameter count and high computational complexity of attention calculations, limiting their application in real-time US image inference tasks. Moreover, while many models focus on segmenting the main body of ROI, they often neglect the attention to boundaries. Thus, achieving precise boundary segmentation of ROI remains a challenging task.

Inspired by the aforementioned literature, this paper introduces a deep learning algorithm named Boundary-aware convolutional attention network (BACANet), designed for the fast and accurate segmentation of liver US images. BACANet consists of three branches: an encoder, a body decoder, and a boundary decoder. The encoder utilizes a lightweight CNN backbone for in-depth feature extraction from complex liver US images. The body decoder is responsible for restoring and progressively refining the spatial position information of the liver area. The boundary decoder focuses on explicitly supervising liver boundaries to enhance the model's understanding of boundary information. At the end of the encoder, we introduced a convolutional attention mechanism to enhance the model's ability to extract multi-scale contextual information. Between the encoder and the body decoder, we introduced a soft attention gate mechanism that aids the body decoder in precisely locating the spatial position and boundary information of the liver region. Through extensive experimental validation across multiple datasets, we have fully demonstrated the outstanding performance of BACANet in the field of liver US image segmentation, effectively achieving a balance between prediction accuracy and inference speed. In summary, the main contributions of this paper are as follows:We designed an explicit liver boundary supervision branch, introduced Selective Large Kernel Convolution Module (SLKCM) to extract liver boundary information from shallow feature maps of the network, and adaptively fused this information to output predictions of liver boundaries.

We introduced Enhanced Attention Gate (EAG) aimed at reducing the semantic gap between the encoder and the body decoder, thereby aiding the body decoder in better recovering the spatial position information and boundary distribution of the liver region.

At the end of the encoder, we introduced Multi-scale Dilated Convolutional Attention Module (MDCAM) to enhance the model's ability to extract global contextual information, which is crucial for extracting features of liver region with varying shapes and sizes.

Based on the proposed enhancement modules, we designed a fast and accurate liver US segmentation algorithm named BACANet. We devised a comprehensive experimental procedure on both public and private liver US datasets and introduced a variety of existing medical image segmentation algorithms for comparison. The experimental results confirmed that BACANet exhibits outstanding performance in liver US segmentation tasks, surpassing comparative algorithms in segmentation performance on the same datasets.

Related works

CNN-based medical image segmentation

In recent years, image semantic segmentation technology based on deep learning has achieved significant breakthroughs in the field of intelligent medical image analysis, greatly advancing the development of this field. In 2015, Long et al.16 introduced the Fully Convolutional Network (FCN), utilizing convolutional layers instead of the fully connected layers typically found at the end of classification networks. This approach generates output feature maps that match the input image resolution, achieving pixel-level image semantic prediction. Inspired by FCN, Ronneberger et al.9 further proposed U-Net, a versatile biomedical image segmentation architecture. They designed unique skip connections to directly merge the high-resolution feature maps produced by the encoder with the high-level semantic feature maps from the decoder, significantly enhancing the network's localization capability and yielding finer segmentation results. However, a large semantic gap between the feature maps can diminish segmentation performance when using direct fusion. Consequently, Zhou et al.17 redesigned the skip connections in U-Net, upgrading them to a series of nested dense connections, reducing the semantic gap between the feature maps generated by the encoder and decoder, and achieving more precise medical image segmentation outcomes. Zhang et al.18 combined ResNet19 with U-Net, effectively merging different semantic information between the shallow and deep layers through residual connections. Moreover, they introduced pyramid dilated convolution to efficiently utilize global contextual information, expanding the receptive field of the network's deeper layers without losing resolution. Zhou et al.20 proposed a lightweight automatic US image segmentation network, LAEDNet, using a lightweight version of CNN as the encoder to enhance feature extraction capabilities for complex US images, achieving an optimal balance between segmentation accuracy and inference efficiency on three public US image datasets. To achieve real-time segmentation of liver US images, Ansari et al.21 proposed DensePSPUNet. They inserted an improved pyramid scene parsing module into the skip connections of U-Net to extract multi-scale features and contextual associations. Conducting liver segmentation experiments on 2400 US images, their method achieved a dice similarity coefficient (DSC) of 0.913 with real-time inference performance of 37 frames per second (FPS). However, their experimental data were limited to eight healthy volunteers and did not include private data.

Attention mechanism

In computer vision, attention mechanism serves as a method to help models differentiate the importance of various regions within feature maps, typically by incorporating learnable weights or parameters. Oktay et al.22 introduced a method called Attention Gate (AG), which automatically focuses on target structures of different shapes and sizes in medical images. It does this by implicitly suppressing irrelevant areas of the input image while highlighting features useful for specific tasks, thus improving model sensitivity and prediction accuracy with minimal computational overhead. Recent research has shown that the Transformer structure, based on the self-attention mechanism, has achieved remarkable performance in visual tasks. Inspired by this, Chen et al.15 proposed a medical image segmentation framework called TransUNet. They first encoded images using a pretrained ResNet to obtain high-level semantic feature maps, which were then converted into 2D sequences and processed by Transformer for global contextual information extraction. This method effectively combines the strengths of CNN and Transformer, achieving optimal performance in abdominal multi-organ segmentation tasks. Following the inspiration from TransUNet, increasing efforts are being made to incorporate the Transformer into the US image segmentation field. Yang et al.23 introduced a method called CSwin-PNet, aimed at the automatic segmentation of breast tumors in US images. They connected CNN with residual modules to Swin Transformer24 as the feature extraction backbone and designed an interactive channel attention module to assign higher weights to tumor-related feature areas. He et al.25 proposed a hybrid CNN-Transformer network called HCTNet to enhance the segmentation of breast lesions in US images. In the encoder of HCTNet, they designed Transformer encoder blocks to learn global contextual information, which were combined with CNN to extract features. In the decoder, they introduced a spatial attention mechanism module to reduce the semantic gap between it and the encoder. Zhang et al.4 presented a model called SEG-LUS for the segmentation of the liver and its associated structures. They combined the advantages of the cross-attention mechanism and the shifted windows computational model to capture image information at different scales and resolutions, allowing the network to retain more comprehensive information during the feature extraction process. They conducted extensive experiments on 7416 liver US images, achieving a DSC of 0.9855 for liver area segmentation. However, their experiments lacked validation on public datasets and did not discuss the model's inference speed.

Boundary supervision

In medical imaging, anatomical structure boundaries often appear as blurry or irregular features, negatively affecting the accuracy of segmentation results. To address this challenge, the introduction of an explicit boundary supervision branch in the segmentation network has been proven to be an effective solution. Lin et al.26 proposed an explicit learning strategy that utilizes boundary detection operators to extract boundary information of foreground objects and generate boundary masks for explicit supervision, guiding the learning process of the decoder. They believe that this strategy is highly beneficial for refining boundaries. Mishra et al.27 adopted a deep supervision approach, applying boundary supervision to the high-resolution feature maps at the shallow layers of the network, while the deeper, lower-resolution feature maps focused on supervising the main body of the segmentation targets. This is because anatomical structure boundary information is most effectively visualized at high resolutions. Wu et al.28 designed a Boundary-Guided Feature Enhancement (BGFE) module that uses multiple convolutional layers to learn boundary maps of breast lesion areas, thereby strengthening the feature maps generated by the network backbone. Their method significantly enhanced the network's overall boundary detection capabilities, allowing weak boundaries in blurred areas to be accurately identified. Sun et al.29 addressed the issue of fuzzy boundary identification of thyroid nodules in US images by proposing a dual-path segmentation network comprising a region path and a shape path. They designed a soft shape supervision block based on cross-path attention mechanism to enhance the feature representation of thyroid nodule boundaries and achieved effective fusion of dual-path information, thereby improving nodule contour prediction and segmentation accuracy. Ji et al.30 proposed a network called BAG-Net for segmenting liver targets with different morphologies in US images. They innovatively introduced a boundary detection module to fuse feature maps from different levels of the encoder and generate high-quality target boundary information. Ablation experiments showed that the boundary detection module significantly enhanced the segmentation performance of the network, leading to an approximately 1.4% increase in liver segmentation IOU score.

Methodology

BACANet overall architecture

Figure 1 illustrates an overview of the proposed BACANet, with liver ultrasound images as input and liver boundary and body as output. BACANet comprises three main branches: an encoder, a boundary decoder, and a body decoder. Additionally, the MDCAM module is located at the end of the encoder, and the EAG structure is situated between the encoder and the body decoder. Specifically, the body part of the encoder adopts ResNet10t31, which is a lightweight version of ResNet with novel structural optimization designed to balance performance and efficiency in feature extraction. The internal structure of ResNet10t mainly consists of an initial module (referred to as stem) and four stages of feature extraction modules (labeled as stage1 to stage4). The stem is composed of three sets of concatenated combinations of ‘3 × 3 convolution + BN + ReLU’. The channel numbers of these three convolutional layers are 24, 32, and 64, with the first layer having a stride of 2, while the other two layers have strides set to 1. The four feature extraction modules are constructed by removing half of the residual units from the convolutional stages Conv2_x, Conv3_x, Conv4_x, and Conv5_x of ResNet1819. The output channel numbers of these four modules are 64, 128, 256, and 512, respectively. The core component of the boundary decoder is SLKCM, which processes the first three groups of high-resolution feature maps from the encoder and generates refined boundary predictions. This is because high-resolution feature maps better preserve the boundary information of the target. The body decoder consists of multiple residual blocks and bilinear up-sampling operations, aiming to restore the spatial resolution of feature maps and locate the liver region. The structure of these residual blocks is consistent with the basic residual units in ResNet10t. The output channel numbers of these residual blocks are set to 256, 128, 64, 32, and 16, respectively. The core structure of MDCAM is a multi-scale convolutional attention mechanism, similar to the vision Transformer based on the self-attention mechanism, which can model long-range dependencies. EAG is a soft attention mechanism that effectively integrates features from three different branches to reduce the semantic gap between the encoder and the body decoder. Next, we will provide a detailed introduction to the core components of BACANet, including SLKCM, EAG, and MDCAM.Fig. 1 Overview of the proposed BACANet. Pink arrow denotes 2 × 2 max pooling, orange arrows denote 2 times bilinear up-sampling, red arrow represents 1 × 1 convolution with a sigmoid function, purple arrow represents 1 × 1 convolution with a softmax function.

Selective large kernel convolution module

In medical image segmentation tasks, predicting the boundary regions of the target is often more challenging than predicting internal pixels. The accuracy of boundary prediction directly affects the final segmentation performance32. Especially in liver US images, distinguishing clear boundaries of liver body is challenging due to low contrast and susceptibility to displacement and deformation. Therefore, we propose a method called SLKCM to explore liver boundary features of different shapes and scales. The detailed structure of SLKCM is illustrated in the Fig. 2. Firstly, we fuse three sets of low-level feature maps of different scales from the early stages of the encoder. This is because low-level feature maps typically have high resolution, enabling them to capture rich boundary information of the target. Next, to overcome the limitation of small receptive fields of common convolution kernels, we use oversized convolution kernels to generate three sets of attention maps. Large convolution kernels can better cover both the liver and background regions simultaneously, helping to capture liver boundary features. Finally, we fuse the three sets of low-level feature maps according to their corresponding attention weights.Fig. 2 Structural diagram of the proposed SLKCM. Orange arrows denote bilinear up-sampling, ⊕  and ⊗  represent element-wise addition and multiplication, respectively.

Assuming H, W, C represent the number of channels, height, and width of the feature maps, respectively. Given three sets of input feature maps F1∈RC×H2×W2, F2∈RC×H4×W4, and F3∈RC×H8×W8, obtained from the first three stages of ResNet10t and already adjusted to have consistent channel numbers through 1 × 1 convolution (set to 64 in this paper). Firstly, we interpolate F1, F2, and F3 using bilinear up-sampling to adjust their spatial resolutions to H × W, denoted as F1′∈RC×H×W, F2′∈RC×H×W, and F3′∈RC×H×W, respectively. Then, we fuse F1′, F2′, and F3′ using element-wise addition, and obtain spatial feature statistics Fgmp∈R1×H×W through global max pooling. Subsequently, we use three 31 × 31 oversized convolutions with sigmoid activation functions to generate boundary attention maps Fbs∈R3×H×W. Finally, we assign the three channels of Fbs to F1′, F2′, and F3′, and fuse them through element-wise addition to generate the feature output Rb∈RC×H×W with advanced boundary semantic properties. Formally, the above process can be represented by the following equations:1 Fgmp=GMPF1′+F2′+F3′

2 Fbs=sigmoid(fconv31×31(Fgmp))

3 Rb=F1′∗Fbs1+F2′∗Fbs2+F3′∗Fbs3

where GMP represents global max pooling, fconv31×31 denotes 31 × 31 convolution, ∗ denotes element-wise multiplication, and Fbsk represents the k-th channel of Fbs. Rb is a set of feature maps that are highly perceptive of liver boundary information. Subsequently, we apply a 1 × 1 convolution and sigmoid activation function to Rb to generate the output of the boundary decoder, which receives supervision from the mask of the liver boundary.

Enhanced attention gate

The encoder downsamples feature maps while extracting image features, which inevitably reduces their resolution and spatial information. This reduction negatively affects the body decoder's ability to reconstruct detailed target information. To address this, UNet innovatively introduced skip connections to directly fuse the high-resolution feature maps from the encoder with the high-level semantic feature maps from the decoder, effectively guiding the decoder to restore the spatial positional information of the target. However, directly fusing feature maps with a large semantic gap through copying and concatenation may reduce the segmentation performance. Therefore, we propose a skip connection method called EAG. EAG combines semantic information from boundary feature maps and high-level feature maps to enrich low-level feature maps semantically. This enhancement improves the body decoder's capability to restore spatial detail information in the feature maps.

The detailed structure of EAG is illustrated in Fig. 3. Assuming H, W, C represent the number of channels, height, and width of the feature maps, respectively. Given three sets of input feature maps Fb∈RCb×H×W, Fh∈RCh×H2×W2, and Fl∈RCl×H×W from the boundary decoder, body decoder, and encoder, respectively. Firstly, we interpolate Fh using bilinear up-sampling to match the resolution of Fb and Fl. Then, we uniformly adjust the number of channels of the three groups of feature maps to Cl by 1 × 1 convolution, and fuse them together using element-wise addition to obtain the feature map Fm∈RCl×H×W. Next, we utilize the mish function33 to update the values in Fm, retaining more favorable feature information for segmentation targets. Finally, we use a 1 × 1 convolution with a sigmoid activation function to generate the soft attention map Fs∈R1×H×W, and multiply it with Fl to obtain the final output Rf∈RCl×H×W of EAG. Formally, the above process can be represented by the following equations:4 Fm=fconv1×1Fb+fconv1×1(fup2Fh)+fconv1×1Fl

5 Fs=sigmoid(fconv1×1(mishFm))

6 Rf=Fl∗Fs

where fconv1×1 denotes 1 × 1 convolution, fup2 represents 2 times bilinear up-sampling, and ∗ denotes element-wise multiplication.Fig. 3 Structural diagram of the proposed EAG. Orange arrow denotes bilinear up-sampling, ⊕  and ⊗  represent element-wise addition and multiplication, respectively.

Multi-scale dilated convolutional attention module

Integrating Transformer into CNN structures addresses limitations in explicitly modeling long-range dependencies, which are often challenging for CNN. However, Transformer neglects the two-dimensional structural characteristics of images, and its attention mechanism entails quadratic computations and memory overheads, making them unsuitable for real-time image processing. To overcome these challenges, Guo et al.34 proposed a novel linear attention method based on depth-wise convolution, achieving similar long-range correlation modeling as the self-attention mechanism. Furthermore, SegNeXt35 demonstrated that convolutional attention is more efficient and effective than self-attention mechanism in Transformer, especially in semantic segmentation tasks. Inspired by these findings, we propose MDCAM to construct channel and spatial attention, enabling the acquisition of multiscale contextual information from local to global scales with fewer parameters.

The structure of MDCAM is depicted in Fig. 4. It resembles VIT14 in form, but we do not employ the self-attention mechanism and layer normalization. Instead, we utilize convolutional attention and image-specific batch normalization. FFN represents a feed forward network composed of a concatenation of 1 × 1 convolution, 3 × 3 depth-wise convolution, GELU36, and another 1 × 1 convolution. Attention represents the convolutional attention module composed of a concatenation of 1 × 1 convolution, GELU, MSDCA, and another 1 × 1 convolution. Here, MSDCA denotes multi-scale dilated convolutional attention, which consists of three parallel branches of depth-wise dilated convolution to capture multi-scale contextual information. Specifically, given the input Fin∈RC×H×W to MSDCA, first, a 3 × 3 depth-wise convolution aggregates local information to generate Fin′∈RC×H×W. Next, three parallel 3 × 3 depth-wise dilated convolutions capture multi-scale contextual information, and the features from all scales are fused using element-wise addition. Finally, attention weights Fatt∈RC×H×W are obtained through 1 × 1 convolution, and the output of MSDCA, Fout∈RC×H×W, is obtained by element-wise multiplication of Fin and Fatt. Formally, MSDCA can be expressed by the following equations:7 Fin′=fdwconv1Fin

8 Fatt=Fin′+fconv1×1∑i=24fdwconviFin

9 Fout=Fin∗Fatt

where fdwconvd represents 3 × 3 depth-wise convolution with dilation rate of d. ∗ denotes element-wise multiplication. In contrast to SegNeXt, we directly employ square convolutions instead of stripe convolutions. This choice is made because the latter may result in imbalanced perceptual capabilities across different orientations. Our aim with MDCAM is to extract more comprehensive and diverse multi-scale contextual features.Fig. 4 Structural diagram of the proposed MDCAM. The light orange square represents 3 × 3 depth-wise convolution and d is dilatation rate. ⊕  and ⊗  represent element-wise addition and multiplication, respectively.

Loss function

In this study, each US image is associated with a binary mask, which serves to supervise the body decoder of BACANet. We utilized the Canny37 edge detector to extract the liver boundary and generate a mask for the liver boundary, which is employed to supervise the boundary decoder of BACANet. Therefore, the total loss function Ltotal consists of two components: the body loss Lbody and the boundary loss Lboundary, formulated as follows:10 Ltotal=Lbody+αLboundary

where α is a weighting factor used to adjust the importance of Lbody and Lboundary. In this paper, the final setting for α is 0.3, determined to yield optimal performance during experimentation. Lbody comprises the combination of the cross-entropy function Lbody_ce and the dice loss function38 Lbody_dl, where the former focuses on the classification of each pixel, and the latter emphasizes the overlap between the predicted results and the mask. Lboundary is composed of the binary cross-entropy function Lboundary_bce and the mean square error loss function39 Lboundary_mse, with the latter encouraging smooth boundary predictions. Formally, the definitions of Lbody and Lboundary are as follows:11 Lbody=θLbodydlGb,Pb+LbodyceGb,Pb

12 Lboundary=LboundarybceGe,Pe+LboundarymseGe,Pe

where Gb and Pb represent the mask of the liver body and the output of the body decoder, respectively. Ge and Pe represent the mask of the liver boundary and the output of the boundary decoder, respectively. θ denotes the weighting factor used to balance Lbody_dl and Lbody_ce, which is set to 0.5 in this paper.

Experiments

Ethics statement

The research involving private data in this study was approved by the Clinical Research Ethics Committee of Taizhou People's Hospital Affiliated to Nanjing Medical University (Approval No.: KY2020-210-01), in accordance with the relevant provisions of the Declaration of Helsinki. This study directly retrieved data from the electronic medical record system. All patients signed a general informed consent form upon admission, and the system ensured the confidentiality of patients' basic information during data retrieval.

Datasets

We collected two sets of liver US datasets, named Dataset A and Dataset B, to comprehensively evaluate the segmentation performance of BACANet.

Dataset A is a publicly available dataset initially utilized and open-sourced by Ansari et al.19. This dataset comprises liver US videos from 8 healthy volunteers undergoing free-breathing motion. These videos exhibit extensive speckle noise, rib and lung shadows, artifacts, and incomplete anatomical boundaries, posing challenges for real-time liver US segmentation. For each US video, 300 frames are extracted, resulting in a total of 2400 liver US images (actually 2372 used), with the liver regions of these images annotated by three experienced radiologists through binarization. Additionally, 600 images from volunteers numbered 6 and 8 were used as the test set, while the remaining 1772 images were used for training. Due to the limited number of volunteers in Dataset A, to further evaluate the proposed method, we conducted additional experiments on BACANet using a fourfold cross-validation strategy.

Dataset B is a private dataset sourced from liver US examinations conducted at the Taizhou People's Hospital Affiliated to Nanjing Medical University between April 2022 and July 2023. During data collection, US scans displaying no apparent liver lesions were included according to the criteria. The US equipment used was Mindray DC-8EXP Diagnostic Ultrasound System. Finally, we successfully collected liver US images from 515 healthy individuals, totaling 1435 images. The binarization annotation of the liver region was performed by a radiologist with 5 years of US diagnostic experience and confirmed by another radiologist with 20 years of US diagnostic experience. For further research and analysis, Dataset B was divided into training (1147 images), validation (144 images), and test set (144 images).

Implementation details

In all experiments conducted in this paper, we utilized a workstation equipped with an NVIDIA RTX 4090 GPU (24 GB) and an Intel(R) Xeon(R) Platinum 8358P CPU @ 2.60 GHz. We implemented all deep learning models using the PyTorch (version 2.0.0) framework. Before training, all images were preprocessed to a resolution of 224 × 224 pixels, and the grayscale images were then duplicated 3 times to generate a 3-channel input format, which is the preferred input format for certain models. US images were resized using bilinear interpolation, while mask images were resized using nearest neighbor interpolation. During the training phase, the initial learning rate was set to 1e-3, and the batch size was set to 10. We employed the AdamW40 optimizer and utilized the Cosine Annealing41 learning rate scheduling strategy (with a period of 10) to dynamically adjust the learning rate. The training epochs for Dataset A and B were set to 50 and 100, respectively, due to the latter being sourced from a larger number of volunteers. We employed various data augmentation techniques to increase the sample size effectively, thereby mitigating the risk of overfitting. These methods were specifically detailed in Table 1 and were applied in real-time to the data in each training batch. Table 1 Data augmentation methods during the training process.

Method	Description	Probability	
Flip	Horizontal or vertical	0.5	
Brightness enhancement	Enhancement range set to 0.2–2.8	0.3	
Contrast enhancement	Enhancement range set to 0.2–2.8	0.3	
Image translation	Translation by ± 10% along the height and width direction	0.5	
Rotation	Rotation angle range set to ± 15°	0.3	

We implemented the data augmentation methods using the Pillow package. Before inputting into the model, all images were first divided by 255, and then normalized by subtracting the mean values of [0.485, 0.456, 0.406] and dividing by the standard deviations of [0.229, 0.224, 0.225] across the 3 channels. In addition, at the end of each training epoch, we performed inference on the validation set (which served as the test set for Dataset A) to record the corresponding DSC, and saved the weights of the best-performing model for the testing phase.

Evaluation metrics

To comprehensively quantify the predictive performance of the models, we selected six commonly used evaluation metrics in the field of image segmentation: Dice Similarity Coefficient (DSC), Intersection Over Union (IOU), Average Surface Distance (ASD), Positive Predictive Value (PPV), True Negative Rate (TNR), and True Positive Rate (TPR). DSC and IOU assess the spatial overlap between predicted results and masks. ASD quantifies the accuracy of segmentation boundaries. PPV reflects which pixels labeled as the liver in the prediction results are correct. TNR and TPR respectively indicate the proportions of background and liver regions correctly predicted, used to evaluate the accuracy of classification. The formulas for the above six evaluation metrics are as follows:13 DSC=2TP2TP+FP+FN

14 IOU=TPTP+FP+FN

15 ASD=12(∑x∈Xminy∈Yd(x,y)X+∑y∈Yminx∈Xd(y,x)Y)

16 PPV=TPTP+FP

17 TNR=TNTN+FP

18 TPR=TPTP+FN

where TN represent the number of pixels correctly predicted as background, FN represent the number of pixels incorrectly predicted as background, TP represent the number of pixels correctly predicted as liver, FP represent the number of pixels incorrectly predicted as liver. X and Y represent the liver region masks and the model's predicted results, respectively. d(x,y) represents the euclidean distance from point x to point y. In the above metrics, a smaller value of ASD indicates a more precise segmentation boundary of predicted results. For other metrics, values range between 0 and 1, where closer proximity to 1 indicates superior segmentation performance of the model.

Results and discussion

We have selected several excellent image segmentation algorithms for comparative experiments, including UNet9, RDeeplab42, TransUNet15, UNeXt43, LWBNAUNet44, and SegNeXt35. RDeeplab is an improved structure based on Deeplabv3+ with ResNet18 as the backbone. Its core component is a multi-scale feature extraction module called Atrous Spatial Pyramid Pooling (ASPP). SegNeXt is the first image segmentation network to utilize the convolutional attention mechanism. The other models mentioned are variations or improvements upon UNet. This section evaluates the segmentation performance of all models from both quantitative and qualitative perspectives, and discusses in depth their differences in liver segmentation tasks. It is important to emphasize that, to ensure a fair comparison of the models' performance, we did not use pre-trained weights from other datasets to fine-tune the models, but instead trained all models from scratch.

Quantitative analysis of Dataset A

This section presents the quantitative analysis results of all experimental models on Dataset A, including an analysis of the training curves and the prediction performance on the test set. Furthermore, we conducted a fourfold cross-validation experiment on BACANet and reported the results for each fold.

Figure 5 illustrates the average DSC changes of all models on the validation set over 50 epochs. Under the same number of training epochs, we can observe differences in the generalization abilities of different models on the validation set. The validation DSC of UNet exhibits the most significant variation, starting at a relatively low value in the early stage of training and experiencing a sharp decline in the mid stage. This indicates that the simple stacking of convolutional modules may struggle to capture the essential features of the liver region. In contrast, other models achieve a validation DSC above 0.75 in the early stage of training and demonstrate more stable fluctuations. RDeeplab extract rich multi-scale features through the powerful ASPP module, which enhanced the expression of liver features, peaking in the 34th epoch. TransUNet introduces the Transformer block for global dependency modeling of feature maps, achieving a peak value of 0.8835 in the 23rd epoch. SegNeXt and LWBNAUNet enhance key features through carefully designed attention mechanism, achieving their optimal values of 0.8870 and 0.8932 in the 5th and 34th epochs, respectively. UNeXt enriches feature representation through tokenized multi-layer perceptron and outperforms comparative models, reaching its highest value of 0.8971 in the 15th epoch. Finally, BACANet, by incorporating explicit boundary supervision and effective convolutional attention mechanism, is able to further capture the detailed information of liver regions and boundaries. Consequently, it achieved a validation DSC surpassing all comparative models, reaching an optimal value of 0.9207 in the 30th epoch. This indicates that BACANet possesses enhanced learning and generalization capabilities, enabling it to effectively understand the characteristics of liver US images and perform precise segmentation of liver anatomical areas.Fig. 5 Validation DSC over epochs for all models on Dataset A.

Table 2 shows the best predictive results of all models on the validation set of Dataset A, with each model achieving a DSC above 0.86. Further analysis reveals that UNet has a low TPR of 0.854, which suggests that it is deficient in understanding liver region features, resulting in some liver pixels being incorrectly predicted as background. TransUNet, RDeeplab, LWBNAUNet, UNeXt, and SegNeXt exhibit DSC and IOU around 0.89 and 0.80, respectively, indicating their strong perceptual ability for locating the liver body region. However, the higher ASD suggests limitations in their ability to refine the boundaries of the liver. DensePSPUNet was the first model to be experimented on dataset A. Under the same training and validation set conditions, it achieved DSC and IOU of 0.913 and 0.841, respectively, outperforming all comparative models. Nevertheless, BACANet exceled in overall performance, achieving a DSC of 0.921, an IOU of 0.854, and an ASD of 3.783, proving its effectiveness in refining liver boundary details. Additionally, BACANet achieved excellent PPV, TNR, and TPR of 0.920, 0.976, and 0.924 respectively, demonstrating its ability to effectively differentiate between liver regions and the background, achieving precise segmentation in liver US images. Table 2 Best validation results for all models on Dataset A, bolded text represents the best indicators.

Model	DSC	IOU	ASD	PPV	TNR	TPR	
UNet	0.862	0.758	6.251	0.876	0.964	0.854	
TransUNet	0.884	0.793	5.617	0.867	0.960	0.904	
RDeeplab	0.890	0.803	5.141	0.879	0.964	0.904	
LWBNAUNet	0.894	0.809	5.116	0.866	0.959	0.926	
UNeXt	0.898	0.817	4.911	0.880	0.963	0.920	
SegNeXt	0.889	0.803	5.761	0.852	0.952	0.932	
DensePSPUNet	0.913	0.841	–	–	0.979	0.929	
BACANet	0.921	0.854	3.783	0.920	0.976	0.924	

Dataset A comprised only eight volunteers, and the validation dataset included data from just two individuals, potentially limiting the comprehensive assessment of the model's performance. To fully utilize all experimental data and holistically evaluate the model, a fourfold cross-validation method was employed. Table 3 shows the amount of training and validation image data for each fold of the experiment, along with the average performance results of BACANet on the validation dataset. The results indicate that BACANet consistently achieved a DSC above 0.91 and an IOU above 0.84 in each fold. The best performance metrics reached were a DSC of 0.941 and an IOU of 0.889. The average DSC and IOU from the fourfold cross-validation were 0.925 and 0.862, respectively, with standard deviations of 0.013 and 0.023. This demonstrates the model's stable performance and excellent generalization capability. Overall, the experimental results on Dataset A robustly confirm BACANet's exceptional performance in precisely segmenting liver US images and detailing liver boundaries. Table 3 The results of the fourfold cross-validation on Dataset A for BACANet.

Fold	Train	Valid	DSC	IOU	ASD	PPV	TNR	TPR	
1	1794	578	0.941	0.889	3.373	0.964	0.986	0.921	
2	1778	594	0.931	0.873	3.837	0.936	0.982	0.929	
3	1772	600	0.913	0.841	4.372	0.896	0.973	0.934	
4	1772	600	0.915	0.845	5.021	0.885	0.956	0.952	
Mean	–	–	0.925	0.862	4.151	0.920	0.974	0.934	
Std	–	–	0.013	0.023	0.709	0.036	0.013	0.013	

Quantitative analysis of Dataset B

This section presents a quantitative analysis of all experimental models on Dataset B, including an analysis of the variations in different losses of BACANet. It also provides the predictive metrics results for all models on both the validation and test set. Furthermore, ablation experiments are conducted to verify the impact of different modules on the overall segmentation performance of BACANet. Finally, a brief discussion on the model’s parameter scale and inference speed is included.

Figure 6 illustrates the variations in different losses of BACANet during the training process. It is observed that all losses gradually decrease and stabilize as the number of training epochs increases. The total loss curve shows periodic fluctuations due to the adoption of a cosine annealing schedule for learning rate adjustments. After 100 training epochs, the body cross-entropy loss and dice loss decreased to their lowest points at 0.0515 and 0.0267, respectively. The boundary binary cross-entropy loss and mean squared error loss decreased to their lowest at 0.0336 and 0.0075, respectively, with the total loss reaching a minimum of 0.0773. These results indicate that each component of the loss contributes to the optimization of the overall total loss. Particularly, the body loss plays a dominant role, while the boundary loss serves a supportive function.Fig. 6 Individual loss variations of BACANet during the training process.

Figure 7 presents the average performance of various models on the validation dataset of Dataset B in the form of a radar chart. It is noteworthy that among the comparative models, the UNet model exhibits the least favorable performance, with a DSC of 0.9079 and an IOU of 0.8384. This indicates that UNet has limited capability in feature extraction and a weaker generalization ability on unfamiliar data. In contrast, SegNeXt performs notably better, effectively capturing multi-scale global feature information, with its DSC and IOU reaching 0.9293 and 0.8726, respectively. In terms of overall performance, the BACANet model stands out as the best, achieving the highest levels of DSC and IOU at 0.9456 and 0.8987, respectively, demonstrating its robustness and effectiveness in segmentation tasks.Fig. 7 Best validation results of different models on Dataset B.

Table 4 illustrates the predictive performance of various experimental models on test set of Dataset B. It is observed that all models achieved a DSC of over 0.90. UNet performed the worst with a DSC of 0.903, an IOU of 0.834, and an ASD of 4.691. The other comparison models all achieved DSC above 0.92 and IOU above 0.85, but their higher ASD values indicate deficiencies in refining liver boundary segmentation. BACANet achieved the best predictive performance across all segmentation metrics, with a DSC of 0.950, an IOU of 0.907, and an ASD of 2.075. Compared to UNet, BACANet showed improvements of 0.047 and 0.073 in DSC and IOU, respectively, and a reduction in ASD by 2.616. These experimental results demonstrate BACANet’s predictive capabilities on unknown liver US images, effectively achieving precise segmentation of both the liver's main body and its boundary. Table 4 Test results for all models on Dataset B, bolded text represents the best indicators.

Model	DSC	IOU	ASD	PPV	TNR	TPR	
UNet	0.903	0.834	4.691	0.922	0.984	0.899	
TransUNet	0.920	0.859	3.964	0.927	0.984	0.924	
RDeeplab	0.922	0.859	3.718	0.938	0.986	0.912	
LWBNAUNet	0.929	0.874	3.182	0.923	0.982	0.945	
UNeXt	0.925	0.866	3.402	0.931	0.985	0.927	
SegNeXt	0.933	0.879	2.778	0.944	0.987	0.930	
BACANet	0.950	0.907	2.075	0.947	0.988	0.957	

Table 5 presents the results of ablation experiments conducted on core modules within BACANet. The baseline configuration refers to a lightweight UNet model using ResNet10t as the backbone, from which the boundary decoder, MDCAM, and EAG were removed. These core modules proposed in the paper were incrementally added to the baseline, and their improved models were trained and evaluated on the test dataset. The results show that after explicitly introducing a boundary supervision branch, the model’s DSC and IOU improved by 0.013 and 0.024 respectively, demonstrating the auxiliary role of the boundary decoder in the final segmentation of the liver main body. With the introduction of MDCAM, the model’s ability to capture global features was enhanced, further increasing the DSC and IOU to 0.939 and 0.889. Finally, by incorporating boundary features into the decoder through the EAG, the liver main body decoder’s ability to locate the liver area and refine liver boundaries was effectively improved, reaching optimal levels of DSC, IOU, and ASD at 0.950, 0.907, and 2.075 respectively. Table 5 Results of ablation experiments on the core module of BACANet.

Model	DSC	IOU	ASD	
Baseline	0.913	0.848	4.253	
Baseline + boundary decoder	0.926	0.872	2.988	
Baseline + boundary decoder + MDCAM	0.939	0.889	2.701	
BACANet	0.950	0.907	2.075	

Table 6 shows the parameter count P (in millions, M) and the inference time per image on CPU (in seconds, s) and GPU (in milliseconds, ms) for different models. For BACANet, only the parameter required for the inference of the liver's main body is listed, as it does not need to perform the forward propagation of the boundary branch during inference. It can be observed that UNet and TransUNet have larger parameter count, at 31.03 M and 105.27 M respectively, with corresponding TCPU of 2.085 s and 3.008 s. Therefore, they may not be suitable for real-time liver US segmentation tasks under resource-constrained conditions. In contrast, the parameter count of BACANet is 7.56 M (around 24% of UNet), and its corresponding TCPU is 0.32 s (around 15% of UNet). It is noteworthy that UNeXt has the fewest parameter at 1.47 M and the shortest TCPU at 0.093 s. However, the overall performance of UNeXt on the test set is still inferior to BACANet, with a lower DSC and coarser boundaries. The experimental results indicate that in the task of liver US segmentation, BACANet outperforms the other models, successfully balancing inference speed and accuracy. Additionally, except for TransUNet and SegNeXt, the inference time per image on GPU for the remaining models is within 5 ms. It is worth noting that TGPU is associated with multiple factors, including the deep learning framework, GPU version, and model architecture. In real-time scenarios, powerful computational resources may not be available, so the listed TGPU values are provided for reference only. Table 6 Parameter count and inference time for different models.

Model	P	TCPU	TGPU	
UNet	31.03	2.085	1.226	
TransUNet	105.27	3.008	12.881	
RDeeplab	15.31	0.465	2.053	
LWBNAUNet	2.95	1.993	4.425	
UNeXt	1.47	0.093	3.241	
SegNeXt	4.22	0.192	10.308	
BACANet	7.56	0.320	4.015	

Qualitative analysis of Dataset B

This section presents the qualitative analysis results of all experimental models on Dataset B. We first visualized the predicted results for the liver's main body, and then discuss the quality of the feature maps generated within the models.

Due to the presence of speckle noise, shadows, and other factors in US images, accurately locating the liver region and refining anatomical boundaries is a challenging task. Furthermore, in liver and gallbladder US examinations, the technician will employ various scanning planes, such as longitudinal, transverse, and oblique views, in order to obtain detailed images of the corresponding locations. As shown in Fig. 8, the first and second rows of images focus on the right and left lobes of the liver, respectively, while the third row of images emphasizes the observation of the gallbladder, resulting in an incomplete depiction of the adjacent liver morphology.Fig. 8 Visualization of the predicted results on test images for different models. The red region indicates the liver mask, the blue region represents the predicted results, and the green area shows the overlap between the two.

Upon further analysis of segmentation results, most models can roughly locate the liver region but show variations in detailed predictions. In the first row of images, the liver occupies a large proportion of the image, with an unclear boundary in the upper right corner, leading to more false negatives (red areas) in comparative models. In contrast, BACANet accurately positions the liver early in the network and employs effective supervision through SLKM, identifying liver features deeply in the network layers, resulting in superior segmentation. In the second row, UNet, TransUNet, and RDeeplab exhibit more false positives (blue areas), possibly due to redundant feature extraction at larger scales, impacting segmentation accuracy. BACANet's explicit supervision strategy effectively addresses the challenging liver boundary in the upper right, leading to the most accurate segmentation. The third row primarily observes the gallbladder, which has distinct morphological characteristics and contrast with the liver, enabling all models to achieve relatively accurate segmentation. Nonetheless, BACANet demonstrates the most precise delineation of the liver boundary, producing the most complete and smooth segmentation outcome.

Figure 9 shows the internal feature maps of different models during the decoding of liver US images, revealing varying levels of understanding of the liver's main body among these models. UNet roughly identifies the liver region but exhibits indistinct boundaries. In contrast, SegNeXt and UNeXt provide more precise localization and clearer boundaries by capturing contextual information from a multi-scale perspective, enhancing their understanding of the liver's characteristics. Furthermore, BACANet not only gathers multi-scale global information through MDCAM but also effectively captures liver boundary features using boundary supervision and the EAG module. As depicted in Fig. 9, BACANet's internal feature maps not only highlight the liver body region but also emphasize liver boundaries, which is absent in the comparative models. Analyzing these internal feature maps provides insights into BACANet's regions of interest for liver segmentation tasks, enhancing interpretability of its predictions.Fig. 9 Visualization of internal feature maps for different models.

Qualitative analysis of liver segmentation results confirms BACANet's superiority in precise liver segmentation, demonstrating superior visual outcomes compared to other models. Observing the internal feature maps further reveals BACANet's deeper understanding of liver boundary information, significantly improving its predictive accuracy. These observations underscore BACANet's effectiveness and validate its superiority in liver segmentation tasks.

Conclusion

In this paper, we have proposed a novel lightweight image segmentation algorithm named BACANet, aimed at addressing liver US segmentation tasks in real-time scenarios. Specifically, we designed the MDCAM to extract multiscale global context features, the SLKCM to capture boundary feature information in high-resolution feature maps, and the EAG to enhance the liver main body and boundary information in the shallow feature maps of the network, aiding the decoder in more accurately locating the liver area and refining liver boundary features. The results on multiple liver US datasets showed that BACANet effectively completes the liver US segmentation task, achieving a balance between inference speed and accuracy. On a private dataset, we achieved a DSC of 0.950, an IOU of 0.907, and a PPV of 0.947. The inference time for a single image on CPU processor is 0.320 s. We believed that the above experimental results demonstrate the advantages of BACANet in performing liver ultrasound segmentation in real-time scenarios. Additionally, we conducted a visual analysis of the internal feature maps of BACANet, enhancing the model’s interpretability. In future research, we plan to deepen our work in three areas: 1. Expanding the liver US dataset to improve the model’s generalization ability; 2. Further reducing the model’s parameter size to increase inference speed; 3. Integrating the model into software systems and applying it to clinical practice for real-time US segmentation.

Author contributions

J.W. carried out the innovative design of the model, wrote the code and conducted all the experiments, and wrote the first draft of the paper; W.S. and P.R. revised and embellished the manuscript; Private data labeled and reconciled by Z.L. and H.H.; R.J., H.H. and R.Z. were involved in the processing and analysis of the experimental data. F.L. and X.Z. reviewed the article in detail and made the necessary changes. All authors reviewed the manuscript.

Funding

This research was supported by the Jiangsu Provincial Graduate Student Research and Innovation Program (Serial No. KYCX23_2965), and by the Unveiling & Leading Project of XZHMU (Grant No. JBGS202204).

Data availability

Publicly available dataset used in this study can be found at https://tinyurl.com/DensePSPUNet. Private dataset used in this study can be requested from the corresponding author upon reasonable request.

Competing interests

The authors declare no competing interests.

Publisher's note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

These authors contributed equally: Jiawei Wu and Fulong Liu.
==== Refs
References

1. Moran CM Thomson AJW Preclinical ultrasound imaging—a review of techniques and imaging applications Front. Phys. 2020 8 124 10.3389/fphy.2020.00124
Moran, C. M. & Thomson, A. J. W. Preclinical ultrasound imaging—a review of techniques and imaging applications. Front. Phys. 8, 124. 10.3389/fphy.2020.00124 (2020).10.3389/fphy.2020.00124
2. Ferraioli G Monteiro LBS Ultrasound-based techniques for the diagnosis of liver steatosis World J. Gastroenterol. 2019 25 6053 6062 10.3748/wjg.v25.i40.6053 31686762
Ferraioli, G. & Monteiro, L. B. S. Ultrasound-based techniques for the diagnosis of liver steatosis. World J. Gastroenterol. 25, 6053–6062. 10.3748/wjg.v25.i40.6053 (2019).31686762 10.3748/wjg.v25.i40.6053
3. Ferrero A Ultrasound-guided laparoscopic liver resections Surg. Endosc. 2015 29 1002 1005 10.1007/s00464-014-3762-9 25135446
Ferrero, A. et al. Ultrasound-guided laparoscopic liver resections. Surg. Endosc. 29, 1002–1005. 10.1007/s00464-014-3762-9 (2015).25135446 10.1007/s00464-014-3762-9
4. Zhang L SEG-LUS: A novel ultrasound segmentation method for liver and its accessory structures based on muti-head self-attention Comput. Med. Imaging Graph. 2024 113 102338 10.1016/j.compmedimag.2024.102338 38290353
Zhang, L. et al. SEG-LUS: A novel ultrasound segmentation method for liver and its accessory structures based on muti-head self-attention. Comput. Med. Imaging Graph. 113, 102338. 10.1016/j.compmedimag.2024.102338 (2024).38290353 10.1016/j.compmedimag.2024.102338
5. Song, Y., Elibol, A. & Chong N. Y. Two-path augmented directional context aware ultrasound image segmentation. In 1923 IEEE International Conference on Mechatronics and Automation (ICMA). 1815–1822. 10.1109/ICMA57826.2023.10215672 (2023).
6. Senthilkumaran N Vaithegi S Image segmentation by using thresholding techniques for medical images Comput. Sci. Eng. 2016 6 1 13 10.5121/cseij.2016.6101
Senthilkumaran, N. & Vaithegi, S. Image segmentation by using thresholding techniques for medical images. Comput. Sci. Eng. 6, 1–13. 10.5121/cseij.2016.6101 (2016).10.5121/cseij.2016.6101
7. Zhou S Wang J Zhang S Liang Y Gong Y Active contour model based on local and global intensity information for medical image segmentation Neurocomputing. 2016 186 107 118 10.1016/j.neucom.2015.12.073
Zhou, S., Wang, J., Zhang, S., Liang, Y. & Gong, Y. Active contour model based on local and global intensity information for medical image segmentation. Neurocomputing. 186, 107–118. 10.1016/j.neucom.2015.12.073 (2016).10.1016/j.neucom.2015.12.073
8. Chen X Pan L A survey of graph cuts/graph search based medical image segmentation IEEE Rev. Biomed. Eng. 2018 11 112 124 10.1109/RBME.2018.2798701 29994356
Chen, X. & Pan, L. A survey of graph cuts/graph search based medical image segmentation. IEEE Rev. Biomed. Eng. 11, 112–124. 10.1109/RBME.2018.2798701 (2018).29994356 10.1109/RBME.2018.2798701
9. Ronneberger, O., Fischer, P. & Brox, T. U-net: Convolutional networks for biomedical image segmentation. In Medical image computing and computer-assisted intervention–MICCAI 2015. 234–241. 10.1007/978-3-319-24574-4_28 (2015).
10. Wang R Chen S Ji C Fan J Li Y Boundary-aware context neural network for medical image segmentation Med. Image Anal. 2022 78 102395 10.1016/j.media.2022.102395 35231851
Wang, R., Chen, S., Ji, C., Fan, J. & Li, Y. Boundary-aware context neural network for medical image segmentation. Med. Image Anal. 78, 102395. 10.1016/j.media.2022.102395 (2022).35231851 10.1016/j.media.2022.102395
11. Yadav N Dass R Virmani J Objective assessment of segmentation models for thyroid ultrasound images J. Ultrasound. 2023 26 673 685 10.1007/s40477-022-00726-8 36195781
Yadav, N., Dass, R. & Virmani, J. Objective assessment of segmentation models for thyroid ultrasound images. J. Ultrasound. 26, 673–685. 10.1007/s40477-022-00726-8 (2023).36195781 10.1007/s40477-022-00726-8
12. Yadav N Dass R Virmani J Deep learning-based CAD system design for thyroid tumor characterization using ultrasound images Multimed. Tools Appl. 2024 83 43071 43113 10.1007/s11042-023-17137-4
Yadav, N., Dass, R. & Virmani, J. Deep learning-based CAD system design for thyroid tumor characterization using ultrasound images. Multimed. Tools Appl. 83, 43071–43113. 10.1007/s11042-023-17137-4 (2024).10.1007/s11042-023-17137-4
13. Vaswani, A. et al. Attention is all you need. 10.48550/arXiv.1706.03762 (2017).
14. Dosovitskiy, A. et al. An image is worth 16x16 words: Transformers for image recognition at scale. 10.48550/arXiv.2010.11929 (2020).
15. Chen, J. et al. Transunet: Transformers make strong encoders for medical image segmentation. 10.48550/arXiv.2102.04306 (2021).
16. Long, J., Shelhamer, E., Darrell, T. Fully convolutional networks for semantic segmentation. In Proceedings of the IEEE conference on computer vision and pattern recognition (CVPR). 3431–3440. 10.1109/CVPR.2015.7298965 (2015).
17. Zhou Z Siddiquee MMR Tajbakhsh N Liang J Unet++: Redesigning skip connections to exploit multiscale features in image segmentation IEEE Trans. Med. Imaging. 2019 39 1856 1867 10.1109/TMI.2019.2959609 31841402
Zhou, Z., Siddiquee, M. M. R., Tajbakhsh, N. & Liang, J. Unet++: Redesigning skip connections to exploit multiscale features in image segmentation. IEEE Trans. Med. Imaging. 39, 1856–1867. 10.1109/TMI.2019.2959609 (2019).31841402 10.1109/TMI.2019.2959609
18. Zhang, Q., Cui, Z., Niu, X., Geng, S. & Qiao, Y. Image segmentation with pyramid dilated convolution based on ResNet and U-Net. In Neural Information Processing: 24th International Conference (ICONIP). 364–372. 10.1007/978-3-319-70096-0_38 (2017).
19. He, K., Zhang, X., Ren, S. & Sun, J. Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 770–778. 10.1109/CVPR.2016.90 (2016).
20. Zhou Q LAEDNet: A lightweight attention encoder–decoder network for ultrasound medical image segmentation Comput. Electr. Eng. 2022 99 107777 10.1016/j.compeleceng.2022.107777
Zhou, Q. et al. LAEDNet: A lightweight attention encoder–decoder network for ultrasound medical image segmentation. Comput. Electr. Eng. 99, 107777. 10.1016/j.compeleceng.2022.107777 (2022).10.1016/j.compeleceng.2022.107777
21. Ansari MY Yang Y Meher PK Dakua SP Dense-PSP-UNet: A neural network for fast inference liver ultrasound segmentation Comput. Biol. Med. 2023 153 106478 10.1016/j.compbiomed.2022.106478 36603437
Ansari, M. Y., Yang, Y., Meher, P. K. & Dakua, S. P. Dense-PSP-UNet: A neural network for fast inference liver ultrasound segmentation. Comput. Biol. Med. 153, 106478. 10.1016/j.compbiomed.2022.106478 (2023).36603437 10.1016/j.compbiomed.2022.106478
22. Oktay, O. et al. Attention u-net: Learning where to look for the pancreas10.48550/arXiv.1804.03999 (2018).
23. Yang H Yang D CSwin-PNet: A CNN-Swin transformer combined pyramid network for breast lesion segmentation in ultrasound images Expert Syst. Appl. 2023 213 119024 10.1016/j.eswa.2022.119024
Yang, H. & Yang, D. CSwin-PNet: A CNN-Swin transformer combined pyramid network for breast lesion segmentation in ultrasound images. Expert Syst. Appl. 213, 119024. 10.1016/j.eswa.2022.119024 (2023).10.1016/j.eswa.2022.119024
24. Liu, Z. et al. Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 10012–10022. 10.1109/ICCV48922.2021.00986 (2021).
25. He Q Yang Q Xie M HCTNet: A hybrid CNN-transformer network for breast ultrasound image segmentation Comput. Biol. Med. 2023 155 106629 10.1016/j.compbiomed.2023.106629 36787669
He, Q., Yang, Q. & Xie, M. HCTNet: A hybrid CNN-transformer network for breast ultrasound image segmentation. Comput. Biol. Med. 155, 106629. 10.1016/j.compbiomed.2023.106629 (2023).36787669 10.1016/j.compbiomed.2023.106629
26. Lin, Y. et al. Rethinking boundary detection in deep learning models for medical image segmentation. In International Conference on Information Processing in Medical Imaging (IPMI). 730–742. 10.1007/978-3-031-34048-2_56 (2023).
27. Mishra D Chaudhury S Sarkar M Soin AS Ultrasound image segmentation: A deeply supervised network with attention to boundaries IEEE Trans. Biomed. Eng. 2018 66 1637 1648 10.1109/TBME.2018.2877577 30346279
Mishra, D., Chaudhury, S., Sarkar, M. & Soin, A. S. Ultrasound image segmentation: A deeply supervised network with attention to boundaries. IEEE Trans. Biomed. Eng. 66, 1637–1648. 10.1109/TBME.2018.2877577 (2018).30346279 10.1109/TBME.2018.2877577
28. Wu Y BGM-Net: Boundary-guided multiscale network for breast lesion segmentation in ultrasound Front. Mol. Biosci. 2021 8 698334 10.3389/fmolb.2021.698334 34350211
Wu, Y. et al. BGM-Net: Boundary-guided multiscale network for breast lesion segmentation in ultrasound. Front. Mol. Biosci. 8, 698334. 10.3389/fmolb.2021.698334 (2021).34350211 10.3389/fmolb.2021.698334
29. Sun J TNSNet: Thyroid nodule segmentation in ultrasound imaging using soft shape supervision Comput. Methods Programs Biomed. 2022 215 106600 10.1016/j.cmpb.2021.106600 34971855
Sun, J. et al. TNSNet: Thyroid nodule segmentation in ultrasound imaging using soft shape supervision. Comput. Methods Programs Biomed. 215, 106600. 10.1016/j.cmpb.2021.106600 (2022).34971855 10.1016/j.cmpb.2021.106600
30. Ji Z Che H Yan Y Wu J BAG-Net: A boundary detection and multiple attention-guided network for liver ultrasound image automatic segmentation in ultrasound guided surgery Phys. Med. Biol. 2024 69 035015 10.1088/1361-6560/ad1cfa
Ji, Z., Che, H., Yan, Y. & Wu, J. BAG-Net: A boundary detection and multiple attention-guided network for liver ultrasound image automatic segmentation in ultrasound guided surgery. Phys. Med. Biol. 69, 035015. 10.1088/1361-6560/ad1cfa (2024).10.1088/1361-6560/ad1cfa
31. Wightman, R., Touvron, H. & Jégou, H. Resnet strikes back: An improved training procedure in timm.10.48550/arXiv.2110.00476 (2021).
32. Huang R Boundary-rendering network for breast lesion segmentation in ultrasound images Med. Image Anal. 2022 80 102478 10.1016/j.media.2022.102478 35691144
Huang, R. et al. Boundary-rendering network for breast lesion segmentation in ultrasound images. Med. Image Anal. 80, 102478. 10.1016/j.media.2022.102478 (2022).35691144 10.1016/j.media.2022.102478
33. Misra D. A self regularized non-monotonic activation function.10.48550/arXiv.1908.08681 (2019).
34. Guo MH Lu CZ Liu ZN Cheng MM Hu SM Visual attention network Comput. Vis. Media. 2023 9 733 752 10.1007/s41095-023-0364-2
Guo, M. H., Lu, C. Z., Liu, Z. N., Cheng, M. M. & Hu, S. M. Visual attention network. Comput. Vis. Media. 9, 733–752. 10.1007/s41095-023-0364-2 (2023).10.1007/s41095-023-0364-2
35. Guo, M. H. et al. Segnext: Rethinking convolutional attention design for semantic segmentation. In Advances in Neural Information Processing Systems 35 (NeurIPS). 1140–1156. 10.48550/arXiv.2209.08575 (2022).
36. Hendrycks, D. & Gimpel, K. Gaussian error linear units (gelus).10.48550/arXiv.1606.08415 (2016).
37. Canny J A computational approach to edge detection IEEE Trans. Pattern Anal. Mach. Intell. 1986 6 679 698 10.1109/TPAMI.1986.4767851
Canny, J. A computational approach to edge detection. IEEE Trans. Pattern Anal. Mach. Intell. 6, 679–698. 10.1109/TPAMI.1986.4767851 (1986).10.1109/TPAMI.1986.4767851
38. Milletari, F., Navab, N. & Ahmadi S. A. V-net: Fully convolutional neural networks for volumetric medical image segmentation. In 2016 Fourth International Conference on 3D Vision (3DV). 565–571. 10.1109/3DV.2016.79 (2016).
39. Kim M Lee BD A simple generic method for effective boundary extraction in medical image segmentation IEEE Access. 2021 9 103875 103884 10.1109/ACCESS.2021.3099936
Kim, M. & Lee, B. D. A simple generic method for effective boundary extraction in medical image segmentation. IEEE Access. 9, 103875–103884. 10.1109/ACCESS.2021.3099936 (2021).10.1109/ACCESS.2021.3099936
40. Loshchilov, I. & Hutter, F. Decoupled weight decay regularization.10.48550/arXiv.1711.05101 (2017).
41. Loshchilov, I. & Hutter, F. Sgdr: Stochastic gradient descent with warm restarts. 10.48550/arXiv.1608.03983 (2016).
42. Murugappan M Bourisly AK Prakash NB Sumithra MG Acharya UR Automated semantic lung segmentation in chest CT images using deep neural network Neural Comput. Appl. 2023 35 15343 15364 10.1007/s00521-023-08407-1 37273912
Murugappan, M., Bourisly, A. K., Prakash, N. B., Sumithra, M. G. & Acharya, U. R. Automated semantic lung segmentation in chest CT images using deep neural network. Neural Comput. Appl. 35, 15343–15364. 10.1007/s00521-023-08407-1 (2023).37273912 10.1007/s00521-023-08407-1
43. Valanarasu, J. M. J. & Patel, V. M. Unext: Mlp-based rapid medical image segmentation network. In Medical Image Computing and Computer Assisted Intervention–MICCAI 2022. 23–33. 10.48550/arXiv.2203.04967 (2022).
44. Sharma P A lightweight deep learning model for automatic segmentation and analysis of ophthalmic images Sci. Rep. 2022 12 8508 10.1038/s41598-022-12486-w 35595784
Sharma, P. et al. A lightweight deep learning model for automatic segmentation and analysis of ophthalmic images. Sci. Rep. 12, 8508. 10.1038/s41598-022-12486-w (2022).35595784 10.1038/s41598-022-12486-w
