
==== Front
Sci Rep
Sci Rep
Scientific Reports
2045-2322
Nature Publishing Group UK London

39300076
70570
10.1038/s41598-024-70570-9
Article
A lightweight parallel attention residual network for tile defect recognition
Lv Cheng
Zhang Enxu
Qi Guowei
Li Fei
Huo Jiaofei 95339323@qq.com

https://ror.org/05xsjkb63 grid.460132.2 0000 0004 1758 0275 School of Mechanical Engineering, Xijing University, Xi’an, 710123 China
19 9 2024
19 9 2024
2024
14 2187221 5 2024
19 8 2024
© The Author(s) 2024
2024
https://creativecommons.org/licenses/by-nc-nd/4.0/ Open Access This article is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License, which permits any non-commercial use, sharing, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if you modified the licensed material. You do not have permission under this licence to share adapted material derived from this article or parts of it. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by-nc-nd/4.0/.
In modern industrial production, permanent magnet motors are an indispensable part of industrial manufacturing. The quality of the magnetic tiles directly affects the working performance of the permanent magnet motors, making the detection of defects on the surface of magnetic tiles critically important. However, due to the small size of defects on the tile image and the reflectivity of the defective surface, the details of image characteristics are not prominently acquired.These problems bring a lot of difficulties for the recognition of magnetic tile defects. In this paper, a magnetic tile defect detection method is proposed for the probAlems of unclear image features and small defects. First, the image is processed using linear variation to enhance the image detail features. Then, by introducing the inverted bottleneck block structure in MobileNetV2, the Attention Parallel Residual Convolution Block (APR) is proposed, and the Lightweight Parallel Attention Residual Network (LPAR-Net) is built. In APR Block, 7 × 7 convolution is introduced so that the model can extract spatial features from a larger range, and weighted fusion of input images by residual structure. In addition, in this paper, CBAM is improved, split into two parts and inserted into APR Block. Finally, the mainstream image classification models and the LPAR-Net proposed in this paper are used for comparison, respectively. The experimental results show that the method achieves 93.63% accuracy on the adopted dataset, which is better than the existing mainstream image classification network models DenseNet, MobileNet, ConvNext and so on. In addition, this paper introduces a strip steel surface defect dataset and compares it with the above image classification model, which verifies that the detection method proposed in this paper still has strong recognition capability.

Keywords

Metal surface defects
Deep learning
Machine vision
Attention mechanism
Subject terms

Mechanical engineering
Computer science
issue-copyright-statement© Springer Nature Limited 2024
==== Body
pmcIntroduction

Permanent magnet motors (PM motors) are an essential part of modern industrial manufacturing, and are often used in new energy vehicles1–4, aerospace5–7, household appliances8–10, and other industries. As a key part of permanent magnet motors, they are mainly made of a special material, magnetic tiles11. The quality of magnetic tiles has a direct impact on the performance and service life of the motors, and even minor surface defects may pose a potential threat to public safety. However, during the production process of magnetic tiles, various defects such as cracks, porosity, wear, chipping and other processing defects are inevitably generated due to the quality of raw materials, production processes, equipment conditions, etc12. These defects can reduce the performance of the motor, increase the noise and vibration of the motor, and shorten the service life of the motor, among other effects. Therefore, the detection of minute defects on the surface of magnetic tiles has always been a challenging task. The existing difficulties include: the surface defects of magnetic tiles are too minute; the defect features are difficult to extract and are restricted by the arcuate magnetic tile material and the data set acquisition environment, which makes the acquired magnetic tile data set images have uneven contrast, resulting in difficult feature discrimination. In modern industrial manufacturing, the accurate identification of small defects in magnetic tiles can ensure the working performance of permanent magnet motors, thus meeting the demand for high precision and high performance in modern manufacturing and reducing unnecessary waste. Therefore, it is of great significance to realize the fast and accurate detection of surface defects on magnetic tiles.

The poor subjectivity of traditional manual visual inspection and the tiny defects of magnetic tiles affect the accuracy and efficiency of magnetic tile defect detection13,14. The gradually increasing labor cost greatly limits the development of enterprises. With the rapid development of computer vision bringing a new direction to defect detection, deep learning-based computer vision algorithms provide a new solution for magnetic tile defect detection. In the early days, machine learning-based defect detection was widely used15. However, traditional machine learning algorithms require manually designed feature extraction methods, which make it difficult to differentiate effectively in the face of the challenge of tiny and complex defects. With the development of deep learning technology, which extracts features by means of a multilevel neural network architecture, it has great advantages in image recognition16,17. Deep learning-based image recognition has become a new way of defect detection for magnetic tiles. Among them, convolutional neural networks (CNNs) have made great progress in defect detection of steel, bearings, turning inserts, etc., by extracting features from the input image through multilevel convolution and pooling operations18,19. CNNs have been widely regarded as one of the best algorithms for image classification tasks20. Therefore, the use of CNNs for detecting defects generated in industrial manufacturing has become an industrial research hotspot for intelligence21. These methods have good performance in detecting large defects on surfaces. However, they do not perform nearly as well for small defects in tiles. This is due to the structural differences of different network models and the inconspicuous features of the magnet tile image, which make the recognition effect have a large difference. This suggests that small defects on the surface of magnetic tiles are not only common, but also have a significant impact on the results.

This paper proposes a novel method for detecting surface defects on magnetic tiles to overcome the above challenges. To address the issue of uneven contrast in magnetic tile images caused by the acquisition environment, data augmentation is applied to the original acquired images to optimize the images and enhance defect details. To solve the problem that neural networks have difficulty in identifying minute defects on the surface of magnetic tiles, a network model is designed specifically for the task of classifying defect images on the surface of magnetic tiles, which is better able to identify minute defects. This improves the accuracy of magnetic tile surface defect detection and achieves better results in practical applications. Even in scenarios with complex image backgrounds and uneven contrast, it can still identify defects on magnetic tiles. Therefore, this paper designs a lightweight parallel attention residual network to identify defects on magnetic tiles.

Our main contributions are summarized below:A data enhancement method is designed to adjust the contrast and brightness in the graph by linear changes to highlight the metal surface defect features in the image.

A new lightweight neural network (LPAR-Net) model is proposed for the recognition of magnetic tile defects. The model is constructed using a well-designed APR to improve the feature extraction capability of the network model for magnetic tile defects.

Attention Parallel Residual Block APR is proposed to reduce the number of model parameters and enable the model to locate the defective region quickly and accurately.

The model performance is evaluated using Magnetic Tile Surface Defects, NEU-DET dataset, and it is verified that the model proposed in this paper has a strong capability of recognizing defects on metal surfaces.

The structure of this paper is summarized as follows: Sect. 2 describes the work related to the detection of defects in magnetic tiles. Then, Sect. 3 describes the materials and methods. Then, Sect. 4 presents the analysis of the experimental results, which verifies the superiority of the methodology of this study. Finally, Sect. 5 summarizes the whole paper.

Related work

In recent years, with the development of science and technology, the traditional manufacturing industry has been transformed to intelligent manufacturing, and many researchers and developers have begun to explore how to effectively carry out defect detection of industrial products. Machine learning techniques are widely used in defect recognition of industrial products.

Huang et al.22 used wavelet packet transform (WPT), linear discriminant analysis (LDA) and support vector machine (SVM) to detect internal defects in magnetic tiles. WPT was applied to extract the normalized features of the image and SVM was optimized to extract the features based on LDA and constraint algorithm to identify the defects in the magnetic tiles. Zhang et al.23 proposed an incremental support vector machine (MEISVM) intelligent fault diagnosis system based on multivariate integration, which can effectively correlate multiple monitoring variables with the corresponding defect types. Chu et al.24 proposed a machine learning with discretized hyperspheres to achieve multi-category classification of steel surface defects using a novel classifier with multiple discretized hyperspheres. Wang et al.25 used a diffusion algorithm to classify and recognize the number of defects, a support vector machine as a detection method to extract the location and shape of surface defects, and a covariance matrix 3D measurement method to calculate the size of defects. However, the recognition rate is not high enough for the classification of large datasets. Yi-Bei Huang et al.26 segmented the collected sample images after preprocessing, extracted gray scale, shape, and HOG feature values, and then used SVM and KNN for classification, with an average recognition rate of 92.6%. Zhao et al.27 used a Gabor filter to extract texture information from images of thermoelectric cooler (TEC) components, while principal component analysis (PCA) was used to select classification features, and then a support vector machine was used to classify TEC defects. Choi et al.28 designed the machine learning iterative filtering (MLIF) algorithm, which performs a comprehensive inspection of domestic gas boilers at the process stage by forming n learners through multiple random undersampling in the majority category of defect-free products. From the above literature, it can be found that traditional machine learning methods mainly extract the defect features in the visible region through the size and shape features of the defects and, finally, classify them through various classifiers.

In recent years, deep learning has made significant development in the field of machine vision, and defect detection using CNN has become a hotspot for the development of intelligent manufacturing. Cui et al.29 proposed a fast and accurate surface defect detection network called SDDNet, which solves the challenging problems of large texture variations and small defect sizes by introducing two modules, mainly the feature-fixed block and the skip-dense connectivity module, and achieves an accuracy of 93.4% in the recognition of defects in magnetic tiles. Hu et al.30 proposed a UPM-DenseNet two-stage detection model, which can eliminate the influence of complex background within the feasible range, and its recognition accuracy can reach a high level after the experiment. Lu et al.31 designed a multimodal fusion convolutional neural network (MMFCNN) for the detection of internal defects in tiles, which extracts features from the generated modal data, fuses multimodal feature maps and analyzes the tiles for the presence of internal defects and introduces a cross-attention mechanism. Liang et al.32 proposed a new framework called feature enhancement and ring fusion convolutional neural network, which enhances shallow features and fuses the features with a ring feature pyramid structure for automatic detection of surface defects on magnetic tiles. Yu et al.33 proposed an efficient scale-aware network (ES-Net) to improve defect detection, which improves the detection of small defects by solving the problems of information loss of tiny targets and the mismatch between the sensing field of the detection head and the scale of the target. Xiao et al.34 proposed a global receptive attention network (GRA-Net) for surface defect detection, which improves the global representation capability, expands the receptive field, and is able to capture tiny targets with improved accuracy over existing models. Wan et al.35 proposed a lightweight semantic-aware multilevel feature interaction network (SMINet) that utilizes two attentional strategies to perceive contextual information to recover more spatial details and validated the network's sophistication on four surface-deficient datasets. Zhu et al.36 proposed for the first time to introduce rotational invariance (RI) into convolution by rotating the feature map, and built a lightweight neural architecture, CNN with RI (RICNN), which obtains better classification performance while requiring much lower computational cost. Yun et al.37 proposed a new Convolutional Variational Autocoder (CVAE) and deep CNN-based defect classification algorithm with a DCNN-based classifier that enables high generalization performance of the generated data and exhibits excellent performance for defect images obtained from real metal production lines. Li et al.38 introduced a novel convolutional retinal attention block containing three modules: a multiresolution module, a global attention aggregation module, and a local attention aggregation module, and verified its effectiveness in multiple backbone networks. Guo et al.39 designed a Dilated Swin Transformer UNet (DSUNET) model with privacy-preserving features to capture global and remote semantic information, introduced a decentralized federated learning framework to protect data privacy, and achieved an accuracy of 97.98% on small target defects in heat sinks.

In order to enhance the network model's attention to input features, attention mechanisms are widely used in deep learning. Zhang et al.40 proposed an edge-guided and differential attention network (EGD-Net) whose differential attention module performs top-down attention, which can effectively eliminate the clutter in the background region caused by connected edge features. Shi et al.41 proposed a deep self-attentive residual convolutional neural network (DSARNet) for automatic identification of surface defects in aluminum profiles, which uses ResNet as the backbone and introduces a self-attention (SA) mechanism to aggregate global deep features with an accuracy of 91.99%. Yu et al. proposed a new deep learning detection network that implements channel attention and bi-directional feature fusion on a fully convolutional single-stage (CABF-FCOS) network, which reduces the loss of feature information with an average accuracy of 76.68%. Zhang et al.42 combined the methodology of the attention module CBAMC3, which utilizes the CBAM and C3 modules, with the analysis of the module BiFPN_concat, which enables the network to extract and fuse defective features separately. Xiang et al.43 improved the mAP (0.5) of the industrial surface defect dataset from 88.3% in Yolov5 to 89.2% by increasing the resolution of the feature layer and introducing a CBAM attention mechanism in the feature aggregation network to better integrate the features at different scales and utilize the information of small defects.

Materials and methods

Image sources and pre-processing

Sample data set

The images used in this study are Magnetic Tile Surface Defects collected by the public dataset website Kaggle. The dataset consists of a total of 1344 grayscale images, including both defective and non-defective images, categorized into six types: Blowhole, Break, Crack, Fray, Uneven, and Free. The number of images for each category is 115, 85, 57, 32, 103, and 952, respectively. Among them, there are 952 non-defective images and 392 defective images. The number of images of the defect-free type is much larger than that of the other five defect types, and the image sizes of each defect type are different. The five different defect types of magnetic tiles are shown in Fig. 1.Fig. 1 Samples of magnetic tiles with five defects and one without defects.

Image preprocessing

In the training of convolutional neural networks, the quality of image datasets significantly impacts the training effectiveness of the model. However, due to factors such as environmental conditions and human operations, the dataset may suffer from uneven illumination, making the features in the images less distinct. This can affect the learning efficiency of the network model during training and further impact the training effectiveness of the model44. To address this issue, this paper proposes an image enhancement method that uses linear transformation to process images. This method can unify the brightness of different images in the dataset and improve the learning efficiency of neural network models. The calculation formula can be expressed as:1 O(r,c)=a∗I(r,c)+b,0≤r<H,0≤c<W

where I is the image, W and H denote the width and height of the input image respectively, O is the output image and a is the threshold value. As shown in Fig. 2.Fig. 2 Example of data enhancement.

Convolutional neural network models often require more samples to effectively learn different categories of image features. However, influenced by realistic image acquisition, there may be some categories in the dataset with fewer samples and some with an excessive number of samples. This will cause the convolutional neural network model to overlearn the categories with a large number of samples, triggering an overfitting phenomenon and making it difficult to correctly recognize the categories with fewer samples. To solve this problem, data enhancement techniques are often used to expand the number of images in the original dataset. Common data enhancement methods include operations such as rotating, scaling, panning, mirror flipping, cropping, and contrasting images45. In this study, data enhancement techniques are used to process the above dataset that has undergone image enhancement to expand the number of defective images for the five categories. The augmented dataset is divided into training set and test set with a ratio of 8:2. This ratio division can meet the needs of training the model, facilitate the model to learn the features and patterns in the data, and avoid overfitting. The training and testing quantities of different types of magnetic tile defects are shown in Fig. 3.Fig. 3 Dataset segmentation of tile defect images.

LPAR-Net model

In this paper, a Lightweight Parallel Attention Residual Network (LPAR-Net) is proposed in order to recognize metal surface defects effectively.The structure of LPAR-Net model is shown in Fig. 4, which mainly consists of APR Block, Down Block and Classifier The model is mainly composed of APR Block, Down Block and Classifier. In this model, the input image is first downsampled using a 3 × 3 convolutional layer and a 3 × 3 maximum pooling layer to reduce the image size and thus reduce the model volume. Then the image features are extracted using APR Block and the image is further downsampled using Down Block. Finally, the extracted features are pooled using Adaptive Mean Pooling to retain the feature distribution in the feature map and a classifier is used to obtain the category output.Fig. 4 Structure of LPAR-Net.

Attention parallel residual block

The APR Block in this paper is mainly improved based on the inverted bottleneck block, whose structure is shown in Fig. 5. The residual structure was first proposed in ResNet networks to effectively minimize overfitting and the gradient explosion scenarios generated46. In addition, in MobileNet, the inverted bottleneck block adopts a residual structure that can effectively preserve the original information of the image and achieve lightweighting47. Figure 5a shows the original inverted bottleneck block structure, and Fig. 5b shows the APR Block proposed in this paper.The inverted bottleneck block was first applied in MobileNetV2, as a replacement of the traditional convolution, which can reduce the amount of computation required for model training and reduce the model parameters. However, although the inverted bottleneck block reduces the amount of computation, its complexity is also relatively high due to the increased number of convolutional layers. In addition, the inverted bottleneck block may sacrifice some of the model's accuracy and expressiveness while reducing the parameters, especially on complex tasks or large-scale datasets, where the performance may not be as good as that of a larger or more complex network structure. For this reason, this paper improves the existing inverted bottleneck block.Fig. 5 Construction of APR module. (a) represents the Bottleneck Residual Block, (b) represents the APR Block.

Firstly, to expand the sensory field of the model, parallel 7 × 7 convolutional layers and BatchhNorma, ReLU activation function and composition of NAPR Block (No-Attention Parallel Residual Block) are added. By this way it can help the model to take multiple global features and enhance the model recognition ability. The large convolutional kernel has a larger sensory field and is able to obtain long distance dependencies at different locations in the image space and more detailed features in the image compared to the traditional 3 × 3 convolutional kernel in the CNN model. In this paper, we are inspired by the multi-scale inverse bottleneck residual network48. Therefore, the APR Block, which is a combination of 3 × 3 convolution and 7 × 7 convolution, is used to replace the Block structure in the original ResNet-18 network model, in order to help the model capture the contextual information and realize the accurate recognition of defects. Then the CBAM attention mechanism is improved as shown in Fig. 5b. The APR block can be divided into two parts, the left side and the right side, to obtain the attention weight of the image. The left side consists of two layers of point-by-point convolution, channel attention mechanism and one layer of depth convolution with ReLU6 activation function and BatchNorm normalization. the channel attention mechanism acquires the image channel weights while helping the image channels to be integrated. The right component is a layer of 7 × 7 large convolutional kernel and spatial attention mechanism with ReLU6 activation function and BatchNorm normalization, which is capable of extracting the spatial features of the image over a large receptive field and obtaining the spatial weights of the image through spatial attention. Finally, the obtained channel weights and spatial weights are weighted and fused to the output image.

Improvements in attention mechanisms

CBAM (Convolutional Block Attention Module) is a type of attention module for convolutional neural networks that infers the attention map along two independent dimensions (channel and space) and then weights the attention map against the input feature map for adaptive feature refinement. This approach can help the model to extract image features efficiently and thus improve the model's recognition performance. In this paper, inspired by this research, two independent dimensions of CBAM are split and used.

Channel attention module

The channel attention mechanism is a compression of the spatial dimensions. As shown in Fig. 6. First, average and maximum pooling of spatial dimensions is performed on the input feature maps to obtain two feature maps.2 Favgc=GAPc(X)

3 Fmaxc=GMPc(X)

where GAP and GMP denote the average pooling and maximum pooling operations in the channel domain, respectively.Fig. 6 Mechanism of channel attention in CBAM.

The results of average pooling and maximum pooling are then fed into a shared multilayer perceptron to learn two feature maps, respectively. The MLP has a number of neurons in the first layer as C/r and an activation function as Relu, and a number of neurons in the second layer as C. The neurons can be denoted as:4 aj=σ(∑i=1nwijxi+bj)

where aj is the output of the neuron, σ is the activation function, wij is the weight between the i-th input and the j-th output, xi is the i-th input, bj is the bias of the j-th output, n is the number of inputs.

Finally, the MLP outputs are subjected to a summing operation, followed by a mapping process with a Sigmoid activation function to finally obtain the channel attention weight matrix Mc.5 Mc(F)=σ(MLP(AvgPool(F))+MLP(MaxPool(F)))=σ(W1(W0(Favgc))+W1W0(Fmaxc)))

where Favgc and Fmaxc denote the average and maximum pooling characteristics, respectively.

Spatial attention module

The spatial attention mechanism is to compress the channel. This is shown in Fig. 7. To compute spatial attention, we first apply average pooling and maximum pooling operations along the channel axes, and then concatenate them to generate a comprehensive feature descriptor.6 Favgs=GAPs(X)

7 Fmaxs=GMPs(X)

where GAP and GMP denote the global average pooling and global maximum pooling operations in the spatial domain, respectively. Then, we apply a convolutional layer on this combined feature descriptor to generate a spatial attention map Ms(F)∈RH×W, which encodes information about locations to be emphasized or suppressed in the feature map. In the following, we perform the detailed operations.Fig. 7 Spatial attention mechanism in CBAM.

Two 2D maps are generated by aggregating the channel information of the feature maps using two pooling operations: Favgs∈R1×H×W and Fmaxs∈R1×H×W, which represent the average pooled features and the maximum pooled features for the entire channel. These two feature maps are then connected and processed through 7 × 7 convolutional layers to generate our 2D spatial attention map. Briefly, the spatial attention is computed as follows:8 Ms(F)=σ(f7×7([Avgpool(F);MaxPool(F)]))=σ(f7×7([Favgs;Fmaxs]))

where σ denotes the sigmoid function and f7×7 denotes the convolution operation with filter size 7 × 7.

We add the attention mechanism (SAM) after the 7 × 7 convolutional layer, and this design not only ensures the lightweight of the network, but also makes the resource allocation more reasonable, which effectively improves the network performance. Moreover, the spatial attention module can highlight important regions in defective images and suppress irrelevant regions.

Flow chart of tile defect identification method

The flowchart of the whole model processing of our study is shown in Fig. 8. In this study, the image recognition model LPAR-Net is constructed based on the CNN model.Because a large number of datasets are required for deep neural network model training, this study uses the publicly available dataset from the Kaggle website as the original dataset. Since the contrast of images in the dataset is inconsistent, image enhancement is performed on the training data using linear variation. Then the dataset is expanded by rotating, panning, mirroring and other processing methods. For the expanded dataset, we divided it into 80% training set and 20% test set. The images in the training and test sets are preprocessed.Fig. 8 Flow chart for identifying tile defects.

Experimental results and discussion

Experimental design

The experimental hardware for this study was Windows 11 operating system and Intel Core i7-12700 processor (2.10 Ghz) with NVIDIA Geforce RTX3060 (12 GB) GPU model. The software environment used python 3.9.13,pytocrh 1.13.1 and cudn11.6 framework. During training, the training parameters for all models were: the learning rate was set to 0.001 (Aadm's default learning rate), the number of iterations was epochs = 100, and the number of input images per batch was batch size = 32.

The experiments are divided into four parts: first, comparing different network models; second, comparing different attentional mechanisms: third, ablation experiments; and fourth, comparing experiments on different datasets.

Evaluation indicators

In this study, Accuracy, Precision, Recall, F1 Score, Params and Flops are utilized to measure the performance of the model for the recognition of magnetic tile defects. The computation of these parameters is shown below:9 Accuracy=TP+TNTP+TN+FP+FN

10 Precision=TPTP+FP

11 Recall=TPTP+FN

12 F1=2×precision×recallprecision+recall

where TP (True Positive): denotes the number of samples correctly identified by the model, FP (False Positive): denotes the number of negative samples incorrectly identified by the model, TN (True Negative): denotes the number of negative samples correctly identified by the model, FN (False Negative): denotes the number of positive samples incorrectly identified by the model.

Comparative experiments with different network models

In this study, in order to effectively evaluate the recognition ability of LPAR-Net model, respectively and ResNet18, ConvNextV2, DenseNet121, MobileNetV2, EfficientNet, ShuffleNetV2, where the large volume model includes ResNet18, ConvNextV2, DenseNet121, and the lightweight model includes MobileNetV2, EfficientNet, ShuffleNetV2. The experimental data is shown in Table 1, the change curve of image accuracy during training is shown in Fig. 9, and the confusion matrix test is shown in Fig. 10.Table 1 Comparison experiment of evaluation parameters of different network models

Model	Accuracy	Precision	Recall	F1-score	Params (M)	Flops (G)	
ResNet18	0.9052	0.9068	0.8963	0.9008	11.18	1.82	
ConvNextV2	0.8973	0.8931	0.8874	0.8902	27.79	4.46	
DenseNet121	0.9273	0.9282	0.9185	0.9233	7.89	2.83	
MobileNetV2	0.8958	0.8946	0.8865	0.8905	3.47	0.30	
EfficientNet	0.9022	0.9016	0.8973	0.8994	5.25	0.39	
ShuffleNetV2	0.9051	0.9047	0.8955	0.9000	2.27	0.15	
LPAR_Net	0.9363	0.9406	0.9306	0.9354	1.66	0.24	

Fig. 9 Accuracy curves for different model comparisons. a shows the comparison of the curves of this study's model with the heavyweight network model, and b shows the comparison of the curves of this study's model with the lightweight model.

Fig. 10 Confusion matrix plots for different network models. a–g ResNet18, ConvNextV2, DenseNet121, MobileNetV2, EfficientNet, ShuffleNetV2, LAPR-Net, respectively.

As shown in Table 1, compared to the heavyweight networks (ResNet18, ConvNextV2, and DenseNet121), the proposed network model is much lower than the three networks mentioned above in terms of Floating Point Operations (Flops) and the number of parameters (Params). The model in this study is improved based on the ResNet18 network architecture, and the four parameters of Accuracy, Precision, Recall, and F1 scores obtained are higher than those of ResNet18.The model parameters (Params) of DenseNet121 are the lightest among the three large volume models, only 7.89 M. However, it improves the classification accuracy of the model by connecting each layer with all the previous layers for feature reuse, which helps the model to capture the background noise information efficiently in complex dataset environments, but the model has 2.83G Flops.LPAR-Net achieves 0.9363, 0.9406 in terms of average classification accuracy, precision, recall, and F1 score, respectively, 0.9306, 0.9354 all achieved the highest values.

Compared to lightweight models such as MobileNet-V2, EfficientNet, and ShuffleNetV2, the Params values of the proposed model in this study are lower than those of other lightweight models, demonstrating better recognition capabilities. MobileNetV2 benefits from its innovative architecture design, which effectively reduces the number of parameters and Flops values, achieving accuracy comparable to larger models while maintaining a lightweight structure. ShuffleNetV2 reduces the model parameters through channel splitting and shuffling, resulting in lower Flops than LPAR-Net, but the Params value is still higher than that of the model proposed in this study. EfficientNet exhibits low Flops values through a rational scaling strategy and a lightweight module structure.

In order to have a clear picture of the recognition effectiveness of the seven models, a confusion matrix was plotted for the dataset of this study as shown in Fig. 10. Blowhole, Break, Crack, Fray, and Uneven represent the five types of defects of the tiles, and Free denotes the defect-free tiles. They are the five different types of defects shown in Fig. 1 and one no defect. Because Uneven defects are more obvious and Free is a defect-free tile, the LPAR-Net model in this study achieves higher accuracy in correctly identifying these two categories. The model has less number of misidentified defects than the other six models in all cases. Since both defects, Crack and Blowhole, produce black defect points, which are only a little different in the shape and size of the defect points and the color shades, this caused some network models to misidentify Crack as Blowhole. Especially in Break defect recognition, some of these defects are less obvious and similar to both defects, Blowhole and Crack, causing some confusion. Due to the complexity of the background of the test set images, this resulted in the failure of the other six networks to achieve more effective recognition. This proves that the network model proposed in this paper has better recognition performance.

Comparison of different attention mechanisms

In order to better test the effects of different attention mechanisms on the model, the method proposed in this paper is compared with two attention mechanisms, CBAM attention and NAM attention, respectively. All three attention mechanisms are lightweight and do not significantly increase the number of parameters of the model. The recognition results are shown in Table 2. It can be clearly seen that the accuracy of the LPAR-Net model with APR module proposed in this Q&A is higher than that with NAM and CBAM modules by 2.14% and 2.11%, respectively, when compared with the models using CBAM module and NAM module. Especially, the accuracy of recognizing defects in small samples of magnetic tiles is higher.Table 2 Comparison of results of different attention mechanisms

Attention mechanism	Accuracy	Precision	Recall	F1-score	Params (M)	
NAM	0.9149	0.9164	0.9080	0.9104	1.43	
CBAM	0.9152	0.9196	0.9065	0.9090	1.44	
APR	0.9363	0.9406	0.9306	0.9354	1.66	

The class activation maps of the Grad-CAM visualization model image the visualization results of the output features of each attention mechanism, and the visualization image of the magnet tile defect image and its heat map superimposed is shown in Fig. 11. The first image shows the Blowole defects in the magnetic tile dataset, and the defects in the image are relatively simple, but the background is dark. The improved APR Block in this paper was able to basically recognize the defect location, while the NAM attention and CBAM attention did not accurately find the defective region of the magnetic tile, resulting in misidentification. In the second image, the defect of the magnetic tile is Crack, which is difficult to detect because the defect is on the far right side of the image and resembles a thin straight line. The image background is complex and has strong interference.NAM attention can recognize where the defect is, but it also focuses on other regions in the image, resulting in misidentification; CBAM attention does not completely focus on the defective region in the image.APR Block effectively improves the anti-interference against the complex background, and has the ability to recognize the defective part of the magnetic tile effectively. The experimental results show that the model in this study can better find the defective parts, reduce the influence of complex background, highlight the defective features more accurately, and help to improve the recognition rate of magnetic tile defects.Fig. 11 Grad-CAM visualization results obtained by the proposed LPAR-Net and two other different fusion methods. (a) is the input image, (b) is the APR module, (c) is NAM, and (d) is CBAM.

Ablation experiments

In order to verify the effect of different methods proposed in this paper on the performance of the network model, ablation experiments are conducted to test the IB Block (Inverted Bottleneck Block), NAPR, and APR proposed in this paper, and the results of the experiments are shown in Table 3.Table 3 Comparison of ablation experiment results

Method	Accuracy (%)	Precision (%)	Recall (%)	F1-score (%)	Params (M)	Flops (G)	
Baseline	90.52	90.68	89.63	89.73	11.18	1.82	
 + IB	91.23	91.68	90.38	90.68	1.39	0.23	
 + NAPR	91.54	91.74	90.82	91.05	1.42	0.24	
 + APR	93.63	94.06	93.06	93.54	1.66	0.24	

The Baseline model used has an average accuracy of 90.52%, which has a low recognition accuracy and has the largest volume. Then, the residual block structure in the Baseline network was changed to IB Block, which increased the accuracy by 0.71% to 91.23% and increased the Precision, Recall, and F1-Score values (+ 1%, + 0.75%, and + 0.95%) compared to the Baseline network.

Based on the IB module, a layer of 7 × 7 convolutional layer in parallel with it is added to form the NAPR module, which is used to realize the acquisition of long-range dependent information on spatial features, with an increase in accuracy of 0.31% and improved Precision, Recall, and F1-Score values (+ 0.06%, + 0.44%, and + 0.37%).

Based on the designed NAPR module, splitting the CBAM attention mechanism, respectively, yields the composition APR module with 2.09% improvement in average accuracy and improved Precision, Recall, and F1-Score values (+ 2.32%, + 2.24%, and + 2.49%). The use of parallel inverted bottleneck block structure, channel attention and spatial attention mechanisms can better improve the average accuracy. The average accuracy is improved by 3.11% using APR structure as compared to Baseline network. The model LPAR-Net for this study was obtained after improvement in the above manner.This model has the ability to extract image features of tile defects and can better recognize smaller defects. The recognition accuracy of this model reaches 93.63%, and the Precision, Recall, and F1-Score values are 94.06%, 93.06%, and 93.54%, respectively, which is the best performance among the recognition experiments of all defect categories, and has a good performance.

Comparison of experiments on different datasets

In order to test the recognition performance of the proposed method in this paper, it is validated on the public dataset NEU-DET. The dataset has 1800 images containing six categories: crazing, inclusion, patches, pitted_surface, rolled-in_scale, and scratches. It is shown in Fig. 12.Fig. 12 Sample plots of different defects in the strip steel dataset. (a–f) are crazing, inclusion, patches, pitted_surface, rolled-in_scale, and scratches defects, respectively.

In order to test the ability of the LPAR-Net model of this study to recognize the images of this dataset, six mainstream network models in Sect. "Attention parallel residual block" were used to compare with it (containing three large-body models and three lightweight models). In the experiments in this section, the learning rate during the training period is set to 0.0001, the number of iterations is epochs = 100, and the number of input images in each batch is batch size = 32. Table 4 shows the evaluation parameters obtained from the tests, and Fig. 13 shows the changes in the accuracy curves of each model. The images in the strip steel dataset have a more complex background, and our method can effectively overcome this difficulty, and the accuracies are all higher than the other six models. Therefore, the experimental results show that the LPAR-Net model performs well on this dataset, which proves that the model has a more excellent recognition ability.Table 4 Comparison of evaluation parameters of different network models using the strip steel dataset

Model	Accuracy	Precision	Recall	F1-score	Params (M)	Flops (G)	
ResNet18	0.9408	0.9460	0.9425	0.9382	11.69	1.82	
ConvNextV2	0.9464	0.9496	0.9478	0.9477	27.79	4.46	
DenseNet121	0.9789	0.9793	0.9786	0.9789	7.89	2.83	
MobileNetV2	0.9388	0.9477	0.9405	0.9393	3.47	0.30	
EfficientNet	0.9494	0.9543	0.9513	0.9503	5.25	0.39	
ShuffleNet-V2	0.9487	0.9521	0.9505	0.9489	2.27	0.15	
LPAR-Net	0.9814	0.9829	0.9814	0.9812	1.66	0.24	

Fig. 13 Plots of average accuracy curves of different models on the strip steel dataset. (a) shows the comparison between the curves of the model of this study and the heavyweight network model, and (b) shows the comparison between the curves of the model of this study and the lightweight model.

Conclusion

In this paper, we propose an innovative method for classifying and recognizing magnetic tile defects, which achieves better results in the testing of magnetic tile defect dataset. It can solve the problem of complex background and small surface defects in images, which is difficult to be solved by the current mainstream network model. Firstly, the image of magnetic tile defects collected under the complex background is enhanced using the linear variation method. Then, the APR Block composition LPAR-Net is proposed for magnetic tile defect detection. Among them, using APR Block to obtain weights in channel and spatial dimensions respectively and weighting the input feature map can reduce the influence of the complex background in the image, which effectively improves the detection accuracy of the defective parts of the magnetic tiles.The LPAR-Net model is trained on the image-enhanced dataset in order to recognize the magnetic tile defects. The experimental results surface that the average recognition accuracy of the model proposed in this paper reaches 93.63%. Compared with several common CNN models, the LPAR-Net model has better recognition rate during the training process. It can recognize the tiny tile surface defects better. It performs well in classifying tiny defects on the surface of magnetic tiles in complex backgrounds, which can meet the demand for high accuracy and performance and reduce unnecessary waste in modern manufacturing industries. In addition, the LPAR-Net model is fully evaluated using different attention mechanism experiments, ablation experiments, and comparison experiments with different datasets, respectively, which show its strong recognition ability. In this paper, a new method for lightweight, fast and accurate identification of magnetic tile defects is explored, which can provide a new research direction for metal surface defect detection. In the future, we will further optimize the model and improve the generalization performance of the algorithm. Although the method proposed in this study can effectively categorize and identify the defects of magnetic tiles, it fails to achieve effective localization of the defective regions in the category activation map, and we follow up with further research on the precise location and identification of the defective regions.

Author contributions

C.L., J.H. and E.Z. wrote the main manuscript text.G.Q. and F.Lprovided the images for this article. All authors reviewed the manuscript.

Data availability

This study did not report any data. The proposed method was evaluated on two public datasets widely used in the field of metal surface defect detection: the Magnetic Tile Surface Defects (https://www.kaggle.com/datasets/alex000kim/magnetic-tile-surface-defects), NEU-DET (https://www.kaggle.com/datasets/kaustubhdikshit/neu-surface-defect-database).

Competing interests

The authors declare no competing interests.

Publisher's note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
==== Refs
References

1. Li Z Khajepour A Song J A comprehensive review of the key technologies for pure electric vehicles Energy 2019 182 824 839 10.1016/j.energy.2019.06.077
Li, Z., Khajepour, A. & Song, J. A comprehensive review of the key technologies for pure electric vehicles. Energy 182, 824–839. 10.1016/j.energy.2019.06.077 (2019).
2. Wang H Zhang C Guo L Li X Novel revolving heat pipe cooling structure of permanent magnet synchronous motor for electric vehicle Appl. Therm. Eng. 2024 236 121641 10.1016/j.applthermaleng.2023.121641
Wang, H., Zhang, C., Guo, L. & Li, X. Novel revolving heat pipe cooling structure of permanent magnet synchronous motor for electric vehicle. Appl. Therm. Eng. 236, 121641. 10.1016/j.applthermaleng.2023.121641 (2024).
3. Taheri F Sauve G Van Acker K Circular economy strategies for permanent magnet motors in electric vehicles: Application of SWOT Proced. CIRP 2024 122 265 270 10.1016/j.procir.2024.01.038
Taheri, F., Sauve, G. & Van Acker, K. Circular economy strategies for permanent magnet motors in electric vehicles: Application of SWOT. Proced. CIRP 122, 265–270. 10.1016/j.procir.2024.01.038 (2024).
4. Li Y Li Q Fan T Wen X Heat dissipation design of end winding of permanent magnet synchronous motor for electric vehicle Energy Rep. 2023 9 282 288 10.1016/j.egyr.2022.10.416
Li, Y., Li, Q., Fan, T. & Wen, X. Heat dissipation design of end winding of permanent magnet synchronous motor for electric vehicle. Energy Rep. 9, 282–288. 10.1016/j.egyr.2022.10.416 (2023).
5. Xu J Modeling and analysis of oil frictional loss in wet-type permanent magnet synchronous motor for aerospace electro-hydrostatic actuator Chin. J. Aeronaut. 2023 36 328 341 10.1016/j.cja.2023.05.026
Xu, J. et al. Modeling and analysis of oil frictional loss in wet-type permanent magnet synchronous motor for aerospace electro-hydrostatic actuator. Chin. J. Aeronaut. 36, 328–341. 10.1016/j.cja.2023.05.026 (2023).
6. Guo H Xu J Kuang X A novel fault tolerant permanent magnet synchronous motor with improved optimal torque control for aerospace application Chin. J. Aeronaut. 2015 28 535 544 10.1016/j.cja.2015.01.008
Guo, H., Xu, J. & Kuang, X. A novel fault tolerant permanent magnet synchronous motor with improved optimal torque control for aerospace application. Chin. J. Aeronaut. 28, 535–544. 10.1016/j.cja.2015.01.008 (2015).
7. Sun J Xing G Zhou X Sun H Static magnetic field analysis of hollow-cup motor model and bow-shaped permanent magnet design Chin. J. Aeronaut. 2022 35 306 313 10.1016/j.cja.2021.10.033
Sun, J., Xing, G., Zhou, X. & Sun, H. Static magnetic field analysis of hollow-cup motor model and bow-shaped permanent magnet design. Chin. J. Aeronaut. 35, 306–313. 10.1016/j.cja.2021.10.033 (2022).
8. Abdalla II Ibrahim T Mohd Nor NB Development and optimization of a moving-magnet tubular linear permanent magnet motor for use in a reciprocating compressor of household refrigerators Int. J. Electr. Power Energy Syst. 2016 77 263 270 10.1016/j.ijepes.2015.11.020
Abdalla, I. I., Ibrahim, T. & Mohd Nor, N. B. Development and optimization of a moving-magnet tubular linear permanent magnet motor for use in a reciprocating compressor of household refrigerators. Int. J. Electr. Power Energy Syst. 77, 263–270. 10.1016/j.ijepes.2015.11.020 (2016).
9. Huang Y Jiang L Lei H Research on cogging torque of the permanent magnet canned motor in domestic heating system Energy Rep. 2021 7 1379 1389 10.1016/j.egyr.2021.09.124
Huang, Y., Jiang, L. & Lei, H. Research on cogging torque of the permanent magnet canned motor in domestic heating system. Energy Rep. 7, 1379–1389. 10.1016/j.egyr.2021.09.124 (2021).
10. Cho SK Jung KH Choi JY Design optimization of interior permanent magnet synchronous motor for electric compressors of air-conditioning systems mounted on EVs and HEVs IEEE Trans. Magn. 2018 54 1 5 10.1109/TMAG.2018.2849078
Cho, S. K., Jung, K. H. & Choi, J. Y. Design optimization of interior permanent magnet synchronous motor for electric compressors of air-conditioning systems mounted on EVs and HEVs. IEEE Trans. Magn. 54, 1–5. 10.1109/TMAG.2018.2849078 (2018).
11. Cao X Chen B He W Unsupervised defect segmentation of magnetic tile based on attention enhanced flexible U-net IEEE Trans. Instrum. Meas. 2022 71 1 10 10.1109/TIM.2022.3170989
Cao, X., Chen, B. & He, W. Unsupervised defect segmentation of magnetic tile based on attention enhanced flexible U-net. IEEE Trans. Instrum. Meas. 71, 1–10. 10.1109/TIM.2022.3170989 (2022).
12. Huang, Y., Qiu, C., Guo, Y., Wang, X. & Yuan, K. in 2018 IEEE 14th International Conference on Automation Science and Engineering (CASE), 612–617.
13. Li Q Internal defects inspection of arc magnets using multi-head attention-based CNN Measurement 2022 202 111808 10.1016/j.measurement.2022.111808
Li, Q. et al. Internal defects inspection of arc magnets using multi-head attention-based CNN. Measurement 202, 111808. 10.1016/j.measurement.2022.111808 (2022).
14. Zhang Y Development of a cross-scale weighted feature fusion network for hot-rolled steel surface defect detection Eng. Appl. Artif. Intell. 2023 117 105628 10.1016/j.engappai.2022.105628
Zhang, Y. et al. Development of a cross-scale weighted feature fusion network for hot-rolled steel surface defect detection. Eng. Appl. Artif. Intell. 117, 105628. 10.1016/j.engappai.2022.105628 (2023).
15. Liu T An adaptive image segmentation network for surface defect detection IEEE Trans. Neural Netw. Learn. Syst. 2022 10.1109/TNNLS.2022.3230426 36279326
Liu, T. et al. An adaptive image segmentation network for surface defect detection. IEEE Trans. Neural Netw. Learn. Syst.10.1109/TNNLS.2022.3230426 (2022).36279326
16. Ye L Xia X Chai B Wang S Yang B Application of deep learning in workpiece defect detection Proced. Comput. Sci. 2021 183 267 273 10.1016/j.procs.2021.02.058
Ye, L., Xia, X., Chai, B., Wang, S. & Yang, B. Application of deep learning in workpiece defect detection. Proced. Comput. Sci. 183, 267–273. 10.1016/j.procs.2021.02.058 (2021).
17. Naveen Venkatesh S Automatic detection of visual faults on photovoltaic modules using deep ensemble learning network Energy Rep. 2022 8 14382 14395 10.1016/j.egyr.2022.10.427
Naveen Venkatesh, S. et al. Automatic detection of visual faults on photovoltaic modules using deep ensemble learning network. Energy Rep. 8, 14382–14395. 10.1016/j.egyr.2022.10.427 (2022).
18. García-Pérez A CNN-based in situ tool wear detection: A study on model training and data augmentation in turning inserts J. Manuf. Syst. 2023 68 85 98 10.1016/j.jmsy.2023.03.005
García-Pérez, A. et al. CNN-based in situ tool wear detection: A study on model training and data augmentation in turning inserts. J. Manuf. Syst. 68, 85–98. 10.1016/j.jmsy.2023.03.005 (2023).
19. Yadav E Chawla VK An explicit literature review on bearing materials and their defect detection techniques Mater. Today Proceed. 2022 50 1637 1643 10.1016/j.matpr.2021.09.132
Yadav, E. & Chawla, V. K. An explicit literature review on bearing materials and their defect detection techniques. Mater. Today Proceed. 50, 1637–1643. 10.1016/j.matpr.2021.09.132 (2022).
20. Ameri R Hsu C-C Band SS A systematic review of deep learning approaches for surface defect detection in industrial applications Eng. Appl. Artif. Intell. 2024 130 107717 10.1016/j.engappai.2023.107717
Ameri, R., Hsu, C.-C. & Band, S. S. A systematic review of deep learning approaches for surface defect detection in industrial applications. Eng. Appl. Artif. Intell. 130, 107717. 10.1016/j.engappai.2023.107717 (2024).
21. Jha SB Babiceanu RF Deep CNN-based visual defect detection: Survey of current literature Comput. Ind. 2023 148 103911 10.1016/j.compind.2023.103911
Jha, S. B. & Babiceanu, R. F. Deep CNN-based visual defect detection: Survey of current literature. Comput. Ind. 148, 103911. 10.1016/j.compind.2023.103911 (2023).
22. Huang Q Yin Y Yin G Automatic classification of magnetic tiles internal defects based on acoustic resonance analysis Mech. Syst. Signal Process. 2015 60–61 45 58 10.1016/j.ymssp.2015.02.018
Huang, Q., Yin, Y. & Yin, G. Automatic classification of magnetic tiles internal defects based on acoustic resonance analysis. Mech. Syst. Signal Process. 60–61, 45–58. 10.1016/j.ymssp.2015.02.018 (2015).
23. Zhang X Wang B Chen X Intelligent fault diagnosis of roller bearings with multivariable ensemble-based incremental support vector machine Knowl. -Based Syst. 2015 89 56 85 10.1016/j.knosys.2015.06.017
Zhang, X., Wang, B. & Chen, X. Intelligent fault diagnosis of roller bearings with multivariable ensemble-based incremental support vector machine. Knowl. -Based Syst. 89, 56–85. 10.1016/j.knosys.2015.06.017 (2015).
24. Chu M Zhao J Liu X Gong R Multi-class classification for steel surface defects based on machine learning with quantile hyper-spheres Chemometr. Intell. Lab. Syst. 2017 168 15 27 10.1016/j.chemolab.2017.07.008
Chu, M., Zhao, J., Liu, X. & Gong, R. Multi-class classification for steel surface defects based on machine learning with quantile hyper-spheres. Chemometr. Intell. Lab. Syst. 168, 15–27. 10.1016/j.chemolab.2017.07.008 (2017).
25. Wang Z Zhu D An accurate detection method for surface defects of complex components based on support vector machine and spreading algorithm Measurement 2019 147 106886 10.1016/j.measurement.2019.106886
Wang, Z. & Zhu, D. An accurate detection method for surface defects of complex components based on support vector machine and spreading algorithm. Measurement 147, 106886. 10.1016/j.measurement.2019.106886 (2019).
26. Huang, Y., Yu, T., Wan, K. & Yuan, J. in 2021 IEEE International Conference on Advances in Electrical Engineering and Computer Applications (AEECA), pp. 983–987.
27. Zhao M Qiu W Wen T Liao T Huang J Feature extraction based on gabor filter and support vector machine classifier in defect analysis of thermoelectric cooler component Comput. Electr. Eng. 2021 92 107188 10.1016/j.compeleceng.2021.107188
Zhao, M., Qiu, W., Wen, T., Liao, T. & Huang, J. Feature extraction based on gabor filter and support vector machine classifier in defect analysis of thermoelectric cooler component. Comput. Electr. Eng. 92, 107188. 10.1016/j.compeleceng.2021.107188 (2021).
28. Choi Y-H Yang J Machine learning iterative filtering algorithm for field defect detection in the process stage Comput. Ind. 2022 142 103740 10.1016/j.compind.2022.103740
Choi, Y.-H. & Yang, J. Machine learning iterative filtering algorithm for field defect detection in the process stage. Comput. Ind. 142, 103740. 10.1016/j.compind.2022.103740 (2022).
29. Cui L SDDNet: A fast and accurate network for surface defect detection IEEE Trans. Instrum. Meas. 2021 70 1 13 10.1109/TIM.2021.3056744 33776080
Cui, L. et al. SDDNet: A fast and accurate network for surface defect detection. IEEE Trans. Instrum. Meas. 70, 1–13. 10.1109/TIM.2021.3056744 (2021).33776080
30. Hu C Liao H Zhou T Zhu A Xu C Online recognition of magnetic tile defects based on UPM-DenseNet Mater. Today Commun. 2022 30 103105 10.1016/j.mtcomm.2021.103105
Hu, C., Liao, H., Zhou, T., Zhu, A. & Xu, C. Online recognition of magnetic tile defects based on UPM-DenseNet. Mater. Today Commun. 30, 103105. 10.1016/j.mtcomm.2021.103105 (2022).
31. Lu H Zhu Y Yin M Yin G Xie L Multimodal fusion convolutional neural network with cross-attention mechanism for internal defect detection of magnetic tile IEEE Access 2022 10 60876 60886 10.1109/ACCESS.2022.3180725
Lu, H., Zhu, Y., Yin, M., Yin, G. & Xie, L. Multimodal fusion convolutional neural network with cross-attention mechanism for internal defect detection of magnetic tile. IEEE Access 10, 60876–60886. 10.1109/ACCESS.2022.3180725 (2022).
32. Liang W Sun Y ELCNN: A deep neural network for small object defect detection of magnetic tile IEEE Trans. Instrum. Meas. 2022 71 1 10 10.1109/TIM.2022.3193175
Liang, W. & Sun, Y. ELCNN: A deep neural network for small object defect detection of magnetic tile. IEEE Trans. Instrum. Meas. 71, 1–10. 10.1109/TIM.2022.3193175 (2022).
33. Yu X Lyu W Zhou D Wang C Xu W ES-Net: Efficient scale-aware network for tiny defect detection IEEE Trans. Instrum. Meas. 2022 71 1 14 10.1109/TIM.2022.3168897
Yu, X., Lyu, W., Zhou, D., Wang, C. & Xu, W. ES-Net: Efficient scale-aware network for tiny defect detection. IEEE Trans. Instrum. Meas. 71, 1–14. 10.1109/TIM.2022.3168897 (2022).
34. Xiao M GRA-Net: Global receptive attention network for surface defect detection Knowl. -Based Syst. 2023 280 111066 10.1016/j.knosys.2023.111066
Xiao, M. et al. GRA-Net: Global receptive attention network for surface defect detection. Knowl. -Based Syst. 280, 111066. 10.1016/j.knosys.2023.111066 (2023).
35. Wan B SMINet: Semantics-aware multi-level feature interaction network for surface defect detection Eng. Appl. Artif. Intelli. 2023 123 106474 10.1016/j.engappai.2023.106474
Wan, B. et al. SMINet: Semantics-aware multi-level feature interaction network for surface defect detection. Eng. Appl. Artif. Intelli. 123, 106474. 10.1016/j.engappai.2023.106474 (2023).
36. Zhu Y Xie L Yin M Yin G Convolution with rotation invariance for online detection of tiny defects on magnetic tile surface IEEE Trans. Instrum. Meas. 2023 72 1 12 10.1109/TIM.2023.3295477 37323850
Zhu, Y., Xie, L., Yin, M. & Yin, G. Convolution with rotation invariance for online detection of tiny defects on magnetic tile surface. IEEE Trans. Instrum. Meas. 72, 1–12. 10.1109/TIM.2023.3295477 (2023).37323850
37. Yun JP Automated defect inspection system for metal surfaces based on deep learning and data augmentation J. Manuf. Syst. 2020 55 317 324 10.1016/j.jmsy.2020.03.009
Yun, J. P. et al. Automated defect inspection system for metal surfaces based on deep learning and data augmentation. J. Manuf. Syst. 55, 317–324. 10.1016/j.jmsy.2020.03.009 (2020).
38. Li J Wang K He M Ke L Wang H Attention-based convolution neural network for magnetic tile surface defect classification and detection Appl. Soft Comput. 2024 159 111631 10.1016/j.asoc.2024.111631
Li, J., Wang, K., He, M., Ke, L. & Wang, H. Attention-based convolution neural network for magnetic tile surface defect classification and detection. Appl. Soft Comput. 159, 111631. 10.1016/j.asoc.2024.111631 (2024).
39. Guo F Zhang Y Lan R Ran S Liang Y Privacy-preserving small target defect detection of heat sink based on DeceFL and DSUNet Neurocomputing 2024 575 127276 10.1016/j.neucom.2024.127276
Guo, F., Zhang, Y., Lan, R., Ran, S. & Liang, Y. Privacy-preserving small target defect detection of heat sink based on DeceFL and DSUNet. Neurocomputing 575, 127276. 10.1016/j.neucom.2024.127276 (2024).
40. Zhang E Ma Q Chen Y Duan J Shao L EGD-Net: Edge-guided and differential attention network for surface defect detection J. Ind. Inf. Integr. 2022 30 100403 10.1016/j.jii.2022.100403
Zhang, E., Ma, Q., Chen, Y., Duan, J. & Shao, L. EGD-Net: Edge-guided and differential attention network for surface defect detection. J. Ind. Inf. Integr. 30, 100403. 10.1016/j.jii.2022.100403 (2022).
41. Shi, Z. et al. in 2021 International Conference on Computer Information Science and Artificial Intelligence (CISAI), pp. 83–87.
42. Zhang S He M Zhong Z Zhu D An industrial interference-resistant gear defect detection method through improved YOLOv5 network using attention mechanism and feature fusion Measurement 2023 221 113433 10.1016/j.measurement.2023.113433
Zhang, S., He, M., Zhong, Z. & Zhu, D. An industrial interference-resistant gear defect detection method through improved YOLOv5 network using attention mechanism and feature fusion. Measurement 221, 113433. 10.1016/j.measurement.2023.113433 (2023).
43. Xiang X Liu M Zhang S Wei P Chen B Multi-scale attention and dilation network for small defect detection Pattern Recognit. Lett. 2023 172 82 88 10.1016/j.patrec.2023.06.010
Xiang, X., Liu, M., Zhang, S., Wei, P. & Chen, B. Multi-scale attention and dilation network for small defect detection. Pattern Recognit. Lett. 172, 82–88. 10.1016/j.patrec.2023.06.010 (2023).
44. Yang K Yi J Chen A Liu J Chen W ConDinet++: Full-scale fusion network based on conditional dilated convolution to extract roads from remote sensing images IEEE Geosci. Remote Sens. Lett. 2022 19 1 5 10.1109/LGRS.2021.3093101
Yang, K., Yi, J., Chen, A., Liu, J. & Chen, W. ConDinet++: Full-scale fusion network based on conditional dilated convolution to extract roads from remote sensing images. IEEE Geosci. Remote Sens. Lett. 19, 1–5. 10.1109/LGRS.2021.3093101 (2022).
45. Zhang L Chen J Chen J Wen Z Zhou X LDD-Net: Lightweight printed circuit board defect detection network fusing multi-scale features Eng. Appl. Artif. Intell. 2024 129 107628 10.1016/j.engappai.2023.107628
Zhang, L., Chen, J., Chen, J., Wen, Z. & Zhou, X. LDD-Net: Lightweight printed circuit board defect detection network fusing multi-scale features. Eng. Appl. Artif. Intell. 129, 107628. 10.1016/j.engappai.2023.107628 (2024).
46. He, K., Zhang, X., Ren, S. & Sun, J. in 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 770–778.
47. Sandler, M., Howard, A., Zhu, M., Zhmoginov, A. & Chen, L. C. in 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 4510–4520.
48. Wang M Image super-resolution via enhanced multi-scale residual network J. Parallel Distrib. Comput. 2021 152 57 66 10.1016/j.jpdc.2021.02.016
Wang, M. et al. Image super-resolution via enhanced multi-scale residual network. J. Parallel Distrib. Comput. 152, 57–66. 10.1016/j.jpdc.2021.02.016 (2021).
