
==== Front
Sci Rep
Sci Rep
Scientific Reports
2045-2322
Nature Publishing Group UK London

39237695
71795
10.1038/s41598-024-71795-4
Article
Intelligent tool wear prediction based on deep learning PSD-CVT model
Si Sumei 1
Mu Deqiang 2
Si Zekai szk980120@gmail.com

13
1 https://ror.org/052pakb34 0000 0004 1761 6995 College of Electromechanical Engineering, Changchun University of Technology, Changchun, 130012 China
2 grid.523788.3 0000 0004 1761 6995 Changchun University of Technology, Changchun, 130012 China
3 Beijing Long March Tianmin High-Tech Co., LTD, Beijing, 100176 China
5 9 2024
5 9 2024
2024
14 207547 3 2024
30 8 2024
© The Author(s) 2024
2024
https://creativecommons.org/licenses/by/4.0/ Open Access This article is licensed under a Creative Commons Attribution 4.0 International License, which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made. The images or other third party material in this article are included in the article's Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article's Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by/4.0/.
To ensure the reliability of machining quality, it is crucial to predict tool wear accurately. In this paper, a novel deep learning-based model is proposed, which synthesizes the advantages of power spectral density (PSD), convolutional neural networks (CNN), and vision transformer model (ViT), namely PSD-CVT. PSD maps can provide a comprehensive understanding of the spectral characteristics of the signals. It makes the spectral characteristics more obvious and makes it easy to analyze and compare different signals. CNN focuses on local feature extraction, which can capture local information such as the texture, edge, and shape of the image, while the attention mechanism in ViT can effectively capture the global structure and long-range dependencies present in the image. Two fully connected layers with a ReLU function are used to obtain the predicted tool wear values. The experimental results on the PHM 2010 dataset demonstrate that the proposed model has higher prediction accuracy than the CNN model or ViT model alone, as well as outperforms several existing methods in accurately predicting tool wear. The proposed prediction method can also be applied to predict tool wear in other machining fields.

Keywords

Convolutional neural network (CNN)
Deep learning
Tool wear prediction
Power spectral density (PSD)
Vision transformer (ViT)
Subject terms

Mechanical engineering
Electrical and electronic engineering
http://dx.doi.org/10.13039/501100011789 Department of Science and Technology of Jilin Province 20200401119GX issue-copyright-statement© Springer Nature Limited 2024
==== Body
pmcIntroduction

Tool wear prediction holds great importance in the field of construction machinery as it significantly contributes to improving machining efficiency, optimizing product quality, and reducing maintenance costs1. Intelligent manufacturing and predictive maintenance have become more popular as a result of the introduction of the Industry 4.0 concept. Deep learning technology, being a key aspect of Industry 4.0, opens up new avenues for tool wear prediction, leading to a surge in research exploring its application2.

In the area of predicting tool wear, there are two commonly used methods: direct measurements3 and indirect measurements4. Direct measurement methods involve acquiring image information of the tool surface using advanced visualization techniques. Through high-resolution image acquisition and detailed image analysis, these methods enable direct observation and assessment of the degree and variation of tool wear. For instance, Nguyen et al.5 employed a digital optical microscope and a scanning electron microscope to examine tool wear at a specific cutting distance. These non-contact direct measurement methods provide researchers with valuable visual references, offering insights into microscopic features and surface topography associated with tool wear. In contrast, indirect measurement methods estimate tool wear by collecting multiple signals generated during machining6. These signals include forces, vibrations, and sounds produced by the tool. With sophisticated sensors and signal processing techniques, researchers can capture and analyze the tool wear characteristics embedded in these signals. By extensively analyzing and mining the spectrum, amplitude, and waveform of the signals, it is possible to determine useful features that have a close connection to the tool wear state. Due to the numerous influencing factors involved in direct measurements during actual production machining, indirect measurement techniques are frequently utilized in tool wear prediction. These indirect methods possess advantages such as non-invasiveness, real-time capability, automation, and intelligence. They provide multidimensional information for accurate wear prediction and monitoring7. Although direct measurement methods still have unique advantages in specific cases, data-driven indirect measurement methods are more commonly utilized and preferred in practical applications due to the production pace and real-world machining environments.

The current state of tool wear prediction relies mainly on statistics and conventional machine learning methods like regression analysis8, Decision Tree9, Random Forest10, Support Vector Machines (SVM)11, Naive Bayes12, Principal Component Analysis (PCA)13, Hidden Markov Model (HMM)14, and Fuzzy Inference (ANFIS)15, etc. Xu et al.16 introduced an intelligent model called the Adaptive Neuro-Fuzzy Inference System (ANFIS) coupled with an improved Particle Swarm Optimization (PSO) algorithm to estimate tool wear.

However, the prediction of tool wear is limited by conventional machine learning techniques, including their inability to fully exploit complex features, resulting in reduced prediction accuracy and stability. These methods typically rely on manually designed feature extraction, which relies on expert knowledge and experience, making it challenging to capture implicit information in the data. Furthermore, the limited feature representation of traditional methods hinders their ability to manage massive amounts of data and intricate relationships. To overcome these challenges, researchers have turned to deep learning techniques, which have become a popular research focus for tool wear prediction in recent years.

Deep learning techniques achieve outstanding results in tool wear prediction, including convolutional neural networks (CNN)17,18 and recurrent neural networks (RNN)19,20. CNNs excel at extracting features from original data through convolutional and pooling layers, enabling the capture of local patterns and features. Pooling layers reduce the feature map size while retaining important features, and fully connected layers learn correlations between different features. Ambadekar et al.21 performed a triple classification task on rear tool face surface texture features to determine the wear status. In addition, based on this architecture of the RNN model, the short-term memory (LSTM)22,23 and the gated recursive unit (GRU)24 are derived, which establish a temporal relationship. Zhou et al.25 combined wear features and operating conditions using LSTM to predict tool life. In recent years, Transformers26–28 have gained attention and have been applied to various engineering fields due to increased computational power. The CNN-Transformer neural network (CTNN) model, proposed by Liu et al.29, aims to estimate wear unsupervised by processing data in parallel and learning variance. Li et al.30 introduced an IE-SBiGRU model that generates long time series feature sequences from multiple signals to achieve global awareness and long-distance parallel operations for tool wear prediction.

However, each of these methods has its drawbacks in engineering applications. CNNs have limitations in capturing global features31, traditional RNNs struggle with the long-term dependency problem32, and LSTM and GRU lack parallel computation capabilities33. Transformers, widely used in natural language processing (NLP)34, excel in processing sequential data with long-range dependencies and parallel computing capabilities. Transformer variants such as GPT35, BERT36, XLNet37, T538, and ViT39 have shown success in other domains. However, using large Transformer-based models in engineering problems, where data is often limited, may lead to overfitting and increased computational resource requirements.

To solve the problem of complementary combination of local and global features in the tool wear prediction process, a new deep learning model PSD-CVT is proposed in this paper, which uses multi-channel sensor signals to generate power spectral density (PSD) maps and enhances the prediction capability by extracting key features from local and global perspectives through CNN and ViT, respectively. Validation conducted on the PHM2010 dataset demonstrates the favorable performance of the proposed model.

The following are the primary contributions of this paper:A PSD-VCT model for tool wear prediction is proposed, which provides an ingenious and effective approach by utilizing PSD for data transformation processing and utilizing the benefits of CNN and ViT for feature extraction.

Tool wear prediction is performed using the novel deep learning model, which provides a new solution in this field.

Results from the experiments on the PHM2010 dataset demonstrate that the model can extract not only key features from a local perspective but also global features by capturing long-distance dependencies in images through a self-attention mechanism, thus achieving high accuracy in wear prediction.

What follows is the remainder of this paper: section “Related work” gives background information on the method and related work on data processing and model building. In section “Research method”, a detailed description of the advantages of the proposed tool wear prediction method, and its model structure is provided. Section “Experiment study” presents the experimental and analytical results obtained from the PHM2010 dataset and the comparison model. Finally, Section “Conclusions and future works” summarizes the research findings and offers a reference for future research.

Related work

Power spectral density (PSD)

The signal is transformed from the time domain to the frequency domain through the application of the Fourier transform. By employing the Fourier transform, a signal can be broken down into a series of sinusoidal or complex exponential components of varying frequencies. The PSD is the square of the amplitude spectrum of the Fourier transform result, and the formula is shown below:1 SXXf=∫-∞∞xt·e-2πiftdt2

where SXXf denotes the PSD of the signal, and xt is the signal in the time domain.

PSD plots play a crucial role in spectral analysis as they effectively depict the energy distribution of a signal at different frequencies. These plots offer an intuitive representation of the frequency domain characteristics of the signal, aiding in the comprehension of its spectral features. In comparison to directly utilizing Fourier transform results, PSD plots emphasize the frequency components, making the spectral characteristics more apparent and facilitating the analysis and comparison of different signals. Additionally, PSD plots are often subjected to smoothing techniques to minimize the impact of noise interference and enhance the readability of the graph. In essence, PSD plots serve as a valuable tool for comprehensively understanding the spectral shape and frequency domain properties of signals. Consequently, they are widely employed in spectral analysis and signal processing endeavors.

Convolutional neural network (CNN)

CNN is a well-known architecture in the field of neural networks, particularly effective for handling data with spatial structures. In the context of analyzing and processing signals related to tool wear processes, where multi-signals are transformed into PSD maps, CNNs prove to be a suitable choice. Features are extracted and analyzed from the input data in the CNN model through a combination of convolutional and pooling layers. In the following section, a brief overview of the CNN architecture is illustrated in Fig. 1.Fig. 1 1-D convolution.

Convolutional Layer: The convolutional layer is a crucial component of CNNs. It applies a learnable convolution kernel, also known as a filter, to the input image through a convolution operation. This process calculates the result of the convolution operation at each position. By convolving the kernel over the entire image, local features are extracted from the input. The convolution operation involves element-wise multiplication of local regions of the input image with the convolution kernel. The products are then summed to obtain an output value. Mathematically, this can be expressed as follows:2 (I∗K)a,b=∑k,lIa+k,b+l·Kk,l

where I is the input feature map, K is the convolution kernel, (I∗K)a,b denotes the elements in the output feature map, k,l are the index of the convolution kernel, and a,b are the index of the output feature map.

Activation Function: An activation function is applied to the output of a convolutional layer to introduce nonlinearity. The ReLU function enhances the expressiveness of the network by performing an element-by-element nonlinear transformation of the output of the convolutional layer. It can be defined as follows:3 fx=max0,x

where fx denotes the result from the function of activation, and max0,x denotes taking the larger value between 0 and x.

Pooling Layer: The output of the convolutional layer is spatially down-sampled using the pooling layer. This downsampling process reduces the number of parameters and computational complexity while extracting more robust features. One commonly used pooling operation is MaxPooling, which divides the feature map into non-overlapping regions using a 2∗2 pooling window. Then, the maximum value within each region is taken as the output. The MaxPooling operation can be expressed as follows:4 MaxPoolingx=maxxa,b,xa+1,b,xa,b+1,xa+1,b+1

where xa,b denotes the elements within the pooling window.

Vision transformer (ViT)

ViT is a computer vision model that adopts the Transformer architecture for tasks like image classification, target detection, and semantic segmentation. Its structure is depicted in Fig. 2. ViT leverages the capabilities of a Transformer and treats an image as a sequence, akin to text sequences in natural language processing.Fig. 2 Network structure of ViT.

Input representation: In ViT, the image is broken up into fixed-size image blocks that are then vectorized and embedded in a lower-dimensional feature space. Consequently, the image is represented as a sequence with a size of N∗D, where N denotes the number of image blocks, and D represents the vector dimension of each block.

Embedding Layer: The embedding layer in ViT employs a simple linear transformation to convert the N∗D input image sequence into a lower-dimensional N∗d embedding sequence. This transformation can be mathematically expressed as follows:5 Emberxi=Wexi+be

where xi is the i-th image block in the input sequence, Emberxi is the corresponding embedding vector, and We and be are the learnable parameters.

Positional Encoding: In the ViT, positional encoding plays a crucial role in associating each embedding vector with its corresponding position in the input image. To achieve this, a commonly used approach involves generating a fixed set of positional encoding vectors through the utilization of sine and cosine functions. The attention mechanism in ViT can be represented by Eq. (6):6 AttentionQ,K,V=softmaxQKTdkV

where Q is a query matrix, K is a key matrix, and V is a value matrix.

After the self-attentive layer, the features at each location undergo a nonlinear transformation using a feedforward neural network. This neural network typically comprises two fully connected layers with an activation function, such as ReLU, and a batch normalization layer placed between these two layers.

Research method

This section presents the model architecture based on PSD-CVT, and the corresponding flowchart is depicted in Fig. 3. The signals captured by each sensor undergo a conversion process, transforming them into a PSD map. Subsequently, a normalization operation is applied to adjust the size of the PSD image to 224∗224. For further processing, the input is subsequently divided into two parts.Fig. 3 Overall structure of PSD-CVT.

In the first part, a 3∗3 convolution kernel is applied to the input image to perform the convolution operation, resulting in a feature map with 32 channels. This operation aims to extract image features. The convolved feature map is subsequently downsampled using a maximum pooling layer, reducing the size by half.

The second part incorporates the ViT module. The ViT-B/16 model, which has been pre-trained for large-scale image tasks, is employed in the ViT module, giving it powerful image feature extraction and generalization capabilities. Its effectiveness has been verified on different tasks and domains through extensive training and validation on numerous datasets.

The 16∗16 size patches are also used because using larger image blocks as input may increase the computational and memory requirements, leading to more complex and inefficient models. To balance computing resources and performance needs, a smaller image block size is chosen. Initially, an image segmentation layer (Patch Embedding) is applied to segment the input image into a set of 16∗16 size image blocks. Each image block is transformed into a vector through a linear transformation to capture block-level features. To preserve positional information, sine, and cosine functions are utilized to generate position encoding. Next, the core part of the ViT model is introduced, which consists of multiple encoder layers. Each encoder layer comprises self-attention40 mechanism sub-layers and feedforward network sub-layers. Self-attention mechanism sub-layer adaptively computes the weight of each patch based on its relationship with other patches, facilitating the capture of global dependencies. The feedforward network sub-layer performs nonlinear transformations on each patch, ensuring a consistent mapping of the output content.

Finally, the convolved feature vectors from the previous part are then concatenated with the feature vectors from the ViT module along the first dimension, resulting in a more comprehensive feature representation. A fully connected layer (FC1) is used to process the concatenated feature vectors, resulting in an output size of 256. The output of FC1 is then passed through the ReLU activation function. Subsequently, the resulting output is fed into another fully connected layer (FC2) with an output size of 1. Finally, the output of FC2 is returned as the prediction result.

The multi-channel sensor signal is converted into a PSD map by a PSD-CVT model to capture the frequency domain characteristics. This conversion allows the model to understand the frequency distribution of the signal and thus analyze its periodicity and frequency characteristics. Secondly, the CNN model is good at extracting detailed local features, while the ViT model is fast at capturing the global relationship between pixels using a self-attention mechanism with powerful global sensing capability. By fusing these architectures, the PSD-CVT model can effectively consider both local and global features to achieve a more comprehensive analysis and processing of signal data.

Experiment study

Introduction to the baseline dataset

The effectiveness and high accuracy of the proposed PSD-CVT model are demonstrated through experimental model training using the PHM2010 dataset, which involved the use of a 3-flute ball-tipped carbide milling cutter on a Roders Tech RFM760 high-digit CNC machine. For the experiment, the following cutting conditions were applied: 10,400 revolutions per minute (rpm) spindle speed, 1555 mm per minute (mm/min) feed rate, 0.2 mm axial depth of cut, 0.125 mm radial width of cut, and 0.001 mm feed per journey. During the machining process, various sensors were employed to measure different signals. A three-way force gauge was used to capture force signals, an acoustic emission sensor was used to record acoustic emission data, and three accelerometers were used to monitor vibration signals. These seven signals were gathered with an NI DAQ data acquisition device at a frequency of 50 kHz. After each machining stroke, the three-edge wear was simultaneously assessed with a LEICA MZ12 microscope. Data sets C1, C4, and C6, which included the data of the entire cycles, were chosen for this experiment and utilized as the training and validation data sets. Each data set contains one “wear” file that lists wear after each cut in 10–3 mm and a folder with 315 individual data acquisition files (one for each cut).

Data processing

The tool generates seven signals in each stroke, as shown in Fig. 4. By calculating the square of the amplitude spectrum of the Fourier transforms result, the PSD image was obtained, which is shown in Fig. 5. Subsequently, the PSD map of each stroke was adjusted to a specified size of 224∗224 and normalized. This processed PSD map served as the input for the subsequent ViT with the CNN part of the model. In terms of the experimental setup, the wear label is selected based on the average value of the wear on the three sides. To create individual datasets for the experiments, the original data set is divided into an 80% training set, a 10% validation set, and a 10% test set. This division ensured that each dataset was appropriately split for training, evaluating, and testing the PSD-CVT model.Fig. 4 The seven signals collected.

Fig. 5 PSD part of the frequency domain detail map.

To ensure a comprehensive assessment and to fully utilize the entire dataset, cross-validation was employed. In this method, the dataset was divided into multiple subsets, and the model was trained and tested multiple times, with each subgroup serving as the test set in one of the iterations. This approach allowed for the use of all data points for both training and evaluation, thereby providing a robust estimate of the performance of the model. As shown in Fig. 6.Fig. 6 Data partitioning.

Parameter and hyper-parameter settings

Table 1. provides the parameters for each layer of the model. The optimizer used for the model is Adam41.In the case of regression problems, the loss function usually selects mean square error (MSE). Because the smaller the value of MSE, the better the model fits. The average value of the MSE on each training batch is calculated as an evaluation metric to assess the performance of the model:7 MSE=1n∑i=1n(yi-y^i)2

where n is the sample size, yi is the true value, and y^i is the predicted value.Table 1 PSD-CVT model parameters.

Layers	Parameters	
ViT	Transformer layers: 4 Attention Heads: 8 Hidden:256	
Convolutional	Input channels: 3, Output channels: 32, Kernel size: 3 × 3, Padding: 1	
Max pooling	Kernel size: 2 × 2, Stride: 2	
Linear	Layer number:2, hidden dim per layer:256,1	

To address the common issue of overfitting in the training of neural network models, this study employs the early stopping strategy42. This strategy is a crucial technique for preventing model overfitting during the training process. The basic idea of this strategy involves setting an early stopping patience value (n). During the training process, if there is an improvement in validation loss within the patience range, the patience value is reset to zero, and the training continues. If the validation loss does not decrease for n consecutive training epochs, it is considered that the model has started to overfit. At this point, the early stopping strategy is triggered to prevent further decline in the performance on the validation set. The StepLR scheduler reduces the learning rate of the optimizer by a factor (gamma) every few epochs (step size), helping the model to converge more efficiently and potentially improving its performance by making the learning rate adjustments more gradual. Table 2 presents the specific values of these hyperparameters.Table 2 Hyperparameter settings.

Hyper-parameters	Learning rate	Step size	Gamma	Epoch number	Batch size	Maximum patience value	
Values	0.001	10	0.5	1000	8	4	

Evaluation metrics

In previous studies, the mean absolute error (MAE) and root mean square error (RMSE) are commonly utilized as performance metrics for prediction problems. These metrics provide quantitative measures of prediction accuracy. The MAE and RMSE can be calculated using the following equations:8 MAE=1n∑i=1n∣yi-y^i∣

9 RMSE=1n∑i=1n(yi-y^i)2

where n is the sample size, yi is the true value, and y^i is the predicted value of the model.

Result discussion and comparison

The trained PSD-CVT model is applied to the raw data for wear prediction, resulting in accurate wear prediction results. Figure 7 illustrates the high agreement observed between the predicted wear curves generated by the model and the actual curves. This demonstrates that the PSD-CVT model is capable of effectively performing the prediction task related to tool wear. The close match between predicted and actual wear curves highlights the accuracy of the model and its ability to accurately capture and predict wear patterns.Fig. 7 The wear prediction results of the PHM2010 testing dataset for the proposed PSD-CVT model: (a) predicted wear of C1 cutter, (b) predicted wear of C4 cutter, (c) predicted wear of C6 cutter.

In this paper, the PSD-CVT model is compared with the CNN model and ViT model alone, as well as some existing models. The experimental results, presented in Table 3, prove the effectiveness of fusing frequency domain features through PSD. By combining the local feature extraction capability of CNN and the global perception capability of ViT, the PSD-CVT model achieves significantly improved prediction accuracy. The CNN model excels in extracting local features and spatial information, while the attention mechanism in the ViT model enables effective modeling of global features. By leveraging both local and global features, the PSD-CVT model demonstrates enhanced accuracy in tool wear prediction tasks. This finding underscores the importance of integrating different components and operations in deep learning models. Comparisons with published papers further support the superior performance of the PSD-CVT model. The experimental results highlight the potential of leveraging PSD data pre-processing and the complementary advantages of ViT and CNN in improving tool wear prediction accuracy. This research provides valuable insights for optimizing model architectures and underscores the practical applications of deep learning models in tool wear prediction and related tasks. In conclusion, the experimental outcomes verify that the proposed PSD-CVT model is effective. And emphasize the significance of integrating different components and operations in deep learning models. This research contributes to the advancement of tool wear prediction models and opens avenues for further exploration and optimization of model architectures.Table 3 Performance analysis of different models.

Model	Datasets	
MAE	RMSE	
C1	C4	C6	C1	C4	C6	
MLP20	24.5	18.0	24.8	31.2	20.0	31.4	
Deep LSTMs20	8.3	8.7	15.2	12.1	10.2	18.9	
TBNN43	4.294	–	7.772	6.116	–	7.835	
CTNN29	3.634	–	7.531	5.358	–	9.209	
IE-SBIGRU30	3.694	5.189	3.398	5.056	6.884	4.527	
CNN	1.234	3.988	2.467	1.600	6.113	3.384	
ViT	23.36	27.71	33.94	27.54	37.67	40.12	
PSD-CVT	0.683	0.734	2.107	0.939	0.979	2.954	

Conclusions and future works

A new PSD-CVT model is proposed in this paper for predicting tool wear, which combines the benefits of PSD, CNN, and ViT architectures to achieve accurate tool wear prediction. The proposed scheme aims to enhance machining efficiency, improve quality, and reduce production costs in the tool wear machining field. By converting force, acceleration, and acoustic emission signals into PSD images and utilizing the CNN and ViT encoder for feature extraction, the PSD-CVT model demonstrates superior performance compared to other researchers and individual CNN or ViT-based approaches. The experimental results of the PHM2010 dataset strongly demonstrate the effectiveness and high accuracy of the scheme in capturing the unique characteristics of tool wear and accurately predicting wear. The contributions of this research are as follows:The proposal of the PSD-CVT scheme introduces a novel approach that intelligently combines PSD, ViT, and CNN techniques for tool wear prediction.

The feasibility of the scheme is experimentally verified, highlighting its superior performance compared to some existing methods.

These findings hold promising prospects for advancing machining intelligence and provide help for future research and advancements in related fields.

The future should focus on further improving the performance of the model in real production environments and exploring practical applications. Overall, the proposed method provides a reliable and innovative approach to tool wear prediction that has implications for a variety of industries and applications.

Acknowledgements

This research acknowledges the financial and equipment support partially provided by the Department of Science and Technology of Jilin Province (20200401119GX).

Author contributions

Zekai Si: Conceptualization, Methodology, Software. Sumei Si: Validation, Visualization, Data curation. Deqiang Mu: Supervision, Project administration.

Data availability

The datasets used and/or analyzed during the current study available. They are come from the PHM 2010 dataset (https://phmsociety.org/phm_competion/2010-phm-society-conference-data-challenge).

Competing intersts

The authors declare no competing interests.

Publisher's note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
==== Refs
References

1. Kuntoğlu M Sağlam H Investigation of progressive tool wear for determining of optimized machining parameters in turning Measurement 2019 140 427 436 10.1016/j.measurement.2019.04.022
Kuntoğlu, M. & Sağlam, H. Investigation of progressive tool wear for determining of optimized machining parameters in turning. Measurement 140, 427–436 (2019).10.1016/j.measurement.2019.04.022
2. Diez-Olivan A Del Ser J Galar D Sierra B Data fusion and machine learning for industrial prognosis: Trends and perspectives towards Industry 4.0 Inf. Fusion 2019 50 92 111 10.1016/j.inffus.2018.10.005
Diez-Olivan, A., Del Ser, J., Galar, D. & Sierra, B. Data fusion and machine learning for industrial prognosis: Trends and perspectives towards Industry 4.0. Inf. Fusion 50, 92–111 (2019).10.1016/j.inffus.2018.10.005
3. Mehta, S., Singh, R. A., Mohata, Y. & Kiran, M. Measurement and analysis of tool wear using vision system. In 2019 IEEE 6th International Conference on Industrial Engineering and Applications (ICIEA), IEEE 45–49 (2019).
4. Wu J-Y Wu M Chen Z Li X Yan R A joint classification-regression method for multi-stage remaining useful life prediction J. Manufact. Syst. 2021 58 109 119 10.1016/j.jmsy.2020.11.016
Wu, J.-Y., Wu, M., Chen, Z., Li, X. & Yan, R. A joint classification-regression method for multi-stage remaining useful life prediction. J. Manufact. Syst. 58, 109–119 (2021).10.1016/j.jmsy.2020.11.016
5. Nguyen D Abdullah MSB Khawarizmi R Kim D Kwon P The effect of fiber orientation on tool wear in edge-trimming of carbon fiber reinforced plastics (CFRP) laminates Wear 2020 450 203213 10.1016/j.wear.2020.203213
Nguyen, D., Abdullah, M. S. B., Khawarizmi, R., Kim, D. & Kwon, P. The effect of fiber orientation on tool wear in edge-trimming of carbon fiber reinforced plastics (CFRP) laminates. Wear 450, 203213 (2020).10.1016/j.wear.2020.203213
6. Wang G Zhang F A sequence-to-sequence model with attention and monotonicity loss for tool wear monitoring and prediction IEEE Trans. Instrument. Meas. 2021 70 1 11 10.1109/TIM.2021.3123218
Wang, G. & Zhang, F. A sequence-to-sequence model with attention and monotonicity loss for tool wear monitoring and prediction. IEEE Trans. Instrument. Meas. 70, 1–11 (2021).10.1109/TIM.2021.3123218
7. Jaen-Cuellar AY Osornio-Ríos RA Trejo-Hernández M Zamudio-Ramírez I Díaz-Saldaña G Pacheco-Guerrero JP Antonino-Daviu JA System for tool-wear condition monitoring in cnc machines under variations of cutting parameter based on fusion stray flux-current processing Sensors 2021 21 8431 10.3390/s21248431 34960525
Jaen-Cuellar, A. Y. et al. System for tool-wear condition monitoring in cnc machines under variations of cutting parameter based on fusion stray flux-current processing. Sensors 21, 8431 (2021).34960525 10.3390/s21248431
8. Soori M Arezoo B Cutting tool wear prediction in machining operations, a review J. New Technol. Mater. 2022 2022 859
Soori, M. & Arezoo, B. Cutting tool wear prediction in machining operations, a review. J. New Technol. Mater. 2022, 859 (2022).
9. Li G Wang Y He J Hao Q Yang H Wei J Tool wear state recognition based on gradient boosting decision tree and hybrid classification RBM Int. J. Adv. Manufact. Technol. 2020 110 511 522 10.1007/s00170-020-05890-x
Li, G. et al. Tool wear state recognition based on gradient boosting decision tree and hybrid classification RBM. Int. J. Adv. Manufact. Technol. 110, 511–522 (2020).10.1007/s00170-020-05890-x
10. Wu D Jennings C Terpenny J Gao RX Kumara S A comparative study on machine learning algorithms for smart manufacturing: Tool wear prediction using random forests J. Manufact. Sci. Eng. 2017 2017 139
Wu, D., Jennings, C., Terpenny, J., Gao, R. X. & Kumara, S. A comparative study on machine learning algorithms for smart manufacturing: Tool wear prediction using random forests. J. Manufact. Sci. Eng. 2017, 139 (2017).
11. Li B Tian X An effective PSO-LSSVM-based approach for surface roughness prediction in high-speed precision milling, Ieee Access 2021 9 80006 80014 10.1109/ACCESS.2021.3084617
Li, B. & Tian, X. An effective PSO-LSSVM-based approach for surface roughness prediction in high-speed precision milling, Ieee. Access 9, 80006–80014 (2021).10.1109/ACCESS.2021.3084617
12. Luo, C. et al. Hob wear state identification method based on SSA and Weighted Naive Bayes. In 2021 4th International Conference on Information Communication and Signal Processing (ICICSP), IEEE 404–408 (2021).
13. Wang G Zhang Y Liu C Xie Q Xu Y A new tool wear monitoring method based on multi-scale PCA J. Intell. Manufact. 2019 30 113 122 10.1007/s10845-016-1235-9
Wang, G., Zhang, Y., Liu, C., Xie, Q. & Xu, Y. A new tool wear monitoring method based on multi-scale PCA. J. Intell. Manufact. 30, 113–122 (2019).10.1007/s10845-016-1235-9
14. Li W Liu T Time varying and condition adaptive hidden Markov model for tool wear state estimation and remaining useful life prediction in micro-milling Mech. Syst. Signal Process. 2019 131 689 702 10.1016/j.ymssp.2019.06.021
Li, W. & Liu, T. Time varying and condition adaptive hidden Markov model for tool wear state estimation and remaining useful life prediction in micro-milling. Mech. Syst. Signal Process. 131, 689–702 (2019).10.1016/j.ymssp.2019.06.021
15. Xu L Huang C Li C Wang J Liu H Wang X Prediction of tool wear width size and optimization of cutting parameters in milling process using novel ANFIS-PSO method Proc. Inst. Mech. Eng. Part B: J. Eng. Manuf. 2022 236 111 122 10.1177/0954405420935262
Xu, L. et al. Prediction of tool wear width size and optimization of cutting parameters in milling process using novel ANFIS-PSO method. Proc. Inst. Mech. Eng. Part B: J. Eng. Manuf. 236, 111–122 (2022).10.1177/0954405420935262
16. Xu L Huang C Li C Wang J Liu H Wang X Estimation of tool wear and optimization of cutting parameters based on novel ANFIS-PSO method toward intelligent machining J. Intell. Manufact. 2021 32 77 90 10.1007/s10845-020-01559-0
Xu, L. et al. Estimation of tool wear and optimization of cutting parameters based on novel ANFIS-PSO method toward intelligent machining. J. Intell. Manufact. 32, 77–90 (2021).10.1007/s10845-020-01559-0
17. Duan J Duan J Zhou H Zhan X Li T Shi T Multi-frequency-band deep CNN model for tool wear prediction Meas. Sci. Technol. 2021 32 065009 10.1088/1361-6501/abb7a0
Duan, J. et al. Multi-frequency-band deep CNN model for tool wear prediction. Meas. Sci. Technol. 32, 065009. 10.1088/1361-6501/abb7a0 (2021).10.1088/1361-6501/abb7a0
18. García-Pérez A Ziegenbein A Schmidt E Shamsafar F Fernández-Valdivielso A Llorente-Rodríguez R Weigold M CNN-based in situ tool wear detection: A study on model training and data augmentation in turning inserts J. Manufact. Syst. 2023 68 85 98 10.1016/j.jmsy.2023.03.005
García-Pérez, A. et al. CNN-based in situ tool wear detection: A study on model training and data augmentation in turning inserts. J. Manufact. Syst. 68, 85–98 (2023).10.1016/j.jmsy.2023.03.005
19. Cheng M Jiao L Yan P Jiang H Wang R Qiu T Wang X Intelligent tool wear monitoring and multi-step prediction based on deep learning model J. Manufact. Syst. 2022 62 286 300 10.1016/j.jmsy.2021.12.002
Cheng, M. et al. Intelligent tool wear monitoring and multi-step prediction based on deep learning model. J. Manufact. Syst. 62, 286–300 (2022).10.1016/j.jmsy.2021.12.002
20. Zhao, R., Wang, J., Yan, R. & Mao, K. Machine health monitoring with LSTM networks. In 2016 10th international conference on sensing technology (ICST), IEEE 1–6 (2016).
21. Ambadekar P Choudhari C CNN based tool monitoring system to predict life of cutting tool SN Appl. Sci. 2020 2 1 11 10.1007/s42452-020-2598-2
Ambadekar, P. & Choudhari, C. CNN based tool monitoring system to predict life of cutting tool. SN Appl. Sci. 2, 1–11 (2020).10.1007/s42452-020-2598-2
22. Hao, G. & Kunpeng, Z. Pyramid LSTM auto-encoder for tool wear monitoring. In 2020 IEEE 16th International Conference on Automation Science and Engineering (CASE), IEEE 190–195 (2020).
23. Si Z Si S Mu D Efficient tool wear prediction in manufacturing: BiLPReS hybrid model with performer encoder Arab. J. Sci. Eng. 2024 10.1007/s13369-024-08943-5
Si, Z., Si, S. & Mu, D. Efficient tool wear prediction in manufacturing: BiLPReS hybrid model with performer encoder. Arab. J. Sci. Eng.10.1007/s13369-024-08943-5 (2024).10.1007/s13369-024-08943-5
24. Wang J Yan J Li C Gao RX Zhao R Deep heterogeneous GRU model for predictive analytics in smart manufacturing: Application to tool wear prediction Comput. Ind. 2019 111 1 14 10.1016/j.compind.2019.06.001
Wang, J., Yan, J., Li, C., Gao, R. X. & Zhao, R. Deep heterogeneous GRU model for predictive analytics in smart manufacturing: Application to tool wear prediction. Comput. Ind. 111, 1–14 (2019).10.1016/j.compind.2019.06.001
25. Zhou J-T Zhao X Gao J Tool remaining useful life prediction method based on LSTM under variable working conditions Int. J. Adv. Manufact. Technol. 2019 104 4715 4726 10.1007/s00170-019-04349-y
Zhou, J.-T., Zhao, X. & Gao, J. Tool remaining useful life prediction method based on LSTM under variable working conditions. Int. J. Adv. Manufact. Technol. 104, 4715–4726 (2019).10.1007/s00170-019-04349-y
26. Wang H Men T Li Y-F Transformer for high-speed train wheel wear prediction with multiplex local–global temporal fusion IEEE Trans. Instrum. Meas. 2022 71 1 12 10.1109/TIM.2022.3216413
Wang, H., Men, T. & Li, Y.-F. Transformer for high-speed train wheel wear prediction with multiplex local–global temporal fusion. IEEE Trans. Instrum. Meas. 71, 1–12 (2022).10.1109/TIM.2022.3216413
27. Vaswani A Attention is all you need Adv. Neural Inf. Process. Syst. 2017 2017 30
Vaswani, A. et al. Attention is all you need. Adv. Neural Inf. Process. Syst. 2017, 30 (2017).
28. Si Z Si S Mu D Precision forecasting of grinding wheel Wear: A TransBiGRU model for advanced industrial predictive maintenance Measurement 2024 2024 114859 10.1016/j.measurement.2024.114859
Si, Z., Si, S. & Mu, D. Precision forecasting of grinding wheel Wear: A TransBiGRU model for advanced industrial predictive maintenance. Measurement 2024, 114859. 10.1016/j.measurement.2024.114859 (2024).10.1016/j.measurement.2024.114859
29. Liu H Liu Z Jia W Zhang D Wang Q Tan J Tool wear estimation using a CNN-transformer model with semi-supervised learning Meas. Sci. Technol. 2021 32 125010 10.1088/1361-6501/ac22ee
Liu, H. et al. Tool wear estimation using a CNN-transformer model with semi-supervised learning. Meas. Sci. Technol. 32, 125010. 10.1088/1361-6501/ac22ee (2021).10.1088/1361-6501/ac22ee
30. Li W Fu H Han Z Zhang X Jin H Intelligent tool wear prediction based on Informer encoder and stacked bidirectional gated recurrent unit Robot. Comput.-Integr. Manufact. 2022 77 102368 10.1016/j.rcim.2022.102368
Li, W., Fu, H., Han, Z., Zhang, X. & Jin, H. Intelligent tool wear prediction based on Informer encoder and stacked bidirectional gated recurrent unit. Robot. Comput.-Integr. Manufact. 77, 102368. 10.1016/j.rcim.2022.102368 (2022).10.1016/j.rcim.2022.102368
31. Li G Yu Y Visual saliency detection based on multiscale deep CNN features IEEE Trans. Image Process. 2016 25 5012 5024 10.1109/TIP.2016.2602079 28113629
Li, G. & Yu, Y. Visual saliency detection based on multiscale deep CNN features. IEEE Trans. Image Process. 25, 5012–5024 (2016).28113629 10.1109/TIP.2016.2602079
32. Lin T Horne BG Tino P Giles CL Learning long-term dependencies in NARX recurrent neural networks IEEE Trans. Neural Netw. 1996 7 1329 1338 10.1109/72.548162 18263528
Lin, T., Horne, B. G., Tino, P. & Giles, C. L. Learning long-term dependencies in NARX recurrent neural networks. IEEE Trans. Neural Netw. 7, 1329–1338 (1996).18263528 10.1109/72.548162
33. Fan J Zhang K Huang Y Zhu Y Chen B Parallel spatio-temporal attention-based TCN for multivariate time series prediction Neural Comput. Appl. 2021 2021 1 10
Fan, J., Zhang, K., Huang, Y., Zhu, Y. & Chen, B. Parallel spatio-temporal attention-based TCN for multivariate time series prediction. Neural Comput. Appl. 2021, 1–10 (2021).
34. Wolf, T. et al. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations 38–45 (2020).
35. Schick, T. et al. Toolformer: Language models can teach themselves to use tools, arXiv preprint arXiv:2302.04761 (2023).
36. Devlin, J., Chang, M.-W., Lee, K. & Toutanova, K. Bert: Pre-training of deep bidirectional transformers for language understanding, arXiv preprint arXiv:1810.04805, 10.48550/arXiv.1810.04805 (2018).
37. Yang Z Dai Z Yang Y Carbonell J Salakhutdinov RR Le QV Xlnet: Generalized autoregressive pretraining for language understanding Adv. Neural Inf. Process. Syst. 2019 2019 32
Yang, Z. et al. Xlnet: Generalized autoregressive pretraining for language understanding. Adv. Neural Inf. Process. Syst. 2019, 32 (2019).
38. Raffel C Shazeer N Roberts A Lee K Narang S Matena M Zhou Y Li W Liu PJ Exploring the limits of transfer learning with a unified text-to-text transformer J. Mach. Learn. Res. 2020 21 5485 5551
Raffel, C. et al. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res. 21, 5485–5551 (2020).
39. Dosovitskiy, A. et al. An image is worth 16x16 words: Transformers for image recognition at scale, arXiv preprint arXiv:2010.11929 (2020).
40. Zhang J Jiang Y Wu S Li X Luo H Yin S Prediction of remaining useful life based on bidirectional gated recurrent unit with temporal self-attention mechanism Reliab. Eng. Syst. Saf. 2022 221 108297 10.1016/j.ress.2021.108297
Zhang, J. et al. Prediction of remaining useful life based on bidirectional gated recurrent unit with temporal self-attention mechanism. Reliab. Eng. Syst. Saf. 221, 108297 (2022).10.1016/j.ress.2021.108297
41. Kingma, D. P. & Ba, J. Adam: A method for stochastic optimization, arXiv preprint arXiv:1412.6980, 10.48550/arXiv.1412.6980 (2014).
42. Prechelt L Orr GB Müller K-R Early stopping—but when? Neural Networks: Tricks of the Trade 1998 Springer 55 69
Prechelt, L. Early stopping—but when? In Neural Networks: Tricks of the Trade (eds Orr, G. B. & Müller, K.-R.) 55–69 (Springer, 1998).
43. Liu H Liu Z Jia W Lin X Zhang S A novel transformer-based neural network model for tool wear estimation Meas. Sci. Technol. 2020 31 065106 10.1088/1361-6501/ab7282
Liu, H., Liu, Z., Jia, W., Lin, X. & Zhang, S. A novel transformer-based neural network model for tool wear estimation. Meas. Sci. Technol. 31, 065106 (2020).10.1088/1361-6501/ab7282
