
==== Front
Sci Rep
Sci Rep
Scientific Reports
2045-2322
Nature Publishing Group UK London

39237633
71742
10.1038/s41598-024-71742-3
Article
The forecasting of surface displacement for tunnel slopes utilizing the WD-IPSO-GRU model
Ma Guoqing 1
Zang Xiaopeng 1
Chen Shitong 2
Zhi Momo zhimomoqnyh@163.com

2
Huang Xiaoming 3
1 https://ror.org/022e9e065 grid.440641.3 0000 0004 1790 0486 College of Civil Engineering, Shijiazhuang Tiedao University, Shijiazhuang, 050043 China
2 https://ror.org/022e9e065 grid.440641.3 0000 0004 1790 0486 Hebei Engineering Innovation Center for Traffic Emergency and Guarantee, Shijiazhuang Tiedao University, Shijiazhuang, 050043 China
3 https://ror.org/04ct4d772 grid.263826.b 0000 0004 1761 0489 School of Transportation, Southeast University, Nanjing, 211189 China
5 9 2024
5 9 2024
2024
14 2071710 4 2024
30 8 2024
© The Author(s) 2024
2024
https://creativecommons.org/licenses/by-nc-nd/4.0/ Open Access This article is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License, which permits any non-commercial use, sharing, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if you modified the licensed material. You do not have permission under this licence to share adapted material derived from this article or parts of it. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by-nc-nd/4.0/.
To quickly assess slope stability based on field displacement monitoring data, this paper constructs a hybrid optimization model that predicts surface displacement during tunnel excavation in base-overburden slopes. The model combines Wavelet Decomposition (WD) with a Gated Recurrent Unit (GRU), and the GRU's hyperparameters are optimized using an Improved Particle Swarm Optimization algorithm (IPSO). The specific steps are as follows: First, the Wavelet Decomposition (WD) technique is applied to decompose the raw displacement data, extracting features at different time–frequency scales. Next, the Dropout technique is incorporated into the GRU model to prevent overfitting. Additionally, nonlinear inertia weight ω improved cognitive factor c1, and social factor c2 are introduced. The PSO algorithm is improved by integrating crossover and mutation concepts from genetic algorithms. Finally, the IPSO is used to optimize the number of neural units hN, HN, LN and dropout rates D1 and D2 in the GRU network architecture. After constructing the WD-IPSO-GRU model, a comprehensive comparison is made with various swarm intelligence algorithms and state-of-the-art models. The experimental results demonstrate that the WD-IPSO-GRU model significantly improves the prediction accuracy of surface displacement in slopes during tunnel excavation. Compared to directly using raw data for prediction, the introduction of the WD preprocessing technique improved the prediction accuracy at measurement points 01 and 02 by 28% and 45.9%, respectively. Additionally, with the model optimized by IPSO, the prediction accuracy at measurement points 01 and 02 increased by 76% and 56.7%, respectively. The WD-IPSO-GRU model effectively addresses the challenges of extracting features from univariate displacement time-series data and determining the parameters of the GRU network. It improves the prediction accuracy of surface displacement in base-overburden type slopes and demonstrates excellent generalization ability and reliability. The research results validate the potential application of the model in geotechnical engineering and provide strong support for assessing slope stability during tunnel excavation.

Keywords

WD
Dropout technique
IPSO
GA
GRU
Analytical prediction of slope displacement
Subject terms

Environmental sciences
Natural hazards
Engineering
issue-copyright-statement© Springer Nature Limited 2024
==== Body
pmcIntroduction

Deformation and failure of slopes often result in catastrophic economic losses and casualties1. In 2023, China experienced 5659 geological disasters, including 3919 landslides, with direct economic losses amounting to 1.5 billion yuan2.Due to the limitations of land transportation routes, mountain underpass tunnels have become an inevitable choice for transportation infrastructure. The impact of tunnel excavation on slope stability cannot be ignored3, 4, as instability of tunnel slopes can severely affect the smooth progress of construction projects and the normal operation of land transportation5, 6. Surface displacement of slopes during tunnel excavation is an important indicator of slope deformation7. Predicting surface displacement can serve as a basis for assessing slope stability. Therefore, evaluating methods for predicting such displacements is particularly important in geotechnical engineering.

In recent years, with the development of computer science, an increasing number of researchers have been using artificial intelligence methods for predictions in the field of geotechnical engineering. The literature8 reviewed the architectures of Machine Learning (ML), Deep Learning (DL), and Ensemble Learning (EL) methods in the field of geotechnical engineering, highlighting that EL techniques outperform the other three methods and exhibit strong reliability when predicting behaviors in geotechnical engineering. RNN is an early model used for time series prediction, and Zheng et al.9. used the RNN model to predict slope displacement. The structural properties of the RNN model allow it to capture dynamic time series behavior, but it suffers from the problem of vanishing or exploding gradients. In light of this, Hochreiter S et al.10 developed the Long Short-Term Memory (LSTM) network based on the RNN model to overcome the typical issues of vanishing and exploding gradients. The LSTM network improves prediction accuracy through several threshold gates based on the classical RNN. This model has been widely used for predicting displacement time-series data7, 11−13. Although prediction models based on the LSTM network can solve the issues of vanishing gradients and long-distance data transmission, a single-layer LSTM network performs poorly when handling univariate time-series data. This necessitates the construction of more complex LSTM network architectures, resulting in significant computational costs. To address this, Chung et al.14. developed the Gated Recurrent Unit (GRU), which simplifies the gating mechanisms of the LSTM network, making gradient propagation easier. This is particularly advantageous when handling large-scale datasets and training tasks that require frequent iterations. Zhang et al.15, 16. applied the GRU to predict nonlinear displacement sequences. Additionally, to address the problem of overfitting in neural networks, Nitish Srivastava et al.17. introduced the Dropout layer. This technique introduces Bernoulli random variables to any layer, effectively extracting a subnetwork from the complex neural network for training, thus reducing the risk of model overfitting.

When raw displacement data is input into artificial intelligence prediction models, the prediction error is often high. Many researchers use wavelet decomposition (WD) to combine signal analysis and reconstruction with deep learning algorithms for prediction18, 19. This approach extracts features at different scales, improving the accuracy of time-series data analysis and prediction.

Neural network models are highly dependent on the selection of hyperparameters, which directly affect the model's fitting ability, training speed, and prediction results. Therefore, to address the difficulty of determining hyperparameters in neural networks, intelligent swarm optimization algorithms can be used to automatically determine them, thus mitigating the timeliness issues associated with manually selecting parameters20. Lu et al.21. used constrained differential evolution to solve the constrained optimization problem of hidden sparse network attacks. However, extreme optimization is more effective in terms of computational cost, memory requirements, and configurable parameters. Lu et al.22. optimized the parameters of a proportional-integral control system using constrained swarm extreme optimization. Particle Swarm Optimization (PSO) is a population-based metaheuristic optimization method that can quickly and easily navigate a large solution space, providing reliable and efficient solutions without requiring prior expert knowledge. However, traditional PSO methods have the drawback of converging too quickly and only finding local optimal solutions. Maghsoudy, S et al.23. combined the Genetic Algorithm (GA) with the PSO algorithm, but the computational cost was significant. Therefore, this paper develops an Improved PSO (IPSO) algorithm based on their approach, introducing appropriate mutation and crossover rates to enhance the diversity of the particles. Additionally, a nonlinear inertia weight ω and adjusted cognitive factor c1 and social factor c2 are chosen to improve the convergence speed and accuracy of the algorithm.

In summary, this study proposes a GRU model based on Wavelet Decomposition (WD) and uses Improved Particle Swarm Optimization (IPSO) to optimize the model's hyperparameters, reducing the difficulty of extracting features from non-stationary displacement sequences. First, WD decomposes the displacement data into subsequences of different scales and positions, analyzing the features at various time–frequency levels. Next, a Dropout layer is added to the GRU network architecture to prevent overfitting, and the extracted data is fed into the GRU network. Finally, the IPSO algorithm is used to optimize the GRU network's parameters, and the predictions of different subsequences are combined. The proposed method significantly improves the prediction accuracy of slope surface displacement sequences during tunnel excavation conditions.

The main contributions and innovations of this study are as follows:A time-series prediction model based on Wavelet Decomposition (WD) is proposed for predicting surface displacement during tunnel excavation in base-overburden type slopes. The WD technique is used to extract displacement features at different time–frequency levels.

By using appropriate mutation and crossover rates, the particles gain greater diversity. The IPSO algorithm, with nonlinear decreasing inertia weight ω, adjusted cognitive factor c1, and social factor c2, is introduced to optimize the GRU model parameters, improving the accuracy and reliability of surface displacement predictions.

The generalization ability of the WD-IPSO-GRU model is validated through comparisons with multiple machine learning models combining different data preprocessing techniques and swarm intelligence algorithms.

Bootstrap validation is used to perform uncertainty analysis on the WD-IPSO-GRU model across subsequences of different datasets, confirming the reliability of the model.

Methodology

Wavelet decomposition

Wavelet Decomposition (WD) is a signal processing technique used to decompose a signal into wavelets with different frequency components and analyze the characteristics of each component at different time points. Wavelet decomposition not only provides frequency information but also time information, making it advantageous for processing non-stationary signals.

When WD processes the data, the signal is decomposed into a set of wavelet basis functions. These basis functions are generated by scaling and shifting an original wavelet function. The main steps of wavelet decomposition are as follows:

Step 1: Select an appropriate wavelet function and the number of decomposition levels. The choice of the wavelet function depends on the characteristics of the signal and the purpose of the analysis. Common wavelet functions include Haar wavelet, Daubechies wavelet, and Coiflets wavelet, among others. The number of decomposition levels determines the extent to which the signal is decomposed, usually based on the signal's frequency range and the required resolution.

Step 2: Through multi-level decomposition of the signal, the signal is decomposed into approximation coefficients (A) and detail coefficients (D) at different resolutions. The approximation coefficients represent the low-frequency components of the signal, while the detail coefficients represent the high-frequency components. The decomposition process can be represented as follows:1 aj+1(t)=H(aj(t))dj+1(t)=G(dj(t))j=1,2,...,J

where aj(t) and dj(t) are the low-frequency and high-frequency information of the original signal at 2-j, respectively. H and G represent the low-frequency and high-frequency signals24, respectively, and j is the number of wavelet decomposition levels.

Because the wavelet decomposition process is based on binary sampling, the length of each sub-sequence is half the length of the preceding sequence. Therefore, it is necessary to reconstruct the decomposed sequences. The reconstruction formula is as follows:2 Aj(t)=(H∗)jaj(t)Dj(t)=(H∗)jG∗dj(t)j=J-1,...,0

where H* and G* are the paired operators of H and G, respectively.

Reconstruct aJ and d1, d2, …to obtain the approximation sequence AJ and detail sequences D1, D2, …, DJ. The sequence S is as follows:3 S=AJ+D1+D2+⋯+DJ

PSO

Particle Swarm Optimization (PSO) is a population-based global optimization algorithm inspired by the behavior of biological groups such as flocks of birds or schools of fish. It is applied to solve various problems in engineering fields25. PSO simulates the social behavior of individuals (particles) within a group, using information exchange among individuals to find the optimal solution. Each particle represents a candidate solution, with its position and velocity continuously updated in the solution space. The algorithm calculates the optimization variables of multidimensional particles and compares the fitness values of the objective function to achieve global optimization, thereby obtaining better model parameters. The velocity and position update formulas of the PSO algorithm are shown in Eq. (4)26.4 vi,jt+1=ωvi,jt+c1r1(pbesti,jt-xi,jt)+c2r2(gbestjt-xi,jt)xi,jt+1=xi,jt+vi,jt+1

In Eq. (4), ω is the inertia weight; cognitive factor is c1 and social factor is c2; r1 and r2 are two independent random numbers distributed in [0,1]; vt i, j、xt i, j、pbestt i, j and gbestt i, j are the velocity component, position component, individual best value, and global best value of particle i in the t-th iteration, respectively.

GRU

GRU can be considered a variant of LSTM, designed to address the gradient vanishing problem of standard recurrent neural networks27. It is widely used in research areas related to natural language processing27, speech recognition28, and time series data29, and can effectively handle long time series data30. Its core structure (Fig. 1) consists of four parts: the update gate, the reset gate, the current memory content, and the final memory at the current time step.Fig. 1 GRU Core Structure.

1. Update Gate: The update gate zt helps the model determine information from previous time steps. When xt is input into the network unit, it is multiplied by its own weight W(z). ht-1 holds the information from the previous t-1 units and is multiplied by its own weight U(z). The two results are added together and then a sigmoid (σ) activation function is applied to control the result between 0 and 1.5 zt=σWzxt+Uzht-1

2. Reset Gate: The reset gate rt determines how much information from the previous hidden state h(t-1) the model should forget. Its calculation formula is the same as that of the update gate zt:6 rt=σWzxt+Uzht-1

3. Current Memory Content: The input xt is multiplied by the weight W, and ht-1 is multiplied by the weight U. The Hadamard product of rt and Uht-1 is calculated to determine which information from the previous hidden state ht-1 should be deleted. The combined information is then passed to the current candidate set ht' through the application of the nonlinear activation function tanh.7 ht′=tanhWxt+rt⊙Uht-1

4. Final Memory at the Current Time Step: As the final step, ht' includes information from the current time step as well as the output from the hidden layer of the previous time step. The update gate zt combines the output from the hidden layer of the previous time step ht-1 with the current candidate set ht' to obtain the hidden layer output at the current time step ht.8 ht=zt⊙ht-1+1-zt⊙ht′

Dropout

Dropout is a regularization technique used to prevent deep learning models (such as neural networks) from overfitting19. Overfitting refers to the phenomenon where a model performs well on training data but poorly on test data or new data. Overfitting usually occurs when the model is too complex, learning the noise in the training data rather than its underlying structure. In a dropout network, the output of each layer is multiplied element-wise by a Bernoulli random variable before being used as the input to the next layer. This method randomly drops out some neurons, preventing overfitting and improving the model's generalization ability.

For example, consider a neural network with L hidden layers. Let l ∈ {1,…, L} index the hidden layers of the network. Let z(l) denote the input vector to the l-th layer, and y(l) denotes the output vector of the l-th layer (y(0) = x as the input). W(l)and b(l) are the weights and biases of the l-th layer, and f is an arbitrary activation function. The forward propagation operation of a standard neural network can be described as follows:9 zi(l+1)=wi(l+1)yl+bi(l+1)yi(l+1)=f(zi(l+1))

For any layer, a Bernoulli random variable r(l) (with probability p of being 1) is introduced. The product of this random variable and the output y(l) generates a sparse output. This sparse output is used as the input for the next layer. Essentially, this means extracting a sub-network from the more complex neural network for training17. After introducing the Dropout technique, the forward propagation operation of the neural network becomes:10 rj(l)∼Bernoulli(p)y~(l)=r(l)∗y(l)zi(l+1)=wi(l+1)y~(l)+bi(l+1)yi(l+1)=f(zi(l+1))

The comparison of the neural network before and after introducing the Dropout technique is shown in Fig. 2A,B.Fig. 2 Dropout Technique Schematic.

IPSO

To address the issue of network structure and hyperparameter selection in the GRU model, this paper employs the PSO algorithm. In the PSO algorithm, the size of the inertia weight represents the particle's ability to maintain its previous motion state. In the original algorithm, the inertia weight is a fixed value, which reduces the global search capability and the convergence speed of the particles, making it easier to fall into local optima early on. This study uses a nonlinear inertia weight ω20 and adjusted cognitive factor c1 and social factor c226 combined with the diversity obtained from the crossover and mutation operations of the genetic algorithm23, to improve the original PSO algorithm (IPSO). The linear inertia weight, nonlinear inertia weight, and fine-tuned c1 and c2 are shown in Eq. (13):11 ω=ωmax-(ωmax-ωmin)∗tk

12 ω=ωmin+(ωmax-ωmin)∗e-7ttmax3

13 c1=c1max-tk(c1max-c1min)c2=c2min+tk(c2max-c2min)

where tmax represents the maximum number of iterations. The particle swarm parameters used in this study are listed in Table 1. Table 1 IPSO Algorithm Parameters.

ωmax	ωmin	c1max	c1min	c2max	c2min	
0.9	0.3	2..25	0.5	2.25	1.25	

To verify the superiority of the IPSO algorithm, this paper selects three test functions: Schaffer, Rastrigin, and Schwefel, to test the original PSO and IPSO20. The specific parameters for the Schaffer, Rastrigin, and Schwefel functions are shown in Table 2. The iteration curves of the two PSO algorithms are shown in Fig. 3. The experimental results indicate that the IPSO algorithm, with a nonlinear decreasing inertia weight ω and fine-tuned c1 and c2, outperforms the original PSO algorithm in terms of overall convergence speed and accuracy. Therefore, after assessing its effectiveness, the IPSO algorithm is used to optimize the hyperparameters of the GRU model. Table 2 Test Functions.

Test function	Function formula	Range	
Schaffer	f(x1,x2)=0.5+(sinx12+x22)2-0.5(1+0.001(x12+x22))2	-10≤x1,x2≤10	
Rastrigin	f(x1,x2)=20+x12+x22-10[cos(2πx1)+cos(2πx2)]	-5.12≤x1,x2≤5.12	
schwefel	f(x1,x2)=-x1∗sin(x1)-x2∗sin(x2)	-500≤x1,x2≤500	

Fig. 3 Test Functions.

Brief overview of GWO, HHO and SMA

The Grey Wolf Optimizer (GWO)31 is a swarm intelligence optimization algorithm based on the hunting behavior of grey wolves. It simulates the social hierarchy and hunting strategies of grey wolves, with the pack led by alpha (α), beta (β), and delta (δ) wolves, representing the best, second-best, and third-best solutions, respectively. By simulating the relative distance between the grey wolves and their prey, the positions are continuously adjusted to find the optimal solution.

The Harris Hawks Optimization (HHO)32 algorithm is an optimization algorithm based on the hunting behavior of Harris hawks. This algorithm achieves optimization by simulating the cooperative and pouncing behaviors of Harris hawks during different stages of hunting. By mimicking the dynamic changes of the prey's escape, the algorithm adjusts the hawks' pursuit strategies to improve the success rate of the hunt.

The Slime Mould Algorithm (SMA)33 is an optimization algorithm based on the path optimization behavior exhibited by slime moulds (such as amoebae) during their food search process. By simulating the expansion and contraction behaviors of slime moulds, the algorithm seeks the optimal solution. It selects paths based on the probability distribution along the routes, choosing the path most likely to approach the optimal solution.

GWO, HHO, and SMA all belong to swarm intelligence optimization methods. They perform global and local searches in the solution space by simulating the cooperative and competitive behaviors among individuals within a group to find the optimal solution to a problem. Each algorithm employs specific rules and mechanisms, leveraging interactions among group members to continuously update and optimize the positions of solutions, ultimately approaching the optimal solution.

The WD-IPSO-GRU predictive model

The WD-IPSO-GRU hybrid prediction model constructed in this paper combines wavelet decomposition signal processing technology, a particle swarm algorithm improved by genetic algorithm concepts, and gated recurrent units to achieve better time series prediction performance. The WD technique decomposes the time series data of tunnel slope surface displacement into approximation sequences and detail sequences, effectively reducing noise while retaining important information and extracting key features, thereby enhancing the analysis and recognition capabilities of the displacement data. The IPSO solves the issues of GRU network structure and training parameter selection, enabling the GRU model to perform better in long-time series prediction tasks.

To more clearly and comprehensively express the WD-IPSO-GRU, the specific steps are more intuitively and concisely presented in Fig. 4. Figure 5 shows the detailed steps of the developed model. In this paper, the decomposition level l of WD is determined based on multiple experiments.Fig. 4 WD-IPSO-GRU Model Flowchart.

Fig. 5 Network Structure of the WD-IPSO-GRU Model.

The overall process of WD-IPSO-GRU is as follows:

Step 1: Select an appropriate wavelet basis function and decomposition level. Perform wavelet decomposition on the time series data X(i) of tunnel slope surface displacement to obtain approximation coefficients AJ and detail sequences D1, D2, …, DJ。

Step 2: Initialize the parameters of the IPSO algorithm, such as population size, number of particles, number of iterations, particle velocity and position, mutation and crossover probabilities, and termination conditions. Use the number of neural units in the three hidden layers hN, HN, LN and the dropout rates D1, D2 in the two Dropout layers as the parameters to be optimized, and set appropriate ranges based on experience.

Step 3: Use Mean Absolute Error (MSE) as the fitness function to calculate the fitness value of each particle.

Step 4: Compare the current fitness value of each particle with its historical fitness values to generate the local optimal solution for each particle; compare the current fitness value of the particle with all historical fitness values to obtain the global optimal solution.

Step 5: Update the position and velocity of each particle according to Eq. (12) and (13).

Step 6: Repeat Steps (3) to (5) until the termination condition of the algorithm is met.

Step 7: Input the parameters output by the IPSO algorithm into the GRU model to obtain the prediction results.

Model calculations

Data selection

Taking a typical case of a base-overlying slope tunnel in the Kangding region of western Sichuan, the slope tunnel exit area is located in a high mountain canyon mudslide terrace, The general topography is high in the northwest and low in the southeast, with the slope inclined 45° to the south-southeast. The area northeast of the tunnel exit is a naturally weathered base-overlying slope. The 3D diagram of the slope is shown in Fig. 6. The overlying soil layer has a thickness of approximately 16–70 m, with a natural slope angle of 35–45°. There are no visible signs of tension cracks, stepped platforms, or bulging deformation on the slope surface. The bedrock is primarily composed of slate and sandstone, with sporadic bedrock outcrops in some areas. Due to tunnel construction, the mountain is being excavated on-site. The tunnel exit direction is due west, forming a 45° angle with the slope inclination. The tunnel exit is located on the south side of the mudslide terrace.Fig. 6 3D Diagram of Slope Monitoring Points.

In the mudslide terrace area of the slope and the natural slope area, GNSS deformation monitoring stations are set up along different profiles. As shown in Fig. 6, monitoring point 01 is located in the mudslide terrace area of the slope, at the tunnel exit profile; monitoring point 02 is located in the natural slope area where anti-slide piles are set up.

Before tunnel excavation, the slope was reinforced and anchored with anti-slide piles and prestressed anchor rods (cables) in different areas. After the slope support was completed, the tunnel excavation was carried out. Numerical simulation results indicate that the excavation of the tunnel and anti-slide piles causes deformation of the slope. Therefore, monitoring points 01 and 02 were selected as important points to monitor the local stability and deformation trends of the slope. Data collection was performed using GNSS monitoring instruments, powered by a solar power system, with a monitoring frequency of once per hour. The monitoring data can be transmitted wirelessly in real-time, and the monitoring points have been in normal operation since December 2021. For this study, 2700 monitoring data points from May to August 2022 were selected for analysis.

Data pre-processing

In this study, the wavelet basis function for WD is Db3, and the decomposition level is 4. The displacement time series data are decomposed and reconstructed, with the wavelet decomposition and reconstruction of the sub-sequences maintaining the same sequence length as the original data. Figure 7 shows the original sequence and the wavelet decomposition results. The approximation sequence (A4) exhibits the same trend and pattern as the displacement time series, while the detail sequences (D1, D2, D3, D4) represent the internal detail information and frequency characteristics of the displacement time series. Since the two types of sequences have different data characteristics, we need to perform IPSO-GRU modeling and prediction for each sub-sequence separately.Fig. 7 WD Decomposition Results.

Evaluation criterion

In this paper, the model performance is evaluated using R2, VAF, WI, MAE, MAPE, and RMSE indicators26. RMSE (Root Mean Square Error) represents the average error between the predicted values and the actual values; MAPE (Mean Absolute Percentage Error) represents the average absolute percentage error between the predicted values and the actual values; MAE (Mean Absolute Error) represents the average absolute error between the predicted values and the actual values; VAF (Variance Accounted For) represents the proportion of variance explained by the model; and WI (Willmott's Index of Agreement) is the Willmott consistency index. The calculation formulas are as follows:14 R2=1-∑(yi-y^i)2∑(yi-y¯)2

15 VAF=(1-var(yi-y¯i)var(yi))∗100

16 WI=1-∑i=1n(yi-y¯i)2∑i=1ny¯i-yavg+yi-yavg2

17 MAE=1n∑i=1ny~i-yi

18 MAPE=1n∑i=1n‖yi-y~i‖‖yi‖

19 MSE=1n∑i=1n(yi-y~i)2

20 RMSE=(∑(yi-y~i)2)n

where y~i is the model's predicted value, and yi is the actual observed value.

The smaller the RMSE, MAPE, and MAE, the better the model's prediction performance. R2 is a measure of the goodness of fit for a linear regression model, with values ranging between 0 and 1. The closer R2 is to 1, the better the model's fit; the closer R2 is to 0, the worse the model's fit. The closer WI is to 1, the better the model's prediction performance; the closer WI is to 0, the worse the model's prediction performance. The higher the VAF value, the larger the proportion of variance explained by the model, indicating better prediction performance.

Comparison before and after introducing VMD and WD

Yang et al.34. have demonstrated that Variational Mode Decomposition (VMD) is superior to Empirical Mode Decomposition (EMD) and Complete Ensemble Empirical Mode Decomposition (CEEMD) when integrated with the LSTM model for predicting time-series data. The GRU model has a simpler structure and higher computational efficiency than the LSTM model. Therefore, to explore the advantages of WD in handling this dataset, this study uses GRU, LSTM, RNN, and XGBoost models for training and predicting at measurement points 01 and 02. Additionally, WD and VMD techniques are introduced for comparison before and after model optimization.

The training, validation, and test sets are divided in chronological order with ratios of 0.8, 0.1, and 0.1, respectively. The time window is set to 24, and the loss function is MSE, with the learning rate determined by the Adam optimizer. The training period is 100 epochs, and the number of neural units in the three hidden layers is 256, 128, and 64. The two Dropout layers have rates of 0.2 and 0.234. The wavelet decomposition (WD) uses Db3 as the basis function with four decomposition levels, while the Variational Mode Decomposition (VMD) has five intrinsic mode functions (IMFs). Additional parameters are detailed in Table 3. The entire training process is implemented using Python 3.9.7 with the PyCharm Professional 2021 development environment. The operating system used is Windows 11, running on two Intel Xeon Scalable Silver 4510 CPUs, with 128 GB of RAM and a 64-bit operating system and 32 GB of memory. Table 3 Model Parameter Settings.

Parameters	Models	Parameters	Models	
	GRU	LSTM	RNN	XGBoost35	
Hidden layer neurons hN	256	256	256	Number of trees n	300	
Hidden layer neurons HN	128	128	128	Learning rate lr	0.2	
Hidden layer neurons LN	64	64	64	Maximum depth of Trees max_d	6	
Iterations k	100	100	100			
Learning rate lr	Adam	Adam	Adam	Minimum sum of child node weights min_c_w	2	
Dropout rateD1	0.3	0.3	0.3			
Dropout rateD2	0.2	0.2	0.2	L2 regularization term weight reg_l	0.5	
Loss function F	MSE	MSE	MSE	Loss function F	MSE	

Figure 8 shows the training and prediction results at Measurement Point 01 using the GRU, LSTM, RNN, and XGBoost models, as well as the comparison of prediction results after introducing the WD and VMD techniques. Based on the fit between the true values and predicted values in the scatter plots, the prediction intervals of the models with WD introduced are closer to the regression line, indicating better prediction performance (Table 4).Fig. 8 Prediction Results at Measurement Point 01 Before and After Introducing VMD and WD.

Table 4 Model Evaluation.

Models	Measurement point 01/mm	Measurement point 02/mm	
R2	VAF	WI	MAE	MAPE	RMSE	R2	VAF	WI	MAE	MAPE	RMSE	
GRU	0.96	96.40	0.99	0.12	4.31	0.15	0.72	75.92	0.92	0.10	1.89	0.15	
LSTM	0.96	96.73	0.99	0.12	4.36	0.15	0.58	75.67	0.90	0.13	2.60	0.18	
RNN	0.95	95.35	0.99	0.13	4.97	0.18	0.73	73.09	0.93	0.09	1.79	0.15	
XGBoost	0.96	96.02	0.99	0.13	4.80	0.16	0.94	94.00	0.98	0.05	1.03	0.07	
WD-GRU	0.98	98.88	1.00	0.08	3.03	0.08	0.96	98.31	0.99	0.04	0.67	0.05	
WD-LSTM	0.98	98.83	1.00	0.08	3.18	0.10	0.96	97.66	0.99	0.04	0.86	0.06	
WD-RNN	0.98	98.58	1.00	0.09	3.20	0.10	0.96	97.71	0.99	0.04	0.78	0.05	
WD-XGBoost	0.97	96.94	0.99	0.11	3.93	0.14	0.83	83.35	0.96	0.08	1.51	0.12	
VMD-GRU	0.97	97.49	0.99	0.10	3.60	0.13	0.79	82.04	0.94	0.09	1.69	0.13	
VMD-LSTM	0.97	97.37	0.99	0.10	3.81	0.14	0.78	81.74	0.94	0.09	1.78	0.13	
VMD-RNN	0.97	96.62	0.99	0.12	4.41	0.15	0.80	81.51	0.95	0.08	1.62	0.13	
VMD-XGBoost	0.97	96.97	0.99	0.11	3.74	0.14	0.91	91.12	0.98	0.06	1.08	0.08	

Figure 9 visually presents a comparison of the evaluation metrics. The average Mean Absolute Error (MAE) for predictions at Measurement Point 01 using raw data is 0.125. After introducing the VMD technique, the average MAE is 0.1075, and after introducing the WD technique, the average MAE is 0.09. Compared to using raw data alone, the prediction accuracy improved by 14% and 28%, respectively. To further validate the advantage of introducing the WD technique, the same model parameters were used for training and prediction at Measurement Point 02, with the prediction results shown in Figs. 10 and 11. The model with the VMD technique has an average MAE of 0.08, while the model with the WD technique has an average MAE of 0.05, improving prediction accuracy by 13.5% and 45.9%, respectively. This further verifies the advantage of introducing the WD technique.Fig. 9 Evaluation Metrics at Measurement Point 01 Before and After Introducing VMD and WD.

Fig. 10 Prediction Results at Measurement Point 02 Before and After Introducing VMD and WD.

Fig. 11 Evaluation Metrics at Measurement Point 01 Before and After Introducing VMD and WD.

Comparison of IPSO, GWO, HHO, and SWA algorithms

To further verify the superiority of the Improved Particle Swarm Optimization (IPSO) algorithm developed in this study when handling this dataset, we also included the Grey Wolf Optimizer (GWO), Harris Hawks Optimization (HHO), and Slime Mould Algorithm (SMA) for comparison. In this study, the inertia weight w of the IPSO is defined in Eq. (12), while c1 and c2 are defined in Eq. (13). The population size for the IPSO, GWO, HHO, and SMA algorithms is 30, with 50 iterations, and the loss function is MSE. The optimization range for the number of hidden layer units is [50, 300], and the optimization range for the Dropout layer is [0.01, 0.3]. The optimal parameters of the GRU model optimized by IPSO, GWO, HHO, and SMA are shown in Table 5. Table 5 Model Parameter Settings.

Optimal Parameters	Measurement point 01/mm	Measurement point 02/mm	
IPSO	GWO	HHO	SMA	IPSO	GWO	HHO	SMA	
Hidden layer neurons hN	283	280	220	276	314	300	210	300	
Hidden layer neurons HN	159	176	120	200	176	141	112	184	
Hidden layer neurons LN	89	83	56	60	76	85	55	98	
Dropout rate D1	0.2	0.12	0.1	0.1	0.2	0.16	0.1	0.1	
Dropout rate D2	0.1	0.11	0.1	0.15	0.1	0.12	0.1	0.1	

Figure 12 shows the training and prediction results at Measurement Point 01 using the optimal parameters. Based on the fit between the true values and predicted values in the scatter plots, the prediction intervals of the IPSO algorithm-optimized model are closer to the regression line, indicating better prediction performance. Figure 13 visually presents a comparison of the evaluation metrics, where the average MAE for predictions at Measurement Point 01 is 0.075 for IPSO, 0.1175 for GWO, 0.115 for HHO, and 0.1175 for SMA. Compared to using raw data directly for prediction, the prediction accuracy improved by 40%, 6%, 8%, and 6%, respectively.Fig. 12 Validation Set Prediction Results at Measurement Point 01 After Algorithm Optimization.

Fig. 13 Evaluation Metrics of Validation Set at Measurement Point 01 After Algorithm Optimization.

To further verify the advantages of IPSO, the same model parameters were used for training and prediction at Measurement Point 02, as shown in Figs. 14 and 15. The average MAE for predictions at Measurement Point 02 is 0.0475 for IPSO, 0.0875 for GWO, 0.09 for HHO, and 0.0825 for SMA. Compared to using raw data directly for prediction, the prediction accuracy improved by 48.6%, 5.4%, 2.7%, and 10.8%, respectively. This further validates the superiority of the IPSO algorithm.Fig. 14 Validation Set Prediction Results at Measurement Point 02 After Algorithm Optimization.

Fig. 15 Evaluation Metrics of Validation Set at Measurement Point 02 After Algorithm Optimization.

To ensure a fair comparison, the WD and VMD techniques were introduced to the models optimized by GWO, HHO, SMA, and IPSO to highlight the advantages of the WD-IPSO-GRU model. Figure 16 shows the training and prediction results at Measurement Point 01 after using the optimal parameters and introducing WD and VMD techniques. Based on the maximum and minimum values, as well as the median line and mean in the box plot, the prediction intervals of the WD-IPSO-GRU model are closer to the actual values, indicating better prediction performance. Figure 17 visually presents a comparison of evaluation metrics. The MAE for WD-IPSO-GRU at Measurement Point 01 is 0.03, indicating a 76% improvement in prediction accuracy compared to using raw data directly. To further verify the advantages of the WD-IPSO-GRU model, the same model parameters were used for training and prediction at Measurement Point 02, where the MAE is 0.04, showing a 56.7% improvement in prediction accuracy compared to using raw data directly. This further validates the superiority of the WD-IPSO-GRU model.Fig. 16 Prediction Results of Optimized Algorithms with WD and VMD.

Fig. 17 Evaluation Metrics of Optimized Algorithms with WD and VMD.

In the previous section, we used IPSO, GWO, HHO, and SMA algorithms to validate three test functions: Schaffer, Rastrigin, and Schwefel (Fig. 3). To more intuitively demonstrate the superiority of IPSO-GRU, the loss curve of the loss function as the number of iterations increases is shown in Fig. 18. At Measurement Point 01, the IPSO algorithm has the fastest convergence speed and the lowest convergence value. At Measurement Point 02, the convergence speed of IPSO is slightly lower than GWO and HHO in the early stages of training, but the convergence speed and accuracy of the IPSO algorithm significantly improve in the middle and later stages. Overall, the IPSO algorithm achieves the best optimization effect on the GRU model for this dataset. Compared to hyperparameters optimized by GWO, HHO, and SMA, the improved IPSO in this paper is more suitable for this dataset, with better generalization ability and predictive performance (Tables 6, 7).Fig. 18 Loss Curve.

Table 6 Model Evaluation.

Models	Measurement point 01/mm	Measurement point 02/mm	
R2	VAF	WI	MAE	MAPE	RMSE	R2	VAF	WI	MAE	MAPE	RMSE	
GWO-GRU	0.96	96.45	0.99	0.10	4.49	0.16	0.72	74.65	0.93	0.08	1.87	0.15	
HHO-GRU	0.97	96.60	0.99	0.11	4.08	0.15	0.76	76.11	0.94	0.08	1.57	0.14	
SMA- GRU	0.96	96.60	0.99	0.10	4.37	0.15	0.74	75.31	0.94	0.09	1.72	0.14	
IPSO- GRU	0.98	98.68	1.00	0.08	2.83	0.10	0.96	96.10	0.99	0.04	0.85	0.06	
GWO-LSTM	0.97	96.70	0.99	0.11	4.32	0.15	0.64	75.34	0.91	0.10	2.31	0.17	
HHO- LSTM	0.97	96.70	0.99	0.11	4.11	0.15	0.77	76.74	0.94	0.08	1.56	0.14	
SMA- LSTM	0.97	96.69	0.99	0.11	4.09	0.15	0.71	76.48	0.93	0.08	1.98	0.15	
IPSO- LSTM	0.98	99.00	1.00	0.07	2.68	0.09	0.94	94.04	0.99	0.05	1.00	0.07	
GWO-RNN	0.91	92.13	0.98	0.15	6.52	0.23	0.75	75.99	0.93	0.09	1.72	0.14	
HHO- RNN	0.96	96.53	0.99	0.12	4.75	0.16	0.37	71.72	0.85	0.12	3.54	0.22	
SMA- RNN	0.95	95.82	0.99	0.12	5.94	0.18	0.65	74.23	0.92	0.08	2.27	0.17	
IPSO- RNN	0.98	98.66	1.00	0.08	3.13	0.10	0.93	93.19	0.98	0.05	0.98	0.08	
GWO-XGBoost	0.97	96.52	0.99	0.11	4.17	0.15	0.75	75.20	0.94	0.08	1.62	0.14	
HHO- XGBoost	0.97	96.71	0.99	0.12	4.27	0.15	0.65	77.13	0.91	0.12	2.31	0.17	
SMA- XGBoost	0.94	95.49	0.99	0.14	5.85	0.19	0.64	75.19	0.91	0.08	2.38	0.17	
IPSO- XGBoost	0.98	98.84	1.00	0.07	2.91	0.09	0.92	92.45	0.98	0.05	1.02	0.08	

Table 7 Model Evaluation.

Models	Measurement point 01/mm	Measurement point 02/mm	
R2	VAF	WI	MAE	MAPE	RMSE	R2	VAF	WI	MAE	MAPE	RMSE	
WD-GWO-GRU	0.98	98.89	1.00	0.08	3.02	0.09	0.95	98.26	0.99	0.06	1.10	0.07	
VMD-GWO-GRU	0.98	97.59	0.99	0.09	3.46	0.12	0.82	82.34	0.95	0.08	1.47	0.12	
WD-HHO-GRU	0.98	98.83	1.00	0.08	3.18	0.09	0.96	98.20	0.99	0.04	0.80	0.05	
VMD-HHO-GRU	0.98	97.58	0.99	0.10	3.55	0.13	0.80	82.39	0.94	0.08	1.61	0.12	
WD-SMA- GRU	0.98	98.87	1.00	0.08	3.10	0.09	0.88	98.21	0.97	0.09	1.76	0.10	
VMD-SMA- GRU	0.98	97.63	0.99	0.09	3.44	0.12	0.82	82.48	0.95	0.08	1.48	0.12	
WD-IPSO- GRU	0.99	99.85	1.00	0.03	1.09	0.04	0.97	98.31	0.99	0.04	0.67	0.05	
VMD-IPSO- GRU	0.98	97.59	0.99	0.10	3.49	0.12	0.82	82.45	0.95	0.08	1.53	0.12	
WD-GWO-LSTM	0.98	98.80	1.00	0.08	3.06	0.09	0.93	97.99	0.98	0.07	1.25	0.07	
VMD-GWO-LSTM	0.97	97.25	0.99	0.11	4.13	0.15	0.81	81.95	0.95	0.08	1.54	0.12	
WD-HHO- LSTM	0.98	98.77	1.00	0.08	3.21	0.09	0.95	97.91	0.99	0.06	1.05	0.07	
VMD-HHO- LSTM	0.97	97.30	0.99	0.10	3.67	0.13	0.81	81.92	0.95	0.08	1.54	0.12	
WD-SMA- LSTM	0.98	98.84	1.00	0.08	3.13	0.10	0.86	98.05	0.97	0.10	1.94	0.11	
VMD-SMA- LSTM	0.97	97.43	0.99	0.10	3.59	0.13	0.80	81.92	0.94	0.08	1.64	0.13	
WD-IPSO- LSTM	0.98	98.68	1.00	0.08	3.23	0.09	0.95	98.09	0.99	0.05	1.01	0.06	
VMD-IPSO- LSTM	0.97	97.37	0.99	0.10	3.85	0.14	0.81	81.91	0.95	0.08	1.53	0.12	
WD-GWO-RNN	0.97	96.58	0.99	0.12	4.18	0.15	0.92	92.17	0.98	0.06	1.20	0.08	
VMD-GWO-RNN	0.88	90.60	0.97	0.19	6.76	0.28	0.80	80.66	0.94	0.08	1.63	0.13	
WD-HHO- RNN	0.98	99.47	1.00	0.07	3.08	0.08	0.98	97.98	0.99	0.03	0.61	0.04	
VMD-HHO- RNN	0.96	96.98	0.99	0.12	4.57	0.15	0.77	81.34	0.93	0.09	1.84	0.14	
WD-SMA- RNN	0.98	98.52	1.00	0.09	3.52	0.11	0.89	94.57	0.98	0.07	1.34	0.09	
VMD-SMA- RNN	0.97	96.96	0.99	0.11	3.94	0.14	0.80	81.90	0.94	0.09	1.68	0.13	
WD-IPSO- RNN	0.98	98.62	1.00	0.08	3.02	0.10	0.93	97.09	0.98	0.06	1.15	0.07	
VMD-IPSO- RNN	0.97	97.35	0.99	0.11	4.12	0.15	0.80	81.15	0.94	0.08	1.59	0.12	
WD-GWO-XGBoost	0.97	96.96	0.99	0.11	3.91	0.14	0.83	83.20	0.96	0.08	1.46	0.12	
VMD-GWO-XGBoost	0.95	95.36	0.99	0.13	4.55	0.17	0.91	90.79	0.97	0.06	1.08	0.09	
WD-HHO- XGBoost	0.97	97.05	0.99	0.10	3.90	0.14	0.83	83.06	0.96	0.08	1.49	0.12	
VMD-HHO- XGBoost	0.95	94.58	0.99	0.14	4.93	0.19	0.91	91.15	0.98	0.05	1.05	0.08	
WD-SMA- XGBoost	0.97	96.93	0.99	0.11	3.91	0.14	0.84	83.89	0.96	0.07	1.43	0.11	
VMD-SMA- XGBoost	0.95	95.46	0.99	0.13	4.68	0.17	0.90	90.51	0.97	0.06	1.09	0.09	
WD-IPSO-XGBoost	0.97	96.93	0.99	0.10	3.85	0.14	0.82	81.72	0.95	0.08	1.49	0.12	
VMD-IPSO-XGBoost	0.95	95.26	0.99	0.13	4.48	0.17	0.91	91.26	0.98	0.06	1.06	0.08	

Uncertainty analysis

In the above article, the advantage of WD was validated by comparing the direct use of raw data with the introduction of WD and VMD techniques. The superiority of IPSO was demonstrated by comparing it with GWO, HHO, and SMA algorithms. To more comprehensively verify the prediction accuracy and stability of the WD-IPSO-GRU model constructed in this paper, the bootstrap method was used for 100 resampling validations on the A4, D4, D3, D2, and D1 subsequences of Measurement Points 01 and 02 (Fig. 7).

Figure 19 illustrates the distribution of Mean Squared Error (MSE) across different experimental groups, generated through the Bootstrap Method and visualized using histograms and a 0.95 confidence interval (red dashed lines), with the confidence interval's lower and upper limits calculated as P2.5 and P97.5, respectively. Specifically, we conducted bootstrap analysis on the subsequences of Measurement Points 01 and 02 by randomly sampling with replacement from the original data to create multiple bootstrap samples and calculating the MSE for each sample. This process was repeated 100 times to generate the bootstrap distribution of the MSE. To better evaluate the generalization ability of the WD-IPSO-GRU model, we used the mean M (Eq. 21) and the standard error se (Eq. 22) as evaluation metrics.21 M=1n∑i=1nyi

22 se=sn=1n-1∑i=1n(yi-y¯)2n

Fig. 19 Bootstrap Validation Frequency Distribution.

The standard error is used to measure the precision of sample statistics, such as the mean. The smaller the standard error, the more precise the estimate of the sample statistic. The evaluation metrics are shown in Table 8, where the MSE mean M for the 01 A4 subsequence is 0.01448, the standard error se is 0.00228, and the confidence interval is 0.00269 to 0.02739. For the 02 A4 subsequence, the MSE mean M is 0.01532, the standard error se is 0.00233, and the confidence interval is 0.00329 to 0.02849. It can be seen that the average error for the 01 and 02A4 subsequences is low, the confidence intervals are narrow, and the accuracy is high. The average error for other subsequences is even lower, with narrower confidence intervals and higher accuracy. Therefore, the WD-IPSO-GRU model constructed in this study exhibits high prediction accuracy across all subsequences, demonstrating strong generalization ability and reliability. Table 8 Bootstrap Validation Evaluation Metrics.

Subsequences	M	P2.5	P97.5	se	
01A4	0.014480	0.002690	0.027390	0.002280	
01D4	0.000311	0.000194	0.000443	0.000027	
01D3	0.000615	0.000487	0.000729	0.000026	
01D2	0.000307	0.000248	0.000369	0.000012	
01D1	0.000664	0.000454	0.000900	0.000042	
02A4	0.015320	0.003290	0.028490	0.002330	
02D4	0.000299	0.000263	0.000335	0.000008	
02D3	0.000366	0.000288	0.000460	0.000016	
02D2	0.000372	0.000259	0.000495	0.000022	
02D1	0.000769	0.000503	0.001040	0.000050	

Measurement Point 01 is located on the mudslide terrace of the slope. Due to tunnel excavation beneath the mudslide terrace, the displacement at this measurement point increases. The WD-IPSO-GRU model predicts that the displacement at Measurement Point 01 will increase from approximately 2.7 mm to around 6.2 mm over the next 245 h. Measurement Point 02 is situated in the natural landslide area of the slope. After the completion of the anti-slide pile construction, the slope remains relatively stable. The WD-IPSO-GRU model predicts that the displacement at Measurement Point 02 will oscillate around 6.3 mm over the next 245 h, consistent with the characteristics of the actual displacement curve of a stable slope. The analysis results indicate that the displacement predicted by the WD-IPSO-GRU model can effectively represent the deformation trend of the slope. The test sets of all models at Measurement Points 01 and 02 are shown in Fig. 20.Fig. 20 Test Set Predictions.

Result and discussion

In the aforementioned section, an AI model for predicting surface displacement of a base-overburden slope during tunnel excavation was thoroughly evaluated. Upon the model's development, six evaluation metrics (R2, VAF, WI, MAE, MAPE, RMSE) were selected to assess the goodness of fit and the generalization ability of the model's predictions. The prediction effects of the model were compared in detail before and after introducing WD and VMD, compared to directly using raw data. The prediction accuracy improved twofold at Measurement Point 01 and 3.4-fold at Measurement Point 02 with WD compared to VMD.

Further optimization of the hyperparameters of GRU, LSTM, RNN, and XGBoost models was conducted using the IPSO, GWO, HHO, and SMA algorithms. Compared to directly using raw data for prediction, the models optimized by swarm intelligence algorithms showed significant improvement in prediction performance at both Measurement Point 01 and Measurement Point 02. Specifically, the improved IPSO algorithm, which incorporates genetic concepts with nonlinear inertia weight ω, and fine-tuned cognitive factor c1 and social factor c2, achieved prediction accuracy improvements of 5.43-fold and 4.3-fold at Measurement Points 01 and 02, respectively. The goodness of fit was 0.99 and 0.97 at these points, which is comparable to the level achieved by Wang et al.11, 36, 37, who used swarm intelligence algorithms to optimize LSTM for predicting landslide displacement. Moreover, the convergence comparison of algorithm iterations showed that IPSO reached superior solutions faster than PSO, GWO, HHO, and SMA. As seen in the iteration curves in Fig. 18, IPSO demonstrated the best overall convergence speed and accuracy, confirming the superiority of the IPSO algorithm in WD-GRU model development.

Additionally, to graphically compare the performance of the models, scatter regression plots, radar charts, and box plots were used to present a visual explanation of the prediction results. The scatter regression plots are used to evaluate the error between the actual and predicted values of the models. The closer the scatter of predicted values is to the regression line, the better the prediction performance. A 0.95 prediction interval was set, and the RMSE of the WD-IPSO-GRU model at Measurement Points 01 and 02 was 0.04 and 0.05, respectively. The radar charts illustrate the various evaluation metrics and the model’s prediction improvement capabilities. To save space in visual graphics, box plots were used to represent the prediction effects and evaluation metrics of the other optimized models (Figs. 16, 17).

Finally, an uncertainty analysis of the optimal model, WD-IPSO-GRU, was conducted by performing 100 resamplings with replacement on the A4, D4, D3, D2, and D1 subsequences at Measurement Points 01 and 02. The frequency and MSE of the prediction results for the newly generated data were visualized (Fig. 19). The average standard errors (se) for all subsequences at Measurement Points 01 and 02 were 0.00048 and 0.00049, respectively, demonstrating the excellent optimization capability of the WD-IPSO-GRU model. In the test set, the model constructed in this study was compared with all models in Tables 4, 6, and 7. The comparison of prediction results is shown in Fig. 20. The predicted displacements effectively reflect the deformation trends and characteristics of the slope, and the prediction performance is the best.

Overall, after decomposing the raw data using the WD technique and using the Dropout technique to prevent overfitting, along with optimizing the GRU network structure using the improved IPSO algorithm, the WD-IPSO-GRU hybrid model exhibits higher generalization ability for predicting surface displacement of base-overburden slopes during tunnel excavation conditions.

Conclusion

To efficiently predict base-overburden engineering slope displacement and assess slope stability, this paper constructs a WD-IPSO-GRU hybrid model. First, the WD technique is used to denoise slope displacement monitoring data from various field monitoring areas. Then, the IPSO-optimized GRU model is applied to predict base-overburden slope displacement under tunnel excavation conditions. The main conclusions are as follows:The Dropout technique reduces the risk of overfitting by introducing Bernoulli random variables r(l) into any layer, effectively extracting a subnetwork from the complex neural network for training. In this study, the prediction performance on the validation and test sets of all models further supports the potential of the Dropout technique to reduce overfitting risk.

The WD preprocessing technique decomposes raw displacement data into wavelets of different frequency components and analyzes the characteristics of each component at different time points. By decomposing and sequentially predicting displacement data at different frequencies, the predicted data is aggregated to obtain the overall prediction. When using the VMD preprocessing technique combined with the GRU model and the WD preprocessing technique combined with the GRU model, the prediction accuracy at Measurement Point 01 improved by 14% and 13.5%, respectively, compared to directly using raw data. At Measurement Point 02, the prediction accuracy improved by 28% and 45.9%, respectively, demonstrating that WD has better applicability than VMD.

This study improves the original PSO algorithm (IPSO) by integrating the concepts of the Genetic Algorithm (GA) with nonlinear inertia weight ω, and fine-tuned cognitive factor c1, and social factor c2. It is tested using the Schaffer, Rastrigin, and Schwefel functions. A comparative analysis was conducted on the parameter search iteration loss curves of the IPSO, GWO, HHO, and SWA algorithms on the model architecture. The analysis results indicate that the IPSO algorithm has a faster overall convergence speed and higher accuracy compared to the original PSO algorithm. Compared to directly using raw data for prediction, the accuracy improved by 76% and 56.7% at Measurement Points 01 and 02, respectively.

The aim of this study is to predict the surface displacement of base-overburden slopes during tunnel excavation conditions and to explore a hybrid model approach that is most suitable for predicting displacement in the base-overburden slopes of mountainous valley regions. The proposed WD-IPSO-GRU model offers higher prediction accuracy and better generalization ability, avoiding the problem of PSO models getting stuck in local optima. Tests at two measurement points in different geological regions validate the generality of the WD-IPSO-GRU model. Overall, WD-IPSO-GRU can effectively predict the surface displacement of base-overburden slopes during tunnel excavation in the mountainous valley regions of Northwest China.

The WD-IPSO-GRU model constructed in this study addresses the prediction problem of univariate time series data, and it has been trained and tested at various measurement points of a specific base-overburden slope, achieving favorable prediction results. Notably, WD-IPSO-GRU can further explore potential applications in geotechnical engineering, such as predicting deep slope deformation, stress, groundwater levels, etc. Additionally, there should be further and more comprehensive research on the accuracy of model evaluation, particularly regarding the advantages of the WD-IPSO-GRU model in predicting multivariate time series data, such as the combined effects of construction impacts and groundwater.

Author contributions

Conceptualization, M.Z. and S.C.; methodology, M.Z. and G.M.; software, G.M.; validation, G.M., M.Z., X.Z. and S.C.; formal analysis, G.M.; investigation, M.Z.; resources and data curation, X.Z.; writing original draft preparation, G.M.; writing review and editing, M.Z.; supervision, M.Z.; funding acquisition, S.C. and X.H.; All authors have read and agreed to the published version of the manuscript.

Data availability

The datasets used and analysed during the current study available from the corresponding author on reasonable request.

Competing interests

The authors declare no competing interests.

Publisher's note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
==== Refs
References

1. Niu W Hu X Lin B Detection and monitoring of potential geological disaster using SBAS-InSAR technology KSCE J. Civ. Eng. 2023 27 4884 4896 10.1007/s12205-023-0759-8
Niu, W. et al. Detection and monitoring of potential geological disaster using SBAS-InSAR technology. KSCE J. Civ. Eng. 27, 4884–4896 (2023).10.1007/s12205-023-0759-8
2. Bureau of Statistics of the People's Republic of China. China Statistical Yearbook [M]. Beijing: China Statistics Press. (2023) (in chinese)
3. Song D Liu X Chen Z Influence of tunnel excavation on the stability of a bedded rock slope: A case study on the mountainous area in southern Anhui, China KSCE J. Civ. Eng. 2021 25 114 123 10.1007/s12205-020-0831-6
Song, D. et al. Influence of tunnel excavation on the stability of a bedded rock slope: A case study on the mountainous area in southern Anhui, China. KSCE J. Civ. Eng. 25, 114–123 (2021).10.1007/s12205-020-0831-6
4. Hu D Hu Y Yi S Prediction method of surface settlement of rectangular pipe jacking tunnel based on improved PSO-BP neural network Sci. Rep. 2023 13 5512 10.1038/s41598-023-32189-0 37015985
Hu, D. et al. Prediction method of surface settlement of rectangular pipe jacking tunnel based on improved PSO-BP neural network. Sci. Rep. 13, 5512 (2023).37015985 10.1038/s41598-023-32189-0
5. Fang K Miao M Tang H Insights into the deformation and failure characteristic of a slope due to excavation through multi-field monitoring: A model test Acta Geotech. 2023 18 1001 1024 10.1007/s11440-022-01627-0
Fang, K. et al. Insights into the deformation and failure characteristic of a slope due to excavation through multi-field monitoring: A model test. Acta Geotech. 18, 1001–1024 (2023).10.1007/s11440-022-01627-0
6. Liu N Liang S Wang S THM model of rock tunnels in cold regions and numerical simulation Sci. Rep. 2024 14 3465 10.1038/s41598-024-53418-0 38342931
Liu, N. et al. THM model of rock tunnels in cold regions and numerical simulation. Sci. Rep. 14, 3465 (2024).38342931 10.1038/s41598-024-53418-0
7. Nava L Carraro E Reyes-Carmona C Landslide displacement forecasting using deep learning and monitoring data across selected sites Landslides 2023 20 2111 2129 10.1007/s10346-023-02104-9
Nava, L. et al. Landslide displacement forecasting using deep learning and monitoring data across selected sites. Landslides 20, 2111–2129 (2023).10.1007/s10346-023-02104-9
8. Yaghoubi E Yaghoubi E Khamees A A systematic review and meta-analysis of artificial neural network, machine learning, deep learning, and ensemble learning approaches in field of geotechnical engineering Neural Comput. Appl. 2024 135 108
Yaghoubi, E. et al. A systematic review and meta-analysis of artificial neural network, machine learning, deep learning, and ensemble learning approaches in field of geotechnical engineering. Neural Comput. Appl. 135, 108 (2024).
9. Zheng T Zhao QH Hu JB An IPSO-RNN machine learning model for soil landslide displacement prediction Arab. J. Geosci. 2021 14 1191 10.1007/s12517-021-07542-0
Zheng, T. et al. An IPSO-RNN machine learning model for soil landslide displacement prediction. Arab. J. Geosci. 14, 1191 (2021).10.1007/s12517-021-07542-0
10. Hochreiter S Jürgen schmidhuber; long short-term memory Neural Comput. 1997 9 8 1735 1780 10.1162/neco.1997.9.8.1735 9377276
Hochreiter, S. Jürgen schmidhuber; long short-term memory. Neural Comput. 9(8), 1735–1780 (1997).9377276 10.1162/neco.1997.9.8.1735
11. Jin A Yang S Huang X Landslide displacement prediction based on time series and long short-term memory networks Bull. Eng. Geol. Environ. 2024 83 264 10.1007/s10064-024-03714-w
Jin, A., Yang, S. & Huang, X. Landslide displacement prediction based on time series and long short-term memory networks. Bull. Eng. Geol. Environ. 83, 264 (2024).10.1007/s10064-024-03714-w
12. Li LM Wang CY Wen ZZ Landslide displacement prediction based on the ICEEMDAN, ApEn and the CNN-LSTM models J. Mt. Sci. 2023 20 1220 1231 10.1007/s11629-022-7606-0
Li, L. M. et al. Landslide displacement prediction based on the ICEEMDAN, ApEn and the CNN-LSTM models. J. Mt. Sci. 20, 1220–1231 (2023).10.1007/s11629-022-7606-0
13. Gao W Li Z Chen Q Modelling and prediction of GNSS time series using GBDT, LSTM and SVM machine learning approaches J. Geod. 2022 96 71 10.1007/s00190-022-01662-5
Gao, W. et al. Modelling and prediction of GNSS time series using GBDT, LSTM and SVM machine learning approaches. J. Geod. 96, 71 (2022).10.1007/s00190-022-01662-5
14. Chung, J., Gulcehre, C., Cho, K., & Bengio, Y. Empirical evaluation of gated recurrent neural networks on sequence modeling. In NIPS 2014 Workshop on Deep Learning, December 2014 (2014).
15. Liu ZQ Guo D Lacasse S Algorithms for intelligent prediction of landslide displacements J. Zhejiang Univ. Sci. 2020 21 412 429 10.1631/jzus.A2000005
Liu, Z. Q. et al. Algorithms for intelligent prediction of landslide displacements. J. Zhejiang Univ. Sci. 21, 412–429 (2020).10.1631/jzus.A2000005
16. Zhang W Li H Tang L Displacement prediction of Jiuxianping landslide using gated recurrent unit (GRU) networks Acta Geotech. 2022 17 1367 1382 10.1007/s11440-022-01495-8
Zhang, W. et al. Displacement prediction of Jiuxianping landslide using gated recurrent unit (GRU) networks. Acta Geotech. 17, 1367–1382 (2022).10.1007/s11440-022-01495-8
17. Srivastava N Hinton G Dropout: A Simple Way to Prevent Neural Networks from Overfitting J. Mach. Learn. Res. 2014 15 1929 1958
Srivastava, N. et al. Dropout: A Simple Way to Prevent Neural Networks from Overfitting. J. Mach. Learn. Res. 15, 1929–1958 (2014).
18. Hu S Liu P Qiao Y PM2.5 concentration prediction based on WD-SA-LSTM-BP model: A case study of Nanjing city Environ. Sci. Pollut. Res. 2022 29 70323 70339 10.1007/s11356-022-20744-7
Hu, S. et al. PM2.5 concentration prediction based on WD-SA-LSTM-BP model: A case study of Nanjing city. Environ. Sci. Pollut. Res. 29, 70323–70339 (2022).10.1007/s11356-022-20744-7
19. Wang S Lyu Tl Luo N Deformation prediction of rock cut slope based on long short-term memory neural network Int. J. Mach. Learn. Cyber. 2024 15 795 805 10.1007/s13042-023-01939-x
Wang, S. et al. Deformation prediction of rock cut slope based on long short-term memory neural network. Int. J. Mach. Learn. Cyber. 15, 795–805 (2024).10.1007/s13042-023-01939-x
20. Huang Y Huang Z Yu J Short-term load forecasting based on IPSO-DBiLSTM network with variational mode decomposition and attention mechanism Appl. Intell. 2023 53 12701 12718 10.1007/s10489-022-04174-z
Huang, Y. et al. Short-term load forecasting based on IPSO-DBiLSTM network with variational mode decomposition and attention mechanism. Appl. Intell. 53, 12701–12718 (2023).10.1007/s10489-022-04174-z
21. Lu K-D Wu Z-G Constrained-differential-evolution-based stealthy sparse cyber-attack and countermeasure in an AC smart grid IEEE Transact. Ind. Inform. 2022 18 8 5275 5285 10.1109/TII.2021.3129487
Lu, K.-D. & Wu, Z.-G. Constrained-differential-evolution-based stealthy sparse cyber-attack and countermeasure in an AC smart grid. IEEE Transact. Ind. Inform. 18(8), 5275–5285 (2022).10.1109/TII.2021.3129487
22. Lu K Zhou W Zeng G Zheng Y Constrained population extremal optimization-based robust load frequency control of multi-area interconnected power system Int. J. Electr. Power Energy Syst. 2019 105 249 271 10.1016/j.ijepes.2018.08.043
Lu, K., Zhou, W., Zeng, G. & Zheng, Y. Constrained population extremal optimization-based robust load frequency control of multi-area interconnected power system. Int. J. Electr. Power Energy Syst. 105, 249–271 (2019).10.1016/j.ijepes.2018.08.043
23. Maghsoudy S Zakerabbasi P Baghban A Connectionist technique estimates of hydrogen storage capacity on metal hydrides using hybrid GAPSO-LSSVM approach Sci. Rep. 2024 14 1503 10.1038/s41598-024-52086-4 38233572
Maghsoudy, S. et al. Connectionist technique estimates of hydrogen storage capacity on metal hydrides using hybrid GAPSO-LSSVM approach. Sci. Rep. 14, 1503 (2024).38233572 10.1038/s41598-024-52086-4
24. Wu X Qian JS Huang CH Short-term coalmine gas concentration prediction based on wavelet transform and extreme learning machine Math Probl Eng. 2014 2014 858260
Wu, X. et al. Short-term coalmine gas concentration prediction based on wavelet transform and extreme learning machine. Math Probl Eng. 2014, 858260 (2014).
25. Kusum D Bansal Jagdish Chand, mean particle swarm optimisation for function optimisation Comput. Intell. 2009 1 72 92
Kusum, D. Bansal Jagdish Chand, mean particle swarm optimisation for function optimisation. Comput. Intell. 1, 72–92 (2009).
26. Bardhan A Samui P ELM-based adaptive neuro swarm intelligence techniques for predicting the California bearing ratio of soils in soaked conditions Appl. Soft Comput. 2021 110 1568 4946 10.1016/j.asoc.2021.107595
Bardhan, A. et al. ELM-based adaptive neuro swarm intelligence techniques for predicting the California bearing ratio of soils in soaked conditions. Appl. Soft Comput. 110, 1568–4946 (2021).10.1016/j.asoc.2021.107595
27. Yurtsever M Unemployment rate forecasting: LSTM-GRU hybrid approach J. Labour. Market Res. 2023 57 18 10.1186/s12651-023-00345-8
Yurtsever, M. Unemployment rate forecasting: LSTM-GRU hybrid approach. J. Labour. Market Res. 57, 18 (2023).10.1186/s12651-023-00345-8
28. Subramanian B Olimov B Naik SM An integrated mediapipe-optimized GRU model for Indian sign language recognition Sci. Rep. 2022 12 11964 10.1038/s41598-022-15998-7 35831393
Subramanian, B. et al. An integrated mediapipe-optimized GRU model for Indian sign language recognition. Sci. Rep. 12, 11964 (2022).35831393 10.1038/s41598-022-15998-7
29. Mehra, S., Ranga, V. & Agarwal, R. A deep learning approach to dysarthric utterance classification with BiLSTM-GRU, speech cue filtering, and log mel spectrograms. J. Supercomput. (2024).
30. Xiang X Li X Zhang Y A short-term forecasting method for photovoltaic power generation based on the TCN-ECANet-GRU hybrid model Sci. Rep. 2024 14 6744 10.1038/s41598-024-56751-6 38509109
Xiang, X. et al. A short-term forecasting method for photovoltaic power generation based on the TCN-ECANet-GRU hybrid model. Sci. Rep. 14, 6744 (2024).38509109 10.1038/s41598-024-56751-6
31. Mirjalili S Mirjalili SM Lewis A Grey wolf optimizer Adv. Eng. Softw. 2014 69 46 61 10.1016/j.advengsoft.2013.12.007
Mirjalili, S., Mirjalili, S. M. & Lewis, A. Grey wolf optimizer. Adv. Eng. Softw. 69, 46–61 (2014).10.1016/j.advengsoft.2013.12.007
32. Heidari AA Mirjalili S Faris H Aljarah I Mafarja M Chen H Harris hawks optimization: Algorithm and applications Future Gener. Comput. Syst. 2019 97 849 872 10.1016/j.future.2019.02.028
Heidari, A. A. et al. Harris hawks optimization: Algorithm and applications. Future Gener. Comput. Syst. 97, 849–872 (2019).10.1016/j.future.2019.02.028
33. Li S Chen H Wang M Heidari AA Mirjalili S Slime mould algorithm: A new method for stochastic optimization Future Gener. Comput. Syst. 2020 111 300 323 10.1016/j.future.2020.03.055
Li, S., Chen, H., Wang, M., Heidari, A. A. & Mirjalili, S. Slime mould algorithm: A new method for stochastic optimization. Future Gener. Comput. Syst. 111, 300–323 (2020).10.1016/j.future.2020.03.055
34. Yang Y Wang Z An effective dimensionality reduction approach for short-term load forecasting Electric Power Syst. Res. 2022 210 108150 10.1016/j.epsr.2022.108150
Yang, Y. et al. An effective dimensionality reduction approach for short-term load forecasting. Electric Power Syst. Res. 210, 108150 (2022).10.1016/j.epsr.2022.108150
35. Joshi A Vishnu C Mohan CK Application of XGBoost model for early prediction of earthquake magnitude from waveform data J. Earth Syst. Sci. 2024 133 5 10.1007/s12040-023-02210-1
Joshi, A. et al. Application of XGBoost model for early prediction of earthquake magnitude from waveform data. J. Earth Syst. Sci. 133, 5 (2024).10.1007/s12040-023-02210-1
36. Wang H Ao Y Wang C A dynamic prediction model of landslide displacement based on VMD–SSO–LSTM approach Sci. Rep. 2024 14 9203 10.1038/s41598-024-59517-2 38649403
Wang, H. et al. A dynamic prediction model of landslide displacement based on VMD–SSO–LSTM approach. Sci. Rep. 14, 9203 (2024).38649403 10.1038/s41598-024-59517-2
37. Taorui Z Hongwei J Qingli L Landslide displacement prediction based on variational mode decomposition and MIC-GWO-LSTM model Stoch. Environ. Res. Risk Assess 2022 36 1353 1372 10.1007/s00477-021-02145-3
Taorui, Z. et al. Landslide displacement prediction based on variational mode decomposition and MIC-GWO-LSTM model. Stoch. Environ. Res. Risk Assess 36, 1353–1372 (2022).10.1007/s00477-021-02145-3
