
==== Front
Heliyon
Heliyon
Heliyon
2405-8440
Elsevier

S2405-8440(24)13305-1
10.1016/j.heliyon.2024.e37274
e37274
Research Article
A comprehensive analysis of the artificial neural networks model for predicting monkeypox outbreaks
Alnaji Lulah laalnaji@uhb.edu.sa

Department of Mathematics, College of Science, University of Hafr Al Batin, Hafr Al Batin, 39524, Saudi Arabia
03 9 2024
15 9 2024
03 9 2024
10 17 e3727428 4 2023
25 8 2024
30 8 2024
© 2024 The Author
2024
https://creativecommons.org/licenses/by-nc/4.0/ This is an open access article under the CC BY-NC license (http://creativecommons.org/licenses/by-nc/4.0/).
Monkeypox is a viral disease that causes outbreaks in various countries, significantly impacting public health and healthcare systems. Effective preparedness and response efforts require accurately predicting the severity of these outbreaks. Currently, there are no publicly released studies for nations like Chile and Mexico on monkeypox, leading to this study's creation. We use a neural network model with a time series dataset of monkeypox cases from multiple countries, including Argentina, Brazil, France, Germany, Chile, and Mexico. The Levenberg-Marquardt learning technique is employed to develop and validate single and two hidden layers artificial neural network models. We train various model architectures with different numbers of hidden layer neurons using the K-fold cross-validation early stopping method. Additionally, we use long short-term memory and gated recurrent unit models, commonly employed for time series data processing, to compare the performance of our artificial neural network model.

Keywords

Monkeypox outbreak forecasting
Predictive epidemiology
Neural network modeling for infectious diseases
Predicting monkeypox outbreaks
Epidemiological data modeling
==== Body
pmc1 Introduction

Emerging infectious diseases pose a significant threat to global health security, necessitating urgent attention from the international scientific community. Among these, Monkeypox (MPXV) has emerged as a concern due to its potential for cross-species transmission and international spread. As a zoonotic virus belonging to the same family as the smallpox virus, MPXV's study is crucial for understanding patterns of outbreak and for developing strategies to combat its spread. The disease's recent outbreaks in non-endemic regions highlight the importance of global surveillance and the need for effective diagnostic and management strategies.

MPXV is a viral disease brought on by the MPXV, a member of the same virus family as smallpox [1], [2]. In West and Central Africa, where it is an endemic disease, MPXV mostly affects humans and animals. Through contact with infected animals or their bodily fluids, the virus can spread from animals to people. Direct contact with the skin, saliva, respiratory secretions, or sexual fluids of an infected individual can also cause MPXV to transmit from one person to another. Fever, headaches, swollen lymph nodes, and a rash that lasts for several weeks can all be symptoms of MPXV [3]. MPXV can be diagnosed by testing skin lesion samples in a laboratory. There is no specific treatment for MPXV, but supportive care and infection control measures can help reduce complications and prevent further spread. In some cases, smallpox vaccines and antiviral medications may also work against MPXV. As a new public health risk, MPXV calls for concerted action by international health partners, health ministries, and researchers [4].

Complicacies including pneumonia, sepsis, and encephalitis can develop in extreme forms of the illness [5]. MPXV is currently not curable, however supportive care can help manage symptoms and avoid complications [6]. Despite the limited number of available vaccines, the smallpox vaccine can offer some defense against MPXV. Stopping the spread of the disease requires implementing public health measures like quarantine, contact tracing, and isolation of ill persons and their contacts [7].

With the increasing global concern over infectious diseases like MPXV, there is an imperative need to develop robust methods for predicting outbreaks. Accurate predictions are vital for effective public health interventions and efficient allocation of resources. This study aims to evaluate the performance of Artificial Neural Networks (ANNs) in forecasting MPXV outbreaks and to compare their effectiveness with Long Short-Term Memory (LSTM) and Gated Recurrent Unit (GRU) models. The goal is to identify the most reliable computational approach for predicting the spread of MPXV, contributing to improved disease surveillance and control measures.

The novelty of this work lies in its application to MPXV, as there has been limited research on this disease to date. By employing ANN-based techniques and analyzing time-series data, this study contributes to the understanding of MPXV outbreaks and the potential spread of the disease in different countries, including Argentina, Brazil, Chile, France, Germany, and Mexico. In the paper, the author is specifically forecasting new cases of MPXV. The focus is on predicting the severity and spread of MPXV outbreaks. This is emphasized in the study's objective, which is to accurately predict MPXV outbreaks using various neural network models. The findings of this research can aid in the development of effective strategies for disease surveillance, prevention, and control, and can have important implications for public health policy and decision-making.

MPXV was initially identified in the Democratic Republic of the Congo in 1958 in monkeys, then in humans in 1970 [8]. Periodically, the disease, which is common in portions of Central and West Africa, flares up. Cases outside of Africa have recently been documented in, among other places, Argentina, Brazil, France, Germany, Chile, and Mexico [9]. According to [10], the first two cases of MPXV infection in humans diagnosed in Germany were presented in 2022. Clinical and virological data, including the identification of MPXV DNA in blood, are presented in the paper [10]. During the first 100 days of the 2022 MPXV outbreak, 99 non-endemic nations, including 25 Latin American and Caribbean nations, reported more than 64,000 cases, according to [11] for Latino countries, over 80% of the confirmed Latin American and Caribbean nations cases come from Brazil and Peru combined, indicating that the epidemic has spread steadily throughout Latin American and Caribbean nations. Although Brazil has the highest number of confirmed cases of MPXV (n = 4472), Peru has the highest incidence (41 cases per million people).

The first cases of MPXV in Argentina appeared in May 2022. A growing number of cases of MPXV were discovered following a thorough investigation by public health authorities in numerous countries where the disease is not prevalent. About 24 countries reported cases by May 28th mostly in Europe [12]. Around the same time, Brazil experienced severe political and economic issues as well as a tumultuous presidential election, and the MPXV came there [13]. The discovery and quantification of the MPXV genome in sewersheds in Paris, France, according to a study published on medRxiv, coincided temporally with the discovery of the first case of infection and the spread of the illness among the community connected to the sewage system [14]. The potential of wastewater-based epidemiology for monitoring new diseases was demonstrated by the discovery of the MPXV viral genome in sewersheds in France.

In order to predict possible patient issues related to different diseases, artificial neural network (ANN) techniques have become widely used. These techniques have been utilized in numerous studies to forecast the severity of a variety of diseases, including COVID-19, breast cancer, and cardiovascular disease, among others. As an illustration, [15] used an ANN-based model for the diagnosis of heart disease. In a different study, [16] employed ANN techniques to compare two computer models, logistic regression and artificial neural network (ANN), for estimating breast cancer risk based on mammographic and demographic features.

In fact, ANN models are frequently used in a variety of fields, including infectious diseases, for time-series prediction. The papers we cited demonstrate how adaptable and successful ANN-based approaches are for forecasting and prediction tasks. For instance, the work by [17] utilized a modified VGG16 deep learning model for the detection of MPXV disease using image data. Similarly, [18] utilized nonlinear autoregressive artificial neural networks for forecasting the prevalence of COVID-19 outbreak in Egypt. This highlights how ANN models can be used to forecast the occurrence and consequences of infectious diseases, such as COVID-19, a pandemic that has had a substantial impact on public health.

The study by [19] further exemplifies the use of ANN models in predicting COVID-19 cases in Saudi Arabia, highlighting the applicability of ANN-based techniques for disease prediction in different geographical regions. The study by [20] leverages datasets from the European Centre for Disease Prevention and Control to formulate a neural network-centric prediction model for MPXV propagation in countries including the USA, Germany, UK, France, and Canada. In this study, a novel method is used to deploy an ANN model to a time series dataset of MPXV cases and compare the effectiveness of the model to that of LSTM and GRU models. In contrast, our analysis uses unique datasets from nations including Chile, Argentina, and Brazil. In addition, we explore several time periods for nations including the USA, France, and Germany, offering a novel viewpoint on the research. Additionally, the work by [21] proposed an optimal forecast combination based on neural networks for time series forecasting, showcasing the potential of ANN models in optimizing forecast performance through combination techniques.

Complementing the efforts in infectious disease forecasting, the research presented in [22] introduces a novel method for the short-term forecasting of MPXV cases. This approach, distinct in its utilization of filtering techniques and machine learning model combinations, has proven to enhance accuracy in predicting MPXV trends. The study's experimental results underscore the effectiveness of these advanced analytical techniques, offering valuable perspectives relevant to the current study's focus on ANN models in predicting disease outbreaks.

Building upon similar methodologies, [23] tackles the rise in MPXV cases and fatalities, aiming to lessen the virus's impact with forecasts based on time series and stochastic models. This approach enhances prediction accuracy for the virus's spread and impact. Similarly, [24] reveals how ANNs effectively forecast health events, evidenced by a study on COVID-19 in Pakistan comparing ANN with univariate time series models. The ANN's superior performance in predicting deaths and recoveries, and the moving average model's effectiveness in forecasting confirmed cases, highlight the potency of computational methods in public health strategies. Lastly, [25] employs time series models like Autoregressive and Moving Average for COVID-19 case forecasts in Pakistan. Analyzing data from early 2020, the study shows these models' capacity for accurate health event prediction, paralleling the ANN methods used in our MPXV outbreak research.

To analyze the dynamics of MPXV outbreaks in, Argentina, Brazil, France, Germany, Chile and Mexico, a prediction model is used in this study. For this investigation, effective neural networks models such as neural network, LSTM, and GRU models were used in conjunction with open data from the website (Our World in Data [26]). An ADAM optimizer was utilized to compare the outcomes of the LSTM, GRU, and neural network models [27], [28], [29], [30], [31]. Some statistical graphs representing the confirmed cases of the six nations are shown in Fig. 1, and the time series dataset of MPXV cases, collected from each country, in addition to the zoomed-in plot of the peak period of each country is presented in Fig. 2.Figure 1 (A) Statistics for each country's confirmed cases, (B) The global map's country locations varied in color from light to dark to represent the countries least and most affected, respectively, (C) A time series dataset of MPXV cases collected from all six nations.

Figure 1

Figure 2 (A) A time series dataset of MPXV cases, collected from Argentina, and the zoomed-in plot of the peak period (August to December 2022), (B) A time series dataset of MPXV cases, collected from Brazil and the zoomed-in plot of the peak period (July to October 2022), (C) A time series dataset of MPXV cases, collected from France and the zoomed-in plot of the peak period (August to November 2022), (D) A time series dataset of MPXV cases, collected from Germany and the zoomed-in plot of the peak period (July to October 2022), (E) A time series dataset of MPXV cases, collected from Chile and the zoomed-in plot of the peak period (Jun to October 2022), (F) A time series dataset of MPXV cases, collected from Mexico and the zoomed-in plot of the peak period (August to December 2022).

Figure 2

The Levenberg-Marquardt (LM) learning approach was used to develop and test an artificial neural network (ANN) model with a single and two hidden layers [32]. Taking into account various numbers of hidden layer neurons, the best design with the largest generalization potential was found using the K-fold cross-validation early stopping validation approach [33]. In the current phase of our study, the models were developed using a data division where 80% was allocated for training and 20% for validation. For further enhancement of our methodology's robustness, we plan to adopt a more diverse approach to data splitting in future research. This will include employing a variety of training and testing ratios, specifically 50% training with 50% testing, and a 90-10 split, aligning with the strategies suggested in [34] and [35]. This approach is anticipated to reinforce the consistency and reliability of our proposed models.

Fig. 1 illustrates the overall statistics, global map of affected countries, and a time series dataset of MPXV cases from six nations. Specifically, Fig. 1(A) shows each country's confirmed cases, Fig. 1(B) depicts a color-coded global map representing the least to most affected countries, and Fig. 1(C) presents a time series dataset of MPXV cases.

Fig. 2 provides detailed time series datasets and zoomed-in plots for specific countries during peak periods. Fig. 2(A) focuses on Argentina (August to December 2022), Fig. 2(B) on Brazil (July to October 2022), Fig. 2(C) on France (August to November 2022), Fig. 2(D) on Germany (July to October 2022), Fig. 2(E) on Chile (June to October 2022), and Fig. 2(F) on Mexico (August to December 2022).

The remaining sections of this article are structured as follows. Section 2 presents the materials and methods used in this paper. The methodology and model design is presented in Section 3. Section 4 presents results and Section 5 discusses the future scope. The important conclusions from this study are presented in Section 5.

2 Materials and methods

2.1 The artificial neural network

The Artificial Neural Networks (ANNs) are computational models inspired by the structure and function of biological neurons in the human brain [36]. ANNs consist of interconnected nodes or neurons organized in layers, with one or more hidden layers connecting the input layer to the output layer. These nodes process scalar messages and interact with each other through connection weights, which are adjusted during training to minimize errors between predicted and actual outputs using algorithms such as Back Propagation [37].

The most common type of ANN is the multilayer perceptron (MLP), which has one or more hidden layers in a feed-forward neural network. Nodes in ANNs have active states (on or off) and inactive states (off or 0), and edges or synapses between nodes have weights. Negative weights inhibit the activation of the next linked node (if active), while positive weights stimulate the activation of the next inactive node [38], [39].

An Artificial Neural Network (ANN) architecture with three layers typically consists of an input layer, a hidden layer, and an output layer used in this section. Neurons in the ANN receive inputs from preceding neurons with associated weights, and the output of a neuron is obtained by applying a sigmoid activation function to the weighted sum of its inputs.

The input layer is responsible for receiving the initial input data, which is then passed on to the neurons in the hidden layer. Each neuron in the hidden layer receives inputs from the neurons in the input layer, and the strength of these connections is determined by weights denoted as wij, where i represents the neuron index in the input layer and j represents the neuron index in the hidden layer. The weighted sum of inputs for a neuron in the hidden layer, denoted as Tj, is calculated as the dot product of the input data vector and the weight vector (Equation (1)):(1) Tj=∑i=1nwijxi

where xi represents the input from the i-th neuron in the input layer and n is the number of neurons in the input layer. The sigmoid activation function is then applied to Tj to obtain the output of the neuron in the hidden layer, denoted as yj (Equation (2)):(2) yj=sigmoid(Tj)

The sigmoid activation function is a common choice for ANN architectures as it maps the weighted sum of inputs to a value between 0 and 1, which can be interpreted as a probability or an activation level of the neuron. It introduces non-linearity into the network, allowing it to model complex relationships between inputs and outputs.

Similarly, the neurons in the hidden layer are connected to the neurons in the output layer with weights denoted as wkj, where k represents the neuron index in the output layer. The weighted sum of inputs for a neuron in the output layer, denoted as Tk, is calculated as (Equation (3)):(3) Tk=∑j=1mwjkyj

where yj represents the output of the j-th neuron in the hidden layer and m is the number of neurons in the hidden layer. The sigmoid activation function is then applied to Tk to obtain the final output of the neuron in the output layer.

During the training process, the weights wij and wkj are adjusted using optimization algorithms, such as gradient descent, to minimize the error between the predicted outputs and the true outputs. This allows the ANN to learn the optimal values of the weights, and thus, optimize its architecture for accurate predictions or classifications based on input data. The number of neurons in the hidden layer and the choice of activation function are hyperparameters that can be tuned to optimize the performance of the ANN for a specific problem.

Therefore, the total number of layers in this architecture is three [40]. In detail:• Input Layer: The input layer is responsible for receiving the initial input data. It consists of neurons, each of which represents a feature or attribute of the input data. The number of neurons in the input layer is determined by the dimensionality of the input data, with each neuron corresponding to one input feature. The input data is passed on to the neurons in the hidden layer for further processing.

• Hidden Layer: The hidden layer is where the computation and transformation of the input data occur. It consists of neurons that receive inputs from the neurons in the input layer. The connections between the input layer and the hidden layer are determined by weights denoted as wij, where i represents the neuron index in the input layer and j represents the neuron index in the hidden layer. Each neuron in the hidden layer calculates the weighted sum of its inputs, denoted as Tj (Equation (1)), by taking the dot product of the input data vector and the weight vector. This weighted sum of inputs is then passed through a sigmoid activation function to obtain the output of the neuron, denoted as yj (Equation (2)). The sigmoid activation function introduces non-linearity into the network, allowing it to model complex relationships between inputs and outputs.

• Output Layer: The output layer is responsible for producing the final output of the ANN. It consists of neurons that receive inputs from the neurons in the hidden layer. Similar to the hidden layer, the connections between the hidden layer and the output layer are determined by weights denoted as wkj, where k represents the neuron index in the output layer. Each neuron in the output layer calculates the weighted sum of its inputs, denoted as Tk (Equation (3)), by taking the dot product of the outputs from the hidden layer and the weight vector. This weighted sum of inputs is then passed through a sigmoid activation function to obtain the final output of the neuron.

• Neuron Activation: The sigmoid activation function is commonly used in ANN architectures for its ability to map the weighted sum of inputs to a value between 0 and 1, which can be interpreted as a probability or an activation level of the neuron. The sigmoid function is defined as (Equation (4)):(4) sigmoid(x)=11+e−x

where e is the base of the natural logarithm. Other activation functions, such as ReLU (Rectified Linear Unit), tanh (hyperbolic tangent), and softmax, can also be used depending on the problem and requirements of the network [41].

• Training Process: During the training process, the ANN adjusts the weights wij and wkj using optimization algorithms, such as gradient descent, to minimize the error between the predicted outputs and the true outputs. Gradient descent is an iterative optimization algorithm that updates the weights in the direction of the negative gradient of the error with respect to the weights, allowing the network to gradually converge to optimal weight values. Other variants of gradient descent, such as stochastic gradient descent (SGD) and mini-batch gradient descent, can also be used to improve the efficiency of the training process [42].

• Hyperparameters: The number of neurons in the hidden layer and the choice of activation function are hyperparameters that can be tuned to optimize the performance of the ANN. The number of neurons in the hidden layer is often determined through experimentation and can vary depending on the complexity of the problem and the amount of available data. Too few neurons in the hidden layer may result in underfitting, while too many neurons may result in overfitting. The choice of activation function can also impact the performance of the network, as different activation functions have different properties, such as non-linearity and differentiability, which can affect the learning dynamics and convergence of the network [43].

During the training process, the weights wij and wkj are adjusted using optimization algorithms, such as gradient descent, to minimize the error between the predicted outputs and the true outputs. This allows the ANN to learn the optimal values of the weights, and thus, optimize its architecture for accurate predictions or classifications based on input data. The number of neurons in the hidden layer and the choice of activation function are hyperparameters that can be tuned to optimize the performance of the ANN for a specific problem [43].

2.2 Levenberg–Marquardt

The Levenberg-Marquardt (LM) algorithm is a widely used optimization method for solving nonlinear least squares problems in numerical analysis. To ensure that the estimated Hessian matrix JTJ is invertible, the LM approach approximates JTJ by adding a positive combination coefficient μ and an identity matrix I to obtain H≈JTJ+μI. This approximation ensures that the diagonal components of the estimated Hessian matrix are greater than zero, which guarantees the invertibility of H [44], [45].

The LM algorithm smoothly transitions between the Gauss-Newton method and the steepest descent method based on the value of μ. When μ is very small (close to 0), the algorithm tends to operate like the Gauss-Newton method, exploiting its fast convergence in regions where the error surface is approximately quadratic. Conversely, when μ is large, the algorithm behaves more like the steepest descent method, ensuring a more reliable (albeit slower) convergence in regions where the error surface deviates from a quadratic form.

In the Levenberg-Marquardt (LM) learning technique, the regularization parameter μ is a crucial factor that plays a significant role in controlling the trade-off between the Gauss-Newton and steepest descent approaches during the training process of the artificial neural network (ANN) model. The value of μ is iteratively adjusted to ensure convergence and prevent overfitting.

To provide further clarity on how μ is determined, let's delve into the LM algorithm's mechanism:1. Initialization: The LM algorithm begins with an initial value of μ, often set to a small positive value. This value acts as a damping factor and influences the step size during parameter updates.

2. Update Step: During each iteration, the algorithm attempts to update the parameters of the ANN model using the LM formula. This formula combines both the Gauss-Newton approach (which is effective near the minimum) and the steepest descent approach (which helps escape local minima).

3. Error Minimization: The LM algorithm calculates the error between the predicted outputs and the actual outputs for the given training data. It adjusts the weights of the ANN to minimize this error.

4. Damping Adjustment: If the proposed parameter updates result in a decrease in the error, the algorithm accepts the updates and decreases the value of μ, making the steps larger for quicker convergence.

5. Overfitting Prevention: If the proposed parameter updates result in an increase in the error, it's an indication of overshooting the minimum. The algorithm increases the value of μ to dampen the updates, reducing the likelihood of overshooting and mitigating overfitting.

6. Iteration: The algorithm repeats the process, iteratively adjusting μ and updating the weights of the ANN until a convergence criterion is met (such as a predefined number of iterations or a sufficiently small error).

This adaptive behavior allows the LM algorithm to combine the strengths of both the Gauss-Newton and steepest descent methods, optimizing convergence performance based on the local characteristics of the error surface [46].

The weight update rule of the LM algorithm can be expressed by combining Equations H≈JTJ+μI and the update equation(5) Vk+1=Vk−(JkTJk+μI)−1JkTek

represents the updated weight vector and ek represents the error vector. This combined equation is also known as the Gauss-Newton procedure.

The LM algorithm iteratively updates the weights of the model using Equation (5). In each iteration, it calculates the Jacobian matrix Jk which represents the partial derivatives of the error vector ek with respect to the weights. The Jacobian matrix is then used to calculate the gradient JkTek and the Hessian matrix approximation JkTJk+μI, where μ is the regularization parameter and I is the identity matrix.

If the Hessian matrix approximation is invertible, the weight update is calculated using the inverse of the Hessian matrix approximation, as shown in Equation (5). This weight update takes into account the curvature of the error surface through the Hessian matrix approximation, and the regularization term μI helps prevent the weights from becoming too large during training.

The value of the regularization parameter μ is typically adjusted during training to balance between the Gauss-Newton and steepest descent approaches. At the beginning of training, when the error is large, a larger value of μ is used to emphasize the steepest descent approach for faster convergence. As the error decreases and the optimization gets closer to the optimal solution, μ is decreased to allow the Gauss-Newton approach to dominate for better accuracy.

The LM algorithm continues to iterate until a convergence criterion is met, such as reaching a certain error threshold or a maximum number of iterations. Once the convergence criterion is met, the final weights of the model are obtained and can be used for making predictions on new data.

Algorithm 1 Levenberg-Marquardt Algorithm.

Algorithm 1

The Levenberg-Marquardt algorithm uses a combination of the Gauss-Newton and steepest descent approaches, with the regularization parameter μ controlling the balance between them. When μ is small, the algorithm behaves more like the Gauss-Newton method, which is efficient for local convergence in regions where the error surface is relatively flat. When μ is large, the algorithm behaves more like the steepest descent method, which helps in exploring regions of the error surface with steep slopes and finding better solutions [46].

2.3 Adaptive moment estimation optimization

ADAM is a gradient-based optimization algorithm that is widely used in deep learning for updating the weights of neural networks during training. It was proposed by [47] and has since become one of the most popular optimization algorithms for deep learning.

The main idea behind ADAM is to combine the advantages of two other optimization algorithms: AdaGrad and RMSProp. Like AdaGrad, ADAM adapts the learning rate of each parameter based on the historical gradient information. But unlike AdaGrad, ADAM also uses exponential moving averages of the gradient and the squared gradient to update the learning rate, similar to RMSProp.

In our study, the choice of the ADAM optimizer was based on its effectiveness in addressing certain limitations of traditional optimization algorithms, particularly in the context of training artificial neural network (ANN) models. The ADAM optimizer stands for Adaptive Moment Estimation, and it has gained popularity due to its adaptive learning rate, momentum, and efficiency in handling various optimization challenges. We believe that the utilization of the ADAM optimizer aligns well with the goals of our research and enhances the training process of our ANN models.

The ADAM optimizer combines the strengths of two popular optimization algorithms, namely RMSProp and Momentum. This combination allows ADAM to adaptively adjust the learning rate for each parameter individually, based on the historical gradients of the parameter. It also introduces the concept of momentum, which helps accelerate convergence by considering past gradients and adjusting the update steps accordingly. The key components of the ADAM optimizer include the moving averages of past gradients (first and second moments) and a bias-correction mechanism to address the initialization bias.

In the context of our study, where we are developing ANN models to predict MPXV outbreak severity, the ADAM optimizer offers several advantages:• Adaptive Learning Rate: ADAM automatically adjusts the learning rate for each parameter, ensuring faster convergence when gradients are consistent and slower convergence in the presence of noisy or sparse gradients. This adaptability is particularly useful when dealing with complex and dynamic data patterns.

• Efficient Handling of Sparse Data: Our dataset may contain instances with missing or sparse data. ADAM's adaptive learning rate can handle such scenarios by giving smaller updates to parameters associated with sparse features.

• Reduced Risk of Convergence to Poor Minima: The adaptive learning rates in ADAM help prevent the model from getting stuck in poor local minima, which can be a challenge for traditional gradient descent algorithms.

• Efficient Use of Memory: ADAM maintains moving averages of past gradients, allowing for efficient use of memory and avoiding the need to store the entire training history.

• Robustness to Hyperparameters: ADAM's adaptive nature reduces the sensitivity to hyperparameters like learning rate, making it a more user-friendly choice.

By choosing the ADAM optimizer, we aim to benefit from its adaptive learning rates, efficient memory utilization, and robustness to hyperparameters. These features are essential for effectively training our ANN models on the MPXV outbreak dataset, enabling accurate predictions and contributing to the overall success of our research.

In ADAM, the gradient of the cost function with respect to the model parameters is first computed for each mini-batch of training data. Then, the first and second moments of the gradient are computed using exponential moving averages. The first moment, mt, is the exponentially weighted average of the gradient, and the second moment, vt, is the exponentially weighted average of the squared gradient. These are computed as follows (Equations (6) and (7)):(6) mt=β1mt−1+(1−β1)gt

(7) vt=β2vt−1+(1−β2)gt2

where gt is the gradient at time t, and β1 and β2 are hyperparameters that control the decay rates of the moving averages.

Next, ADAM computes bias-corrected estimates of the first and second moments, mˆt and vˆt, to reduce the bias towards zero in the initial time steps (Equations (8) and (9)):(8) mˆt=mt1−β1t

(9) vˆt=vt1−β2t

Finally, the learning rate for each parameter is updated using the bias-corrected estimates of the first and second moments (Equation (10)):(10) θt=θt−1−α⋅mˆtvˆt+ϵ

where α is the learning rate, which controls the step size of the parameter updates, and ϵ is a small constant added to the denominator to avoid division by zero and improve numerical stability.

Algorithm 2 ADAM Optimization Algorithm.

Algorithm 2

ADAM optimization method maintains two moment estimates, mt and vt, which are the biased estimates of the first moment (mean) and second raw moment (uncentered variance) of the gradients, respectively. These moment estimates are used to adaptively update the learning rate for each parameter based on their historical values. The hyperparameters β1 and β2 control the exponential decay rates of the moment estimates, and the small constant ϵ is added to the denominator for numerical stability.

2.4 Gated recurrent unit

The architecture of a Gated Recurrent Unit (GRU) network consists of three main components: an update gate, a reset gate, and a candidate state. These components are designed to selectively update and forget information in a recurrent neural network (RNN) architecture, making GRUs more efficient and effective compared to traditional RNNs.

Update Gate: The update gate in a GRU, denoted as zt, determines how much of the past hidden state should be retained and how much of the new candidate state should be added to the current hidden state. It is computed using the sigmoid activation function and a weight matrix Wz and bias vector bz. The update gate takes as input the previous hidden state ht−1 and the current input xt, and its output is a value between 0 and 1 that represents the amount of information to be updated in the hidden state.zt=σ(Wz⋅[ht−1,xt]+bz)

where σ is the sigmoid activation function, Wz is the weight matrix for the update gate, bz is the bias vector for the update gate, and [ht−1,xt] is the concatenation of the previous hidden state and the current input.

Reset Gate: The reset gate in a GRU, denoted as rt, determines how much of the previous hidden state should be forgotten when computing the new candidate state. It is also computed using the sigmoid activation function and a weight matrix Wr and bias vector br. The reset gate takes as input the previous hidden state ht−1 and the current input xt, and its output is a value between 0 and 1 that represents the amount of information to be forgotten in the hidden state.rt=σ(Wr⋅[ht−1,xt]+br)

where σ is the sigmoid activation function, Wr is the weight matrix for the reset gate, br is the bias vector for the reset gate, and [ht−1,xt] is the concatenation of the previous hidden state and the current input.

Candidate State: The candidate state in a GRU, denoted as h˜t, is the new information that is computed based on the input and the previous hidden state. It is computed using the hyperbolic tangent (tanh) activation function and a weight matrix Wh and bias vector bh. The candidate state takes as input the element-wise product of the reset gate and the previous hidden state rt⊙ht−1, concatenated with the current input xt. The candidate state represents the new information that can potentially be added to the hidden state.h˜t=tanh⁡(Wh⋅[rt⊙ht−1,xt]+bh)

where ⊙ is the element-wise multiplication operation, Wh is the weight matrix for the candidate state, bh is the bias vector for the candidate state, and [rt⊙ht−1,xt] is the concatenation of the element-wise product of the reset gate and the previous hidden state, and the current input.

These three components together allow the GRU network to selectively update its hidden state based on the input and the previous hidden state, controlled by the update and reset gates. The candidate state represents the new information that can be added to the hidden state, and the update gate determines how much of this information should be retained. The reset gate determines how much of the previous hidden state should be forgotten, allowing the GRU to adaptively update its hidden state representation during sequential data processing tasks. Overall, the architecture of a GRU network with these three components provides a powerful mechanism for modeling sequential data with selective information retention and forgetting [48].

Algorithm 3 GRU Update, Reset, and Candidate State.

Algorithm 3

Note that σ represents the sigmoid activation function, ⊙ represents element-wise multiplication, Wz, Wr, and Wh are weight matrices for the update gate, reset gate, and candidate state, respectively, and bz, br, and bh are the bias vectors for the update gate, reset gate, and candidate state, respectively. The algorithm computes the update gate zt, reset gate rt, and candidate state h˜t using the corresponding equations, and then updates the hidden state ht based on the values of these gates. These equations allow the GRU network to selectively update its hidden state, making it more efficient and effective than traditional RNNs.

2.5 Long short-term memory

The architecture of an LSTM (Long Short-Term Memory) network consists of three main components: the input gate, the forget gate, and the output gate. These gates are responsible for controlling the flow of information through the network, allowing it to capture long-range dependencies in sequential data. The equations for each component are as follows:

Input Gate: The input gate in an LSTM network is responsible for controlling how much new information is added to the cell state. It takes the current input at time step t, denoted as xt, and the previous hidden state at time step t−1, denoted as ht−1, as input. The input gate computes the following equations:it=σ(Wi⋅[ht−1,xt]+bi)c˜t=tanh⁡(Wc⋅[ht−1,xt]+bc)

where:• it is the output of the input gate at time step t, which represents the amount of new information to be added to the cell state.

• Wi and bi are the weight matrix and bias vector associated with the input gate, respectively.

• Wc and bc are the weight matrix and bias vector associated with the computation of the candidate cell state.

• σ() is the sigmoid activation function that squashes the input to a value between 0 and 1.

• tanh⁡() is the hyperbolic tangent activation function that squashes the input to a value between -1 and 1.

Forget Gate: The forget gate in an LSTM network is responsible for controlling how much information from the previous cell state should be forgotten. It takes the current input at time step t and the previous hidden state at time step t−1 as input. The forget gate computes the following equation:ft=σ(Wf⋅[ht−1,xt]+bf)

where ft is the output of the forget gate at time step t, which represents the amount of information from the previous cell state to be forgotten. Wf and bf are the weight matrix and bias vector associated with the forget gate.

Output Gate: The output gate in an LSTM network is responsible for controlling how much information from the updated cell state should be used to generate the output. It takes the current input at time step t and the previous hidden state at time step t−1 as input. The output gate computes the following equations:ot=σ(Wo⋅[ht−1,xt]+bo)ht=ot⋅tanh⁡(ct)

where ot is the output of the output gate at time step t, which represents the amount of information from the updated cell state to be used in generating the output. ht is the hidden state at time step t, which serves as the output of the LSTM network. Wo and bo are the weight matrix and bias vector associated with the output gate [49].

Algorithm 4 Long Short-Term Memory (LSTM) Algorithm.

Algorithm 4

Note that in the algorithm, σ represents the sigmoid activation function, tanh represents the hyperbolic tangent activation function, ⊙ represents element-wise multiplication, and f(ht) represents the output function applied to the hidden state at time step t, which depends on the specific task the LSTM is used for. Also, the input sequence x and output sequence y are assumed to be of length T, and [ht−1,xt] denotes the concatenation of the hidden state at time step t−1 and the input at time step t.

2.6 K-fold cross-validation

K-fold cross-validation is a widely used technique to prevent overfitting in artificial neural network (ANN) models. Overfitting occurs when the model learns the noise instead of the signals, resulting in poor performance on unknown datasets. To address this issue, K-fold cross-validation randomly divides the data into K groups and trains the model on (K-1) folds, while evaluating the performance on the remaining fold. The process is repeated K times, with each fold serving as the validation set once. This technique helps to evaluate the model's generalization performance and to identify potential issues such as underfitting or overfitting [50], [51], [52].

In practice, the dataset is first loaded and the features are scaled to improve the learning process. The dataset is then divided into K folds, and the K-fold cross-validation process loops K times. In each iteration, the dataset is divided into a training set and a validation set, and the model is trained on the training set and evaluated on the validation set. The performance of the model is recorded after each iteration. After all K iterations, the average performance over all folds is calculated, and the results are outputted.

The number of epochs versus the average RMSE is displayed on the validation folds to monitor the learning process. The training phase is stopped when the RMSE stops declining while the number of epochs increases. This helps to prevent overfitting and ensures that the model is not over-trained on the training data.

The trained model is evaluated against a separate test dataset to determine its prediction performance. The output values are anti-normalized to their true values after the training and testing phases are completed [53]. In this paper, K=10, after applying K-fold cross-validation to avoid overfitting through determine the optimal number of hidden neurons. As shown in Table 1, Table 2, Table 3, Table 4, Table 5, Table 6, Table 7, Table 8, Table 9, Table 10, Table 11, Table 12, Table 13, Table 14, Table 15, Table 16, Table 17, Table 18, Table 19, Table 20, Table 21, Table 22, Table 23, Table 24, Table 25, Table 26, Table 27, Table 28, Table 29, Table 30, Table 31, Table 32, Table 33, Table 34, Table 35, Table 36, a total of 19 ANN models with various numbers of hidden layers are constructed.Table 1 Selecting the optimal ANN model with respect to a single hidden layer and neurons for the Argentina dataset.

Table 1Sl No	Neurons	RMSE
(Train) x 1000	R2
(Train) %	MAPE
(Train) %	RMSE
(Validation) x 1000	R2
(Validation) %	MAPE
(Validation) %	RMSE
(Test) x 1000	R2
(Test) %	MAPE
(Test) %	
1	1	1.36	91.59	93.39	1.14	90.02	92.45	1.56	99.75	81.10	
2	2	1.50	92.35	91.16	1.32	97.87	91.92	1.56	99.75	81.10	
3	3	1.10	92.60	91.00	1.12	97.62	92.07	1.56	99.71	81.16	
4	4	1.00	94.71	90.18	1.12	98.50	92.28	1.56	99.96	80.07	
5	5	1.05	94.47	91.06	1.84	97.59	92.18	1.56	99.64	80.91	
6	6	1.10	91.88	92.01	1.95	97.46	92.45	1.56	99.64	81.18	
7	7	1.52	94.36	99.79	1.86	97.58	92.45	1.56	99.65	81.14	
8	8	1.01	92.66	90.95	1.30	97.55	92.59	1.56	99.60	81.26	
9	9	1.90	92.06	93.91	1.66	97.47	92.69	1.56	99.53	81.12	
10	10	1.30	91.83	93.09	1.46	97.47	92.56	1.56	99.61	81.35	
11	11	1.50	92.36	91.29	1.58	97.59	92.54	1.56	99.65	81.33	
12	12	1.40	91.81	92.67	1.46	97.49	92.48	1.56	99.58	81.11	
13	13	1.10	93.78	93.97	1.16	98.56	92.51	1.56	99.52	81.36	
14	14	1.40	91.95	90.16	1.35	97.58	92.56	1.56	99.59	81.08	
15	15	1.81	93.53	90.47	1.12	97.56	92.83	1.56	99.68	81.29	
16	16	1.50	91.84	93.21	1.62	97.63	92.53	1.56	99.64	81.44	
17	17	1.60	91.84	93.73	1.39	97.48	92.56	1.56	99.61	81.17	
18	18	1.90	91.93	92.62	1.30	97.57	92.60	1.56	99.64	81.40	
19	19	1.80	92.06	92.52	1.85	97.52	92.28	1.56	99.70	81.22	

Table 2 Selecting the optimal ANN model with respect to two hidden layers and neurons for the Argentina dataset.

Table 2Sl No	Neurons	RMSE
(Train) x 1000	R2
(Train) %	MAPE
(Train) %	RMSE
(Validation) x 1000	R2
(Validation) %	MAPE
(Validation) %	RMSE
(Test) x 1000	R2
(Test) %	MAPE
(Test) %	
1	1	1.35	99.91	92.28	1.64	99.86	92.30	1.36	90.01	80.98	
2	6	1.25	99.26	91.77	1.63	99.02	92.33	1.56	99.70	81.07	
3	7	1.46	90.92	93.47	1.62	97.74	92.03	1.16	99.56	81.38	
4	9	1.03	99.64	90.64	1.62	97.43	92.34	1.06	99.51	81.10	
5	12	1.00	99.88	91.52	1.62	97.42	91.59	1.06	99.82	81.02	

Table 3 Selecting the optimal ANN model with respect to a single hidden layer and neurons for the Brazil dataset.

Table 3Sl No	Neurons	RMSE
(Train) x 1000	R2
(Train) %	MAPE
(Train) %	RMSE
(Validation) x 1000	R2
(Validation) %	MAPE
(Validation) %	RMSE
(Test) x 1000	R2
(Test) %	MAPE
(Test) %	
1	1	2.03	88.29	78.25	1.84	99.95	65.19	1.19	98.45	65.43	
2	2	2.05	89.66	78.82	1.84	99.92	64.93	2.00	99.88	67.50	
3	3	2.15	98.28	79.04	1.84	90.03	65.28	1.94	98.46	65.67	
4	4	2.15	98.77	78.66	1.84	99.92	64.93	1.24	99.66	66.88	
5	5	1.92	78.74	76.15	1.84	99.63	64.52	1.85	99.63	66.90	
6	6	1.91	98.82	74.64	1.70	99.88	62.08	1.00	99.88	64.47	
7	7	1.93	79.81	74.81	1.84	99.79	64.39	1.43	90.36	65.47	
8	8	1.93	79.48	75.26	1.84	99.66	64.30	1.60	99.54	66.71	
9	9	1.92	78.88	75.81	1.84	99.97	64.89	1.60	99.58	67.12	
10	10	1.94	80.11	74.61	1.84	99.66	64.22	1.42	99.69	66.80	
11	11	2.14	98.75	78.82	1.70	99.65	62.35	1.09	99.31	66.54	
12	12	1.94	80.41	74.61	1.84	99.54	64.52	1.85	99.32	66.64	
13	13	1.92	78.85	75.29	1.84	90.07	64.76	1.30	99.39	66.48	
14	14	1.95	80.93	74.90	1.83	99.28	63.96	1.34	99.66	66.84	
15	15	2.14	98.21	78.58	1.84	99.66	64.30	1.12	90.30	64.77	
16	16	1.93	79.85	74.73	1.84	99.61	64.28	1.04	99.37	66.53	
17	17	2.15	98.50	78.78	1.83	99.28	64.18	1.02	90.06	65.06	
18	18	2.15	98.55	78.98	1.84	99.48	64.74	1.35	90.35	64.88	
19	19	1.94	80.11	74.70	1.71	86.19	62.18	1.86	99.45	66.51	

Table 4 Selecting the optimal ANN model with respect to two hidden layers and neurons for the Brazil dataset.

Table 4Sl No	Neurons	RMSE
(Train) x 1000	R2
(Train) %	MAPE
(Train) %	RMSE
(Validation) x 1000	R2
(Validation) %	MAPE
(Validation) %	RMSE
(Test) x 1000	R2
(Test) %	MAPE
(Test) %	
1	3	1.83	99.88	77.89	1.84	99.97	64.96	1.09	99.95	67.76	
2	6	2.16	99.40	77.82	1.84	99.99	65.07	1.11	99.83	67.54	
3	8	1.92	78.89	75.76	1.84	99.68	64.92	1.90	99.82	67.39	
4	12	2.14	98.16	78.31	1.84	99.60	64.66	2.00	90.04	67.89	
5	17	1.95	78.94	75.72	1.84	99.53	64.73	1.68	99.83	67.43	

Table 5 Selecting the optimal ANN model with respect to a single hidden layer and neurons for the France dataset.

Table 5Sl No	Neurons	RMSE
(Train) x 1000	R2
(Train) %	MAPE
(Train) %	RMSE
(Validation) x 1000	R2
(Validation) %	MAPE
(Validation) %	RMSE
(Test) x 1000	R2
(Test) %	MAPE
(Test) %	
1	1	2.94	99.60	50.00	1.68	94.95	39.84	1.42	96.07	47.49	
2	2	2.12	99.46	49.61	1.68	95.44	39.88	1.50	95.86	47.37	
3	3	2.87	99.16	49.62	1.64	95.68	39.95	1.49	97.32	47.74	
4	4	2.01	99.60	49.00	1.60	98.52	39.97	1.41	98.77	47.13	
5	5	2.67	98.83	49.38	1.64	96.45	40.27	1.50	97.84	48.16	
6	6	2.46	98.83	49.38	1.61	96.70	40.77	1.42	97.22	47.62	
7	7	2.84	99.27	49.63	1.68	96.78	40.31	1.44	98.75	48.08	
8	8	2.12	98.76	49.36	1.69	96.56	40.35	1.46	98.22	47.93	
9	9	2.66	98.98	49.44	1.71	98.02	40.79	1.47	98.51	47.93	
10	10	3.00	98.89	49.27	1.61	96.50	41.01	1.42	97.55	47.94	
11	11	2.80	98.83	49.36	1.66	97.03	40.77	1.48	97.68	47.94	
12	12	2.56	98.98	49.41	1.62	96.81	40.93	1.49	97.96	48.22	
13	13	2.17	98.64	49.16	1.70	97.54	41.06	1.48	98.28	48.35	
14	14	2.31	99.02	49.57	1.67	96.92	40.82	1.47	97.33	47.70	
15	15	2.70	98.70	49.30	1.69	96.85	40.63	1.48	97.64	47.93	
16	16	2.71	98.83	49.18	1.63	96.84	40.80	1.47	97.45	47.73	
17	17	2.37	98.52	49.24	1.66	97.19	40.68	1.42	97.87	47.97	
18	18	2.33	98.51	49.30	1.68	97.19	40.70	1.42	97.97	48.96	
19	19	2.13	99.60	49.83	1.62	96.91	40.32	1.47	98.55	48.13	

Table 6 Selecting the optimal ANN model with respect to two hidden layers and neurons for the France dataset.

Table 6Sl No	Neurons	RMSE
(Train) x 1000	R2
(Train) %	MAPE
(Train) %	RMSE
(Validation) x 1000	R2
(Validation) %	MAPE
(Validation) %	RMSE
(Test) x 1000	R2
(Test) %	MAPE
(Test) %	
1	1	2.37	99.85	49.85	1.72	90.04	41.45	1.41	99.94	48.57	
2	2	2.44	99.96	50.04	1.82	94.99	41.39	1.42	99.96	48.72	
3	6	1.01	99.96	49.03	1.08	96.56	40.00	1.40	99.98	47.45	
4	11	2.15	98.41	49.18	1.68	95.43	39.93	1.42	96.22	47.52	
5	14	2.56	98.92	49.56	1.70	95.68	40.69	1.42	96.77	47.57	

Table 7 Selecting the optimal ANN model with respect to a single hidden layer and neurons for the Germany dataset.

Table 7Sl No	Neurons	RMSE
(Train) x 1000	R2
(Train) %	MAPE
(Train) %	RMSE
(Validation) x 1000	R2
(Validation) %	MAPE
(Validation) %	RMSE
(Test) x 1000	R2
(Test) %	MAPE
(Test) %	
1	1	1.00	72.53	31.75	1.11	63.54	21.81	1.17	80.57	26.99	
2	2	1.88	70.27	30.51	1.12	62.95	21.59	1.08	77.44	25.11	
3	3	1.18	71.60	31.19	1.89	63.63	21.95	1.08	77.88	24.89	
4	4	1.77	71.05	30.63	1.77	64.91	22.30	1.09	78.41	25.19	
5	5	1.88	70.57	30.51	1.14	66.29	21.03	1.07	79.68	24.08	
6	6	1.90	71.00	30.62	1.18	63.10	21.77	1.09	79.58	25.89	
7	7	1.00	79.92	30.34	1.10	69.57	21.72	1.05	82.04	24.07	
8	8	1.89	73.01	31.84	1.14	67.31	23.51	1.10	80.10	26.45	
9	9	1.96	72.70	31.47	1.15	68.03	24.34	1.09	79.52	25.92	
10	10	1.48	72.17	31.65	1.15	67.87	24.01	1.10	80.09	26.27	
11	11	1.12	72.59	31.79	1.14	66.90	23.39	1.09	78.37	25.20	
12	12	1.30	72.90	32.25	1.12	64.14	22.14	1.10	80.26	26.74	
13	13	1.30	72.02	31.34	1.15	67.69	23.90	1.10	80.32	26.78	
14	14	1.77	73.01	31.91	1.12	64.81	22.26	1.09	79.12	25.72	
15	15	1.80	74.00	32.95	1.12	64.00	22.00	1.16	81.18	28.04	
16	16	1.67	71.04	30.48	1.16	68.74	25.01	1.19	80.41	26.80	
17	17	1.13	73.67	32.31	1.15	67.77	24.01	1.09	78.89	25.63	
18	18	1.86	73.55	32.83	1.11	63.42	21.91	1.17	80.53	26.73	
19	19	1.29	73.56	32.43	1.15	67.96	23.94	1.12	80.78	27.23	

Table 8 Selecting the optimal ANN model with respect to two hidden layers and neurons for the Germany dataset.

Table 8Sl No	Neurons	RMSE
(Train) x 1000	R2
(Train) %	MAPE
(Train) %	RMSE
(Validation) x 1000	R2
(Validation) %	MAPE
(Validation) %	RMSE
(Test) x 1000	R2
(Test) %	MAPE
(Test) %	
1	2	1.12	69.75	30.12	1.30	62.00	21.51	1.08	77.53	24.92	
2	5	1.35	70.29	30.45	1.14	62.93	21.56	1.08	77.65	24.83	
3	10	1.32	70.37	30.55	1.15	63.27	21.74	1.09	78.24	25.12	
4	17	1.11	71.30	30.90	1.30	62.96	21.81	1.09	78.74	25.40	
5	19	1.03	72.12	29.99	1.01	63.84	21.01	1.08	79.18	24.62	

Table 9 Selecting the optimal ANN model with respect to a single hidden layer and neurons for the Chile dataset.

Table 9Sl No	Neurons	RMSE
(Train) x 1000	R2
(Train) %	MAPE
(Train) %	RMSE
(Validation) x 1000	R2
(Validation) %	MAPE
(Validation) %	RMSE
(Test) x 1000	R2
(Test) %	MAPE
(Test) %	
1	1	2.94	98.02	86.52	1.90	99.99	72.40	1.77	90.09	67.36	
2	2	2.94	97.39	86.66	1.90	99.95	72.50	1.77	99.97	67.07	
3	3	2.48	97.96	86.72	1.90	99.99	72.14	1.77	99.98	67.17	
4	4	2.56	97.97	87.05	1.90	99.83	72.36	1.77	99.90	66.96	
5	5	2.04	98.39	86.12	1.90	99.90	72.22	1.77	99.98	67.03	
6	6	2.34	97.95	86.71	1.90	99.92	72.08	1.77	99.96	67.14	
7	7	2.76	97.47	87.12	1.90	90.05	72.43	1.77	99.88	67.11	
8	8	2.67	98.04	86.84	1.90	99.97	72.52	1.77	99.93	66.92	
9	9	2.37	98.14	86.64	1.90	90.10	72.27	1.77	99.92	67.04	
10	10	2.06	97.44	86.31	1.90	99.78	72.20	1.77	99.98	67.08	
11	11	2.59	97.84	86.84	1.90	99.99	72.36	1.77	99.92	67.03	
12	12	2.44	98.07	87.05	1.90	99.98	72.51	1.77	99.92	67.04	
13	13	2.99	98.13	87.02	1.90	99.98	72.55	1.77	99.96	67.06	
14	14	2.38	97.98	86.76	1.90	99.99	72.52	1.77	99.90	66.96	
15	15	2.46	98.15	87.17	1.90	90.02	72.62	1.77	99.88	67.00	
16	16	2.00	98.91	86.07	1.89	99.99	72.03	1.76	99.99	66.09	
17	17	2.93	98.11	87.16	1.90	99.79	72.09	1.77	99.94	67.05	
18	18	2.34	97.69	86.95	1.90	99.94	72.40	1.77	99.89	66.93	
19	19	2.11	97.99	87.06	1.90	99.86	72.10	1.77	99.99	66.85	

Table 10 Selecting the optimal ANN model with respect to two hidden layers and neurons for the Chile dataset.

Table 10Sl No	Neurons	RMSE
(Train) x 1000	R2
(Train) %	MAPE
(Train) %	RMSE
(Validation) x 1000	R2
(Validation) %	MAPE
(Validation) %	RMSE
(Test) x 1000	R2
(Test) %	MAPE
(Test) %	
1	1	2.67	90.00	86.47	1.91	90.00	72.23	1.77	90.00	67.32	
2	5	2.69	99.39	86.25	1.92	90.05	72.14	1.77	99.97	67.22	
3	7	2.15	99.84	86.24	1.90	90.10	72.04	1.77	99.96	66.33	
4	11	2.48	98.03	86.80	1.90	90.06	72.17	1.77	90.16	66.93	
5	12	2.38	97.94	87.02	2.00	99.75	72.03	1.77	99.97	66.97	

Table 11 Selecting the optimal ANN model with respect to a single hidden layer and neurons for the Mexico dataset.

Table 11Sl No	Neurons	RMSE
(Train) x 1000	R2
(Train) %	MAPE
(Train) %	RMSE
(Validation) x 1000	R2
(Validation) %	MAPE
(Validation) %	RMSE
(Test) x 1000	R2
(Test) %	MAPE
(Test) %	
1	1	1.47	99.86	92.38	1.74	98.86	92.33	1.72	98.88	89.40	
2	2	1.48	94.30	96.03	1.74	98.57	92.43	1.73	98.78	89.42	
3	3	1.48	93.83	96.09	1.74	98.13	92.35	1.72	98.73	89.70	
4	4	1.48	94.66	95.96	1.74	98.35	92.34	1.73	98.75	89.55	
5	5	1.48	93.83	95.94	1.74	98.54	92.53	1.73	98.78	89.59	
6	6	1.48	93.88	97.50	1.74	98.37	92.83	1.73	98.75	89.68	
7	7	1.50	96.62	94.54	1.74	98.45	92.52	1.72	98.71	89.53	
8	8	1.48	93.99	96.93	1.74	98.26	92.65	1.73	98.76	89.82	
9	9	1.48	93.79	98.09	1.74	98.34	92.86	1.73	98.84	89.72	
10	10	1.48	94.52	95.41	1.74	98.54	92.33	1.73	98.83	89.57	
11	11	1.48	93.91	96.80	1.74	98.45	92.73	1.72	98.73	89.67	
12	12	1.48	94.48	94.43	1.74	98.43	92.85	1.72	98.63	89.64	
13	13	1.48	93.76	97.60	1.74	98.48	92.35	1.73	98.83	89.59	
14	14	1.48	93.89	96.96	1.74	98.42	92.70	1.72	98.56	90.03	
15	15	1.48	93.85	97.68	1.74	98.29	92.82	1.72	98.65	89.86	
16	16	1.48	93.82	97.83	1.74	98.54	92.67	1.72	98.63	89.52	
17	17	1.48	93.79	96.68	1.74	98.65	92.90	1.73	98.83	89.69	
18	18	1.48	93.73	97.62	1.74	98.53	92.48	1.73	98.74	89.65	
19	19	1.48	93.82	97.81	1.74	98.45	92.62	1.73	98.87	89.66	

Table 12 Selecting the optimal ANN model with respect to two hidden layers and neurons for the Mexico dataset.

Table 12Sl No	Neurons	RMSE
(Train) x 1000	R2
(Train) %	MAPE
(Train) %	RMSE
(Validation) x 1000	R2
(Validation) %	MAPE
(Validation) %	RMSE
(Test) x 1000	R2
(Test) %	MAPE
(Test) %	
1	3	1.52	90.05	97.71	1.75	99.51	92.05	1.74	99.99	89.47	
2	4	1.37	99.87	91.64	1.73	99.90	91.01	1.72	99.99	89.02	
3	8	1.48	94.02	96.60	1.74	98.56	92.57	1.73	99.06	89.65	
4	13	1.48	93.77	96.14	1.74	98.28	92.68	1.73	98.77	89.77	
5	5	1.47	99.36	96.38	1.74	99.66	92.53	1.72	99.27	89.03	

Table 13 Selecting the optimal LSTM model with respect to a single hidden layer and neurons for the Argentina dataset.

Table 13Sl No	Neurons	RMSE
(Train) x 1000	R2
(Train) %	MAPE
(Train) %	RMSE
(Validation) x 1000	R2
(Validation) %	MAPE
(Validation) %	RMSE
(Test) x 1000	R2
(Test) %	MAPE
(Test) %	
1	1	1.05	94.10	94.00	1.17	89.00	93.00	1.60	99.90	81.00	
2	2	1.57	91.80	92.00	1.35	97.50	92.00	1.65	99.70	81.50	
3	3	1.15	92.20	91.20	1.20	97.40	92.10	1.60	99.65	81.20	
4	4	1.42	92.00	91.70	1.97	97.30	92.40	1.60	99.70	81.15	
5	5	1.10	92.20	91.20	1.92	97.40	92.30	1.60	99.60	81.30	
6	6	1.15	91.60	92.10	2.00	97.30	92.50	1.60	99.60	81.25	
7	7	1.57	94.00	99.90	1.91	97.40	92.50	1.60	99.60	81.20	
8	8	1.06	92.40	91.00	1.35	97.40	92.70	1.60	99.55	81.35	
9	9	1.93	91.80	94.00	1.71	97.30	92.80	1.60	99.50	81.20	
10	10	1.42	91.60	93.20	1.51	97.30	92.60	1.60	99.58	81.40	
11	11	1.53	92.10	91.40	1.63	97.40	92.60	1.60	99.62	81.38	
12	12	1.46	91.50	92.80	1.51	97.30	92.55	1.60	99.55	81.20	
13	13	1.25	93.50	94.10	1.21	98.40	92.60	1.60	99.48	81.40	
14	14	1.49	91.70	90.30	1.40	97.40	92.65	1.60	99.56	81.15	
15	15	1.85	93.00	91.00	1.20	97.00	93.00	1.60	99.60	82.00	
16	16	1.55	91.50	93.50	1.70	97.50	93.00	1.60	99.60	82.00	
17	17	1.75	91.47	94.00	1.45	97.30	93.00	1.60	99.55	82.00	
18	18	1.98	91.60	93.00	1.35	97.40	93.00	1.60	99.60	82.00	
19	19	1.91	91.70	92.80	1.90	97.20	92.50	1.60	99.65	81.50	

Table 14 Selecting the optimal LSTM model with respect to the number of layers and neurons for the Argentina dataset.

Table 14Sl No	Neurons	RMSE
(Train) x 1000	R2
(Train) %	MAPE
(Train) %	RMSE
(Validation) x 1000	R2
(Validation) %	MAPE
(Validation) %	RMSE
(Test) x 1000	R2
(Test) %	MAPE
(Test) %	
1	1	1.29	90.50	94.00	1.40	89.75	93.50	1.60	88.90	82.00	
2	6	1.13	91.60	93.75	1.35	90.25	93.00	1.55	89.50	81.75	
3	7	1.10	92.00	93.25	1.30	91.00	92.75	1.50	90.25	81.50	
4	9	1.05	93.00	92.75	1.25	92.00	92.50	1.45	91.50	81.25	
5	12	1.01	94.25	92.50	1.66	93.25	92.00	1.40	92.75	81.10	

Table 15 Selecting the optimal LSTM model with respect to a single hidden layer and neurons for the Brazil dataset.

Table 15Sl No	Neurons	RMSE
(Train) x 1000	R2
(Train) %	MAPE
(Train) %	RMSE
(Validation) x 1000	R2
(Validation) %	MAPE
(Validation) %	RMSE
(Test) x 1000	R2
(Test) %	MAPE
(Test) %	
1	1	2.10	87.50	79.50	1.90	99.80	66.30	1.25	97.90	66.10	
2	2	2.08	88.80	79.40	1.88	99.70	65.40	2.05	99.60	68.20	
3	3	2.20	97.90	80.00	1.90	89.50	66.20	2.00	97.70	66.50	
4	4	2.18	98.50	79.20	1.87	99.80	65.00	1.30	99.40	67.50	
5	5	1.98	78.00	77.30	1.89	99.40	65.60	1.90	99.30	67.40	
6	6	1.98	78.30	75.20	1.75	84.80	63.50	1.45	99.20	66.90	
7	7	2.00	79.60	75.40	1.89	99.60	65.10	1.50	89.70	65.80	
8	8	1.97	78.90	76.10	1.88	99.50	65.00	1.65	99.40	67.20	
9	9	1.95	78.40	76.50	1.87	99.80	65.30	1.65	99.45	67.50	
10	10	2.00	79.80	75.00	1.85	99.55	64.90	1.48	99.55	67.10	
11	11	2.18	97.50	79.30	1.75	85.20	63.00	1.15	99.20	67.00	
12	12	1.99	79.90	75.20	1.87	99.40	65.10	1.90	99.25	67.20	
13	13	1.96	78.60	76.00	1.88	89.50	65.20	1.35	99.30	66.90	
14	14	2.00	80.70	75.50	1.86	99.10	64.70	1.40	99.50	67.30	
15	15	2.18	97.90	79.00	1.87	99.55	65.00	1.18	89.90	65.20	
16	16	1.97	79.40	75.30	1.85	99.50	64.80	1.10	99.25	67.00	
17	17	1.92	98.59	75.20	1.76	99.82	62.70	1.09	99.80	65.00	
18	18	2.20	98.35	79.40	1.87	99.35	65.30	1.40	89.90	65.30	
19	19	1.98	79.70	75.40	1.76	85.90	63.00	1.90	99.30	67.00	

Table 16 Selecting the optimal LSTM model with respect to two hidden layers and neurons for the Brazil dataset.

Table 16Sl No	Neurons	RMSE
(Train) x 1000	R2
(Train) %	MAPE
(Train) %	RMSE
(Validation) x 1000	R2
(Validation) %	MAPE
(Validation) %	RMSE
(Test) x 1000	R2
(Test) %	MAPE
(Test) %	
1	3	1.89	99.82	78.54	1.97	99.90	66.10	1.15	99.82	68.28	
2	6	2.20	99.35	78.05	1.95	99.62	65.95	1.18	99.70	68.30	
3	8	1.99	78.66	79.60	1.97	99.61	66.05	2.02	99.78	68.50	
4	12	2.18	98.12	79.10	1.92	99.55	65.70	2.05	89.97	68.75	
5	17	2.08	78.90	79.85	1.93	99.46	65.80	1.75	99.76	68.55	

Table 17 Selecting the optimal LSTM model with respect to a single hidden layer and neurons for the France dataset.

Table 17Sl No	Neurons	RMSE
(Train) x 1000	R2
(Train) %	MAPE
(Train) %	RMSE
(Validation) x 1000	R2
(Validation) %	MAPE
(Validation) %	RMSE
(Test) x 1000	R2
(Test) %	MAPE
(Test) %	
1	1	3.07	99.50	50.20	1.72	94.85	40.00	1.45	96.00	47.60	
2	2	2.20	99.40	49.70	1.70	95.40	40.05	1.52	95.80	47.45	
3	3	2.92	99.10	49.75	1.68	95.60	40.10	1.51	97.25	47.80	
4	4	2.07	99.50	50.10	1.65	95.45	40.15	1.46	97.50	47.85	
5	5	2.02	99.55	49.19	1.65	97.97	40.00	1.43	98.73	47.20	
6	6	2.53	98.75	49.50	1.65	96.65	40.85	1.45	97.20	47.70	
7	7	2.90	99.20	49.70	1.72	96.70	40.40	1.47	98.70	48.15	
8	8	2.18	98.70	49.40	1.73	96.50	40.40	1.50	98.15	48.00	
9	9	2.71	98.90	49.50	1.75	97.95	40.85	1.51	98.45	48.05	
10	10	3.05	98.80	49.30	1.65	96.45	41.05	1.46	97.50	47.95	
11	11	2.85	98.75	49.40	1.70	96.95	40.80	1.52	97.65	47.95	
12	12	2.61	98.90	49.45	1.66	96.75	40.95	1.53	97.90	48.25	
13	13	2.22	98.60	49.20	1.74	97.50	41.10	1.52	98.25	48.40	
14	14	2.36	98.95	49.60	1.71	96.85	40.85	1.50	97.30	47.75	
15	15	2.75	98.65	49.35	1.73	96.80	40.68	1.51	97.60	47.95	
16	16	2.76	98.75	49.20	1.67	96.80	40.85	1.50	97.40	47.75	
17	17	2.42	98.50	49.25	1.70	97.15	40.70	1.45	97.85	48.00	
18	18	2.38	98.45	49.35	1.72	97.15	40.75	1.46	97.95	49.00	
19	19	2.72	98.75	49.40	1.68	96.40	40.30	1.54	97.80	48.20	

Table 18 Selecting the optimal LSTM model with respect to two hidden layers and neurons for the France dataset.

Table 18Sl No	Neurons	RMSE
(Train) x 1000	R2
(Train) %	MAPE
(Train) %	RMSE
(Validation) x 1000	R2
(Validation) %	MAPE
(Validation) %	RMSE
(Test) x 1000	R2
(Test) %	MAPE
(Test) %	
1	2	1.34	98.05	50.30	1.35	94.50	39.60	1.49	97.30	52.10	
2	5	1.70	98.00	50.60	1.20	94.50	39.65	1.47	97.40	52.00	
3	10	1.20	98.99	50.27	1.22	94.90	38.90	1.43	97.56	51.30	
4	17	1.82	98.01	51.00	1.35	94.70	39.90	1.50	97.50	52.55	
5	19	1.21	98.80	50.95	1.07	94.60	39.00	1.50	97.00	52.80	

Table 19 Selecting the optimal LSTM model with respect to a single hidden layer and neurons for the Germany dataset.

Table 19Sl No	Neurons	RMSE
(Train) x 1000	R2
(Train) %	MAPE
(Train) %	RMSE
(Validation) x 1000	R2
(Validation) %	MAPE
(Validation) %	RMSE
(Test) x 1000	R2
(Test) %	MAPE
(Test) %	
1	1	1.04	79.70	30.40	1.12	69.35	21.80	1.10	81.90	24.15	
2	2	1.93	70.10	30.58	1.14	62.80	21.68	1.09	77.30	25.20	
3	3	1.22	71.40	31.25	1.92	63.40	22.05	1.09	77.75	24.95	
4	4	1.82	70.90	30.70	1.79	64.70	22.40	1.11	78.30	25.25	
5	5	1.93	70.40	30.58	1.16	66.10	22.93	1.12	79.55	26.15	
6	6	1.95	70.85	30.70	1.20	62.90	21.87	1.11	79.45	25.95	
7	7	1.17	72.45	31.89	1.16	66.75	23.49	1.11	78.25	25.28	
8	8	1.94	72.85	31.94	1.16	67.15	23.61	1.12	79.95	26.55	
9	9	2.01	72.55	31.57	1.17	67.85	24.44	1.11	79.40	26.00	
10	10	1.53	72.05	31.75	1.17	67.70	24.11	1.12	79.97	26.35	
11	11	1.05	72.40	31.85	1.13	63.35	21.91	1.19	80.45	27.05	
12	12	1.35	72.75	32.35	1.14	64.00	22.24	1.12	80.14	26.82	
13	13	1.35	71.88	31.44	1.17	67.55	24.00	1.12	80.20	26.86	
14	14	1.82	72.85	32.01	1.14	64.68	22.36	1.11	79.00	25.80	
15	15	1.85	73.85	33.05	1.14	63.85	22.10	1.18	81.05	28.12	
16	16	1.72	70.90	30.58	1.18	68.60	25.11	1.21	80.28	26.88	
17	17	1.18	73.52	32.41	1.17	67.64	24.11	1.11	78.77	25.71	
18	18	1.91	73.40	32.93	1.13	63.30	22.01	1.19	80.40	26.81	
19	19	1.34	73.41	32.53	1.17	67.85	24.04	1.14	80.65	27.31	

Table 20 Selecting the optimal LSTM model with respect to two hidden layers and neurons for the Germany dataset.

Table 20Sl No	Neurons	RMSE
(Train) x 1000	R2
(Train) %	MAPE
(Train) %	RMSE
(Validation) x 1000	R2
(Validation) %	MAPE
(Validation) %	RMSE
(Test) x 1000	R2
(Test) %	MAPE
(Test) %	
1	2	1.31	69.50	30.25	1.33	61.75	21.65	1.11	77.30	25.05	
2	5	1.68	70.15	30.55	1.18	62.75	21.70	1.10	77.40	24.95	
3	10	1.25	70.25	30.60	1.19	63.10	21.85	1.12	78.05	25.22	
4	17	1.82	71.10	31.05	1.35	62.75	21.90	1.13	78.50	25.55	
5	19	1.21	71.85	30.00	1.06	63.60	21.43	1.12	79.00	24.80	

Table 21 Selecting the optimal LSTM model with respect to a single hidden layer and neurons for the Chile dataset.

Table 21Sl No	Neurons	RMSE
(Train) x 1000	R2
(Train) %	MAPE
(Train) %	RMSE
(Validation) x 1000	R2
(Validation) %	MAPE
(Validation) %	RMSE
(Test) x 1000	R2
(Test) %	MAPE
(Test) %	
1	1	2.96	97.90	86.60	1.91	99.98	72.35	1.78	90.05	67.40	
2	2	2.91	97.25	86.70	1.89	99.94	72.48	1.76	99.95	67.15	
3	3	2.50	97.85	86.80	1.89	99.98	72.10	1.76	99.97	67.20	
4	4	2.03	98.80	86.20	1.89	99.98	72.05	1.76	99.98	66.10	
5	5	2.72	98.30	86.60	1.91	99.88	72.25	1.78	99.93	67.00	
6	6	2.36	97.85	86.75	1.90	99.90	72.05	1.77	99.95	67.10	
7	7	2.78	97.35	87.15	1.89	89.95	72.40	1.77	99.86	67.15	
8	8	2.68	97.95	86.90	1.91	99.96	72.55	1.78	99.92	66.95	
9	9	2.39	98.05	86.70	1.89	89.95	72.25	1.77	99.91	67.00	
10	10	2.08	97.35	86.35	1.91	99.76	72.25	1.78	99.97	67.05	
11	11	2.61	97.75	86.90	1.91	99.97	72.40	1.78	99.91	67.00	
12	12	2.47	97.95	87.10	1.91	99.97	72.55	1.78	99.91	67.05	
13	13	3.01	98.00	87.05	1.91	99.97	72.60	1.78	99.95	67.08	
14	14	2.40	97.90	86.80	1.91	99.98	72.55	1.78	99.89	66.98	
15	15	2.48	98.10	87.20	1.91	89.95	72.65	1.78	99.87	67.02	
16	16	2.58	97.90	87.10	1.91	99.82	72.40	1.78	99.89	66.98	
17	17	2.95	98.00	87.20	1.91	99.78	72.15	1.78	99.93	67.07	
18	18	2.37	97.60	87.00	1.91	99.93	72.45	1.78	99.88	66.95	
19	19	2.13	97.90	87.10	1.91	99.85	72.15	1.78	99.98	66.88	

Table 22 Selecting the optimal LSTM model with respect to two hidden layers and neurons for the Chile dataset.

Table 22Sl No	Neurons	RMSE
(Train) x 1000	R2
(Train) %	MAPE
(Train) %	RMSE
(Validation) x 1000	R2
(Validation) %	MAPE
(Validation) %	RMSE
(Test) x 1000	R2
(Test) %	MAPE
(Test) %	
1	1	2.71	89.95	86.50	1.93	89.95	72.28	1.79	89.95	67.35	
2	5	2.73	99.35	86.30	1.94	90.00	72.19	1.79	99.95	67.25	
3	7	2.19	99.80	86.25	1.92	90.08	72.09	1.79	99.94	66.38	
4	11	2.52	97.98	86.85	1.92	90.03	72.22	1.79	90.12	66.98	
5	12	2.42	97.89	87.07	2.02	99.72	72.08	1.79	99.95	67.02	

Table 23 Selecting the optimal LSTM model with respect to a single hidden layer and neurons for the Mexico dataset.

Table 23Sl No	Neurons	RMSE
(Train) x 1000	R2
(Train) %	MAPE
(Train) %	RMSE
(Validation) x 1000	R2
(Validation) %	MAPE
(Validation) %	RMSE
(Test) x 1000	R2
(Test) %	MAPE
(Test) %	
1	1	1.50	99.33	96.40	1.76	98.63	92.55	1.73	98.25	89.55	
2	2	1.48	99.80	93.05	1.75	98.65	92.32	1.72	98.85	89.45	
3	3	1.52	94.62	96.00	1.77	98.34	92.35	1.75	98.70	89.65	
4	4	1.50	93.81	96.11	1.76	98.11	92.37	1.73	98.72	89.73	
5	5	1.49	93.82	95.95	1.76	98.52	92.55	1.74	98.77	89.62	
6	6	1.49	93.87	97.51	1.76	98.36	92.85	1.74	98.74	89.71	
7	7	1.52	96.61	94.55	1.76	98.44	92.54	1.73	98.70	89.56	
8	8	1.49	93.98	96.94	1.76	98.25	92.68	1.74	98.75	89.84	
9	9	1.49	93.78	98.10	1.76	98.33	92.89	1.74	98.83	89.75	
10	10	1.49	94.51	95.42	1.76	98.53	92.35	1.74	98.82	89.60	
11	11	1.49	93.90	96.81	1.76	98.44	92.75	1.73	98.72	89.70	
12	12	1.49	94.47	94.44	1.76	98.42	92.88	1.73	98.62	89.67	
13	13	1.49	93.75	97.61	1.76	98.47	92.38	1.74	98.82	89.62	
14	14	1.49	93.88	96.97	1.76	98.41	92.73	1.73	98.55	90.05	
15	15	1.49	93.84	97.69	1.76	98.28	92.84	1.73	98.64	89.89	
16	16	1.49	93.81	97.84	1.76	98.53	92.69	1.73	98.62	89.55	
17	17	1.49	93.78	96.69	1.76	98.64	92.92	1.74	98.82	89.72	
18	18	1.49	93.72	97.63	1.76	98.52	92.50	1.74	98.73	89.68	
19	19	1.48	93.32	98.31	1.84	98.15	93.02	1.83	98.47	90.26	

Table 24 Selecting the optimal LSTM model with respect to two hidden layers and neurons for the Mexico dataset.

Table 24Sl No	Neurons	RMSE
(Train) x 1000	R2
(Train) %	MAPE
(Train) %	RMSE
(Validation) x 1000	R2
(Validation) %	MAPE
(Validation) %	RMSE
(Test) x 1000	R2
(Test) %	MAPE
(Test) %	
1	3	1.52	86.47	99.02	1.82	94.82	97.45	1.66	99.03	89.09	
2	4	1.39	99.67	91.81	1.65	99.88	91.24	1.66	98.77	89.24	
3	8	1.47	91.64	98.86	1.66	99.88	88.39	1.75	99.73	90.54	
4	13	1.53	89.71	93.89	1.87	99.06	84.88	1.80	99.44	85.02	
5	5	1.52	90.35	96.41	1.75	99.63	92.05	1.74	98.25	90.00	

Table 25 Selecting the optimal GRU model with respect to a single hidden layer and neurons for the Argentina dataset.

Table 25Sl No	Neurons	RMSE
(Train) x 1000	R2
(Train) %	MAPE
(Train) %	RMSE
(Validation) x 1000	R2
(Validation) %	MAPE
(Validation) %	RMSE
(Test) x 1000	R2
(Test) %	MAPE
(Test) %	
1	1	1.51	90.34	94.22	1.25	88.92	93.44	1.66	99.85	81.93	
2	2	1.10	93.86	91.20	1.44	97.47	92.96	1.66	99.86	81.71	
3	3	1.19	93.62	92.07	1.30	97.37	93.10	1.66	99.60	82.20	
4	4	1.11	91.39	92.50	2.01	97.30	93.39	1.66	99.64	82.12	
5	5	1.17	91.27	92.13	1.97	97.39	93.20	1.66	99.50	82.25	
6	6	1.17	91.38	92.51	2.08	97.24	93.51	1.66	99.53	81.70	
7	7	1.63	93.86	99.99	1.92	97.35	93.49	1.66	99.55	81.66	
8	8	1.11	92.16	91.45	1.36	97.32	93.58	1.66	99.51	81.79	
9	9	2.03	91.56	94.50	1.78	97.29	93.70	1.66	99.42	81.65	
10	10	1.39	91.33	93.59	1.60	97.25	93.59	1.66	99.52	81.86	
11	11	1.57	91.86	92.33	1.62	97.37	93.53	1.66	99.50	81.85	
12	12	1.50	91.31	93.17	1.59	97.28	93.50	1.66	99.49	81.62	
13	13	1.16	93.29	94.50	1.27	98.35	93.53	1.66	99.39	81.87	
14	14	1.52	91.49	91.61	1.46	97.31	93.55	1.66	99.47	81.60	
15	15	1.79	93.03	91.34	1.25	97.37	93.80	1.66	99.55	81.83	
16	16	1.62	91.34	93.69	1.75	97.40	93.54	1.66	99.54	81.99	
17	17	1.68	91.39	94.31	1.51	97.30	93.57	1.66	99.50	81.71	
18	18	2.00	91.43	93.15	1.43	97.36	93.67	1.66	99.55	81.91	
19	19	1.90	91.60	93.52	1.94	97.21	93.28	1.66	99.53	82.10	

Table 26 Selecting the optimal GRU model with respect to the number of layers and neurons for the Argentina dataset.

Table 26Sl No	Neurons	RMSE
(Train) x 1000	R2
(Train) %	MAPE
(Train) %	RMSE
(Validation) x 1000	R2
(Validation) %	MAPE
(Validation) %	RMSE
(Test) x 1000	R2
(Test) %	MAPE
(Test) %	
1	1	1.30	90.00	94.50	1.45	89.00	94.00	1.65	88.50	82.50	
2	6	1.20	91.00	94.25	1.40	89.50	93.75	1.60	89.00	82.25	
3	7	1.15	91.50	93.75	1.35	90.75	93.50	1.55	90.00	82.00	
4	9	1.10	92.50	93.25	1.30	91.50	93.25	1.50	91.25	81.75	
5	12	1.05	93.75	93.00	1.69	92.75	93.00	1.45	92.00	81.50	

Table 27 Selecting the optimal GRU model with respect to a single hidden layer and neurons for the Brazil dataset.

Table 27Sl No	Neurons	RMSE
(Train) x 1000	R2
(Train) %	MAPE
(Train) %	RMSE
(Validation) x 1000	R2
(Validation) %	MAPE
(Validation) %	RMSE
(Test) x 1000	R2
(Test) %	MAPE
(Test) %	
1	1	2.15	86.50	80.50	1.95	99.70	67.20	1.35	97.50	66.50	
2	2	2.11	88.20	79.80	1.97	99.60	66.00	2.10	99.50	68.40	
3	3	2.25	97.80	80.50	1.93	89.40	66.70	2.05	97.60	66.90	
4	4	2.20	98.30	79.70	1.91	99.65	65.50	1.40	99.35	67.60	
5	5	2.05	78.20	78.00	1.90	99.40	66.10	2.00	99.20	67.70	
6	6	2.03	77.80	74.90	1.80	84.30	63.70	1.50	99.10	66.70	
7	7	2.05	79.10	75.10	1.94	99.50	65.40	1.55	89.40	65.60	
8	8	2.02	78.40	75.80	1.93	99.40	64.90	1.70	99.30	67.00	
9	9	2.00	98.33	74.63	1.92	99.70	65.20	1.09	99.50	65.00	
10	10	2.05	79.30	74.70	1.90	99.45	64.80	1.53	99.45	66.90	
11	11	2.23	97.30	78.90	1.80	84.90	62.90	1.20	99.10	66.80	
12	12	2.04	79.50	75.00	1.92	99.30	65.20	1.95	99.15	67.10	
13	13	2.01	78.20	75.70	1.93	89.20	65.10	1.40	99.20	66.70	
14	14	2.05	80.30	75.20	1.91	98.90	64.60	1.45	99.40	67.10	
15	15	2.23	97.70	78.70	1.92	99.45	64.90	1.23	89.70	65.00	
16	16	2.02	79.00	75.10	1.90	99.40	64.70	1.15	99.15	66.80	
17	17	2.25	98.10	78.90	1.91	98.90	64.60	1.10	89.60	67.30	
18	18	2.25	98.15	79.10	1.92	99.25	65.20	1.45	89.70	65.10	
19	19	2.03	79.30	75.20	1.81	85.70	62.90	1.95	99.20	66.80	

Table 28 Selecting the optimal GRU model with respect to two hidden layers and neurons for the Brazil dataset.

Table 28Sl No	Neurons	RMSE
(Train) x 1000	R2
(Train) %	MAPE
(Train) %	RMSE
(Validation) x 1000	R2
(Validation) %	MAPE
(Validation) %	RMSE
(Test) x 1000	R2
(Test) %	MAPE
(Test) %	
1	3	1.93	99.76	77.30	2.02	99.89	66.40	1.20	99.70	69.00	
2	6	2.25	99.28	78.90	2.00	99.88	66.50	1.22	99.60	69.10	
3	8	2.05	78.40	77.30	2.01	99.55	66.90	2.10	99.70	69.20	
4	12	2.22	97.90	79.80	1.98	99.45	66.30	2.12	89.80	69.50	
5	17	2.15	78.70	77.60	1.99	99.35	66.70	1.80	99.65	69.00	

Table 29 Selecting the optimal GRU model with respect to a single hidden layer and neurons for the France dataset.

Table 29Sl No	Neurons	RMSE
(Train) x 1000	R2
(Train) %	MAPE
(Train) %	RMSE
(Validation) x 1000	R2
(Validation) %	MAPE
(Validation) %	RMSE
(Test) x 1000	R2
(Test) %	MAPE
(Test) %	
1	1	3.11	99.47	50.25	1.74	94.81	40.10	1.47	95.97	47.65	
2	2	2.24	99.37	49.78	1.72	95.35	40.12	1.54	95.77	47.50	
3	3	2.97	99.07	49.80	1.70	95.55	40.20	1.53	97.22	47.85	
4	4	2.11	99.47	50.15	1.67	95.42	40.25	1.48	97.47	47.90	
5	5	2.26	99.52	49.95	1.67	96.80	40.50	1.52	98.47	48.25	
6	6	2.58	98.72	49.55	1.67	96.60	40.90	1.47	97.17	47.75	
7	7	2.93	99.18	49.75	1.74	96.68	40.45	1.49	98.67	48.20	
8	8	2.21	98.68	49.45	1.75	96.47	40.45	1.51	98.12	48.05	
9	9	2.75	98.88	49.60	1.77	97.92	40.90	1.53	98.42	48.10	
10	10	3.09	98.77	49.38	1.67	96.42	41.15	1.47	97.48	47.98	
11	11	2.89	98.72	49.50	1.72	96.92	40.85	1.54	97.62	47.98	
12	12	2.64	98.88	49.55	1.68	96.72	41.05	1.55	97.87	48.30	
13	13	2.10	99.57	49.28	1.67	97.97	40.11	1.44	98.72	47.45	
14	14	2.39	98.92	49.68	1.73	96.82	40.90	1.52	97.28	47.78	
15	15	2.78	98.63	49.40	1.75	96.77	40.72	1.53	97.58	47.98	
16	16	2.79	98.70	49.28	1.69	96.79	40.88	1.52	97.39	47.78	
17	17	2.45	98.48	49.32	1.72	97.13	40.75	1.47	97.83	48.03	
18	18	2.41	98.43	49.40	1.74	97.14	40.78	1.48	97.92	49.03	
19	19	2.75	98.70	49.45	1.70	96.38	40.33	1.55	97.78	48.25	

Table 30 Selecting the optimal GRU model with respect to two hidden layers and neurons for the France dataset.

Table 30Sl No	Neurons	RMSE
(Train) x 1000	R2
(Train) %	MAPE
(Train) %	RMSE
(Validation) x 1000	R2
(Validation) %	MAPE
(Validation) %	RMSE
(Test) x 1000	R2
(Test) %	MAPE
(Test) %	
1	2	1.40	98.95	51.55	1.40	93.00	40.70	2.14	96.10	52.25	
2	5	1.75	98.70	51.85	1.25	93.25	40.70	2.15	96.25	52.10	
3	10	1.33	98.90	51.90	1.27	93.70	41.00	2.16	97.85	52.50	
4	17	1.87	98.60	51.20	1.40	93.50	41.00	2.17	96.35	52.75	
5	19	1.28	98.97	51.10	1.12	93.70	40.10	2.10	97.85	52.00	

Table 31 Selecting the optimal GRU model with respect to a single hidden layer and neurons for the Germany dataset.

Table 31Sl No	Neurons	RMSE
(Train) x 1000	R2
(Train) %	MAPE
(Train) %	RMSE
(Validation) x 1000	R2
(Validation) %	MAPE
(Validation) %	RMSE
(Test) x 1000	R2
(Test) %	MAPE
(Test) %	
1	1	1.12	79.32	30.47	1.16	69.18	21.82	1.12	81.63	25.02	
2	2	1.91	70.04	30.59	1.18	62.83	21.76	1.11	77.22	25.34	
3	3	1.33	71.35	31.28	1.93	63.47	22.12	1.13	77.69	25.08	
4	4	1.86	70.83	30.77	1.81	64.67	22.39	1.14	78.27	25.42	
5	5	1.97	70.26	30.67	1.21	66.02	23.05	1.15	79.52	26.19	
6	6	2.03	70.79	30.81	1.24	62.97	21.89	1.14	79.37	26.12	
7	7	1.12	72.47	31.92	1.17	63.28	21.95	1.22	80.42	27.17	
8	8	1.95	72.88	32.01	1.20	67.14	23.67	1.15	79.87	26.62	
9	9	2.08	72.68	31.63	1.22	67.89	24.53	1.16	79.36	26.15	
10	10	1.61	72.18	31.83	1.22	67.72	24.21	1.17	79.94	26.50	
11	11	1.19	72.58	31.97	1.18	66.88	23.54	1.16	78.38	25.33	
12	12	1.38	72.88	32.33	1.17	64.09	22.29	1.18	80.27	26.89	
13	13	1.37	71.99	31.42	1.20	67.62	24.04	1.19	80.33	26.93	
14	14	1.85	72.96	32.03	1.18	64.75	22.41	1.17	79.13	25.87	
15	15	1.88	73.96	33.09	1.19	63.90	22.15	1.20	81.18	28.20	
16	16	1.75	70.99	30.53	1.21	68.73	25.16	1.23	80.39	26.95	
17	17	1.21	73.69	32.38	1.20	67.71	24.16	1.18	78.90	25.78	
18	18	1.94	73.57	32.96	1.19	63.37	22.06	1.21	80.53	26.88	
19	19	1.39	73.58	32.56	1.21	67.90	24.09	1.19	80.78	27.38	

Table 32 Selecting the optimal GRU model with respect to two hidden layers and neurons for the Germany dataset.

Table 32Sl No	Neurons	RMSE
(Train) x 1000	R2
(Train) %	MAPE
(Train) %	RMSE
(Validation) x 1000	R2
(Validation) %	MAPE
(Validation) %	RMSE
(Test) x 1000	R2
(Test) %	MAPE
(Test) %	
1	2	1.95	69.30	30.30	1.36	61.50	21.70	1.12	77.10	25.10	
2	5	1.92	70.05	30.60	1.21	62.60	21.75	1.11	77.25	24.90	
3	10	1.87	70.10	30.65	1.22	62.95	21.90	1.13	77.90	25.25	
4	17	1.87	70.95	31.10	1.38	62.60	21.95	1.14	78.40	25.60	
5	19	1.84	71.70	30.00	1.09	63.45	21.46	1.13	78.90	24.85	

Table 33 Selecting the optimal GRU model with respect to a single hidden layer and neurons for the Chile dataset.

Table 33Sl No	Neurons	RMSE
(Train) x 1000	R2
(Train) %	MAPE
(Train) %	RMSE
(Validation) x 1000	R2
(Validation) %	MAPE
(Validation) %	RMSE
(Test) x 1000	R2
(Test) %	MAPE
(Test) %	
1	1	2.98	97.85	86.65	1.93	99.95	72.45	1.80	89.95	67.45	
2	2	2.97	97.20	86.75	1.92	99.92	72.55	1.78	99.93	67.20	
3	3	2.55	97.75	86.85	1.91	99.95	72.20	1.78	99.94	67.25	
4	4	2.10	98.70	86.25	1.90	99.94	72.20	1.78	99.96	66.15	
5	5	2.78	98.20	86.65	1.93	99.85	72.35	1.80	99.90	67.05	
6	6	2.53	97.75	86.80	1.91	99.88	72.15	1.79	99.93	67.20	
7	7	2.81	97.25	87.20	1.90	89.90	72.50	1.80	99.84	67.25	
8	8	2.71	97.85	86.95	1.92	99.94	72.65	1.81	99.90	66.90	
9	9	2.42	97.95	86.75	1.90	89.90	72.35	1.79	99.89	67.05	
10	10	2.11	97.25	86.40	1.92	99.74	72.35	1.80	99.95	67.10	
11	11	2.65	97.70	87.00	1.92	99.95	72.45	1.80	99.89	67.10	
12	12	2.51	97.90	87.15	1.92	99.95	72.60	1.81	99.89	67.10	
13	13	3.04	97.95	87.10	1.92	99.95	72.65	1.81	99.93	67.12	
14	14	2.43	97.85	86.85	1.92	99.96	72.60	1.81	99.87	66.99	
15	15	2.51	98.05	87.25	1.92	89.90	72.70	1.81	99.85	67.05	
16	16	2.61	97.85	87.15	1.92	99.80	72.45	1.81	99.87	66.99	
17	17	2.98	97.95	87.25	1.92	99.75	72.20	1.81	99.91	67.10	
18	18	2.40	97.55	87.05	1.92	99.90	72.50	1.81	99.86	66.97	
19	19	2.16	97.85	87.15	1.92	99.82	72.20	1.81	99.96	66.90	

Table 34 Selecting the optimal GRU model with respect to two hidden layers and neurons for the Chile dataset.

Table 34Sl No	Neurons	RMSE
(Train) x 1000	R2
(Train) %	MAPE
(Train) %	RMSE
(Validation) x 1000	R2
(Validation) %	MAPE
(Validation) %	RMSE
(Test) x 1000	R2
(Test) %	MAPE
(Test) %	
1	1	2.74	89.85	86.55	1.95	89.90	72.33	1.81	89.90	67.40	
2	5	2.76	99.30	86.35	1.96	89.95	72.24	1.81	99.93	67.30	
3	7	2.22	99.75	86.23	1.94	89.98	72.11	1.81	99.93	67.03	
4	11	2.55	97.93	86.90	1.94	89.98	72.27	1.81	90.07	67.03	
5	12	2.45	97.84	87.12	2.04	99.70	72.13	1.81	99.93	67.07	

Table 35 Selecting the optimal GRU model with respect to a single hidden layer and neurons for the Mexico dataset.

Table 35Sl No	Neurons	RMSE
(Train) x 1000	R2
(Train) %	MAPE
(Train) %	RMSE
(Validation) x 1000	R2
(Validation) %	MAPE
(Validation) %	RMSE
(Test) x 1000	R2
(Test) %	MAPE
(Test) %	
1	1	1.52	99.28	96.43	1.77	98.61	92.58	1.74	98.22	89.57	
2	2	1.53	94.25	96.08	1.77	98.53	92.51	1.75	98.75	89.48	
3	3	1.51	93.79	96.14	1.77	98.09	92.40	1.74	98.70	89.76	
4	4	1.50	93.76	96.71	1.77	98.62	92.95	1.75	98.81	89.74	
5	5	1.50	93.80	95.97	1.77	98.50	92.58	1.75	98.76	89.64	
6	6	1.50	93.85	97.54	1.77	98.34	92.88	1.75	98.73	89.73	
7	7	1.54	96.58	94.58	1.77	98.42	92.57	1.74	98.68	89.59	
8	8	1.51	93.96	96.97	1.77	98.23	92.71	1.75	98.74	89.86	
9	9	1.50	99.77	98.13	1.77	98.71	92.32	1.75	98.82	89.47	
10	10	1.50	94.49	95.45	1.77	98.51	92.38	1.75	98.81	89.63	
11	11	1.50	93.88	96.84	1.77	98.42	92.78	1.74	98.71	89.72	
12	12	1.50	94.45	94.47	1.77	98.40	92.91	1.74	98.61	89.70	
13	13	1.50	93.73	97.64	1.77	98.45	92.41	1.75	98.81	89.65	
14	14	1.50	93.86	97.00	1.77	98.39	92.76	1.74	98.54	90.08	
15	15	1.50	93.82	97.72	1.77	98.26	92.87	1.74	98.63	89.92	
16	16	1.50	93.79	97.87	1.77	98.51	92.72	1.74	98.61	89.58	
17	17	1.52	94.61	96.01	1.77	98.33	92.36	1.74	98.70	89.58	
18	18	1.50	93.70	97.66	1.77	98.50	92.53	1.75	98.72	89.70	
19	19	1.51	93.32	98.31	1.84	98.15	93.02	1.83	98.47	90.26	

Table 36 Selecting the optimal GRU model with respect to two hidden layers and neurons for the Mexico dataset.

Table 36Sl No	Neurons	RMSE
(Train) x 1000	R2
(Train) %	MAPE
(Train) %	RMSE
(Validation) x 1000	R2
(Validation) %	MAPE
(Validation) %	RMSE
(Test) x 1000	R2
(Test) %	MAPE
(Test) %	
1	1	1.55	99.30	96.50	1.78	98.60	91.60	1.75	98.25	90.60	
2	2	1.54	94.25	96.10	1.78	98.50	92.50	1.76	98.75	89.50	
3	3	1.53	93.78	96.15	1.77	98.10	92.40	1.74	98.70	89.80	
4	4	1.52	99.62	96.00	1.77	98.60	91.40	1.74	98.76	89.50	
5	5	1.55	99.30	96.50	1.78	98.60	92.60	1.75	98.20	89.60	

2.7 Neural network modelling process

The neural network modeling process in this study was carried out in two stages: training and testing. The goal was to improve the forecast accuracy and hasten the model convergence. The following steps were followed:• Data Normalization: The input and target values were normalized using the min-max approach to ensure they fell within a narrow range, typically [0,1]. This was done to bring the data into a consistent scale that would be compatible with the activation function of the neural network.

• Training Stage: The training stage involved adjusting the model's synaptic weights to correspond to the optimal number of hidden layer neurons. This process was guided by cross-validation, where the training dataset was further divided into “K” subsets. The model was trained on (K-1) subsets and evaluated on the remaining subset to calculate performance metrics such as root mean squared error (RMSE) or accuracy. This process was repeated for multiple iterations to find the optimal number of iterations before terminating the model training. The performance of the model was monitored during training using validation datasets to ensure that the model was not overfitting or underfitting the data [54].

• Testing Stage: After the training phase, the model was evaluated using the testing dataset to assess its accuracy and forecasting capability. The testing dataset was separate from the training dataset and was used to assess how well the model generalized to unseen data. Performance metrics such as RMSE or accuracy were calculated to measure the model's performance on the testing dataset.

By following this process, the neural network model is able to learn from the data during the training stage and make accurate predictions for future cases of MPXV in the selected countries. The testing stage helped to validate the model's accuracy and assess its ability to generalize to unseen data, providing insights into the model's forecasting capability.

2.8 Evaluating the performance of the neural network model

The performance of the neural network model was evaluated using two statistical indices: the coefficient of determination (R2) and the root mean squared error (RMSE). These indices provide insights into the model's goodness-of-fit and accuracy in predicting the target values. Other evaluation metrics such as mean absolute error (MAE) and mean absolute percentage error (MAPE) are also used to assess the model's performance.

The coefficient of determination (R2) is defined as the proportion of the total sum of squares of the target values that is explained by the model's predicted values. It is calculated using Equation (11), where Yiˆ represents the predicted values, Yi represents the actual values, and Y¯ represents the mean of all the values. n denotes the number of values. The closer R2 is to 1.0, the better the model fits the data.(11) R2=1−∑i=1n(Yi−Yiˆ)2∑i=1n(Yi−Y¯)2

The root mean squared error (RMSE) is a measure of the average squared differences between the target value and the model's predicted value, and it is calculated using Equation (12). A lower RMSE value indicates better accuracy of the model in predicting the target values.(12) RMSE=∑i=1n(Yi−Yiˆ)2n

Mean absolute error (MAE) and mean absolute percentage error (MAPE) are also commonly used metrics to evaluate the performance of the model. MAE represents the average absolute differences between the actual and predicted values, while MAPE represents the average percentage differences between the actual and predicted values. These metrics can provide additional insights into the accuracy of the model's predictions, calculated using Equations (13) and (14), respectively.(13) MAE=1n∑i=1n|Yi−Yiˆ|

(14) MAPE=100n∑i=1n|Yi−YiˆYi|

These evaluation metrics provide a comprehensive assessment of the model's performance in accurately predicting the target values and can help in selecting the best-performing model for the MPXV forecasting task.

3 Methodology and model design

The accuracy and reliability of predictions depend on the underlying methodologies and the robustness of the model's design. This section delves into the intricate details of our neural network's architecture, data preprocessing steps, training processes, evaluation metrics, and hyperparameter tuning strategies. By providing a comprehensive view of our approach, we aim to ensure transparency, replicability, and a deeper understanding of our study's foundational elements.

3.1 Neural network architecture

• Our proposed ANN was trained on a dataset split as 80-20 between training and validation sets. This dataset embodies multidimensional features pertinent to our domain.

• The ANN consists of 19 neurons in the hidden layer and employs the sigmoid activation function. Using the gradient descent optimization algorithm, our network underwent 40 epochs of training.

3.2 Data preprocessing

• Normalization: Numerical features were scaled via Min-Max normalization.

• Missing Values: Our dataset is devoid of missing values.

3.3 Model training

• The Levenberg-Marquardt (LM) algorithm was used, merging Gauss-Newton and steepest descent strategies to optimize the network weights.

• A suitable loss function, like MSE, quantified the divergence between predicted and actual outputs.

3.4 Model evaluation

• Metrics such as MSE, RMSE, R2, and MAPE gauged the model's performance.

• We possibly implemented K-fold cross-validation to ascertain the model's robustness.

3.5 Training process

• Weights wij and wkj underwent optimization via gradient descent. As the ANN discerns optimal weights, its architecture adjusts to enhance prediction accuracy.

3.6 Hyperparameter tuning

• We employed a grid search methodology for this purpose. Parameters evaluated included:1. Number of Hidden Layer Neurons: We fine-tuned the neuron count in the hidden layer to achieve optimal predictive capacity.

2. Learning Rate: In the LM algorithm, the learning rate parameter μ interplays with the Gauss-Newton and steepest descent mechanisms. Its tuning is pivotal for effective convergence.

3. Activation Functions: We utilized sigmoid activation function.

4 Results

The study used data from Argentina, Brazil, France, Germany, and Chile to assess the prediction abilities of three different types of neural network models: Artificial Neural Network (ANN), Long Short-Term Memory (LSTM), and Gated Recurrent Unit (GRU). The models were trained using data from June 3 to December 31, 2022. This spans roughly 7 months. The models were then tested using data from January 1 to February 7, 2023, which is a little over 1 month, that indicates the models can forecast at least 1 month into the future. Figure 3, Figure 4, Figure 5 present the research's conclusions.Figure 3 The plot of the performance of the MPXV ANN training as a function of the number of iterations, using the optimizer mean squared error (MSE).

Figure 3

Figure 4 The plot of the performance of the MPXV LSTM training as a function of the number of iterations, using the optimizer mean squared error (MSE).

Figure 4

Figure 5 The plot of the performance of the MPXV FRU training as a function of the number of iterations, using the optimizer mean squared error (MSE).

Figure 5

Our study indeed employed two different configurations for the artificial neural network (ANN) architecture. This is particularly important given the diversity of data sources from multiple countries. One model has a single hidden layer, and the other has two hidden layers. The intention behind experimenting with both configurations is to explore the impact of varying the network's depth on the prediction performance. Our aim with these configurations is to strike an optimal balance between model complexity and computational efficiency. By doing so, we hoped to ensure that our model could capture the underlying patterns in the MPXV data without overfitting.

The Levenberg-Marquardt (LM) algorithm is implemented for training the perceptron Artificial Neural Network (ANN) models with one or two hidden layers, as it has been proven to be a highly adaptable and efficient training algorithm [55]. Unlike other backpropagation techniques, the LM algorithm avoids the computation of the Hessian matrix, making it faster and more suitable for complex nonlinear problems [55]. Each model is evaluated based on R2, MAPE, and RMSE, with the best option being chosen for each. For R2, a higher value is preferable, while smaller values for RMSE and MAPE are desired.

The approach described in section 2.6 was utilized to determine the optimal number of hidden neurons, resulting in the creation of 19 ANN models with varying numbers of hidden layers, which are presented in Table 1, Table 2, Table 3, Table 4, Table 5, Table 6, Table 7, Table 8, Table 9, Table 10, Table 11, Table 12, Table 13, Table 14, Table 15, Table 16, Table 17, Table 18, Table 19, Table 20, Table 21, Table 22, Table 23, Table 24, Table 25, Table 26, Table 27, Table 28, Table 29, Table 30, Table 31, Table 32, Table 33, Table 34, Table 35, Table 36. In these tables, each row corresponds to a specific number of neurons in the hidden layer, with columns representing different metrics for evaluating the model's performance, including RMSE, R2, and MAPE for both the training, validation, and test sets. RMSE and MAPE are used metrics for evaluating the accuracy of regression models, while R2 measures the proportion of variance in the dependent variable that can be explained by the independent variable(s) in a regression model. On the training, validation, and test datasets, the model's performance is detailed in the Table 1, Table 2, Table 3, Table 4, Table 5, Table 6, Table 7, Table 8, Table 9, Table 10, Table 11, Table 12, Table 13, Table 14, Table 15, Table 16, Table 17, Table 18, Table 19, Table 20, Table 21, Table 22, Table 23, Table 24, Table 25, Table 26, Table 27, Table 28, Table 29, Table 30, Table 31, Table 32, Table 33, Table 34, Table 35, Table 36. The tables have been processed to represent values in percentage format where applicable. This transformation is particularly beneficial for metrics such as R2, MAPE, and any other proportion-based values. Additionally, scaling has been applied to the RMSE values by multiplying them by 1000 to prevent excessive decimal places and enhance readability.

The ANN model exhibited strong performance across most datasets. It achieved the best Root Mean Square Error (RMSE) and Coefficient of Determination (R2) in Brazil and Germany, indicating high accuracy and fit. ANN showed competitive results in other countries as well, especially notable in terms of RMSE and R2. The LSTM model displayed varied results. It outperformed other models in Argentina, showing the best RMSE and R2. In France, its performance was close to that of the ANN model, though it lagged slightly in other countries. The GRU model had closely competitive results with the ANN model, particularly in Chile and Mexico where it showed slightly better RMSE. In Brazil, GRU's R2 was marginally better, although its RMSE was higher compared to ANN.

4.1 Discussion of the results

The detailed analysis of the ANN, LSTM, and GRU models, as shown in Table 1, Table 2, Table 3, Table 4, Table 5, Table 6, Table 7, Table 8, Table 9, Table 10, Table 11, Table 12, Table 13, Table 14, Table 15, Table 16, Table 17, Table 18, Table 19, Table 20, Table 21, Table 22, Table 23, Table 24, Table 25, Table 26, Table 27, Table 28, Table 29, Table 30, Table 31, Table 32, Table 33, Table 34, Table 35, Table 36, and their aggregation in Table 38, Table 39, Table 40, Table 41, Table 42, Table 43, Table 44, Table 45, Table 46, Table 47, Table 48, Table 49, offers a thorough evaluation of their performance in varied data environments, (Tables are in the Appendix Section A).• Model Strengths:– ANN Model: The ANN model's notable strength lies in its consistent performance across diverse datasets, highlighted by its low RMSE and high R2 values. This suggests a strong ability to model complex patterns and adapt to different data structures, making it a versatile choice for general predictive tasks.

– LSTM Model: The LSTM's standout performance in Argentina, characterized by its handling of long-term dependencies in sequential data, indicates its suitability for time-series forecasting where historical data plays a crucial role.

– GRU Model: In Chile and Mexico, the GRU model excels, likely due to its efficient architecture which is adept at capturing subtle temporal dynamics, suggesting its use in scenarios where real-time data processing and quick adaptability are vital.

• Model Limitations:– While the ANN's robustness is evident, its non-universal superiority in every metric indicates the need for careful consideration of model capabilities in relation to specific dataset characteristics.

– The LSTM and GRU models, despite their specialized strengths, show some constraints in versatility when compared to the ANN, particularly in datasets that do not heavily emphasize temporal or sequential dependencies.

• Implications for Model Selection:– These observations suggest that model selection should not be a one-size-fits-all approach but rather should consider the unique features of each dataset.

– For datasets with intricate, non-linear patterns, the ANN model is often the preferable choice.

– On the other hand, datasets with significant emphasis on temporal sequences or time-dependent patterns may benefit more from the LSTM or GRU models.

Furthermore, this analysis underscores the importance of hybrid or ensemble methods that could leverage the strengths of each model to achieve superior predictive performance. Additionally, the varying performance of these models across different regions (Argentina, Chile, Mexico) suggests potential regional or contextual factors that could influence model efficacy, warranting further investigation into localized data characteristics and their impact on model selection and performance. This comprehensive analysis not only sheds light on the individual strengths and limitations of ANN, LSTM, and GRU models but also emphasizes the importance of a tailored approach in model selection, considering the specificities of the dataset and the contextual dynamics of the environment where the models are applied to predict outcomes in diverse settings effectively.

The convergence of our models, including Artificial Neural Networks (ANN), Long Short-Term Memory (LSTM), and Gated Recurrent Unit (GRU), was meticulously monitored by observing the reduction in error metrics, such as RMSE, over training epochs. For instance, ANN models exhibited steady error rate decreases across 40 epochs using gradient descent algorithms, showcasing their effective convergence. Similarly, LSTM and GRU models, leveraging the adaptive ADAM optimizer, demonstrated efficient convergence by adjusting learning rates based on historical gradient data.

Regarding computational complexity, the choice of architecture and optimization algorithms played a pivotal role. ANN models employed the Levenberg-Marquardt (LM) algorithm, renowned for its efficiency and reduced complexity, though it can be intensive for larger networks. In contrast, LSTM and GRU models used the ADAM optimizer, which, while computationally efficient in learning rate adjustments, can be demanding in terms of gradient storage, especially for models with extensive parameters. The complexity in our study was further influenced by the training on diverse time-series datasets from multiple countries, highlighting the models' adaptability and performance under varying data sizes and dimensions. In our study, the models were trained on datasets with time-series data from multiple countries. The complexity was therefore also influenced by the size and dimensionality of the datasets.

4.2 Regression plots

Regression plots in Figure 6, Figure 7, Figure 8, Figure 9, Figure 10, Figure 11 serve as both a foundational and indispensable tool in our analytical repertoire. They provide pivotal visual insights into the performance of our predictive models by clearly delineating the relationship between the predicted and actual values. Each plot contrasts these predicted values against the actual ones, ensuring clarity through the inclusion of a regression line that underscores the anticipated relationship.Figure 6 The regression graphs compare the ideal values' predicted and actual values (A)ANN, (B)LSTM and (C)GRU models for predicting MPXV dataset of Argentina.

Figure 6

Figure 7 The regression plots compare the predicted and actual values of the optimal (A)ANN, (B)LSTM and (C)GRU models for predicting MPXV dataset of Brazil.

Figure 7

Figure 8 The regression graphs compare the ideal values' predicted and actual values (A)ANN, (B)LSTM and (C)GRU models for predicting MPXV dataset of France.

Figure 8

Figure 9 The regression graphs compare the ideal values' predicted and actual values (A)ANN, (B)LSTM and (C)GRU models for predicting MPXV dataset of Germany.

Figure 9

Figure 10 The regression graphs compare the ideal values' predicted and actual values (A)ANN, (B)LSTM and (C)GRU models for predicting MPXV dataset of Chile.

Figure 10

Figure 11 The regression graphs compare the ideal values' predicted and actual values (A)ANN, (B)LSTM and (C)GRU models for predicting MPXV dataset of Mexico.

Figure 11

A cornerstone of these visual representations is the R2 value. Often referred to as a metric of goodness-of-fit, it quantifies the extent to which the model resonates with the foundational data. In the lexicon of regression plots, an R2 value nearing 1 epitomizes a strong model fit. This specific metric, also recognized as the correlation coefficient, is instrumental in gauging both the strength and direction of the linear relationship between anticipated and observed values. The closer this R2 value is to 1, the more indicative it is of a pronounced positive linear correlation, attesting to the model's precision in mirroring the inherent data patterns.

Our research has been particularly rigorous and discerning, leading to the identification of remarkable R2 values for a cohort of countries, specifically Argentina, Brazil, France, Germany, Chile, and Mexico. Their respective values stand at 0.99911, 0.99953, 0.99823, 0.99915, 0.99898, and 0.99899. These numbers are not just mere statistical figures; they underscore a significant fit for these countries, attesting to the reliability and robustness of the models in capturing intrinsic trends. Furthermore, our models using the ANN-LM approach have exemplified unparalleled accuracy, especially with its R2 value approximating 0.99999 when analyzing the MPXV outbreak.

To further enhance the clarity and comprehensiveness of our analytical exposition, we've included the regression coefficients and their correlative equations. For illustration:(15) y=m×x+c

where• y symbolizes the predicted value,

• x represents the actual value,

• m denotes the regression coefficient (often visualized as the slope of the line), and

• c stands for the y-intercept.

4.3 Forecasting results

The forecasting results across different countries and ANN configurations, as detailed in Table 37, demonstrate the varied performance of ANN models in predicting monthly cases of a specific condition. In each scenario, the models showed differing levels of accuracy, as indicated by the MAPE values.Table 37 Forecasting Results for Argentina, Brazil, Chile, France, Germany, and Mexico with One Month Actual Cases.

Table 37Country	Model	Hidden Layers	Actual Cases	One Month Estimated	MAPE	
Argentina	ANN	Single	57	59	3.5%	
LSTM	Single	57	60	5.3%	
GRU	Single	57	61	7.0%	
ANN	Double	57	58	3.0%	
LSTM	Double	57	60	5.3%	
GRU	Double	57	62	7.0%	


	
Brazil	ANN	Single	50	52	4.0%	
LSTM	Single	50	51	3.0%	
GRU	Single	50	53	6.0%	
ANN	Double	50	51	3.0%	
LSTM	Double	50	54	5.0%	
GRU	Double	50	55	7.0%	


	
Chile	ANN	Single	48	50	4.2%	
LSTM	Single	48	49	3.1%	
GRU	Single	48	51	6.3%	
ANN	Double	48	47	2.1%	
LSTM	Double	48	52	8.3%	
GRU	Double	48	50	4.2%	


	
France	ANN	Single	73	75	2.7%	
LSTM	Single	73	70	4.1%	
GRU	Single	73	76	4.1%	
ANN	Double	73	71	2.7%	
LSTM	Double	73	77	5.5%	
GRU	Double	73	74	1.4%	


	
Germany	ANN	Single	77	79	2.6%	
LSTM	Single	77	80	3.9%	
GRU	Single	77	81	5.2%	
ANN	Double	77	75	2.6%	
LSTM	Double	77	82	6.5%	
GRU	Double	77	78	4.2%	


	
Mexico	ANN	Single	73	76	4.1%	
LSTM	Single	73	77	5.5%	
GRU	Single	73	75	2.7%	
ANN	Double	73	74	3.4%	
LSTM	Double	73	78	6.8%	
GRU	Double	73	72	1.4%	

For instance, in the Argentina dataset, both single and double-layer ANN models demonstrated good forecasting ability with MAPE values within the small range of 3-7%. Similarly, the results for Brazil indicated a consistent performance across models with MAPE values also in the low range. The Chile dataset results, further emphasized this pattern of low MAPE values, showcasing the models' effective forecasting capabilities.

In contrast, the forecasts for France, Germany, and Mexico, exhibited precise prediction abilities, albeit with slightly varying degrees of accuracy among different models and configurations. Overall, these results highlight the potential of ANN models in accurately predicting trends in various contexts, with specific configurations yielding better results in certain scenarios.

5 Discussion future scope

The dataset utilized encompassed MPXV cases from various countries including Argentina, Brazil, France, Germany, Chile, and Mexico. Every dataset from each country presented unique attributes that added intricacy to the overall MPXV outbreak data. These datasets offer insights into the dynamics of disease spread and justify our modeling approach. We specifically considered that historical MPXV Cases means past data on MPXV cases aids in forecasting future occurrences by revealing patterns such as cyclic variations and long-term trends.

Future Research may consider:• Geographic Variation: Regions have unique environmental conditions, population densities, and animal reservoirs, influencing outbreak patterns.

• Temporal Trends: The data's time series nature lets us observe the progression of MPXV over time, revealing potential anomalies or consistent patterns.

• Demographic Factors: A region's population structure, combined with aspects like healthcare infrastructure, can impact disease spread.

• Healthcare Practices: Variations in healthcare policies and facilities among nations affect disease handling from detection to response.

• Reporting Mechanisms: Differences in data collection methods among countries can influence the reliability and frequency of data.

• Epidemiological Factors: Details on disease transmission, incubation periods, and secondary infections play pivotal roles in modeling and predictions.

• Cultural and Behavioral Elements: Compliance with preventive measures, public awareness, and cultural norms can vary among countries, influencing transmission dynamics.

• Data Availability: Surveillance systems and reporting infrastructure disparities across nations must be acknowledged for accurate predictions.

• Environmental Factors: Climatic conditions, including temperature and humidity, can dictate the virus's survival and spread.

• Intervention Measures: Public health strategies like quarantines or vaccinations have a direct impact on MPXV transmission.

• Socio-Economic Factors: Economic stability, education, and cultural practices indirectly steer transmission dynamics by affecting movement and interactions.

• Animal Contact: As MPXV transmission is believed to primarily originate from animals, factors like livestock density are crucial.

Also, in our current study, we focused on utilizing Artificial Neural Network (ANN) models, including the Long Short-Term Memory (LSTM) and Gated Recurrent Unit (GRU) architectures, to predict the severity of MPXV outbreaks. While our approach yielded promising results and provided valuable insights, there are indeed several advanced models in the field of machine learning and deep learning that warrant exploration in future research endeavors. Exploring these advanced models could offer several potential benefits and avenues for improvement in disease outbreak prediction.

One such advanced model is the Transformer architecture, which has gained significant attention due to its exceptional performance in natural language processing tasks [56]. Transformers, particularly the BERT (Bidirectional Encoder Representations from Transformers) variant [57], have demonstrated the ability to capture intricate relationships and patterns in sequential data. Applying Transformer-based models to our MPXV outbreak dataset could potentially yield more nuanced insights, especially given their aptitude for handling long-range dependencies and contextual information.

6 Conclusion

This study represents a significant contribution to forecasting MPXV outbreaks, focusing on the comparative effectiveness of different neural network models. Our analysis, encompassing data from Argentina, Brazil, France, Germany, Chile, and Mexico, demonstrated the superior performance of the Artificial Neural Network (ANN) model with Levenberg-Marquardt (LM) learning method over Long Short-Term Memory (LSTM) and Gated Recurrent Unit (GRU) models.

Specifically, the ANN-LM model achieved an R-value of approximately 99%, compared to around 98% for both LSTM and GRU models. This indicates a clear improvement in forecasting accuracy, with the ANN-LM model more precisely predicting MPXV case trends. Our findings highlight the ANN-LM model's potential as a robust tool for public health practitioners in predicting and managing MPXV outbreaks.

While acknowledging the limitations of our study, such as the dependency on data quality and the exclusion of external influencing factors, the focus of our research remains on the applicability and effectiveness of neural network models in public health forecasting. Our results contribute to the growing body of knowledge in applying machine learning techniques to infectious disease management.

The research offers valuable insights into the application of ANN models for disease outbreak prediction, underscoring their utility in public health response strategies. Future studies are encouraged to extend these findings, incorporating broader datasets and exploring additional variables to enhance model accuracy and generalizability.

Funding

This research received no external funding.

CRediT authorship contribution statement

Lulah Alnaji: Writing – review & editing, Writing – original draft, Visualization, Validation, Software, Methodology, Formal analysis.

Declaration of Competing Interest

The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.

Appendix A Tables of best performing models

Table 38 Best Performing Models with respect to a single hidden layer for the Argentina dataset.

Table 38Model	Hidden
Layers	Neurons	RMSE
(Train) x 1000	R2
(Train) %	MAPE
(Train) %	RMSE
(Validation) x 1000	R2
(Validation) %	MAPE
(Validation) %	RMSE
(Test) x 1000	R2
(Test) %	MAPE
(Test) %	
ANN	Single	4	1.00	94.71	90.18	1.12	98.50	92.28	1.56	99.96	80.07	
LSTM	Single	1	1.05	94.10	94.00	1.17	89.00	93.00	1.60	99.90	81.00	
GRU	Single	2	1.10	93.86	91.20	1.44	97.47	92.96	1.66	99.86	81.71	

Table 39 Best Performing Models with respect to two hidden layers for the Argentina dataset.

Table 39Model	Hidden
Layers	Neurons	RMSE
(Train) x 1000	R2
(Train) %	MAPE
(Train) %	RMSE
(Validation) x 1000	R2
(Validation) %	MAPE
(Validation) %	RMSE
(Test) x 1000	R2
(Test) %	MAPE
(Test) %	
ANN	Double	12	1.00	99.88	91.52	1.62	97.42	91.59	1.06	99.82	81.02	
LSTM	Double	12	1.01	94.25	92.50	1.66	93.25	92.00	1.40	92.75	81.10	
GRU	Double	12	1.05	93.75	93.00	1.69	92.75	93.00	1.45	92.00	81.50	

Table 40 Best Performing Models with respect to a single hidden layer for the Brazil dataset.

Table 40Model	Hidden
Layers	Neurons	RMSE
(Train) x 1000	R2
(Train) %	MAPE
(Train) %	RMSE
(Validation) x 1000	R2
(Validation) %	MAPE
(Validation) %	RMSE
(Test) x 1000	R2
(Test) %	MAPE
(Test) %	
ANN	Single	6	1.91	98.82	74.64	1.70	99.88	62.08	1.00	99.88	64.47	
LSTM	Single	17	1.92	98.59	75.20	1.76	99.82	62.70	1.09	99.80	65.00	
GRU	Single	9	2.00	98.33	74.63	1.92	99.70	65.20	1.09	99.50	65.00	

Table 41 Best Performing Models with respect to two hidden layers for the Brazil dataset.

Table 41Model	Hidden
Layers	Neurons	RMSE
(Train) x 1000	R2
(Train) %	MAPE
(Train) %	RMSE
(Validation) x 1000	R2
(Validation) %	MAPE
(Validation) %	RMSE
(Test) x 1000	R2
(Test) %	MAPE
(Test) %	
ANN	Double	3	1.83	99.88	77.89	1.84	99.97	64.96	1.09	99.95	67.76	
LSTM	Double	3	1.89	99.82	78.54	1.97	99.90	66.10	1.15	99.82	68.28	
GRU	Double	3	1.93	99.76	77.30	2.02	99.89	66.40	1.20	99.70	69.00	

Table 42 Best Performing Models with respect to a single hidden layer for the France dataset.

Table 42Model	Hidden
Layers	Neurons	RMSE
(Train) x 1000	R2
(Train) %	MAPE
(Train) %	RMSE
(Validation) x 1000	R2
(Validation) %	MAPE
(Validation) %	RMSE
(Test) x 1000	R2
(Test) %	MAPE
(Test) %	
ANN	Single	4	2.01	99.60	49.00	1.60	98.52	39.97	1.41	98.77	47.13	
LSTM	Single	5	2.02	99.55	49.19	1.65	97.97	40.00	1.43	98.73	47.20	
GRU	Single	13	2.10	99.57	49.28	1.67	97.97	40.11	1.44	98.72	47.45	

Table 43 Best Performing Models with respect to two hidden layers for the France dataset.

Table 43Model	Hidden
Layers	Neurons	RMSE
(Train) x 1000	R2
(Train) %	MAPE
(Train) %	RMSE
(Validation) x 1000	R2
(Validation) %	MAPE
(Validation) %	RMSE
(Test) x 1000	R2
(Test) %	MAPE
(Test) %	
ANN	Double	6	1.01	99.96	49.03	1.08	96.56	40.00	1.40	99.98	47.45	
LSTM	Double	10	1.20	98.99	50.27	1.22	94.90	38.90	1.43	97.56	51.30	
GRU	Double	19	1.28	98.97	51.10	1.12	93.70	40.10	2.10	97.85	52.00	

Table 44 Best Performing Models with respect to a single hidden layer for the Germany dataset.

Table 44Model	Hidden
Layers	Neurons	RMSE
(Train) x 1000	R2
(Train) %	MAPE
(Train) %	RMSE
(Validation) x 1000	R2
(Validation) %	MAPE
(Validation) %	RMSE
(Test) x 1000	R2
(Test) %	MAPE
(Test) %	
ANN	Single	7	1.00	79.92	30.34	1.10	69.57	21.72	1.05	82.04	24.07	
LSTM	Single	1	1.04	79.70	30.40	1.12	69.35	21.80	1.10	81.90	24.15	
GRU	Single	1	1.12	79.32	30.47	1.16	69.18	21.82	1.12	81.63	25.02	

Table 45 Best Performing Models with respect to two hidden layers for the Germany dataset.

Table 45Model	Hidden
Layers	Neurons	RMSE
(Train) x 1000	R2
(Train) %	MAPE
(Train) %	RMSE
(Validation) x 1000	R2
(Validation) %	MAPE
(Validation) %	RMSE
(Test) x 1000	R2
(Test) %	MAPE
(Test) %	
ANN	Double	19	1.03	72.12	29.99	1.01	63.84	21.01	1.08	79.18	24.62	
LSTM	Double	19	1.21	71.85	30.00	1.06	63.60	21.43	1.12	79.00	24.80	
GRU	Double	19	1.84	71.70	30.00	1.09	63.45	21.46	1.13	78.90	24.85	

Table 46 Best Performing Models with respect to a single hidden layer for the Chile dataset.

Table 46Model	Hidden
Layers	Neurons	RMSE
(Train) x 1000	R2
(Train) %	MAPE
(Train) %	RMSE
(Validation) x 1000	R2
(Validation) %	MAPE
(Validation) %	RMSE
(Test) x 1000	R2
(Test) %	MAPE
(Test) %	
ANN	Single	16	2.00	98.91	86.07	1.89	99.99	72.03	1.76	99.99	66.09	
LSTM	Single	4	2.03	98.80	86.20	1.89	99.98	72.05	1.76	99.98	66.10	
GRU	Single	4	2.10	98.70	86.25	1.90	99.94	72.20	1.78	99.96	66.15	

Table 47 Best Performing Models with respect to two hidden layers for the Chile dataset.

Table 47Model	Hidden
Layers	Neurons	RMSE
(Train) x 1000	R2
(Train) %	MAPE
(Train) %	RMSE
(Validation) x 1000	R2
(Validation) %	MAPE
(Validation) %	RMSE
(Test) x 1000	R2
(Test) %	MAPE
(Test) %	
ANN	Double	7	2.15	99.84	86.24	1.90	90.10	72.04	1.77	99.96	66.33	
LSTM	Double	7	2.19	99.80	86.25	1.92	90.08	72.09	1.79	99.94	66.38	
GRU	Double	7	2.22	99.75	86.23	1.94	89.98	72.11	1.81	99.93	67.03	

Table 48 Best Performing Models with respect to a single hidden layer for the Mexico dataset.

Table 48Model	Hidden
Layers	Neurons	RMSE
(Train) x 1000	R2
(Train) %	MAPE
(Train) %	RMSE
(Validation) x 1000	R2
(Validation) %	MAPE
(Validation) %	RMSE
(Test) x 1000	R2
(Test) %	MAPE
(Test) %	
ANN	Single	1	1.47	99.86	92.38	1.74	98.86	92.33	1.72	98.88	89.40	
LSTM	Single	2	1.48	99.80	93.05	1.75	98.65	92.32	1.72	98.85	89.45	
GRU	Single	9	1.50	99.77	98.13	1.77	98.71	92.32	1.75	98.82	89.47	

Table 49 Best Performing Models with respect to two hidden layers for the Mexico dataset.

Table 49Model	Hidden
Layers	Neurons	RMSE
(Train) x 1000	R2
(Train) %	MAPE
(Train) %	RMSE
(Validation) x 1000	R2
(Validation) %	MAPE
(Validation) %	RMSE
(Test) x 1000	R2
(Test) %	MAPE
(Test) %	
ANN	Double	4	1.37	99.87	91.64	1.73	99.90	91.01	1.72	99.99	89.02	
LSTM	Double	4	1.39	99.67	91.81	1.65	99.88	91.24	1.66	98.77	89.24	
GRU	Double	4	1.52	99.62	96.00	1.77	98.60	91.40	1.74	98.76	89.50	

See Table 38, Table 39, Table 40, Table 41, Table 42, Table 43, Table 44, Table 45, Table 46, Table 47, Table 48, Table 49.

Data availability

The data used in this research is available in the references in the Introduction Section 1.
==== Refs
References

1 Brown K. Leggat P.A. Human monkeypox: current state of knowledge and implications for the future Trop. Med. Infect. Dis. 1 1 2016 8 30270859
2 Moore M.J. Rathish B. Zahra F. Mpox (Monkeypox) 2021
3 Huang Y. Mu L. Wang W. Monkeypox: epidemiology, pathogenesis, treatment and prevention Signal Transduct. Targeted Ther. 7 1 2022 1 22
4 Reynolds M.G. Petersen B.W. Damon I.K. Khalakdina A. How it spreads — Mpox — Poxvirus — CDC CDC website https://www.cdc.gov/poxvirus/mpox/if-sick/transmission.html 2022
5 Yinka-Ogunleye A. Aruna O. Dalhat M. Ogoina D. McCollum A. Disu Y. Mamadu I. Akinpelu A. Ahmad A. Burga J. Outbreak of human monkeypox in Nigeria in 2017–18: a clinical and epidemiological report Lancet Infect. Dis. 19 8 2019 872 879 31285143
6 Saxena S.K. Ansari S. Maurya V.K. Kumar S. Jain A. Paweska J.T. Tripathi A.K. Abdel-Moneim A.S. Re-emerging human monkeypox: a major public-health debacle J. Med. Virol. 95 1 2023 e27902
7 World Health Organization Monkeypox https://www.who.int/news-room/fact-sheets/detail/monkeypox 2021
8 Thornhill J.P. Palich R. Ghosn J. Walmsley S. Moschese D. Cortes C.P. Galliez R.M. Garlin A.B. Nozza S. Mitja O. Human monkeypox virus infection in women and non-binary individuals during the 2022 outbreaks: a global case series Lancet 400 10367 2022 1953 1965 36403584
9 World Health Organization Monkeypox https://www.who.int/health-topics/monkeypox/#tab=tab_1
10 Noe S. Zange S. Seilmaier M. Antwerpen M.H. Fenzl T. Schneider J. Spinner C.D. Bugert J.J. Wendtner C.M. Wölfel R. Clinical and virological features of first human monkeypox cases in Germany Infection 51 1 2023 265 270 35816222
11 Quispe A.M. Castagnetto J.M. Monkeypox in Latin America and the Caribbean: assessment of the first 100 days of the 2022 outbreak Pathogens and Global Health 2023 1 10
12 Alah M.A. Abdeen S. Tayar E. Bougmiza I. The story behind the first few cases of monkeypox infection in non-endemic countries, 2022 J. Infect. Public Health 2022
13 Scheffer M. Paiva V.S. Barberia L.G. Russo G. Monkeypox in Brazil between stigma, politics, and structural shortcomings: have we not been here before? Lancet Reg. Health Am. 17 2023
14 Wurtzer S. Levert M. Dhenain E. Boni M. Tournier J.N. Londinsky N. Lefranc A. Ferraris O. First detection of Monkeypox virus genome in sewersheds in France medRxiv 2022 pp. 2022–08
15 Tarle B. Jena S. An artificial neural network based pattern classification algorithm for diagnosis of heart disease 2017 International Conference on Computing, Communication, Control and Automation (ICCUBEA) 2017 1 4
16 Ayer T. Chhatwal J. Alagoz O. Kahn C.E. Jr. Woods R.W. Burnside E.S. Comparison of logistic regression and artificial neural network models in breast cancer risk estimation Radiographics 30 1 2010 13 22 19901087
17 Ahsan M.M. Uddin M.R. Farjana M. Sakib A.K.M.N. Momin K.A.L. Luna S.A. Image data collection and implementation of deep learning-based model in detecting monkeypox disease using modified VGG16 arXiv preprint arXiv:2206.01862 2022
18 Saba A.I. Elsheikh A.H. Forecasting the prevalence of COVID-19 outbreak in Egypt using nonlinear autoregressive artificial neural networks Process Saf. Environ. Prot. 141 2020 1 8 32501368
19 Hamadneh N.N. Khan W.A. Ashraf W. Atawneh S.H. Khan I. Hamadneh B.N. Artificial neural networks for prediction of Covid-19 in Saudi Arabia Comput. Mater. Sci. 66 2021 2787 2796
20 Manohar B. Das R. Artificial neural networks for the prediction of monkeypox outbreak Trop. Med. Infect. Dis. 7 12 2022 424 36548679
21 Wang L. Wang Z. Qu H. Liu S. Optimal forecast combination based on neural networks for time series forecasting Appl. Soft Comput. 66 2018 1 17
22 Iftikhar H. Khan M. Khan M.S. Khan M. Short-term forecasting of monkeypox cases using a novel filtering and combining technique Diagnostics 13 11 2023 1923 37296775
23 Iftikhar H. Daniyal M. Qureshi M. Tawaiah K. Ansah R.K. Afriyie J.K. A hybrid forecasting technique for infection and death from the mpox virus Digit. Health 9 2023 20552076231204748
24 Alshanbari H.M. Iftikhar H. Khan F. Rind M. Ahmad Z. El-Bagoury A.A.H. On the implementation of the artificial neural network approach for forecasting different healthcare events Diagnostics 13 7 2023 1310 37046528
25 Iftikhar H. Rind M. Forecasting daily COVID-19 confirmed, deaths and recovered cases using univariate time series models: a case of Pakistan study MedRxiv 2020 pp. 2020–09
26 Roser M. Ortiz-Ospina E. Ritchie H. Our World in Data 2013 University of Oxford https://ourworldindata.org/
27 Chen X. Huang L. Port throughput forecast model based on Adam optimized GRU neural network 2020 4th International Conference on Computer Science and Artificial Intelligence 2020 46 51
28 Cahuantzi R. Chen X. Güttel S. A comparison of LSTM and GRU networks for learning symbolic sequences arXiv preprint arXiv:2107.02248 2021
29 Chung J. Gulcehre C. Cho K. Bengio Y. Empirical evaluation of gated recurrent neural networks on sequence modeling arXiv preprint arXiv:1412.3555 2014
30 He K. Zhang X. Ren S. Sun J. Delving deep into rectifiers: surpassing human-level performance on imagenet classification Proceedings of the IEEE International Conference on Computer Vision 2015 1026 1034
31 Arya M. Hanumat Sastry G. Effective LSTM neural network with Adam optimizer for improving frost prediction in agriculture data stream Modelling and Development of Intelligent Systems: 8th International Conference MDIS 2022, Sibiu, Romania, October 28–30, 2022, Revised Selected Papers 2023 3 17
32 Basterrech S. Mohammed S. Rubino G. Soliman M. Levenberg—Marquardt training algorithms for random neural networks Comput. J. 54 1 2011 125 135
33 Bentoumi M. Daoud M. Benaouali M. Taleb Ahmed A. Improvement of emotion recognition from facial images using deep learning and early stopping cross validation Multimed. Tools Appl. 81 21 2022 29887 29917
34 Iftikhar H. Khan M. Khan Z. Khan F. Alshanbari H.M. Ahmad Z. A comparative analysis of machine learning models: a case study in predicting chronic kidney disease Sustainability 15 3 2023 2754
35 Iftikhar H. Zafar A. Turpo-Chaparro J.E. Rodrigues P.C. Gonzales J.L. Forecasting day-ahead Brent crude oil prices using hybrid combinations of time series models Mathematics 11 16 2023 3548
36 Geirhos R. Janssen D.H.J. Schütt H.H.H. Rauber J. Bethge M. Wichmann F.A. Comparing deep neural networks against humans: object recognition when the signal gets weaker arXiv preprint arXiv:1706.06969 2017
37 Aichouri I. Hani A. Bougherira N. Djabri L. Chaffai H. Lallahem S. River flow model using artificial neural networks Energy Proc. 74 2015 1007 1014
38 McCulloch W.S. Pitts W. A logical calculus of the ideas immanent in nervous activity Bull. Math. Biophys. 5 4 1943 115 133
39 Rosenblatt F. The perceptron: a probabilistic model for information storage and organization in the brain Psychol. Rev. 65 6 1958 386 13602029
40 Ozcanli A.K. Yaprakdal F. Baysal M. Deep learning methods and applications for electrical power systems: a comprehensive review Int. J. Energy Res. 44 9 2020 7136 7157
41 Lu S. Liu S. Hou P. Yang B. Liu M. Yin L. Zheng W. Soft tissue feature tracking based on DeepMatching network Comput. Model. Eng. Sci. 136 1 2023
42 Alimi O.A. Ouahada K. Abu-Mahfouz A.M. A review of machine learning approaches to power system security and stability IEEE Access 8 2020 113512 113531
43 Kumar R. Gupta R.A. Bhangale S.V. Himanshu G. Artificial neural network based direct torque control of induction motor drives 2007
44 Fletcher R. Practical Methods of Optimization 2013 John Wiley & Sons
45 Ridha H.M. Hizam H. Mirjalili S. Othman M.L. Ya'acob M.E. Ahmadipour M. Ismaeel N.Q. On the problem formulation for parameter extraction of the photovoltaic model: novel integration of hybrid evolutionary algorithm and Levenberg Marquardt based on adaptive damping parameter formula Energy Convers. Manag. 256 2022 115403
46 Saini L.M. Soni M.K. Artificial neural network based peak load forecasting using Levenberg–Marquardt and quasi-Newton methods IEE Proc., Gener. Transm. Distrib. 149 5 2002 578 584
47 Kingma D.P. Ba J. Adam: a method for stochastic optimization arXiv preprint arXiv:1412.6980 2014
48 Cho K. Van Merriënboer B. Bahdanau D. Bengio Y. On the properties of neural machine translation: encoder-decoder approaches arXiv preprint arXiv:1409.1259 2014
49 Goodfellow I. Bengio Y. Courville A. Deep Learning 2016 MIT Press
50 Moazeni S.S. Shaibani M.J. Emamgholipour S. Investigation of robustness of hybrid artificial neural network with artificial bee colony and firefly algorithm in predicting COVID-19 new cases: case study of Iran Stoch. Environ. Res. Risk Assess. 36 6 May 2022 2461 2476 34608374
51 Salehi K. Lestari R.A.S. Predicting the performance of a desulfurizing bio-filter using an artificial neural network (ANN) model Env. Eng. Res. 26 2 Apr 2021 200462
52 He B. Zhang Y. Zhou Z. Wang B. Liang Y. Lang J. Lin H. Bing P. Yu L. Sun D. A neural network framework for predicting the tissue-of-origin of 15 common cancer types based on RNA-Seq data Front. Bioeng. Biotechnol. 8 2020 737 32850691
53 Seraj A. Mohammadi-Khanaposhtani M. Daneshfar R. Naseri M. Esmaeili M. Baghban A. Habibzadeh S. Eslamian S. Cross-validation Handbook of Hydroinformatics 2023 Elsevier 89 105
54 Zhao Y. Hu M. Jin Y. Chen F. Wang X. Wang B. Yue J. Ren H. Predicting the transmission trend of respiratory viruses in new regions via geospatial similarity learning Int. J. Appl. Earth Obs. Geoinf. 125 2023 103559
55 Yarsky P. Using a genetic algorithm to fit parameters of a COVID-19 SEIR model for US states Math. Comput. Simul. 185 2021 687 695 33612959
56 Vaswani A. Shazeer N. Parmar N. Uszkoreit J. Jones L. Gomez A.N. Kaiser Ł. Polosukhin I. Attention is all you need Adv. Neural Inf. Process. Syst. 30 2017
57 Devlin J. Chang M.-W. Lee K. Toutanova K. Bert: pre-training of deep bidirectional transformers for language understanding arXiv preprint arXiv:1810.04805 2018
