
==== Front
Heliyon
Heliyon
Heliyon
2405-8440
Elsevier

S2405-8440(24)12872-1
10.1016/j.heliyon.2024.e36841
e36841
Research Article
Predicting compressive strength of hollow concrete prisms using machine learning techniques and explainable artificial intelligence (XAI)
Bin Inqiad Waleed WaleedBinInqiad@gmail.com
a
Dumitrascu Elena Valentina elena.dumitrascu97@stud.etti.upb.ro
b
Dobre Robert Alexandru robert.dobre@upb.ro
c⁎
Khan Naseer Muhammad nmkhan@mce.nust.edu.pk
de⁎⁎
Hammood Abbas Hussein bbas.h.hammood@almamonuc.edu.iq
f
Henedy Sadiq N. ci.sadiq@mpu.edu.iq
g
Khan Rana Muhammad Asad masadkhan87@gmail.com
h
a Military College of Engineering (MCE), National University of Science and Technology (NUST), Islamabad, 44000, Pakistan
b Computers Department, Faculty of Automatic Control and Computer Science, National University of Science and Technology Politehnica Bucharest, Bucharest, Romania
c Electronic Technology and Reliability Department, National University of Science and Technology Politehnica Bucharest, 061071, Bucharest, Romania
d Department of Sustainable Advanced Geomechanical Engineering, Military College of Engineering (MCE), National University of Science and Technology (NUST), Islamabad, 44000, Pakistan
e MEU Research Unit, Middle East University, Amman, 11831, Jordan
f Department of Computer Techniques Engineering/ Al-Ma'moon University College /Al Washash, Baghdad, Iraq
g Department of Civil Engineering, Mazaya University College, Nasiriya City, 64001, Iraq
h Pak-Austria Fachhochschule Institute of Applied Sciences and Technology, Pakistan
⁎ Corresponding author. robert.dobre@upb.ro
⁎⁎ Corresponding author. Department of Sustainable Advanced Geomechanical Engineering, Military College of Engineering (MCE), National University of Science and Technology (NUST), Islamabad, 44000, Pakistan. nmkhan@mce.nust.edu.pk
27 8 2024
15 9 2024
27 8 2024
10 17 e3684123 5 2024
21 8 2024
22 8 2024
© 2024 The Authors
2024
https://creativecommons.org/licenses/by-nc/4.0/ This is an open access article under the CC BY-NC license (http://creativecommons.org/licenses/by-nc/4.0/).
The design of masonry structures requires accurate estimation of compressive strength (CS) of hollow concrete masonry prisms. Generally, the CS of masonry prisms is determined by destructive laboratory testing which results in time and resource wastage. Thus, this study aims to provide machine learning-based predictive models for CS of hollow concrete masonry blocks using different algorithms including Multi Expression Programming (MEP), Random Forest Regression (RFR), and Extreme Gradient Boosting (XGB) etc. A dataset of 159 experimental results was collected from published literature for this purpose. The collected dataset consisted of four input parameters including strength of masonry units (fb), height-to-thickness ratio (h/t), strength of mortar (fm), and ratio of fm/fb and only one output parameter i.e., CS. Out of all the algorithms employed in current study, only MEP and GEP expressed their output in the form of an empirical equation. The accuracy of developed models was assessed using root mean squared error (RMSE), objective function (OF), and R2 etc. Among all algorithms assessed, XGB turned out to be the most accurate having R2 = 0.99 and least OF value of 0.0063 followed by AdaBoost, RFR, and other algorithms. The developed XGB model was also used to conduct different explainable artificial intelligence (XAI) analysis including sensitivity and shapley analysis and the results showed that strength of masonry unit (fb) is the most significant variable in predicting CS. Thus, the ML-based predictive models presented in this study can be utilized practically for determining CS of hollow concrete masonry prisms without requiring expensive and time-consuming laboratory testing.

Graphical abstract

Image 1

Keywords

Concrete masonry prisms
Machine learning
Multi expression programming
Compressive strength
Explainable artificial intelligence (XAI)
==== Body
pmc1 Introduction

Masonry, made from naturally occurring materials is one of the oldest building systems which has been used for more than 6000 years. It is the most widely used type of construction due to its low cost, easily available raw materials, and the aesthetics it provides in modern buildings. Also, low-rise masonry houses are known to have good seismic performance and impressive thermal insulation properties [1]. Masonry structures are made up of masonry units and mortar. The masonry units in turn are made up of different materials and maybe hollow or solid. Different types of masonry units which are commonly used in construction include but are not limited to burnt clay bricks, soft mud bricks, and pressed earth bricks etc. Generally, a weak interface joins two material phases in masonry structures and thus it is weak in tension. Therefore, masonry structures are designed to resist compressive forces only [2]. The prevalent design practices of masonry emphasizes that they should only be designed against compressive forces, thus the accurate estimation of masonry CS holds paramount importance [3]. The general configuration of hollow concrete masonry blocks and applied compressive force is given in Fig. 1.Fig. 1 Configuration of hollow concrete masonry blocks.

Fig. 1

The technical and economic aspects of masonry structures are greatly influenced by the CS of hollow concrete masonry [4]. It is an important factor which greatly affects the economy and safety of structures because many other mechanical properties are linked to CS [5]. Despite its significance, the determination of CS is particularly hard due to the complex behaviour of masonry components such as reinforced and unreinforced hollow interlocking blocks, the mortar used to join them and their interfaces etc [6]. There have been some theoretical investigations to study the strength of masonry prisms under axially compressive forces [[7], [8], [9], [10]]. Also in past, different analytical methods have been devised based on conditions of equilibrium and compatibility equations to estimate the CS. Eurocode 6 provides a mechanism to predict CS based on different input parameters such as compressive strength of mortar, blocks etc [11]. There have also been some experimental studies to determine CS [[12], [13], [14], [15]]. These studies mostly use compressive strength of mortar and hollow concrete blocks to estimate CS of hollow concrete masonry prisms but the dataset used to develop these models was very limited. Thus, the developed empirical models using this data have limited applicability and their reliability needs to be improved by incorporation of new data.

The strength of hollow concrete masonry block (fb) and mortar compressive strength (fm) are the most commonly used parameters in previous studies. However, the dimensions of prism and mortar also affects the CS of hollow blocks [16]. The effect of other varibles such as ratio of prism height and masonry width, volume ratio of bed joints to mortar, and volume fraction of masonry unit was studied as input parameters by the study conducted by Thaickavil et al. [16]. The results showed that the developed models with these new input parameters exhibited 88 % correlation between actual and predicted CS. However, this study was conducted using one particular masonry unit having specified dimensions and mortar joint thickness. Also it was highlighted that prediction of CS by virtue of empirical models with complex input variables is particularly challenging [17]. Thus, it is crucial to have some empirical prediction models for estimating CS of concrete masonry prisms to design resilient structures. In this regard, the use of cutting-edge machine learning (ML) techniques can be useful for developing empirical prediction models as discussed in coming sections. The subsequent sections of the paper will first discuss the previous relevant studies regarding prediction of different properties of hollow concrete blocks followed by the explanation of working mechanism of the employed ML models. After the explanation of ML techniques, section 3.2 will discuss the database formation used for development of the predictive models and the results of data analysis techniques which were performed on the collected dataset. In section 3.3, the criteria for assessment of the ML models will be discussed while the model development and results of algorithms will be discussed in sections 4.1 till 4.8. Finally, in section 4.9, the explanatory analyses technqiues employed in current study will be highlighted.

1.1 Overview of relevant literature

During the last few years, civil engineering industry is making a shift towards the sustainable materials and data-driven decision-making tools like every other industry. In the field of civil engineering, artificial Intelligence (AI) is mainly used to optimize material utilization, foster sustainable construction, and improve the overall efficiency and accuracy of construction processes by using various machine learning techniques. Machine learning (ML) is a subset of AI which refers to the process in which machines learn from vast amounts of data without any human intervention. ML techniques learn patterns from the data and use that to make informed decisions for future data. Deep learning is in turn a subset of ML which refers to the use of neural networks to make predictions on unseen data.

Recently, the prediction of various properties of new concrete composites using ML models has captured the interest of researchers. This is due to the simplicity, accuracy, cost, and time effectiveness of the ML techniques [18]. ML algorithms are basically computational methods designed to solve complex problems like humans. It has been reported that ML algorithms can discover the hidden patterns in the data and make predictions based on the learning from the data [19]. ML techniques have been successfully applied to solve problems related to finance, biology, engineering etc. In the domain of civil engineering, ML techniques have found their use in predicting various properties of different concrete and cement composites, concrete columns, soil classification and compaction factors etc [[20], [21], [22], [23], [24], [25], [26], [27], [28], [29], [30], [31], [32], [33], [34], [35], [36], [37], [38], [39], [40], [41], [42], [43], [44], [45], [46], [47]]. The increased applicability of ML techniques to predict different material properties is due to the fact that they can reliably handle multivariate datasets, efficiently detect patterns in the data, and use the learning from the data to make predictions about the future unseen data.

Although various ML techniques like Artificial Neural Networks (ANNs), Random Forest (RF), Gene Expression Programming (GEP) had widely been utilized to predict different concrete and soil properties, their application to predict CS of hollow concrete masonry is severly limited. A study conducted by Zhou et al. [48] predicted CS of concrete masonry prisms using ANN and adapted neuro-fuzzy inference systems (ANFIS). The study considered three input variables including height-to-thickess ratio and compressive strengths of mortar and concrete blocks. The results exhibited higher accuracy and minimal errors indicating the robustness of ML techniques to predict CS of hollow concrete masonry. Similarly, another study conducted by Zhang et al. [49] developed an automated model for predicting cracking patterns of masonry wallets loaded vertically. Also, Garzo'n-Roca et al. utilized fuzzy logic and ANN algorithms for prediction of compressive strength of brick masonry [[50], [51], [52]]. In the same way, Turco et al. [53] leveraged ANN for prediction of compressive and tensile strength of compressed earth blocks containing natural fibre-reinforcement. The natural fibres included banana fibres, oil palm fibres, sugarcane bagasse, and coconut coir etc. The authors trained two ANN-based predictive models for prediction of CS and tensile strength using datasets of 332 and 130 experimental results respectively. The results revealed that ANN can predict mechanical strength properties of compressed earth blocks with great accuracy. The correlation coefficient (R) was 0.97 for CS prediction and 0.91 for tensile strength prediction.

In another study conducted by Asteris et al. [1], the compressive strength of brick masonry was estimated using a back propagation neural network and the analysis showed that developed neural network-based model exhibited good correlation with the actual values. In another study conducted by Fakharian et al. [54], the authors utilized several ML algorithms to predict strength of hollow concrete masonry including ANN and ANFIS. The ANN model exhibited higher accuracy having correlation value (R) between experimental and predicted values as 0.956. The authors also performed sensitivity analysis on the developed ANN-based predictive mode which highlighted that compressive strength of hollow block is the most important parameter for predicting CS. Furthermore, a study by Lan et al. [55] utilized ANFIS and neural networks to predict CS of masonry concrete blocks having three input parameters including ratio of height to thickness, compressive strength of mortar, and masonry blocks. The study leveraged 72 experimental data points for this purpose and the predictive models built using ANFIS and ANN exhibited remarkable accuracy to predict CS of concrete masonry blocks.

Some previously conducted studies also used ML algorithms like namely fuzzy logic and fuzzy set for prediction of CS hollow masonry blocks using data from the experimental works [56,57]. The data used in these studies included both direct experimental data and data from non destructive testing like rebound hammer and ultrasonic pulse velocity tests. More importantly, these studies also conducted sensitivity analysis on the developed ML models which showed that the compression strength of masonry unit is the deciding factor in determining CS of hollow concrete masonry prisms. Also, in previous studies, different XAI approaches like class activation maps (CAMs) have been utilized for different purposes like recognizing defects in existing bridges for efficient risk management [58]. However, the use of more recently developed XAI approaches like shap and individual conditional expectation analysis for interpretation of results of CS prediction of hollow concrete blocks is comparatively limited.

2 Research significance

The CS is one of the most significant parameters in masonry construction and the performance of masonry blocks is largely dependent on its compressive strength. The general practice of determining CS of hollow masonry blocks is by means of destructive compression strength test in the laboratory, which is time consuming and also results in resource wastage. Also, it is very difficult to accurately determine CS due to the complex behaviour of masonry structures. To overcome these limitations, there have been several empirical methods developed for CS prediction, however it was indicated that these empirical models were built using very limited dataset and are difficult to be implemented widely for CS prediction with complex variables. Therefore, some ML-based predictive models have been developed for this purpose mainly using ANN, ANFIS, and fuzzy logic etc. However, there is a lack of work focusing on the utilization of more recently developed techniques like multi expression programming (MEP), extreme gradient boosting (XGB) and gene expression programming (GEP) etc. for CS estimation of hollow concrete blocks. Thus, this study aims to bridge this gap in the literature by providing predictive models based on XGB, MEP, random forest regression (RFR), GEP, AdaBoost and k-nearest neighbor (KNN). The motive behind using MEP and GEP along with other techniques stems from their grey-box nature as compared to black-box nature of XGB and other ML algorithms. Both MEP and GEP can give their output as an empirical equation which provides transparency to the model prediction process and allows the widespread use of developed empirical models [59]. This significance of this study also lies in the fact that it compares some of the most commonly utilized algorithms in the realm of civil engineering ranging from boosting techniques (XGB and AdaBoost), genetic programming techniques (MEP and GEP), with the most widely used tree-based algorithm (random forest). Moreover, this study also tends to utilize different XAI approaches including individual conditional expectation analysis and shapley analysis for getting useful insights into the developed models and understanding the logic behind the prediction process. The employment of these XAI approaches will provide transparency to the black-box ML models and provide information about the importance of different input parameters in predicting the output.

3 Research methodology

The overall methodology of the study is illustrated in Fig. 2. The most important step in the development of ML models is the collection of data. For this purpose, a thorough review of the previous studies was performed, and 159 data points were collected. After data collection, next step involved extracting features from the data which means specifying input and output variables. Thus, based on earlier literature recommendations [48,[60], [61], [62], [63], [64]], four input variables were chosen to predict the CS of concrete masonry prisms. The next step involved carrying out different statistical analyses on the collected dataset which involved investigating correlation and distribution of the data by using heatmap and scatter matrix. Then, the dataset was split into two subsets (training and testing) to be used for training of ML models and their validation. After that, the ML models were trained on the training data and error assessment was performed for both training and testing phases to examine their accuracy. After error assessment and identification of most accurate algorithm, several explanatory analyses were performed on the most accurate model to get further insights into the prediction process.Fig. 2 Research methodology followed in the study.

Fig. 2

3.1 Prediction models

3.1.1 Gene expression programming (GEP)

GEP is an AI technique which uses evolutionary algorithms to find solution to a problem. It is an extension of Genetic Programming (GP) and based on Charles Darwin's principles of natural selection. GEP finds solution to the problem using genes and performing evolutionary rules on these genes [65]. The genes are build using various mathematical functions called primitive functions. These functions range from simple arithmetic functions like addition, subtraction to more complex like square root, sine, cosine etc [66]. These genes which can be expressed as expression trees represent a small piece of code and are combined in various ways using crossover and mutation to create a full program which represents the solution to the given problem [67]. Crossover involves combining two expressions to develop a new one and mutation means changing an existing expression [68]. These processes in GEP are similar to the genetic mutation and crossover in humans. The pictorial representation of the processes of crossover and mutation is given in Fig. 3.Fig. 3 Schematic representation of expression tree and genetic processes.

Fig. 3

The prediction process starts with the formation of a random population of computer programs called chromosomes. These chromosomes are evaluated for their ability to solve the given problem using a pre-defined fitness function. The chromosomes with good performance based on the fitness function are selected for the next generation while the worst performing ones are deleted from the population [69]. This process repeats several times until a population of chromosomes with the highest accuracy is reached [70]. The flowchart of the overall process is given in Fig. 4.Fig. 4 Flowchart representation of MEP and GEP.

Fig. 4

3.1.2 Multi expression programming (MEP)

MEP is a recently developed sub technique of GP. It was developed by Oltean and Dumitrescu [71]. It is a method based on the use of genetic algorithm which creates a population of computer programs and refines the programs to accurately solve the problem. MEP differs from GEP in the way that it uses linear representation of chromosomes having the ability to encode multiple solutions in a single chromosome. The fitness of the individuals is measured, and the most accurate solution is selected as a representation of the chromosome. This technique offers the significant advantage of not rising the complexity of the resulting equation unlike GEP and other variants of GP who can only encode a single solution in one chromosome [72]. This ability is due to the fact that MEP doesn't require any prior assumptions, but it creates and evolves a population of mathematical expressions to solve the problem [73].

The MEP algorithm also starts by creating a random population of mathematical expressions. These expressions are then evaluated using a fitness function. The chromosomes having good accuracy are selected for creating the next generation of chromosomes using mutation and crossover [74]. This iterative process allows MEP to explore new solutions thus discovering potentially better and accurate solutions. This process of creating a new population of chromosomes from already existing chromosomes using evolutionary rules continues until a stopping point such as maximum number of generations is reached. The flowchart of MEP algorithm is shown in Fig. 4.

3.1.3 Extreme gradient boosting (XGB)

Extreme Gradient Boosting algorithm commonly abbreviated as XGBoost or XGB is an ensemble ML algorithm which involves traditional DT algorithm with gradient boosting approach [75]. DTs are very useful to gain insights about the significance of inputs in predicting the response and XGB algorithm magnifies this property by providing user enhanced capabilities to learn these insights. This ability is achieved by iteratively training DTs by leveraging gradient boosting technique. Boosting technique refers to the process in which multiple weak learners such as decision trees are iteratively produced in a sequence to produce a strong learning model. The purpose of adding new learners (decision trees) in a sequence is to reduce the residual (difference between actual and predicted value) of the previously constructed learner [76]. This technique increases the flexibility of algorithm by creating new learners that are robust to reduce the gradient of the loss function [77]. It has found many applications in various domains due to its superior performance. XGB algorithm effectively reduces overfitting and improves accuracy to develop a less complex model. Thus, it achieves higher accuracy than other tree boosting algorithms by effectively avoiding over-fitting of the algorithm to the training data. It has been reported that XGB outperforms other boosting algorithms when it is used to handle large data sets. It is because XGB employ a loss function and a regularization function to check model's fitness [78]. The overall prediction process followed by XGB is shown in Fig. 5. The final XGB model represented by F is given in Equation (1). In Equation (1), yi′ is the predicted value for ith sample, fk denotes the kth decision tree (weak learner) while the total number of trees in the model are represented by k.(1) yi′=F(Xi)=∑k=1kfk(Xi)

Fig. 5 Flowchart representation of XGB and AdaBoost prediction process.

Fig. 5

3.1.4 Adaptive gradient boosting (AdaBoost)

AdaBoost was one of the first boosting methods introduced by Freund [79] to overcome the predictive limitations of a single DT as a weak learner. In this algorithm, several weak learners are combined and “boosted” to create a strong learner. The weak learners are used in ensembles which allow for a strong approximation. Also, the training data in AdaBoost is sampled many times with replacement for training the weak learners. The performance of the weak learner is measured, and the sampling weights are updated. This process allows to designate more weight to the less accurate learners to increase their accuracy and this process is repeated several times [80]. The algorithm also introduces the training diversity by training the subsequent learners on the previously less accurate examples and the final prediction of AdaBoost is the weighted combination of predictions from all the learners.

After AdaBoost's impressive implementation to solve classification problems, its use has been extended to solve regression problems too because its modular nature allows to improve the regressors to adapt to dynamics of specific problems [81]. For solving a regression problem, AdaBoost first creates a sample from the training data by using the weight vector. Then, the base learning algorithm uses this sample to create a learner which links input data to the output and this learner is utilized to calculate predictions for all of the data points. Then the difference between actual and learner predicted values is calculated and compared with a threshold value ∅. This threshold value ∅ then classifies the predicted values as satisfactory or not. Then in the next step, the sampling weight is increased for the unsatisfactory predictions so that the algorithm pays more attention to them [82]. The representation of the AdaBoost prediction process is given in Fig. 5.

3.1.5 Random forest regression (RFR)

RFR is an ensemble ML algorithm consisting of two components i.e., DT and bagging algorithm and it can be utilized to solve both regression and classification problems based on the given dataset. It can swiftly relate input and output variables to make accurate predictions and has the ability to simulate complex and multivariate relationships between different dependent and independent variables. The algorithm runs by iteratively partitioning the feature space into several sub-regions and this partitioning continues until a stopping criterion is met. Although the decision tree algorithm was robust for solving regression problems in civil engineering, Breiman proposed RFR algorithm to overcome the learning limitations of DTs [83] and since its development, RFR has proven to be useful in solving various civil and geotechnical related problems [[84], [85], [86], [87], [88], [89], [90], [91], [92], [93], [94], [95]]. RFR treats DT as base learner and uses bagging technique on them to make predictions. It generates a number of DTs that are correlated by using different samples from the bagging technique and the results of all DTs are averaged to make the final predictions to overcome the overfitting issue. It is reported that RFR has several advantages over the traditional DT algorithm including ability to seamlessly handle large data sets, robustness towards outliers and over-fitting [96]. Also, it can simultaneously deal with various variables without skipping any of them and is simpler than other algorithms like ANN etc. The architecture of the RFR algorithm is illustrated in Fig. 6.Fig. 6 Flowchart of RFR algorithm.

Fig. 6

3.1.6 K-nearest neighbor (KNN)

KNN is a widely used ML technique which can solve both classification and regression problems [97]. The concept of KNN is based on the principle that if the k closest neighboring samples to a given sample in the space belong to a specific category, then the given sample should also belong to that category. In other words, the KNN algorithm is used to decide the category of an unknown sample by looking at the category of its k closest neighboring samples. For solving a regression problem, the KNN algorithm computes the distance between an input unknown sample and all other samples present in the dataset and selects the first k samples with the shortest distances. KNN algorithm computes the value of unknown data samples by computing the average of relevant variables and comparing the results with the k-nearest samples in the training set [98]. The algorithm's effectiveness is highly dependent on the k value [99], which determines the number of samples to be considered in the regression process. A smaller value of k may result in a higher level of noise in the process, while a larger k value may lead to over smoothing and lower accuracy. Therefore, choosing an appropriate k value is crucial for optimization the KNN performance. The distance between neighboring points used for the regression problems is calculated commonly by using Minkowski function as given in Equation (2). In Equation (2), xi and yi are the ith dimensions, and q indicates the order between the points x and y.(2) F(mi)=(∑i=0f(|xi−yi|)q)q1

The KNN method has several benefits over other ML algorithms. Firstly, it is simple and easy to understand. Secondly, it can easily understand nonlinear decision boundaries of both regression and classification problems. Moreover, the flexibility of the KNN approach is enhanced by the ability to adjust the k value to provide a suitable choice limit.

3.2 Data collection

This study was conducted using 159 data instances conducted from internationally published literature [48,60,61,[100], [101], [102], [103], [104], [105], [106]]. The collected database has four input parameters including strengths of mortar and the concrete masonry block (fm and fb respectively), height-to-thickness ratio, and the fm/fb ratio. These four parameters will be utilized towards the prediction of CS of concrete masonry prisms. The selection of these input parameters was based on the previous literature recommendations [48,102]. The input variables and geometry of hollow concrete prisms is shown in Fig. 7.Fig. 7 Geometry and input variables of hollow concrete blocks.

Fig. 7

3.2.1 Data description and splitting

The collection of a reliable dataset represents one of the most important steps in the process of ML models development. Thus, an extensive dataset of 159 points was collected to be used for this study and different statistical parameters of the database like mean, maximum, standard deviation, and minimum values of input and output variable(s) are given in Table 1. Also, according to the recommendations from previous studies, the whole database has been split into two subsets namely training set and testing set [107,108]. The training set comprises 70 % of the data and will be used for training the ML algorithms while testing set has 30 % of the data to be used for validating the performance of the trained ML models. The splitting of data like this makes sure that the algorithms perform good on both training and unseen data and are not overfitted to the training set [109].Table 1 Statistical description of dataset.

Table 1	Mortar Compressive Strength	Block Compressive Strength	height-to-thickness ratio	Ratio of fm/fb	Concrete Prism Compressive Strength	
Units	MPa	MPa	–	–	MPa	
Symbol	fm	fb	h/t	fm/fb	CS	
Maximum	26.5	40.5	5.2	1.48	31	
Minimum	3.9	9.0	1.80	0.10	6.63	
Mean	13.62	23.34	3.165	0.616	17.83	
Standard Deviation	6.011	7.295	0.989	0.322	5.163	
25 %	9.0	19.7	2.10	0.35	13.93	
50 %	14.2	22	3.1	0.60	17.5	
75 %	17.5	26.1	4.2	0.827	21.2	

3.2.2 Data correlation and distribution

The Pearson's correlation matrix or commonly referred to as heatmap developed for the dataset used in current study is given in Fig. 8. It is widely used to access the strength of regressions between all possible input and output combinations by using coefficient of correlation (R) [110]. The value of R ranges from −1 to 1 with 1 indicating a perfect linear positive correlation between two variables while −1 indicating a negative linear correlation. A positive correlation implies that increase in one variable also causes the other variable to increase while a negative correlation means that increasing one variable will lead to a decrease in other variable's values. However, the value of R being closer to zero implies the existence of little or no correlation between two variables. The correlation matrix given in Fig. 8 is symmetric about its diagonal meaning that corresponding values in upper and lower triangles are same. Generally, the R value greater than 0.8 indicates the presence of a good correlation between two variables while values greater than 0.9 highlights the presence of strong correlation [111]. It is important to investigate the correlation between different variable combinations before starting the actual model training process. This is because if the input variables are highly correlated with one another, they might give rise to an issue known as “multi-collinearity” during model development and affect the accuracy of resulting prediction models [108]. Thus, it is advised to keep the correlation values between different input variables less than 0.9 to avoid the potential risk of multi-collinearity [112]. Notice from the correlation matrix for this study given in Fig. 8 that the correlations between all the variables are less than the recommended value of 0.9. Thus, it implies that the problem of multi-collinearity will not arise while training the algorithms.Fig. 8 Correlation matrix of input and output variable(s).

Fig. 8

In addition to the data description and correlation, another important factor which must be considered before training the algorithms is the distribution of the data because the literature suggests having wide range of input and output variables to make widely applicable and robust prediction models [113]. The models developed using such type of data will have the ability to be utilized for a variety of data combinations and configurations. Thus, the distribution of data used in this study has been shown by means of a scatter matrix in Fig. 9. It is a useful statistical tool that can be used to visualize bivariate relationships between different input and output variable(s) and offer important insights about the regressions between variables [114]. It can be seen from Fig. 9 that the scatter matrix is a grid of scatter plots developed between input and output variable(s) while the diagonal shows the frequency distribution curves for variables. It can be seen from the frequency distribution curves and scatter plots that the variables are spread across a wide range in a continuous manner. The fm values range from 3 MPa to over 25 MPa. Similarly, fb values 9–40 MPa. In the same way, the height-to-thickness ratio, and output values are across a wide range. This type of data distribution will make sure that the developed models can be used for multiple data values and configurations. Moreover, the scatter matrix can serve as a tool for outlier detection since it helps to visualize the relationship between variables in the form of scatter plots [115].Fig. 9 Data distribution of input and output variable(s).

Fig. 9

3.3 Performance assessment

It is crucial to assess the predictive abilities of ML models using different statistical metrices to make sure that the developed models are robust in predicting the output without significant errors [116]. Thus, this study employs six widely used error metrices for evaluating the performance of models in both training and testing phases. The error metrices utilized, their mathematical formulae, significance, and recommended values for the models to be deemed acceptable are given in Table 2.Table 2 Summary of the error evaluation criteria.

Table 2No	Metric	Abbreviation	Formula	Criteria	Significance	Reference	
1.	Mean absolute error	MAE	Σ|x−y|n	Close to zero	Used to show average difference between the actual and predicted values.	[116]	
2.	Root mean square error	RMSE	∑(x−y)2n	Close to zero	Used as an indicator of large errors.	[117]	
3.	Coefficient of determination	R2	1−∑(x−y)2∑(y−ymean)2	R2>0.8	It is the most commonly used indicator of model's general accuracy and it allows to measure the correlation between predicted and actual values.	[118]	
4.	a20-index	a20	n20n	Close to 1	A recently developed metric used to indicate the proportion of predictions which deviate more than ±20 % from the actual values.	[119,120]	
5.	Performance index	PI	RRMSE1+R	PI<0.2	It considers the combined effect of correlation coefficient (R) and relative root mean square error (RRMSE).	[121]	
6.	Objective Function	OF	(nTraining−nTestingn)PITraining+2(nTestingn)PITesting	OF<0.2	It utilizes RRMSE, R and number of datapoints in both testing and training tests to access the accuracy of the model as a whole.	[122]	

4 Results and discussion

4.1 ML models development

4.1.1 GEP model development

This section discusses in detail the model development procedure adopted in this study. The GEP model presented in this study was developed using a software called Genexpro Tools. After loading the dataset into the software and splitting it into two sets, the next step is to specify some GEP fitting parameters which directly affect the accuracy of the output. Among these parameters, the most important one is the set of functions which will be used to create the resultant equation. Thus, some simple arithmetic functions like addition, subtraction, square root etc. were chosen to make the GEP equation. The next step is to specify the number of chromosomes (size of the population) which is an indicator of the number of programs in the solution. The presence of large number of programs (large number of chromosomes) generally increases the time required by the algorithm to complete the training process [123]. It is important to mention that increasing number of chromosomes in the population increases the accuracy of the model but only to a certain extent, after which overfitting starts. Thus, it is very important to wisely chose the appropriate value of number of chromosomes [124]. In current study, the value of chromosome number along with number of genes was chosen through a trial-and-error method and previous literature recommendations [125,126]. The values of these parameters were changed until a model with maximum possible accuracy was reached. The set of GEP hyperparameters used in current study is given in Table 3.Table 3 Hyperparameter settings of ML models.

Table 3Parameters	MEP	GEP	XGB	AdaBoost	RFR	KNN	
Constants per Gene	–	10	–	–	–	–	
No. of Genes	–	4	–	–	–	–	
Linking Function	–	Addition	–	–	–	–	
Head Size	–	15	–	–	–	–	
No. of Chromosomes	–	30	–	–	–	–	
Functions	+, -, × , ÷, sqrt	+, -, × , ÷, sqrt, cube root, x2,x3,x4,x5	–	–	–	–	
Number of Generations	500	–	–	–	–	–	
Subpopulation Size	200	–	–	–	–	–	
Runs	10	–	–	–	–	–	
Crossover Probability	0.9	–	–	–	–	–	
No. of Subpopulations	500	–	–	–	–	–	
Code Length	40	–	–	–	–	–	
Max. depth	–	–	10	20	30	–	
n_estimators	–	–	100	200	50	–	
Learning rate	–	–	0.05	0.09	–	–	
n_neighbors	–	–	–	–	–	5	

4.1.2 MEP model development

Similar to the GEP model, MEP algorithm was also employed with the help of MEPX 2021.05.18.0. software which allowed to change different MEP fitting parameters and observe the change in model's accuracy with change in these parameters. The most important MEP fitting parameters include code length, subpopulation size, crossover probability, and functions etc. and these were also chosen using recommendations from the relevant studies and a trial-based method [[127], [128], [129]]. The model development by MEP algorithm was started by using an initial value of subpopulation size as 10 and it was changed until the algorithm yielded maximum accuracy. As a general rule, increasing its value increases the accuracy of model but it can also cause a surge in the model complexity and computing power to execute the algorithm. Thus, its value must be carefully chosen because a higher value may lead to overfitting of the model [130]. The parameter crossover probability signifies the number of individuals that would go under the genetic process of crossover, and it has been chosen as 0.9 for current study meaning that 90 % of the individuals will go though the crossover procedure. The parameter code length and number of generations are also very important factors which directly affect the model's output. Code length is directly related to the length of the resulting equation and number of generations show the number of iterations performed by the MEP algorithm. The set of MEP parameters for this study is given in Table 3.

4.1.3 XGB, AdaBoost, RFR, and KNN model development

Contrary to the MEP and GEP development, the other ML algorithms (AdaBoost, XGB, RFR, and KNN) were employed using python computer language in Jupyter library of Anaconda software. The data loading, defining input and output parameters, and data splitting were all done with the help of python codes. However, the parameter optimization of the four aforementioned algorithms was done using grid search approach. It is one of the most widely used technique for parameter optimization [131]. It is an exhaustive framework to test all combinations of hyperparameters involved in model development. In this approach, the value of one parameter is changed across a range of possible values while keeping all other variables constant. This process is iteratively done for every variable until the optimal combination of hyperparameters yielding the maximum algorithm accuracy is found. The representation of a two-dimensional grid search between two parameters is given in Fig. 10. In current study, the ‘GridSearchCV’ function available in sklearn [132] was used to find the optimal algorithm hyperparameters. The hyperparameters to tune for the RFR model were n-estimators and maximum depth which represent the total number of DTs in the forest and depth of each tree respectively and the optimal values of these two parameters which resulted in the maximum model accuracy are given in Table 3. For XGB and AdaBoost algorithm, the two parameters are the same as for RFR model i.e., n-estimators and maximum depth along with an additional parameter known as learning rate. As mentioned previously in section 3.1.3 that XGB is an ensemble technique that leverages gradient boosting technique on DTs and adds more DTs to the algorithm in a sequence to reduce the disparity between actual and predicted values by the previous tree [133]. This technique allows XGB to quickly fit the training data but if this fitting is done too quicky, it can also lead to the algorithm being overfitted to the training dataset [134]. Thus, to limit the learning rate of algorithm by addition of subsequent trees, a factor known as learning rate is applied to the predictions of every new tree. The value of learning rate ranges from 0.01 to 0.1 and the literature recommends having a lesser value of learning rate so that the residual reduction by addition of new trees is done slowly [135]. For this reason, the learning rate is chosen as 0.05 for XGB and 0.09 for AdaBoost. Furthermore, the n-neighbors which represent the number of nearest neighbors that are averaged for making a prediction for KNN model were also chosen by grid search approach and the optimal number of neighbors turned out to be 5 as shown in Table 3.Fig. 10 Representation of 2D gird search between two hyperparameters.

Fig. 10

4.2 GEP result

The result of GEP model is shown as an empirical equation relating the CS of concrete masonry prisms with the four input parameters in Equation (3). Notice that the GEP equation is made up of only simple arithmetic operations specified in Table 3 like +, -, ÷, sqrt, cube root, x2,x5 etc. It was done to keep the equation as simple as possible. Also, it is important to mention that the GEP algorithm doesn't directly gave Equation (3) as an output, rather it provided four expression trees which were decoded and linked by the chosen pre-determined linking function (addition) to get the final resultant equation. For the ease of users, the equation has been split into different smaller and comparatively simpler subexpressions represented by A, B,C, and D. The user can first calculate the values of these simpler expressions and then join them to get the final result by GEP equation.

The results of GEP to predict CS are shown in the form of scatter plots for both data subsets in Fig. 11 along with the R2 value. These scatter plots serve as valuable tools for visualizing that how close the algorithm predicted values lie to the actual values. It can be seen from Fig. 11 that some points lie close to the linear fit line while others are scattered away. It indicates that GEP predicted some values with greater accuracy but also failed to predict accurately at some points. It can be well understood by looking at the error evaluation metrices of GEP training and testing sets given in Table 4, Table 5 respectively. First of all, the R2 value for GEP training is 0.80 which is exactly equal to the lower threshold for a model to be acceptable as per the criteria of literature already explained in section 3.3. However, the value increases to 0.83 for testing set. Thus, the GEP model is acceptable to be used for predicting CS. In the same way, the average error between real and predicted values by GEP is 1.75 and 1.95 for training and testing respectively which indicates that there is almost a difference of 2 MPa between actual and GEP predictions. Moreover, the a20 value for GEP training and testing are very close to 1 which indicates that most of the predictions made by the GEP algorithm doesn't deviate more than ±20 % from actual values. Furthermore, the PI and OF values of GEP, which are an indicator of model's overall efficiency and accuracy are also less than the upper limit of 0.2 and in fact very close to 0 which shows that overall GEP model can be used for effectively predicting CS of concrete masonry prisms.(3) CS=A+B+C+D

whereA=[fmfb]+[{(0.6612.156−ht)×(ht+fmfb)fmht×ht}+2fb]

B=[fm−{(1.268×(fmfb)5)(fm)}]−ht−0.379ht

C={1.9509−(ht−fm)}×fm2.854+[{(fmfb−5.471)+(fb+fmfb)}×{fmfb−fmfb−1.5580}]

D=[(fmfb×fm3.559)−3.7663]−[−6.2713+3.3953.559ht]

Fig. 11 Scatter plots between actual and predicted values of developed models.

Fig. 11

Table 4 Error metrices of training phase.

Table 4	MEP	GEP	XGB	AdaBoost	RFR	KNN	
MAE	1.802	1.751	0.222	0.445	0.472	1.458	
RMSE	2.351	2.111	0.348	0.587	0.6974	1.886	
a20	0.810	0.9189	1	1	1	0.936	
R2	0.810	0.808	0.9954	0.986	0.984	0.86	
PI	0.070	0.062	0.0097	0.0162	0.0193	0.053	
OF	0.0657	0.0626	0.0063	0.0330	0.0495	0.0728	
							

Table 5 Error metrices of testing phase.

Table 5	MEP	GEP	XGB	AdaBoost	RFR	KNN	
MAE	1.788	1.952	0.235	1.160	1.190	2.05	
RMSE	2.210	2.36	0.351	1.4817	1.6422	2.317	
R2	0.834	0.839	0.995	0.92	0.9135	0.81	
a20	0.833	0.8125	1	1	0.895	0.875	
PI	0.063	0.0678	0.0098	0.04417	0.0484	0.0711	

4.3 MEP result

The MEP algorithm also represented its output as an empirical equation like GEP and is given by Equation (4). This equation was also formulated only using simple mathematical functions for simplicity purposes. It is crucial to highlight that MEP algorithm also didn't directly yield the output equation rather it generated a computer interpretable C++ code which was decoded to get the final equation. Notice that the MEP equation is more compact and computationally simple than the GEP equation. It is because MEP has the ability to encode multiple solutions in a single chromosome which greatly improves accuracy and limits the length of the equation. The variables X and Y given in the equation can be computed separately and then put in Equation (4) to get the final MEP resultant efficiently.

The MEP predictive abilities are also given in the form of scatter plots in Fig. 11 for better visualization. Similar to the GEP plots, the data points in MEP scatter plots for both training and testing sets are scattered around the linear fit line. However, their distance from the linear fit line is less as compared to GEP which indicates the superior performance of MEP. Notice from Table 4, Table 5 that the R2 value for MEP training is 0.81 and 0.834 for testing set. Thus, MEP satisfies this criterion of being an acceptable ML predictive model. Also, the average error values (MAE values) of MEP model are closer to zero. Moreover, the a20-index values are 0.810 and 0.833 respectively which again highlights that MEP model gave predictions closer to the actual values. In the same way, the OF and PI values of MEP being less than 0.2 also verify its robustness.(4) CS=fb−(Y/fb−x3X)−(fmX/fm−(fmfb)X)(h/t)+2(fmfb)2(fmfb)−2(fmfb)+X

whereX=fm+2(fmfb)2(fmfb)−2(fmfb)

Y=(fmfb)X×(fb−(fmfb)X)+2(fmfb)

4.4 KNN result

The KNN algorithm does not expressed its output as an empirical equation. Nevertheless, the predictive ability of KNN for both training and testing set is shown as a scatter plot in Fig. 11. Same as GEP and MEP model, the KNN algorithm also gave some predictions with large errors while predicted other values with greater accuracy as shown by the data points around the linear fit line of training and testing plots. Also notice that the R2 value of KNN training is 0.86 which is significantly larger than MEP and GEP values. It indicates that KNN exhibited good performance as compared to GEP and MEP. The other error metrices of KNN are also shown in Table 4, Table 5 for training and testing sets respectively. Firstly, notice that the training MAE of KNN is 1.45 which is significantly lesser than the MEP and GEP algorithms. However, the performance of KNN dropped a little in testing set as depicted by the MAE value of 2.05. Similarly, the RMSE value also increased from 1.86 for training to 2.31 for testing demonstrating that testing dataset has larger errors in predictions as compared to training set. Despite a little drop in accuracy in testing set, the PI values of both datasets still lie below 0.2 and very close to 0 which indicates that the KNN model can still be used as a reliable model for prediction. The OF value being 0.0728 also suggests the same.

4.5 RFR result

The results of RFR-based predictive model are shown in Fig. 11 in the form of scatter plots and the error metrices are shown in Table 4 for training set and Table 5 for testing set. Notice from scatter plots that the data points lie significantly closer to the linear fit line as compared to KNN, MEP, and GEP. The R2 value of RFR training algorithm is 0.98 which indicates there is 98 % correlation between actual and predicted values. This value is significantly larger than the R2 values of MEP and GEP models which shows that RFR is much better than them. The value drops to 0.91 when used against testing set which indicates slight overfitting of RFR model to the training data. However, it is still larger than MEP and GEP values. Similarly, the MAE values of RFR are less than the other algorithms. The a20-index value of RFR training is perfectly 1 meaning that there are no predictions made by the RFR algorithm which deviate more than ±20 % from the actual values while this value decreases a little to 0.895 in testing phase. The RMSE values of RFR being close to zero also suggests that there aren't any abnormally large errors made by the RFR algorithm. Moreover, the PI and OF values are also very close to zero which suggests the overall robustness of the algorithm to predict CS of concrete masonry prisms.

4.6 AdaBoost result

The results of AdaBoost algorithm are also shown as scatter plots in Fig. 11 and it can be seen that the data points in AdaBoost training and testing scatter plots lie on the linear fit line or very close to it with only four to five points deviating slightly from the line. This fitting is better than all of the previous algorithms and it shows the excellent predictive capabilities of AdaBoost. The training and testing error metrices for the same algorithm are also shown in Table 4, Table 5 respectively. It can be seen that AdaBoost has training and testing R2 values of 0.98 and 0.92 which is significantly larger than the previous algorithms. Similarly, MAE value is 0.44 for training and 1.160 for testing which also shows the great accuracy of AdaBoost. Notably, the a20 value of AdaBoost is perfectly 1 for training and testing sets which shows that the predictions made by the algorithm are very close to the actual values. Although the algorithm suffered a slight loss in accuracy when tested against unseen testing data, but the PI values of both training and testing sets (0.016 and 0.044 respectively) suggests that the algorithm is feasible to calculate the desired output and it satisfies the criteria suggested by the literature. It can also be seen from the OF value being extremely close to zero i.e., 0.033.

4.7 XGB result

The results of XGB algorithm are also depicted as scatter plots in Fig. 11 since it does not yield an empirical equation like MEP and GEP algorithms. The scatter plots suggest that the XGB predicted values coincide almost 99 % with the actual values because the data points practically lie on the line of ideal fit. It is also evident from training and testing R2 values given in Table 4, Table 5 that XGB has predicted values with over 99 % accuracy. Additionally, the average error between actual and predictive values is less than 0.3 MPa for both training and testing sets of XGB. The RMSE values being very close to zero and a20 index values being perfectly 1 in both data subsets also suggests the same. Notably, XGB maintained its accuracy and exhibited the same performance in both training and testing subsets. The PI values of training data (0.0097) and testing data (0.0098) are a justification of this argument. Moreover, notice from Table 4 that OF value of XGB i.e., 0.0063 is significantly lesser than the OF values of all other algorithms employed in current study. Thus, it can be inferred that XGB proved to be the most accurate algorithm used in this study to predict CS of hollow concrete masonry blocks.

4.8 Error assessment and selection of best model

The commonly used error metrices are calculated for the developed ML models and results are given in Table 4, Table 5 for training data and testing data respectively. However, it is necessary to compare the models based on ability to predict the desired output accurately and specify a single algorithm which surpasses the other algorithms in terms of accuracy. Thus, a comparison has been drawn between the developed algorithms based on their residuals (difference between actual and predicted values) and the results are given in Fig. 12. A model having the smallest range of residuals will be considered the most accurate one and vice versa. First of all, notice that the GEP residuals are spread across the widest range of −6 to 6 with almost 95 % lying between −4 and 4. It shows that the maximum error that can occur between actual and GEP predicted values can range anywhere from −4 to 4 for most of the cases. In case of MEP, the range of more than 90 % residuals is somewhat compacted from −2 to 4 while there are very few residuals which go beyond this range. It shows that MEP-based predictions have lesser chances of resulting in large errors compared to GEP and thus is more accurate than GEP. Moving forward, some of the KNN residuals even go beyond 6 MPa but majority of them are concentrated in the range of −3 to 4. It also shows that KNN mostly underestimates the predicted value resulting in more positive residuals. In regard to RFR model, the range of residuals showed a significant decrease with more than 95 % residuals lying between −2 and 2. Also very few of the RFR residuals are in the range of 4 and 5 MPa which indicates that RFR is accurate than MEP, GEP, and KNN and the predictions based on RFR predictive model will be close to the actual values having the error between actual and predicted values in the range of −2 to 2 mostly. Additionally, the AdaBoost algorithm showed greater accuracy than the RFR model because most of the AdaBoost residuals are between −1 and 1 MPa with only two residuals extending beyond 2 MPa. The range of residuals further reduces in case of XGB which exhibits excellent accuracy having more than 99 % residuals in the range of −0.5 to 0.5 and only three to four residuals go beyond this range. It demonstrates that when using the XGB-based predictive model developed in current study for estimating CS of concrete masonry prisms, the error between actual and predicted value will always be close to 0.5 and won't go beyond that. Therefore, the overall order of algorithm accuracy based on error evaluation in Table 4, Table 5 and residual comparison is: XGB > AdaBoost > RFR > KNN > MEP > GEP. Moreover, to further verify and visualize the performance of employed algorithms for predicting CS across training and testing sets, the line plots drawn between actual and predicted CS values for all algorithms are given in Fig. 13. The line plots in Fig. 13 also confirm the excellent accuracy of XGB for predicting CS since the predicted curve by XGB is lying directly on the actual values curve at all points. In line plots of other algorithms like MEP, GEP, and KNN etc. it can be seen that the algorithms failed to predict CS accurately on several points, since the predicted curve deviates from the actual one. But the XGB algorithm proved to be effective for predicting CS with remarkable accuracy in both training and testing phases. Thus, it can be concluded that XGB is the most accurate algorithm employed in current study which has the least value of error between real and estimated value in both training and testing phases. Therefore, further analysis will be performed on the XGB model.Fig. 12 Residual comparison of developed models.

Fig. 12

Fig. 13 Series plot-based comparison of developed models.

Fig. 13

4.9 External validation of models

To further validate the performance of models after error evaluation by means of statistical error metrices, this study employs some external validation checks suggested by the literature. Table 6 provides the summary of all validation checks used in current study, their mathematical formulas, the suggested criteria for a model to be acceptable, and the calculated values of the algorithms used the whole dataset of 159 instances. Each of the employed external validation parameter has its own significance. For instance, K and K′ depict the slope of regression lines passing through the origin. Ro2 shows the correlation between real and predicted values while Rm is a measure of absolute difference between Ro2 and R2. It can be seen from Table 6 that overall, all the algorithms satisfied almost all external validation checks and their values lie within the recommended range. However, XGB again proved its superior performance and had values most closely related to the recommended criteria. The correlation coefficient (R) value od XGB is the highest at 0.9987 followed by AdaBoost at 0.9837, RFR at 0.978, KNN at 0.921, followed by GEP and lastly MEP. In the same way, Rm which simultaneously considers Ro2 and R2 is the highest for XGB (0.853), AdaBoost, RFR, and other algorithms. This concludes that the developed predictive models in this study using diverse ML techniques are robust and suitable to be used in practical scenarios.Table 6 External Validation of models.

Table 6Reference	Expression	Criteria	MEP	GEP	KNN	RFR	AdaBoost	XGB	
[136]	K=∑i=1n(xi×yi)∑i=1n(xi2)	0.85 < K < 1.15	0.969	0.969	0.974	0.988	0.987	0.997	
[136]	K′=∑i=1n(xi×y)∑i=1n(yi2)	0.85 < K′ < 1.15	0.106	1.01	1.00	1.007	0.999	1.013	
[137]	Rm=R2×(1−|R2−Ro2|)	Rm>0.5	0.526	0.536	0.534	0.763	0.800	0.853	
[138]	Ro2=1−∑i=1n(yi−xio)2∑i=1n(yi−yio)2 , xio=K×yi	Ro2≈1	0.945	0.98	0.810	0.986	0.968	0.9978	
[138]	Ro′2=1−∑i=1n(xi−yio)2∑i=1n(yi−yio)2 , yio=K′×xi	Ro′2≈1	0.984	0.93	0.80	0.949	0.960	0.999	
[139]	R=(n∑y−(∑x)(∑y))(n∑x2−(∑x)2)(n∑y2−(∑y)2)	R>0.8	0.903	0.907	0.921	0.978	0.9837	0.9987	

4.10 Explainable artificial intelligence (XAI)

Post-hoc explanations are used to advance the prediction-making process by bringing transparency to the traditional ML models. With the increasing use of ML models to predict different properties, the stress on increasing the transparency and interpretability of ML models have also increased [120,140].Since in the development of ML models, it is not always clear to know that how the algorithm calculates the predictions and it can serve as potential drawback to limit the widespread use of ML models in predicting critical parameters in materials and structural engineering. Thus, XAI approaches are increasingly getting attention to provide transparency to the ML model's inner working mechanisms. If the insights from the post-hoc explanations coincide with the observations of previous experimental studies and real-world scenarios, the developed ML models can be deemed acceptable [141]. The XAI approaches are used to reveal information about the internal working mechanisms of the ML models and are computationally simple and easy to implement. Also, these techniques does not need the information about the whole methodology of the algorithm rather the explanation is based on how the trained ML models reaches a specific prediction. Some of the most common forms of data-based explanatory analysis techniques used in the field of engineering for providing transparency to the black-box models are SHAP, ICE, Partial dependence plots (PDP) analysis, permutation feature importance, and local interpretable model-agnostic explanations. The relative strengths and weaknesses of some of the most commonly used XAI techniques are given in Table 7. Based on the comparison given in Table 7, this study implements SHAP, ICE, and sensitivity analysis (SA) on the XGB model since it was mentioned earlier in section 4.8 that XGB is the most accurate of all the algorithms employed in current study, the three types of XAI analyses will be performed on the trained XGB model. These three approaches are selected specifically since they offer significant advantages over other techniques as evident from Table 7. Also, by using these techniques, almost all type of explanations will be yielded from the XGB model. The SA will provide global feature importance based on its equations. SHAP will provide both local and global interpretations in an interactive manner to expose the hidden prediction mechanism of XGB while ICE analysis will give insights into the feature relationship with the outcome.Table 7 Summary of strengths and weaknesses of different XAI techniques.

Table 7Reference	Technique	Abbreviation	Benefits	Limitations	
[142]	Shapely Additive Explanatory Analysis	SHAP	• A unified measure of input feature importance across models

• Strong theoretical foundation based on the coalitional game theory

• Yields interpretations at both local and global level

• Easy to implement

• Provides feature importance in the form of additive feature attributions

• Results can be easily represented using a variety of plots

	• Computationally challenging when dealing with large datasets

• Different plots convey different information which can lead to confusion

	
[143]	Local Interpretable Model-agnostic Explanations	LIME	• A model agnostic approach

• Provides information about local interpretability of models

• Easy implementation

	• Doesn't provide theoretical explanation like SHAP

• The model agnostic explanations can sometimes be unstable across diverse data configurations

	
[144]	Partial Dependence Plots	PDP	• Gives global interpretation

• Helps to visualize the average effect of features on output

	• Inability to provide local interpretability

• Only applicable to numerical and categorical inputs/outputs

	
[145]	Permutation Feature Importance	–	• Model agnostic approach

• Gives local interpretability

• Easy to interpret

	• Computationally extensive

• Results can be affected by feature correlation

• Cannot provide local interpretation

	
[146]	Individual Conditional Expectation	ICE	• Very useful tool to check global interpretability of models

• Depicts detailed view of feature variation on the predicted output

• Provides the relationship between different inputs and outcome in the form of a curve

• Useful for capturing heterogeneity in feature effects

• Easy to interpret and implement

	• Can be visually challenging to interpret for complex feature relations

• Provides only visual analysis

• Computationally challenging for large datasets

	
[147]	Sensitivity Analysis	SA	• One of the oldest and most commonly used type of XAI technique

• Easy to implement

• Helps in identification of influential features

• Can be applied to almost every type of ML-based model

	• Cannot provide local interpretation

• Results can be affected by feature correlation, number of inputs, data size etc

• Can be computationally extensive for large datasets

	

4.10.1 Sensitivity analysis (SA)

The first XAI technique applied on the trained XGB model is the SA which tells about the importance of each input parameter in determining the output. It is an effective way to know about the sensitivity of the developed ML model towards the input data and parameters on which it has been trained [148]. Larger value of SA for a particular variable indicates the greater extent to which the output depends upon it and vice versa [149]. The SA result also depend upon other factors like data used to train the model and number of input variables considered in the study etc. [150]. The SA was done on the XGB model using Equations (5), (6). In Equations (5), (6), ei is the itℎ input variable with all other variables constant while fmax(ei) and fmin(ei) represent maximum predicted output and minimum predicted output respectively.(5) Rangei=fmax(ei)−fmin(ei)

(6) SA=Rangei∑nj=1Rangej

The result of SA is given in Fig. 14. It can be seen from Fig. 14 that the most important variable for predicting CS of hollow concrete masonry prisms is fb (compressive strength of masonry unit). After fb, the second most important variable is the height-to-thickness ratio followed by fm. However, the contribution of ratio of fm and fb is the least in output prediction.Fig. 14 Sensitivity analysis of XGB model.

Fig. 14

4.10.2 Shapley additive explanatory analysis (SHAP)

The employment of various explanatory techniques holds paramount importance for empirical studies due to the fact that they increase the transparency of black-box ML algorithms [151]. Thus, shapley analysis was conducted using the XGB-based predictive model because it proved to be the most accurate algorithm employed in current study as explained in section 4.8. SHAP analysis takes inspiration from the coalitional game theory and it assumes the feature values as coalition members [152]. The SHAP explanations of a ML model are of two broad types known as SHAP local interpretation and SHAP global interpretation.

The global interpretation plots provide information about the whole model highlighting the importance of input features as a whole for predicting the output. The SHAP global interpretation of a ML model is generally given in the form of mean absolute shap plot and beeswarm shap plot. The idea of mean absolute shap plot revolves around the fact that variables with high mean absolute values are more important to predict the outcome and vice versa. Fig. 15 (a) represents the mean absolute SHAP plot in which the input variables are listed according to their importance from top to bottom. Notice from Fig. 15 (a) that strength of masonry block (fb) has the highest mean absolute SHAP value (7.38). After fb, the ratio h/t has the largest shap absolute value of 2.19 followed by 0.77 for fm while the mean shap absolute value is least for fm/fb (0.25).

Although Fig. 15 (a) is a useful tool to know about the importance of inputs based on their mean absolute shap values, it fails to highlight the positive and negative correlations of input variables with the outcome. Thus, Fig. 15 (b) shows Shap summary plot to overcome this shortcoming of mean absolute shap plot. In the summary plot, the dots are colour coded such that blue colour shows lower values while red indicates higher values. The input variables are listed according to their importance from top to bottom in this plot too and the most important input features are the ones with the widest range of shapley values as indicated by the red and blue dots on the plot. Notice from the summary plot in Fig. 15 (b) that fb has the broadest range of both red and blue points indicating that it is most important variable to predict CS of hollow concrete masonry. Next to fb is the variable h/t who also has a wide range of SHAP values as indicated by the blue and red dots in the beeswarm plot. The range of fm is almost the same as of h/t while the SHAP values indicated by red and blue dots on the beeswarm plot are more concentrated around the zero line for the parameter fm/fb which shows that it has the least importance.

Although Fig. 15 is useful for getting a holistic overview of the most important inputs to predict the outcome, it is also important to investigate the local interpretation of developed model to further highlight that process used by the algorithm for calculating individual predictions. Thus, the shapley force plot in Fig. 16 (a) provides insight into the model prediction process. The plot is for the predicted CS value of 11.88 MPa and is centred at base value, which is basically an average value of all predictions made by the model; 15.5 in our case as shown in the plot. Each arrow represents an input feature which contributed to estimation of the output while the length of arrow specifies its magnitude while its direction indicates its effect on the output (positive or negative). It can be seen from Fig. 16 (a) that fb contributes the most in prediction as indicated by the longest length of its arrow followed by other variables. It is because fb is the most significant variable in outcome determination as previously indicated by the SHAP global interpretation.Fig. 15 SHAP global interpretation; (a) Mean absolute plot; (a) Beeswarm plot.

Fig. 15

In addition to the identification of most crucial input features, it is important to investigate the relationship between various inputs and their SHAP values to demonstrate that how the change in value of one variable affects its SHAP value and SHAP values of other variables. This is done by using partial dependence plots in Fig. 16 (b). In the partial dependence plots, the variable values are given along the bottom axis and right-hand side while the corresponding SHAP values for variable in x-axis are given on the left-hand side. Thus, these plots depict the change in shap values of a variable with the change in values of other variables. Overall, the SHAP values of fb, h/t, and fm vary greatly as compared to fm/fb. It is because fb, h/t, and fm are the most prominent variables in predicting the CS of hollow concrete masonry prisms. Notice that the SHAP values of fm vary from −2 to 2 as the ratio fm/fb changes from 0.2 to 1. Similarly, the SHAP values of fb exhibited a change from −8 all the way up to 6 with increase in h/t value from 2 to 5. The change in SHAP values of fb across a wide range also explains why it is the most prominent factor in output prediction. Moving forward, the SHAP values of h/t showed a decrease from 2 to almost −4 when the ratio fm/fb was varied 0.2 to 1.0. Finally, the SHAP values of fm/fb seem to be unbothered by change in h/t and remained almost zero with change in h/t values across a range of 2.0–5.0.Fig. 16 SHAP local interpretation; (a) Force plot; (a) Partial dependence plots.

Fig. 16

4.10.3 Individual conditional expectation (ICE) analysis

The ICE plots for XGB model developed in current study are shown in Fig. 17. These plots are an interactive and efficient way to demonstrate the variation in XGB-based output value with change in values of four input variables used to build the XGB model. To build an ICE plot, the values of all variables were set to their average values and only one variable was changed across its range. Then, the change in XGB predictions was plotted against the variable values. Each curve in the ICE plot corresponds to one of the data samples while the average is represented by the thick blue curve. Notice from Fig. 17 that increase in masonry block strength (fb) from 10 MPa to 40 MPa results in a stepwise increase in the XGB predictions from 10 to over 25 MPa. This is because fb is the most significant variable in output prediction as indicated previously by sensitivity and shapley analysis. In contrast, the increase in h/t value across a range of 1.5–5.5 results in a slight decrease in output from 19 MPa to almost 15 MPa. The relationship between fm and XGB predictions however does not follow a specific trend and the output initially remains constant at a value of around 16 MPa when fm changes from 5 to almost 8 MPa. The output then changes to almost 18 MPa when fm values reaches up to 10 MPa followed by a slight decrease until the output value increases again to almost 19 MPa before slightly decreasing again. The output value does not show an appreciable increase with increase in fm/fb value since it was already established that it is the least contributing variable. The XGB-based output value increases only from 16 MPa to only over 18 MPa when fm/fb changes from 0 to 1.6.Fig. 17 ICE analysis of XGB model.

Fig. 17

The insights from the above three XAI approaches described above are also in line with findings of previous literature. It was indicated consistently by sensitivity analysis, shapley global explanation, and ICE analysis that fb proved to be the most significant variable in predicting the outcome, which is in line with the results obtained in previous studies (both experimental and modelling) [54,153,154]. The second most important variable according to all three XAI approaches was the h/t ratio whose significance was also highlighted in a study conducted by Sathiparan et al. [153]. Similarly, the slight decrease in CS of masonry prisms with increase in h/t value as indicated by the ICE plot in Fig. 17 and the lesser overall effect of fm/fb in predicting CS also aligns with findings of literature [61]. Thus, it can be concluded that XGB-based prediction model developed in this study can be effectively used to compute CS of hollow concrete masonry prisms because it proved to be more accurate than all other algorithms based on the error evaluation and the insights into the XGB prediction process also align with the previous studies.

5 Practical implications of the study

It is important to highlight the practical implications of such empirical studies to further highlight their importance. It has already been discussed in introduction section that the accurate determination of CS of hollow concrete blocks is inevitable for design of resilient infrastructure. However, the accurate determination of CS of hollow concrete blocks is difficult using traditional predictive techniques due to various input factors involved in determination of strength. Also, the lack of a reliable database in previous studies greatly limits their use. The widespread use of hollow concrete blocks will help to achieve sustainability goals by providing a suitable alternative for expensive steel and concrete construction. Thus, this study was conducted in an attempt to provide empirical methods which can be effectively used by professionals in the construction industry to estimate CS of hollow concrete masonry prior to construction. The models developed in this study are also important for predicting the CS of existing structures made from concrete masonry prisms. The equations provided by MEP and GEP algorithms will help the professionals to effectively calculate the CS values using the set of inputs. Moreover, the information provided by the shapley analysis is important for highlighting the importance of different input variables in predicting CS. The use of sensitivity, shapley, and ICE analysis highlight the importance the using explainable AI approaches on black-box ML models for their interpretation and widespread use.

6 Conclusions

This study aimed to utilize multiple ML algorithms for prediction of CS of hollow concrete masonry prisms. A database of 159 experimental data points was collected from published literature for this purpose. The whole database was split into training and testing subsets to be used for training and testing of the algorithms including MEP, XGB, KNN, GEP, RFR, and AdaBoost. The main findings drawn from the study are.• The developed ML models were assessed for their accuracy using several error metrices including average error, a20-index, objective function, and R2 etc. and the results showed that all models satisfy the criteria suggested in the literature for being an acceptable predictive model.

• Only MEP and GEP expressed their output in the form of an empirical equation which can be used to calculate CS of concrete masonry blocks while all other algorithms failed to do so.

• Among the ML algorithms used in current study, XGB turned out to be the best and most accurate for prediction of CS (R2 = 0.99, a20-index = 1, and OF = 0.0063).

• The residual comparison of developed models revealed that the order of accuracy of algorithms used in this study is: XGB > AdaBoost > RFR > KNN > MEP > GEP.

• The outcomes of three XAI approaches employed in this study demonstrated that strength of masonry unit (fb) is the most significant variable in predicting CS of masonry blocks followed by height-to-thickness ratio (h/t). However, fm/fb contributes the least to the output prediction.

7 Recommendations

Although this study provides an extensive framework for CS prediction of hollow concrete masonry prisms using various ML algorithms to add to the knowledge and practical application of this subject. It is necessary to highlight potential limitations and provide avenues for future research.• The main focus of the study was to predict CS of hollow concrete masonry prism by using strengths of mortar and masonry unit along with height-to-thickness ratio as inputs. However, it is advisable to conduct future research by considering other input factors like different dimensions of masonry units and different mortar and masonry classifications etc.

• The predictive models presented in current study were developed using a limited dataset. It is imperative to consider much larger datasets in future studies for developing widely applicable and robust predictive models.

• Future studies should consider using other types of boosting based ML algorithms like CatBoost and Gradient Boosting Regressor etc. Also, the feasibility of newly developed physics-based approaches should be explored in future studies for prediction purposes.

Data availability

The raw data supporting the conclusions of this study is available with the corresponding author and will be furnished upon reasonable request.

CRediT authorship contribution statement

Waleed Bin Inqiad: Writing – original draft, Software, Methodology, Formal analysis, Data curation, Conceptualization. Elena Valentina Dumitrascu: Software, Resources, Project administration, Methodology, Investigation, Data curation. Robert Alexandru Dobre: Visualization, Validation, Supervision, Project administration, Funding acquisition, Data curation. Naseer Muhammad Khan: Writing – review & editing, Methodology, Investigation, Data curation, Conceptualization. Abbas Hussein Hammood: Validation, Supervision, Software, Resources, Project administration, Investigation, Data curation. Sadiq N. Henedy: Validation, Supervision, Software, Resources, Project administration, Methodology, Investigation. Rana Muhammad Asad Khan: Visualization, Validation, Supervision, Software, Resources.

Declaration of competing interest

The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.

Appendix A Supplementary data

The following is the Supplementary data to this article:Multimedia component 1

Multimedia component 1

Acknowledgements

This work was supported by a grant from the National Program for Research of the National Association of Technical Universities - GNAC ARUT 2023 .

Appendix A Supplementary data to this article can be found online at https://doi.org/10.1016/j.heliyon.2024.e36841.
==== Refs
References

1 P. G. Asteris et al.,“Masonry Compressive Strength Prediction Using Artificial Neural Networks.”.
2 H. B. Kaushik, ; Durgesh, C. Rai, S. K. Jain, and M. Asce, “Stress-Strain Characteristics of Clay Brick Masonry under Uniaxial Compression”, doi: 10.1061/ASCE0899-1561200719:9728.
3 Shu Z. Reinforced moment-resisting glulam bolted connection with coupled long steel rod with screwheads for modern timber frame structures Earthq. Eng. Struct. Dynam. 52 4 Apr. 2023 845 864 10.1002/EQE.3789
4 Lourenço P. J. P.-H.-C. & structures, and undefined Validation of Analytical and Continuum Numerical Methods for Estimating the Compressive Strength of Masonry 2006 Elsevier [Online]. Available: https://www.sciencedirect.com/science/article/pii/S0045794906002367
5 Li H. Yang Y. Wang X. Tang H. Effects of the position and chloride-induced corrosion of strand on bonding behavior between the steel strand and concrete Structures 58 Dec. 2023 10.1016/J.ISTRUC.2023.105500
6 Köksal H.O. Karakoç C. Yildirim H. Compression behavior and failure mechanisms of concrete masonry prisms J. Mater. Civ. Eng. 17 1 Feb. 2005 107 115 10.1061/(ASCE)0899-1561(2005)17:1(107)
7 Atkinson R.H. Noland J.L. Abrams D.P. McNary S. Deformation failure theory for stack-bond brick masonry prisms in compression [Online]. Available: https://experts.illinois.edu/en/publications/deformation-failure-theory-for-stack-bond-brick-masonry-prisms-in 1985
8 Khoo C.L. Hendry A.W. Strength tests on brick and… - Google Scholar.” [Online]. Available: https://scholar.google.com/scholar?hl=en&as_sdt=0%2C5&q=C.L.+Khoo%2C+A.W.+Hendry%2C+Strength+tests+on+brick+and+mortar+under+complex++stresses+for+the+development+of+a+failure+criterion+for+brickwork+in++compression%2C+in%3A+Proc.+British+Cer.+Soc.+21.+England%3A+Stoke-on-Trent%2C+1973%2C++pp.+51%E2%80%9366.&btnG=
9 Zeng H. Performance evolution of low heat cement under thermal cycling fatigue: a comparative study with moderate heat cement and ordinary Portland cement Construct. Build. Mater. 412 Jan 2024 10.1016/J.CONBUILDMAT.2024.134863
10 Wang M. Xi X. Guo Q. Pan J. Cai M. Yang S. Sulfate diffusion in coal pillar: experimental data and prediction model Int J Coal Sci Technol 10 1 Dec. 2023 1 12 10.1007/S40789-023-00575-8/FIGURES/12
11 He H. A general and simple method to disperse 2D nanomaterials for promoting cement hydration Construct. Build. Mater. 427 May 2024 10.1016/J.CONBUILDMAT.2024.136217
12 Thaickavil N.N. Thomas J. Behaviour and strength assessment of masonry prisms Case Stud. Constr. Mater. 8 2018 23 38
13 Lu D. Zhou X. Du X. Wang G. 3D dynamic elastoplastic constitutive model of concrete within the framework of rate-dependent consistency condition J. Eng. Mech. 146 11 Nov. 2020 10.1061/(asce)em.1943-7889.0001854
14 Soleimani F. Si G. Roshan H. Zhang J. Numerical modelling of gas outburst from coal: a review from control parameters to the initiation process Int J Coal Sci Technol 10 1 Dec. 2023 10.1007/S40789-023-00657-7
15 Zhang C. Wang P. Wang E. Chen D. Li C. Characteristics of coal resources in China and statistical analysis and preventive measures for coal mine accidents Int J Coal Sci Technol 10 1 Dec. 2023 1 13 10.1007/S40789-023-00582-9/FIGURES/13
16 Thaickavil N.N. Thomas J. Behaviour and strength assessment of masonry prisms Case Stud. Constr. Mater. 8 Jun. 2018 23 38 10.1016/J.CSCM.2017.12.007
17 Gao Q. Ding Z. Liao W.H. Effective elastic properties of irregular auxetic structures Compos. Struct. 287 May 2022 10.1016/J.COMPSTRUCT.2022.115269
18 Rahman S.K. Al-Ameri R. Experimental investigation and artificial neural network based prediction of bond strength in self-compacting geopolymer concrete reinforced with basalt FRP bars Appl. Sci. 11 11 Jun. 2021 10.3390/app11114889
19 He H. Exploring green and efficient zero-dimensional carbon-based inhibitors for carbon steel: from performance to mechanism Construct. Build. Mater. 411 Jan 2024 10.1016/J.CONBUILDMAT.2023.134334
20 Bin Inqiad W. Siddique M.S. Alarifi S.S. Butt M.J. Najeh T. Gamil Y. Comparative analysis of various machine learning algorithms to predict 28-day compressive strength of Self-compacting concrete Heliyon 9 11 Nov. 2023 e22036 10.1016/j.heliyon.2023.e22036
21 Qi Q. Yue X. Duo X. Xu Z. Li Z. Spatial prediction of soil organic carbon in coal mining subsidence areas based on RBF neural network Int J Coal Sci Technol 10 1 Dec. 2023 1 13 10.1007/S40789-023-00588-3/TABLES/4
22 Alyami M. Application of metaheuristic optimization algorithms in predicting the compressive strength of 3D-printed fiber-reinforced concrete Developments in the Built Environment 17 Mar. 2024 100307 10.1016/j.dibe.2023.100307
23 Mustapha I.B. Comparative analysis of gradient-boosting ensembles for estimation of compressive strength of quaternary blend concrete Int J Concr Struct Mater 18 1 Dec. 2024 10.1186/s40069-023-00653-w
24 Yao X. AI-based performance prediction for 3D-printed concrete considering anisotropy and steam curing condition Construct. Build. Mater. 375 Apr 2023 10.1016/J.CONBUILDMAT.2023.130898
25 Long X. hui Mao M. xiong Su T. tai Su Y. ke Tian M. Machine learning method to predict dynamic compressive response of concrete-like material at high strain rates Defence Technology 23 May 2023 100 111 10.1016/J.DT.2022.02.003
26 Liu C. Cui J. Zhang Z. Liu H. Huang X. Zhang C. The role of TBM asymmetric tail-grouting on surface settlement in coarse-grained soils of urban area: field tests and FEA modelling Tunn. Undergr. Space Technol. 111 May 2021 10.1016/J.TUST.2021.103857
27 Lu D. Meng F. Zhou X. Zhuo Y. Gao Z. Du X. A dynamic elastoplastic model of concrete based on a modeling method with environmental factors as constitutive variables J. Eng. Mech. 149 12 Dec. 2023 10.1061/JENMDT.EMENG-7206
28 Huang H. Yuan Y. Zhang W. Zhu L. Property assessment of high-performance concrete containing three types of fibers Int J Concr Struct Mater 15 1 Dec. 2021 10.1186/s40069-021-00476-7
29 Huang H. Huang M. Zhang W. Pospisil S. Wu T. Experimental investigation on rehabilitation of corroded RC columns with BSP and HPFL under combined loadings J. Struct. Eng. 146 8 Aug. 2020 10.1061/(ASCE)ST.1943-541X.0002725
30 Zhang H. Xiang X. Huang B. Wu Z. Chen H. Static homotopy response analysis of structure with random variables of arbitrary distributions by minimizing stochastic residual error Comput. Struct. 288 Nov 2023 10.1016/J.COMPSTRUC.2023.107153
31 Cao J. He H. Zhang Y. Zhao W. Yan Z. Zhu H. Crack detection in ultrahigh-performance concrete using robust principal component analysis and characteristic evaluation in the frequency domain Struct. Health Monit. 23 2 Mar. 2024 1013 1024 10.1177/14759217231178457
32 Ali Z. Karakus M. Nguyen G.D. Amrouch K. Effect of loading rate and time delay on the tangent modulus method (TMM) in coal and coal measured rocks Int J Coal Sci Technol 9 1 Dec. 2022 1 13 10.1007/S40789-022-00552-7/FIGURES/16
33 Wang G. Research and practice of intelligent coal mine technology systems in China Int J Coal Sci Technol 9 1 Dec. 2022 1 17 10.1007/S40789-022-00491-3/FIGURES/13
34 Yao Z. Wang Y. Shen J. Niu Y. Yang J.F. Wei X. Synergistic CO2 mineralization using coal fly ash and red mud as a composite system International Journal of Coal Science & Technology 11 1 May 2024 1 10 10.1007/S40789-024-00672-2 2024 11:1
35 Zheng H. Jiang B. Wang H. Zheng Y. Experimental and numerical simulation study on forced ventilation and dust removal of coal mine heading surface Int J Coal Sci Technol 11 1 Dec. 2024 1 17 10.1007/S40789-024-00667-Z/FIGURES/20
36 Nematzadeh M. Shahmansouri A.A. Zabihi R. Innovative models for predicting post-fire bond behavior of steel rebar embedded in steel fiber reinforced rubberized concrete using soft computing methods Structures 31 Jun. 2021 1141 1162 10.1016/J.ISTRUC.2021.02.015
37 Ashrafian A. Shahmansouri A.A. Akbarzadeh Bengar H. Behnood A. Post-fire behavior evaluation of concrete mixtures containing natural zeolite using a novel metaheuristic-based machine learning method Arch. Civ. Mech. Eng. 22 2 May 2022 1 25 10.1007/S43452-022-00415-7/METRICS
38 Ghanbari S. Shahmansouri A.A. Akbarzadeh Bengar H. Jafari A. Compressive strength prediction of high-strength oil palm shell lightweight aggregate concrete using machine learning methods Environ. Sci. Pollut. Control Ser. 30 1 Jan. 2023 1096 1115 10.1007/S11356-022-21987-0/METRICS
39 Pap E. Park C. Saadati R. Additive σ-RANDOM operator inequality and rhom-derivations in fuzzy banach algebras U.P.B. Sci. Bull., Series A 82 2 2020
40 Huang F. Slope stability prediction based on a long short-term memory neural network: comparisons with convolutional neural networks, support vector machines and random forest models Int J Coal Sci Technol 10 1 Dec. 2023 1 14 10.1007/S40789-023-00579-4/FIGURES/5
41 Zhao Y. He X. Jiang L. Wang Z. Ning J. Sainoki A. Influence analysis of complex crack geometric parameters on mechanical properties of soft rock Int J Coal Sci Technol 10 1 Dec. 2023 10.1007/S40789-023-00649-7
42 Liu B. Yao J. Sun T. Numerical analysis of water-alternating-CO2 flooding for CO2-EOR and storage projects in residual oil zones Int J Coal Sci Technol 10 1 Dec. 2023 10.1007/S40789-023-00647-9
43 Lu S. Zhao J. Song J. Chang J. Shu C.M. Apparent activation energy of mineral in open pit mine based upon the evolution of active functional groups Int J Coal Sci Technol 10 1 Dec. 2023 10.1007/S40789-023-00650-0
44 Yang Y. Simulation study of hydrogen sulfide removal in underground gas storage converted from the multilayered sour gas field Int J Coal Sci Technol 10 1 Dec. 2023 10.1007/S40789-023-00631-3
45 Tao K. Dang W. Liao X. Li X. Experimental study on the slip evolution of planar fractures subjected to cyclic normal stress Int J Coal Sci Technol 10 1 Dec. 2023 10.1007/S40789-023-00654-W
46 Zou Q. Zhan J. Wang X. Huang Z. Influence of nanosized magnesia on the hydration of borehole-sealing cements prepared using different methods Int J Coal Sci Technol 10 1 Dec. 2023 10.1007/S40789-023-00605-5
47 Yang G. Chen Y. Liu X. Yang R. Zhang Y. Zhang J. Stability analysis of a slope containing water-sensitive mudstone considering different rainfall conditions at an open-pit mine Int J Coal Sci Technol 10 1 Dec. 2023 10.1007/S40789-023-00619-Z
48 Zhou Q. Wang F. Zhu F. Estimation of compressive strength of hollow concrete masonry prisms using artificial neural networks and adaptive neuro-fuzzy inference systems Construct. Build. Mater. 125 Oct. 2016 417 426 10.1016/j.conbuildmat.2016.08.064
49 Zhang Y. Zhou G.C. Xiong Y. Rafiq M.Y. Techniques for predicting cracking pattern of masonry wallet using artificial neural networks and cellular automata J. Comput. Civ. Eng. 24 2 Feb. 2010 161 172 10.1061/(ASCE)CP.1943-5487.0000021
50 Garzón-Roca J. Marco C.O. Adam J.M. Compressive strength of masonry made of clay bricks and cement mortar: Estimation based on Neural Networks and Fuzzy Logic Eng. Struct. 48 2013 21 27
51 Garzón-Roca J. Adam J.M. Sandoval C. Roca P. Estimation of the axial behaviour of masonry walls based on artificial neural networks Comput. Struct. 125 2013 145 152
52 Ma D. Duan H. Zhang J. Bai H. A state-of-the-art review on rock seepage mechanism of water inrush disaster in coal mines International Journal of Coal Science & Technology 9 1 Jul. 2022 1 28 10.1007/S40789-022-00525-W 2022 9:1
53 Turco C. Funari M.F. Teixeira E. Mateus R. Artificial neural networks to predict the mechanical properties of natural fibre-reinforced compressed earth blocks (Cebs) Fibers 9 12 Dec. 2021 10.3390/fib9120078
54 Fakharian P. Rezazadeh Eidgahee D. Akbari M. Jahangir H. Ali Taeb A. Compressive strength prediction of hollow concrete masonry blocks using artificial intelligence algorithms Structures 47 Jan. 2023 1790 1802 10.1016/J.ISTRUC.2022.12.007
55 Lan G. Wang Y. Zeng G. Zhang J. Compressive strength of earth block masonry: estimation based on neural networks and adaptive network-based fuzzy inference system Compos. Struct. 235 Mar 2020 10.1016/j.compstruct.2019.111731
56 Mishra M. Bhatia A.S. Maity D. Predicting the compressive strength of unreinforced brick masonry using machine learning techniques validated on a case study of a museum through nondestructive testing J Civ Struct Health Monit 10 3 Jul. 2020 389 403 10.1007/S13349-020-00391-7
57 Mishra M. Bhatia A.S. Maity D. A comparative study of regression, neural network and neuro-fuzzy inference system for determining the compressive strength of brick–mortar masonry by fusing nondestructive testing data Eng. Comput. 37 1 Jan. 2021 77 91 10.1007/S00366-019-00810-4
58 Huang H. Guo M. Zhang W. Zeng J. Yang K. Bai H. Numerical investigation on the bearing capacity of RC columns strengthened by HPFL-BSP under combined loadings J. Build. Eng. 39 Jul 2021 10.1016/J.JOBE.2021.102266
59 Shahin M. Jaksa M.B. Maier H.R. Physical modeling of rolling dynamic compaction view project artificial neural networks-pile capacity prediction view project [Online]. Available: https://www.researchgate.net/publication/228364758 2008
60 Ramamurthy K. Sathish V. Ambalavanan R. Compressive strength prediction of hollow concrete block masonry prisms Struct. J. 97 1 2000 61 67
61 Ho L.S. Tran V.Q. Evaluation and estimation of compressive strength of concrete masonry prism using gradient boosting algorithm PLoS One 19 3 March Mar. 2024 10.1371/journal.pone.0297364
62 Mohamad G. Lourenço P. R.-C H. Mechanics of hollow concrete block masonry prisms under compression: review and prospects C. Composites, and undefined 2007 ElsevierG Mohamad, PB Lourenço, HR RomanCement and Concrete Composites, 2007•Elsevier 2006 10.1016/j.cemconcomp.2006.11.003
63 Wu H. Chen Y. Lv H. Xie Q. Chen Y. Gu J. Stability analysis of rib pillars in highwall mining under dynamic and static loads in open-pit coal mine Int J Coal Sci Technol 9 1 Dec. 2022 10.1007/S40789-022-00504-1
64 Animah F. Keles C. Reed W.R. Sarver E. Effects of dust controls on respirable coal mine dust composition and particle sizes: case studies on auxiliary scrubbers and canopy air curtain Int J Coal Sci Technol 11 1 Dec. 2024 1 16 10.1007/S40789-024-00688-8/FIGURES/6
65 Modiga A. Eterigho-Ikelegbe O. Bada S. Extractability and mineralogical evaluation of rare earth elements from Waterberg Coalfield run-of-mine and discard coal Int J Coal Sci Technol 11 1 Dec. 2024 65 10.1007/S40789-024-00702-Z
66 Gandomi A.H. Alavi A.H. Mirzahosseini M.R. Nejad F.M. Nonlinear genetic-based models for prediction of flow number of asphalt mixtures J. Mater. Civ. Eng. 23 3 Mar. 2011 248 263 10.1061/(asce)mt.1943-5533.0000154
67 J. R. Koza and M. Jacks Hall, “SURVEY OF GENETIC ALGORITHMS AND GENETIC PROGRAMMING.” [Online]. Available:http://www-cs-faculty.stanford.edu/∼koza/.
68 Huang H. Li M. Yuan Y. Bai H. Theoretical analysis on the lateral drift of precast concrete frame with replaceable artificial controllable plastic hinges J. Build. Eng. 62 Dec 2022 10.1016/J.JOBE.2022.105386
69 Gholampour A. Gandomi A.H. Ozbakkaloglu T. New formulations for mechanical properties of recycled aggregate concrete using gene expression programming Construct. Build. Mater. 130 Jan. 2017 122 145 10.1016/j.conbuildmat.2016.10.114
70 Zhang Y. Research on coal-rock identification method and data augmentation algorithm of comprehensive working face based on FL-Segformer Int J Coal Sci Technol 11 1 Dec. 2024 10.1007/S40789-024-00704-X
71 Oltean M. Multi expression programming [Online]. Available: www.cs.ubbcluj.ro/∼molteanwww.mep.cs.ubbcluj.ro 2006
72 Crina M.O. Gros¸an G. A comparison of several linear GP techniques A comparison of several linear genetic programming techniques [Online]. Available: www.mep.cs.ubbcluj.ro 2003
73 Xiao C. Additive manufacturing of high solid content lunar regolith simulant paste based on vat photopolymerization and the effect of water addition on paste retention properties Addit. Manuf. 71 Jun 2023 10.1016/J.ADDMA.2023.103607
74 Mousavi S.M. Alavi A.H. Gandomi A.H. Esmaeili M.A. Gandomi M. A Data Mining Approach to Compressive Strength of CFRP-Confined Concrete Cylinders 2010
75 Chen T. Guestrin C. XGBoost: a scalable tree boosting system Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining Aug. 2016 Association for Computing Machinery 785 794 10.1145/2939672.2939785
76 Wang C. Xu S. Yang J. Adaboost algorithm in artificial intelligence for optimizing the IRI prediction accuracy of asphalt concrete pavement Sensors 21 17 Sep. 2021 10.3390/s21175682
77 Le L.T. Nguyen H. Zhou J. Dou J. Moayedi H. Estimating the heating load of buildings for smart city planning using a novel artificial intelligence technique PSO-XGBoost Appl. Sci. 9 13 2019 10.3390/APP9132714
78 Wang G. Influences of clean fracturing fluid viscosity and horizontal in-situ stress difference on hydraulic fracture propagation and morphology in coal seam International Journal of Coal Science & Technology 11 1 May 2024 1 17 10.1007/S40789-024-00692-Y 2024 11:1
79 Freund Y. R. S.-J. of computer and system sciences, and undefined “A decision-theoretic generalization of on-line learning and an application to boosting,” ElsevierY Freund, RE SchapireJournal of computer and system sciences, 1997•Elsevier [Online]. Available: https://www.sciencedirect.com/science/article/pii/S002200009791504X 1997
80 Freund Y. Schapire R.E. A short introduction to boosting J. Jpn. Soc. Artif. Intell. 14 5 1999 771 780 [Online]. Available: www.research.att.com/fyoav
81 G. Ridgeway,“The State of Boosting *.”.
82 Luo T. Wang J. Chen L. Sun C. Liu Q. Wang F. Quantitative characterization of the brittleness of deep shales by integrating mineral content, elastic parameters, in situ stress conditions and logging analysis Int J Coal Sci Technol 11 1 Dec. 2024 1 13 10.1007/S40789-023-00637-X/FIGURES/14
83 L. Breiman,“Random Forests,” 2001.
84 Qi C. Fourie A. Du X. Tang X. Prediction of open stope hangingwall stability using random forests Nat. Hazards 92 2 Jun. 2018 1179 1197 10.1007/s11069-018-3246-7
85 Shanmugasundar G. Vanitha M. Čep R. Kumar V. Kalita K. Ramachandran M. A comparative study of linear, random forest and adaboost regressions for modeling non-traditional machining Processes 9 11 Nov. 2021 10.3390/pr9112015
86 Wei J. Ying H. Yang Y. Zhang W. Yuan H. Zhou J. Seismic performance of concrete-filled steel tubular composite columns with ultra high performance concrete plates Eng. Struct. 278 Mar. 2023 10.1016/J.ENGSTRUCT.2022.115500
87 Cui D. Wang L. Zhang C. Xue H. Gao D. Chen F. Dynamic splitting performance and energy dissipation of fiber-reinforced concrete under impact loading Materials 17 2 Jan. 2024 10.3390/ma17020421
88 Ghasemi M. Zhang C. Khorshidi H. Zhu L. Hsiao P.C. Seismic upgrading of existing RC frames with displacement-restraint cable bracing Eng. Struct. 282 May 2023 10.1016/J.ENGSTRUCT.2023.115764
89 Jiang Y. Mechanical properties and acoustic emission characteristics of soft rock with different water contents under dynamic disturbance International Journal of Coal Science & Technology 11 1 May 2024 1 14 10.1007/S40789-024-00682-0 2024 11:1
90 Wang S. Guo J. Yu Y. Shi P. Zhang H. Quality evaluation of land reclamation in mining area based on remote sensing Int J Coal Sci Technol 10 1 Dec. 2023 1 10 10.1007/S40789-023-00601-9/TABLES/6
91 Huang F. Slope stability prediction based on a long short-term memory neural network: comparisons with convolutional neural networks, support vector machines and random forest models Int J Coal Sci Technol 10 1 Dec. 2023 1 14 10.1007/S40789-023-00579-4/FIGURES/5
92 Ma D. Duan H. Li Q. Wu J. Zhong W. Huang Z. Water–rock two-phase flow model for water inrush and instability of fault rocks during mine tunnelling Int J Coal Sci Technol 10 1 Dec. 2023 10.1007/S40789-023-00612-6
93 Sun M. The release and migration mechanism of arsenic during pyrolysis process of Chinese coals Int J Coal Sci Technol 11 1 Dec. 2024 10.1007/S40789-024-00715-8
94 Asiwaju L. Mustapha K.A. Abdullah W.H. Sia S.G. Hakimi M.H. Geochemistry of Cenozoic coals from Sarawak Basin, Malaysia: implications for paleoclimate, depositional conditions, and controls on petroleum potential Int J Coal Sci Technol 11 1 Dec. 2024 10.1007/S40789-024-00690-0
95 Lu S. A review of coal permeability models including the internal swelling coefficient of matrix Int J Coal Sci Technol 11 1 Dec. 2024 10.1007/S40789-024-00701-0
96 liang Jin X. Estimation of wheat agronomic parameters using new spectral indices PLoS One 8 8 Aug. 2013 10.1371/journal.pone.0072736
97 Wu X. Top 10 algorithms in data mining Knowl. Inf. Syst. 14 1 2008 1 37 10.1007/s10115-007-0114-2
98 Qian Y. Zhou W. Yan J. Li W. Han L. Comparing machine learning classifiers for object-based land cover classification using very high resolution imagery Rem. Sens. 7 1 2015 153 168 10.3390/rs70100153
99 Akbulut Y. Sengur A. Guo Y. Smarandache F. NS-k-NN: neutrosophic set-based k-nearest neighbors classifier Symmetry (Basel) 9 9 Sep. 2017 10.3390/sym9090179
100 Sarhat S.R. Sherwood E.G. The prediction of compressive strength of ungrouted hollow concrete block masonry Construct. Build. Mater. 58 May 2014 111 121 10.1016/j.conbuildmat.2014.01.025
101 J. Thompson, N. Lang, and T. Witthuhn,“RECALIBRATION OF THE UNIT STRENGTH METHOD FOR VERIFYING COMPLIANCE WITH THE SPECIFIED COMPRESSIVE STRENGTH OF CONCRETE MASONRY.”.
102 Andolfato R. J. C.-… materials journal, and undefined “Brazilian results on structural masonry concrete blocks,” search.proquest.comRP Andolfato, JS Camacho, MA RamalhoACI materials journal, 2007•search.proquest.com [Online]. Available: https://search.proquest.com/openview/9f6e03879a53a258729e3f5496076033/1?pq-origsite=gscholar&cbl=37076 2007
103 Huang H. Huang M. Zhang W. Guo M. Liu B. Progressive collapse of multistory 3D reinforced concrete frame structures after the loss of an edge column Structure and Infrastructure Engineering 18 2 2022 249 265 10.1080/15732479.2020.1841245
104 Shu Z. Reinforced moment-resisting glulam bolted connection with coupled long steel rod with screwheads for modern timber frame structures Earthq. Eng. Struct. Dynam. 52 4 Apr. 2023 845 864 10.1002/EQE.3789
105 Wang L. Effect of long reaction distance on gas composition from organic-rich shale pyrolysis under high-temperature steam environment Int J Coal Sci Technol 11 1 Dec. 2024 1 18 10.1007/S40789-024-00689-7/TABLES/2
106 Saberi F. Hosseini-Barzi M. Effect of thermal maturation and organic matter content on oil shale fracturing Int J Coal Sci Technol 11 1 Dec. 2024 1 19 10.1007/S40789-024-00666-0/FIGURES/16
107 Bin Inqiad W. Dumitrascu E.V. Dobre R.A. Forecasting residual mechanical properties of hybrid fibre-reinforced Self-compacting concrete (HFR-SCC) exposed to elevated temperatures Heliyon 10 12 Jun. 2024 e32856 10.1016/J.HELIYON.2024.E32856
108 Bin Inqiad W. Soft computing models for prediction of bentonite plastic concrete strength Sci. Rep. 14 1 Aug. 2024 18145 10.1038/s41598-024-69271-0
109 Bin Inqiad W. Siddique M.S. Ali M. Najeh T. Predicting 28-day compressive strength of fibre-reinforced self-compacting concrete (FR-SCC) using MEP and GEP Sci. Rep. 14 1 Jul. 2024 17293 10.1038/s41598-024-65905-5
110 Li Z. Gao X. Lu D. Correlation analysis and statistical assessment of early hydration characteristics and compressive strength for multi-composite cement paste Construct. Build. Mater. 310 2021 125260
111 Sarveghadi M. Gandomi A.H. Bolandi H. Alavi A.H. Development of Prediction Models for Shear Strength of SFRCB Using a Machine Learning Approach Jul. 01, 2019 Springer London 10.1007/s00521-015-1997-6
112 Probability, statistics, and decision for civil engineers - Jack R Benjamin, C. Allin cornell - Google Books [Online]. Available: https://books.google.com.pk/books?hl=en&lr=&id=Gqm-AwAAQBAJ&oi=fnd&pg=PP1&dq=Probability,+Statistics+and+Decision+for+Civil+Engineers%3B+Courier+Cooperation,+Dover+Publication,+Mineola:+New+York,+NY,+USA,+2014%3B+p.+244.&ots=6cVO1xCc4F&sig=lFzwIUMnqpjNdyfRdnIGFNWfIi0&redir_esc=y#v=onepage&q&f=false
113 Khan M.A. Geopolymer concrete compressive strength via artificial neural network, adaptive neuro fuzzy interface system, and gene expression programming with K-fold cross validation Front Mater 8 May 2021 10.3389/fmats.2021.621163
114 Ware C. Interacting with visualizations Inf Vis 2021 359 392 10.1016/B978-0-12-812875-6.00010-4
115 Bin Inqiad W. Javed M.F. Siddique M.S. Alarifi S.S. Alabduljabbar H. A comparative analysis of boosting and genetic programming techniques for predicting mechanical properties of soilcrete materials Mater. Today Commun. 40 Aug. 2024 109920 10.1016/j.mtcomm.2024.109920
116 Rostami A. Raef A. Kamari A. Totten M.W. Abdelwahhab M. Panacharoensawad E. Rigorous framework determining residual gas saturations during spontaneous and forced imbibition using gene expression programming J. Nat. Gas Sci. Eng. 84 Dec. 2020 10.1016/J.JNGSE.2020.103644
117 Despotovic M. Nedic V. Despotovic D. Cvetanovic S. Evaluation of empirical models for predicting monthly mean horizontal diffuse solar radiation Renew. Sustain. Energy Rev. 56 Apr. 2016 246 260 10.1016/J.RSER.2015.11.058
118 Li H. Yang Y. Wang X. Tang H. Effects of the position and chloride-induced corrosion of strand on bonding behavior between the steel strand and concrete Structures 58 Dec. 2023 10.1016/J.ISTRUC.2023.105500
119 Asteris P.G. Koopialipoor M. Armaghani D.J. Kotsonis E.A. Lourenço P.B. Prediction of cement-based mortars compressive strength using machine learning techniques Neural Comput. Appl. 33 19 Oct. 2021 13089 13121 10.1007/s00521-021-06004-8
120 Gao Q. Method for rock fracture prediction and early warning: insight from fusion of multi-physics field information Heliyon 10 10 May 2024 e30660 10.1016/j.heliyon.2024.e30660
121 Gandomi A.H. Roke D.A. Assessment of artificial neural network and genetic programming as predictive tools Adv. Eng. Software 88 Jun. 2015 63 72 10.1016/j.advengsoft.2015.05.007
122 Bin Inqiad W. Comparison of boosting and genetic programming techniques for prediction of tensile strain capacity of Engineered Cementitious Composites (ECC) Mater. Today Commun. 39 Jun. 2024 109222 10.1016/j.mtcomm.2024.109222
123 Chu H.H. Sustainable use of fly-ash: use of gene-expression programming (GEP) and multi-expression programming (MEP) for forecasting the compressive strength geopolymer concrete Ain Shams Eng. J. 12 4 2021 3603 3617 10.1016/j.asej.2021.03.018
124 Davarpanah T.Q A. Masoodi A.R. Gandomi A.H. Unveiling the potential of an evolutionary approach for accurate compressive strength prediction of engineered cementitious composites Case Stud. Constr. Mater. 19 2023 10.1016/j.cscm.2023.e02172
125 Jalal F.E. Indirect estimation of swelling pressure of expansive soil: GEP versus MEP modelling Adv. Mater. Sci. Eng. 2023 2023 1 25 10.1155/2023/1827117
126 Shahmansouri A.A. Akbarzadeh Bengar H. Ghanbari S. Compressive strength prediction of eco-efficient GGBS-based geopolymer concrete using GEP method J. Build. Eng. 31 Sep 2020 10.1016/J.JOBE.2020.101326
127 Amin M.N. Forecasting compressive strength of RHA based concrete using multi-expression programming Materials 15 11 Jun. 2022 10.3390/ma15113808
128 Althoey F. Crack width prediction of self-healing engineered cementitious composite using multi-expression programming J. Mater. Res. Technol. 24 May 2023 918 927 10.1016/j.jmrt.2023.03.036
129 Khan M. Ali M. Najeh T. Gamil Y. Computational prediction of workability and mechanical properties of bentonite plastic concrete using multi-expression programming Sci. Rep. 14 1 2024 6105 10.1038/s41598-024-56088-0 38480772
130 Iqbal M.F. Sustainable utilization of foundry waste: forecasting mechanical properties of foundry sand based concrete using multi-expression programming Sci. Total Environ. 780 Aug 2021 10.1016/j.scitotenv.2021.146524
131 Zöller M.-A. Huber M.F. Benchmark and Survey of Automated Machine Learning Frameworks 2021
132 Pedregosa Fabianpedregosa F. Scikit-learn: machine learning in Python gaël varoquaux bertrand thirion vincent dubourg alexandre passos PEDREGOSA, VAROQUAUX, GRAMFORT ET AL. Matthieu perrot [Online]. Available: http://scikit-learn.sourceforge.net 2011
133 Hoang N.D. A novel ant colony-optimized extreme gradient boosting machine for estimating compressive strength of recycled aggregate concrete Multiscale and Multidisciplinary Modeling, Experiments and Design 2023 10.1007/s41939-023-00220-6
134 Al-Taai S.R. Azize N.M. Thoeny Z.A. Imran H. Bernardo L.F.A. Al-Khafaji Z. XGBoost prediction model optimized with bayesian for the compressive strength of eco-friendly concrete containing ground granulated blast furnace slag and recycled coarse aggregate Appl. Sci. 13 15 Aug. 2023 10.3390/app13158889
135 Cui L. Chen P. Wang L. Li J. Ling H. Application of extreme gradient boosting based on grey relation analysis for prediction of compressive strength of concrete Adv. Civ. Eng. 2021 2021 10.1155/2021/8878396
136 Roy P.P. Roy K. On some aspects of variable selection for partial least squares regression models QSAR Comb. Sci. 27 3 Mar. 2008 302 313 10.1002/qsar.200710043
137 Ali Khan M. Zafar A. Akbar A. Javed M.F. Mosavi A. Application of Gene Expression Programming (GEP) for the Prediction of Compressive Strength of Geopolymer Concrete 2021 10.3390/ma14051106
138 Chu Mohsin Ali H.-H.K. Javed Muhammad Faisal Zafar Adeel Khan M. Ijaz Alabduljabbar Hisham Qayyum Sumaira Sustainable use of fly-ash: use of gene-expression programming (GEP) and multi-expression programming (MEP) for forecasting the compressive strength geopolymer concrete Ain Shams Eng. J. 12 4 2021 3603 3617 10.1016/j.asej.2021.03.018
139 Pang B. Ultraductile waterborne epoxy-concrete composite repair material: epoxy-fiber synergistic effect on flexural and tensile performance Cem. Concr. Compos. 129 May 2022 10.1016/J.CEMCONCOMP.2022.104463
140 Sagi O. Rokach L. Explainable decision forest: transforming a decision forest into an interpretable tree Inf. Fusion 61 2020 124 138 10.1016/j.inffus.2020.03.013
141 Thisovithan P. Aththanayake H. Meddage D.P.P. Ekanayake I.U. Rathnayake U. A novel explainable AI-based approach to estimate the natural period of vibration of masonry infill reinforced concrete frame structures using different machine learning techniques Results in Engineering 19 Sep 2023 10.1016/j.rineng.2023.101388
142 Cakiroglu C. Aydın Y. Bekdaş G. Geem Z.W. Interpretable predictive modelling of basalt fiber reinforced concrete splitting tensile strength using ensemble machine learning methods and SHAP approach Materials 16 13 Jul. 2023 10.3390/ma16134578
143 Visani G. Bagli E. Chesani F. Poluzzi A. Capuzzo D. Statistical Stability Indices for LIME: Obtaining Reliable Explanations for Machine Learning Models Jan. 2020 10.1080/01605682.2020.1865846
144 Karim R. Islam M.H. Datta S.D. Kashem A. Synergistic effects of supplementary cementitious materials and compressive strength prediction of concrete using machine learning algorithms with SHAP and PDP analyses Case Stud. Constr. Mater. 20 Jul 2024 10.1016/j.cscm.2023.e02828
145 Sun G. Kong G. Liu H. Amenuvor A.C. Vibration velocity of X-section cast-in-place concrete (XCC) pile–raft foundation model for a ballastless track 54 9 2017 1340 1345 10.1139/CGJ-2015-0623 10.1139/cgj-2015-0623
146 Goldstein A. Kapelner A. Bleich J. Pitkin E. Peeking inside the black box: visualizing statistical learning with plots of individual conditional expectation [Online]. Available: http://arxiv.org/abs/1309.6392 Sep. 2013
147 Dimopoulos T. Bakas N. Sensitivity analysis of machine learning models for the mass appraisal of real estate. Case study of residential units in Nicosia, Cyprus Rem. Sens. 11 24 Dec. 2019 10.3390/rs11243047
148 Jagadesh P. de Prado-Gil J. Silva-Monteiro N. Martínez-García R. Assessing the compressive strength of self-compacting concrete with recycled aggregates from mix ratio using machine learning approach J. Mater. Res. Technol. 24 May 2023 1483 1498 10.1016/j.jmrt.2023.03.037
149 De-Prado-gil J. Palencia C. Jagadesh P. Martínez-García R. A comparison of machine learning tools that model the splitting tensile strength of self-compacting recycled aggregate concrete Materials 15 12 Jun. 2022 10.3390/ma15124164
150 Alaskar A. Comparative study of genetic programming-based algorithms for predicting the compressive strength of concrete at elevated temperature Case Stud. Constr. Mater. 18 Jul 2023 10.1016/j.cscm.2023.e02199
151 Iqbal M.F. Sustainable utilization of foundry waste: forecasting mechanical properties of foundry sand based concrete using multi-expression programming Sci. Total Environ. 780 Aug 2021 10.1016/j.scitotenv.2021.146524
152 Lundberg S. Erion G. Chen H. A. D.-N. machine, and undefined From local explanations to global understanding with explainable AI for trees,” nature.comSM Lundberg, G Erion, H Chen, A DeGrave, JM Prutkin, B Nair, R Katz, J HimmelfarbNature machine intelligence, 2020•nature.com [Online]. Available: https://www.nature.com/articles/s42256-019-0138-9 2020
153 Sathiparan N. Jeyananthan P. Prediction of masonry prism strength using machine learning technique: effect of dimension and strength parameters Mater. Today Commun. 35 Jun 2023 10.1016/j.mtcomm.2023.106282
154 Lu D. Wang G. Du X. Wang Y. A nonlinear dynamic uniaxial strength criterion that considers the ultimate dynamic strength of concrete Int. J. Impact Eng. 103 May 2017 124 137 10.1016/J.IJIMPENG.2017.01.011
