
==== Front
PLoS One
PLoS One
plos
PLOS ONE
1932-6203
Public Library of Science San Francisco, CA USA

10.1371/journal.pone.0309869
PONE-D-24-05709
Research Article
Medicine and Health Sciences
Medical Conditions
Cardiovascular Diseases
Cardiovascular Disease Risk
Medicine and Health Sciences
Cardiology
Cardiovascular Medicine
Cardiovascular Diseases
Cardiovascular Disease Risk
Medicine and Health Sciences
Medical Conditions
Metabolic Disorders
Metabolic Syndrome
Medicine and Health Sciences
Epidemiology
Medical Risk Factors
Computer and Information Sciences
Artificial Intelligence
Machine Learning
Support Vector Machines
Biology and Life Sciences
Physiology
Physiological Parameters
Body Weight
Obesity
Computer and Information Sciences
Artificial Intelligence
Machine Learning
Medicine and Health Sciences
Vascular Medicine
Blood Pressure
People and Places
Geographical Locations
Asia
Bangladesh
Metabolic syndrome predictive modelling in Bangladesh applying machine learning approach
Metabolic Syndrome prediction in Bangladesh using machine learning
https://orcid.org/0000-0002-3105-5915
Hossain Md Farhad Conceptualization Investigation Methodology Validation Writing – original draft 1 2
Hossain Shaheed Formal analysis Methodology Software Writing – original draft 2 *
Akter Mst. Nira Data curation Visualization 2
Nahar Ainur Data curation Project administration 2
https://orcid.org/0000-0002-6918-9871
Liu Bowen Supervision Validation Writing – review & editing 1
Faruque Md Omar Visualization Writing – review & editing 3
1 Division of Computing, Analytics and Mathematics, Department of Mathematics and Statistics, School of Science and Engineering, University of Missouri, Kansas City, MO, United States of America
2 Department of Statistics, Comilla University, Cumilla, Bangladesh
3 Division of Energy, Matter and Sciences, School of Science and Engineering, University of Missouri, Kansas City, MO, United States of America
Klisic Aleksandra Editor
University of Montenegro-Faculty of Medicine, MONTENEGRO
Competing Interests: There is no conflict of interest among the authors.

* E-mail: shaheedhossaincou98@stud.cou.ac.bd
5 9 2024
2024
19 9 e030986915 2 2024
12 8 2024
https://creativecommons.org/publicdomain/zero/1.0/ This is an open access article, free of all copyright, and may be freely reproduced, distributed, transmitted, modified, built upon, or otherwise used by anyone for any lawful purpose. The work is made available under the Creative Commons CC0 public domain dedication.

Metabolic syndrome (MetS) is a cluster of interconnected metabolic risk factors, including abdominal obesity, high blood pressure, and elevated fasting blood glucose levels, that result in an increased risk of heart disease and stroke. In this research, we aim to identify the risk factors that have an impact on MetS in the Bangladeshi population. Subsequently, we intend to construct predictive machine learning (ML) models and ultimately, assess the accuracy and reliability of these models. In this particular study, we utilized the ATP III criteria as the basis for evaluating various health parameters from a dataset comprising 8185 participants in Bangladesh. After employing multiple ML algorithms, we identified that 27.8% of the population exhibited a prevalence of MetS. The prevalence of MetS was higher among females, accounting for 58.3% of the cases, compared to males with a prevalence of 41.7%. Initially, we identified the crucial variables using Chi-Square and Random Forest techniques. Subsequently, the obtained optimal variables are employed to train various models including Decision Trees, Random Forests, Support Vector Machines, Extreme Gradient Boosting, K-nearest neighbors, and Logistic Regression. Particularly we employed the ATP III criteria, which utilizes the Waist-to-Height Ratio (WHtR) as an anthropometric index for diagnosing abdominal obesity. Our analysis indicated that Age, SBP, WHtR, FBG, WC, DBP, marital status, HC, TGs, and smoking emerged as the most significant factors when using Chi-Square and Random Forest analyses. However, further investigation is necessary to evaluate its precision as a classification tool and to improve the accuracy of all classifiers for MetS prediction.

The author(s) received no specific funding for this work. Data AvailabilityData set is publicly available in https://datadryad.org/stash/dataset/doi:10.5061/dryad.zkh18937f.
Data Availability

Data set is publicly available in https://datadryad.org/stash/dataset/doi:10.5061/dryad.zkh18937f.
==== Body
pmcIntroduction

Metabolic syndrome (MetS) is a cluster of conditions that, when combined, increase the vulnerability of a patient to several life-threatening diseases such as coronary heart disease, diabetes, stroke, and various other consequential health issues [1]. Approximately 422 million individuals globally suffer from diabetes, with the majority residing in low and middle-income nations [2]. The disease is directly responsible for 1.5 million fatalities annually. Over the past few decades, there has been a steady rise in both the number of cases and the incidence of diabetes. Metabolic syndrome, also known as insulin resistance syndrome and syndrome X, encompasses a range of risk factors related to blood pressure, glucose, and plasma lipid levels. Hypertension was identified by a systolic blood pressure of at least 130 mm Hg, a diastolic blood pressure of at least 85 mm Hg, or the use of antihypertensive medications. The risk of sudden cardiac death was found to be 70% higher in those with the MetS. According to the United States National Heart Lung and Blood Institute, every 1 in 3 Americans is suffering from MetS [3]. A 2018 non-communicable disease risk factor evaluation revealed that 15.5% of Bangladeshis aged 40–69 are at risk for cardiovascular diseases [4]. According to a systematic review and meta-analysis [5], 37.0% of Bangladeshi people were found to have MetS. The investigation of potential risk factors is necessary due to the high prevalence of MetS, which is a serious public health concern. According to the National Institute of Health, with advancing years comes an increased risk of MetS [6]. The risk of MetS may increase due to lifestyle choices, being inactive, poor diet, lack of quality sleep, smoking, excessive alcohol consumption, poor socioeconomic standing, and working irregular shifts. The chances of MetS in adults with a sleeping duration of less than six hours per day were roughly five times higher than the odds of MetS in adults who slept often and for more hours per day [OR: 4.62; 95% CI: (1.02, 20.98)] [7]. Genetic and family history and obesity also worsen the condition as these conditions can reduce the “good” HDL cholesterol and increase the “bad” LDL cholesterol while impacting badly on blood triglycerides, and blood pressure [8]. Additionally, certain medical conditions such as polycystic ovary syndrome (PCOS) and insulin resistance can also increase the risk of developing MetS [9]. Pregnancy-related overweight and obesity can increase the child’s chance of developing MetS [10]. The impacts of socioeconomic conditions on health are demonstrated in a study where they showed that rural populations have significantly higher tobacco use (45.2%), inadequate fruit/vegetable intake (92.1%), and higher daily salt intake (9.0 g) compared to urban populations [11]. Socioeconomic condition, age, sex, obesity, hypertension, wealth, and living conditions all had impacts on the prevalence of diabetes according to the Bangladesh Demographic and Health Survey 2017–18 [12]. Of those with diabetes, 61.5% were not reported of the condition, 35.2% were receiving regular treatment, and 30.4% had it under control [13]. Predictive healthcare strategies, utilizing machine learning (ML), are crucial in predicting and mitigating the impact of metabolic diseases, particularly in countries like Bangladesh with rising prevalence of diabetes, obesity, and cardiovascular diseases. ML significantly enhances data-driven research efficiency, reducing manual inspection burden and enabling the development of new models for optimal operation. ML uses feature selection techniques which can be carried out via random forests, chi-square tests, and correlation plots to identify key factors in large datasets with numerous less significant variables. Our study aims to identify the potential risk factors associated with metabolic diseases and to generate a predictive model of such diseases using anML approach. ML is preferred for accurate prediction due to its ability to recognize patterns and relationships, analyze larger data sets, and generate predictions quickly, making it more efficient than traditional methods, which are time-consuming and vulnerable to bias [14]. According to a nationwide cross-sectional survey carried out in 2018, 12.3% of adult Bangladeshis reported engaging in insufficient physical activity, with women (14.8%) and urban groups (14.1%) exhibiting higher frequencies. Additionally, the study found that the prevalence of overweight and obesity was 25.9% and that it was considerably higher in the groups of women (33.7%), urban (34.3%) and wealthiest (34.3%) [11].

In this research, we aim to identify the risk factors that have an impact on MetS. Subsequently, we intend to construct predictive ML models and ultimately, assess the accuracy and reliability of these predictive ML models.

Related work

In the last 30 years, there has been a significant rise in the global prevalence of MetS [15]. The International Diabetes Federation (IDF), the World Health Organization (WHO), the European Group for the Study of Insulin Resistance (EGIR), and the US National Cholesterol Education Program Adult Treatment Panel III (NCEP ATP III) have defined and published separate clinical criteria for MetS [16].

Mohammad Ziaul Islam Chowdhury et al. proposed an examination of the MetS prevalence in Bangladesh using meta-analysis [5]. The findings indicated that the prevalence of MetS in females (32%) is higher than in males (25%), although this difference is not statistically significant (p = 0.434). When the modified NCEP III criteria were utilized, the highest occurrence of metabolic syndrome was observed at 37%. Conversely, the prevalence decreased to its lowest level of 20% when the WHO criteria were applied. The studies indicated that geographical factors played a significant role in the variation observed [5].

Another method proposed by Suparno Datta et al. utilized ML to detect MetS at an early stage [15], which relies exclusively on non-invasive features such as height, weight, waist circumference (WC), triglycerides (TGs), blood sugar, and HDL levels. The ensemble learning approach demonstrates superior performance, with GBMs and RF closely trailing behind. The findings indicate that machine learning can effectively forecast MetS, eliminating the need for invasive biomarkers, and enhancing the convenience of early detection [15].

Guadalupe Obdulia Gutiérrez-Esparza et al. proposed an ML algorithm that predicts MetS in the Mexican population [17]. Random Forest was employed to prioritize health parameters. The key prognostic factors for MetS, based on their significance, included FPG, TGs, WHtR, HDL-C, and BMI. The data was analyzed using the Random Forest and chi-squared methods, which showed that WHtR had the highest values and was the most significant factor. Additionally, when evaluated with the C45 and JRip algorithms, WHtR demonstrated better performance in terms of balanced accuracy, sensitivity, and specificity. The RF model, which utilized ATP III variables, demonstrated the highest performance with a balanced accuracy of 0.875, closely trailed by JRip [17].

Shu-Jie-Xia et al. developed a diagnostic model for MetS that can be developed by incorporating symptoms into a physiochemical index [18]. Their selected cohort was compared to three traditional machine learning methods: Decision tree (DT), Support vector machine (SVM), and Random Forest (RF). Comparison among the three models indicated that the RF model exhibited superior performance, boasting the highest average accuracy (0.942 on average, with a 95% CI of [0.925, 0.958]) and sensitivity (0.993 on average, with a 95% CI of [0.990, 0.996]), when compared to SVM. The significance of the TGM indexes in predicting MetS was clearly stated in this study [18].

DarkoIvanovic et al. proposed an Artificial Neural Network (ANN) that can be utilized to predict the diagnosis of MetS by solely relying on non-invasive, cost-effective, and readily available diagnostic methods [16]. They included gender, age, body mass index (BMI), waist-to-height ratio, and systolic and diastolic blood pressure as the input vectors. The outcome of this study demonstrated that the implementation of ANN effectively predicts both positive and negative cases of MetS, thereby aiding in the early prevention of metabolic syndrome. The highest positive predictive value (PPV) was found to be 0.858, while the negative predictive value (NPV) was close to PPV at 0.832 [16].

Hui Zhang et al. also used ML in a retrospective cohort study to predict the probability of adults developing MetS within a 4-year period. Three ML techniques were selected, namely ANN, classification, regression tree, and SVM. All models, except for the classification and regression tree model in internal validation, had discrimination values greater than 0.7. In external validation, the Logistic regression model showed the highest discrimination. Furthermore, both external validation (0.780) and internal validation (0.788) demonstrated satisfactory calibration for the ANN model [19].

Mohammad Salim Hossain et al. examined the MetS among individuals with diabetes who reside in aBangladeshi coastal area [20]. It was discovered that approximately 47.00% of patients diagnosed with type 2 diabetes mellitus were afflicted with MetS. The prevalence rate of MetS was higher in females, with 58.60%, compared to males, who had a rate of 36.14%. Females showed higher rates of obesity and hypertriglyceridemia, along with lower levels of HDL. Also, the age group of 55–64 showed the highest occurrence of MetS [20].

Suresh Mehata et al. evaluated the occurrence and factors influencing MetS in Nepalese adults based on a study that represents the entire nation. The most common combination was low HDL-C, abdominal obesity, and high blood pressure, making up 8.18% of cases. Close behind was abdominal obesity, low HDL-C, and high triglyceride levels, accounting for 8% of cases. Only a small fraction, specifically less than two percent, of the participants exhibited all five components of the syndrome, while a significant portion, 19%, did not display any of the components. The prevalence of the syndrome consistently increased as the age group advanced, with adults between the ages of 45 and 69 having the highest prevalence, ranging from 28% to 30% [21].

Hayat Ali Shah et al. used deep neural networks to generate feature representations of metabolic pathways, which are then fed into random forests for pathway prediction. The DeepRF model accurately predicts both known and unknown metabolic pathways in organisms. It has been tested on a dataset of over 318,016 instances, showing high accuracy (>97%), recall (>95%), and precision (>99%). When compared to other methods, DeepRF consistently provides more reliable results [22].

Methodology

Materials

Data source

The dataset in this research was taken from the survey “National STEPS Survey for Non-communicable Diseases [NCDs] Risk Factors in Bangladesh 2018” which was conducted by the National Institute of Preventive and Social Medicine (NIPSOM) under the World Health Organization (WHO) [11]. Dataset link (STATA): https://extranet.who.int/ncdsmicrodata/index.php/access_licensed/download/1763/5374,(CSV):https://extranet.who.int/ncdsmicrodata/index.php/access_licensed/download/1763/5375. The study involved 8185 respondents, with 3804 male (46.5%) and 4381 female (53.5%) participants aged between 18–69 years. The dataset encompasses various risk factors for metabolic diseases including lifestyle habits, clinical and anthropometric measurements, and biomedical evaluation. A national cross-sectional population-based survey utilized a multi-stage cluster sampling design to select households and eligible adult men and women (aged 18–69) for an interview and physical examination. The physical examination consisted of anthropometry, blood pressure measurement, blood glucose, cholesterol, and a urine sample for salt analysis. The WHO NCD STEPS instrument version 3.2 was utilized to carry out the survey. The questionnaire is comprised of three STEPS aimed at assessing the NCD risk factors. Each step encompassed a range of core, expanded, and country-specific questions that were adjusted to cater to the local requirements. In Bangladesh, all core modules and optional modules, namely oral health, and cervical cancer screening, were included. The questionnaire was translated into Bengali, and the validation of the translated questionnaire was conducted through back translation. In the initial phase, personal information from participants in STEP 1, which included documenting their height, weight, and hip and waist measurements was collected. These measurements were taken from individuals who agreed to move on to STEP 2. After completing data collection in STEP 1 and STEP 2 at selected households, the following day, biochemical assessments were carried out at specified locations for each Primary Sampling Unit (PSU). These assessments involved analyzing blood samples for glucose and total cholesterol levels, which were taken from venous blood samples. Plasma samples were also used to measure the concentrations of glucose, total cholesterol, and HDL cholesterol. Fasting blood samples were specifically collected to identify elevated blood glucose levels. The subjects were classified as having type 2 diabetes if they reported being informed by their doctor about the disease (provided the diagnosis was made after the age of 25 and not due to pregnancy), if they reported using insulin or a hypoglycemic medication, or if their fasting blood sugar level exceeded 100 mg/dL [23].

Habits and lifestyles

Three STEPS were used in validated questionnaires [11] to measure the risk factors for NCDs. These questionnaires were used to collect data on lifestyle variables such as alcohol and smoking consumption, physical activity levels, and salt intake before meals.

Clinical and anthropometric measurements

The measurements of the diastolic and systolic blood pressure were taken following the JNC-established standard protocol [24], WC, height, weight, and BMI were determined using the formula weight/height2. The WHtR was computed by dividing the WC by height (waist/height).

Biochemical evaluation

The laboratory tests that were acquired were fasting blood glucose (FBG),TGs, and HDL cholesterol (HDL-C). Blood samples were obtained after a 12-hour overnight fast.

Diagnostic criteria

MetS encompasses a combination of significant risk factors, lifestyle-related risk factors, and emerging risk factors. These factors include abdominal obesity, atherogenic dyslipidemia (elevated triglyceride levels, small LDL particles, low HDL cholesterol), high blood pressure, insulin resistance (with or without glucose intolerance), and prothrombotic and proinflammatory states [25]. The clinical diagnosis criteria for MetS based on the ATP III criteria [26] was used. It is shown in Table 1.

10.1371/journal.pone.0309869.t001 Table 1 The clinical diagnosis criteria for metabolic syndrome based on the ATP III criteria.

Risk Factors	Criteria	
MetS	Three or more of the following criteria	
Waist Circumference	Male: >102 cm, Female: >88 cm	
Systolic Blood Pressure	≥130 mmHg	
Diastolic Blood Pressure	≥85 mmHg	
TGs	≥150 mg/dL	
Fasting Blood Glucose	≥100 mg/dL	
Body Mass Index	≥30 kg/m2	
HDL	<40 mg/dL	

Methods

Decision tree

A decision tree is a tree-structured classifier in which the features of a dataset are represented by internal nodes, the decision rules are represented by branches, and the conclusion is represented by each leaf node. It is a graphical tool that shows all the options for solving a problem or making a decision given certain parameters [27]. The method is non-parametric, effective for large datasets, and can be divided into training and validation datasets for optimal decision tree model construction [28].

Random forest

Adele Cutler and Breiman [29] introduced Random Forest which is a prediction technique that generates a set of CART classification trees and assigns the class to the instance based on a majority vote. This approach outperforms individual classification trees in terms of prediction accuracy and can be used for a variety of prediction situations [30]. In cases of regression or classification, Random Forest offers a technique called Variable Importance Measures (VIMs) to rank the importance of variables.

Support vector machine (SVM)

A support vector machine (SVM) is an ML algorithm that uses supervised learning models to solve complex classification, regression, and outlier detection problems by performing optimal data transformations that determine boundaries between data points based on predefined classes, labels, or outputs. SVMs are widely adopted across disciplines such as healthcare, natural language processing, signal processing applications, and speech & image recognition fields [31].

Extreme gradient boosting (XGBoost)

An ensemble ML technique that uses decision trees to provide a gradient boosting framework is called Xtreme Gradient Boosting (XGBoost). To reach the final prediction, XGBoost creates new models that predict the residuals of the earlier models [32].

K-nearest neighbors (KNN)

The k-nearest neighbors technique transforms Big Data into Smart Data, free from noise, redundant information, and missing values [33]. This approach is crucial for accurate data mining and revealing insightful information.

Logistic regression

A logistic regression model examines the relationship between one or more independent variables that are already present in order to predict a dependent variable in the data. Multiple input criteria can be considered by the model. Logistic regression is used in the field of ML as a key technique. The algorithms improve at classifying data sets as more pertinent data becomes available.

Feature selection criteria

The process of feature selection holds great significance in the realm of ML model development it allows the identification of crucial variables from extensive datasets which have a significant influence on the model in comparison to other variables. In our study, we utilized chi-square and random forest methodologies to identify these important variables.

Statistical analysis

To reveal the characteristics of objects under study we exert a chi-square test employing statistical packages for social science by using SPSS software version 28.0.We also utilized Python 3.0 with Jupyter Notebook where Pandas, NumPy, Scikit-learn, Seaborn, and Matplotlib libraries were used to reveal the results.

Metrices

In ML, various performance metrics are used to evaluate the accuracy and effectiveness of a classification model. Some of the key metrics include:

Precision: The proportion of true positive predictions among all positive predictions made by the model. It measures how accurate the model’s positive predictions are [34].

Sensitivity: Also known as recall, it is the proportion of true positive predictions among all actual positive cases. It measures how well the model identifies all the positive cases [34].

SENS=TPTP+FN

Specificity: It is the proportion of true negative predictions among all actual negative cases. Specificity measures how well the model identifies all the negative cases [34]salma.

SPC=TNFP+TN

Balanced accuracy (AUC-ROC): The area under the receiver operating characteristic (ROC) curve, which measures the model’s ability to distinguish between positive and negative cases. A higher AUC-ROC value indicates a better model performance [35]. BACC=12TPP+TNN

where P = Positive, N = Negative, TP = True Positive, FN = False Negative, TN = True Negative and FP = False Positive, respectively.

F1 score: The harmonic means of precision and recall; it measures the model’s overall accuracy by balancing both precision and recall. A higher F1 score indicates better model performance [36].

F1=2×precision×recallprecision+recall

Results and analysis

In our research, we utilized the dataset obtained from the "National STEPS Survey for Non-communicable Diseases Risk Factors in Bangladesh 2018" [11]. The ATP III criteria were employed to identify the crucial cardiovascular risk factors, and subsequently, the participants were categorized into two groups: MetS Group and Normal Group. The variables in our study can be classified into four categories: Lifestyle variables, Anthropometric variables, Clinical variables, and Biochemical variables [26, 37]. Lifestyle variables encompass factors such as smoking, alcohol consumption, and marital status. Anthropometric variables include age, weight, height, BMI, WHtR, waist circumference (WC), and hip circumference (HC). Clinical variables consist of systolic and diastolic blood pressure measurements. Lastly, biochemical variables encompass fasting blood glucose (FBG), high-density lipoprotein (HDL), and triglycerides (TGs).

In Fig 1, the process of constructing our proposed model is depicted. Initially, the crucial variables are identified through the utilization of Chi-Square and Random Forest techniques [17]. Subsequently, the obtained optimal variables are employed to train various models including Decision Trees, Random Forests, Support Vector Machines, Extreme Gradient Boosting, K-nearest neighbors, and Logistic Regression. The validity of these models is then assessed, and their performance is compared based on metrics such as Precision, Sensitivity, Specificity, Balanced Accuracy, AUC score, Recall, and F1 score. Furthermore, the performance of these models is visualized through the illustration of the ROC curve.

10.1371/journal.pone.0309869.g001 Fig 1 Diagram to show the process of building the model.

Prevalence of MetS

In our study, we found that the prevalence of MetS was 27.8% among the 8185 participants. Notably, there were significant variations between males and females, with males accounting for 41.7% of the MetS group and females comprising 58.3%. Fig 2 further highlights the dominance of the female population in both the MetS group and the Normal group.

10.1371/journal.pone.0309869.g002 Fig 2 Bar chart of showing prevalence of MetS regarding sex.

Table 2 displays the overall attributes of the participants in relation to the Mets group and Normal group. The Chi-Square test was employed in SPSS version 28.0 to analyze the data.

10.1371/journal.pone.0309869.t002 Table 2 General descriptions of the participants regarding MetS group and normal group using chi square test.

Variables	MetS Group (n = 2275) Count (%)	Normal Group (n = 5910) Count (%)	Total (n = 8185) Count (%)	P-value	
Sex				0.000	
Female	1327 (58.3)	3054 (51.7)	4381 (53.5)		
Male	948 (41.7)	2856 (48.3)	3804 (46.5)		
Age				0.000	
18–24	92 (4.0)	934 (15.8)	1026 (12.5)		
25–39	723 (31.8)	2766 (46.8)	3489 (42.6)		
40–54	924 (40.6)	1579 (26.7)	2503 (30.6)		
55–69	536 (23.6)	631 (10.7)	1167 (14.3)		
Residence				0.000	
Urban	1239 (54.5)	2763 (46.8)	4002 (48.9)		
Rural	1036 (45.5)	3147 (53.2)	4183 (51.1)		
Marital Status				0.000	
Never Married	41 (1.8)	449 (8.4)	535 (6.5)		
Currently Married	2062 (90.6)	5188 (87.8)	7250 (88.6)		
Separated	11 (0.5)	27 (0.5)	38 (0.5)		
Divorced	9 (0.4)	20 (0.3)	29 (0.4)		
Widowed	152 (6.7)	181 (3.1)	333 (4.1)		
Smoking				0.000	
Yes	411 (18.1)	1512 (25.6)	1923 (23.5)		
No	1864 (81.9)	4398 (74.4)	6262 (76.5)		
Alcohol				0.011	
Yes	149 (6.5)	477 (8.1)	626 (7.6)		
No	2126 (93.5)	5433 (91.9)	7559 (92.4)		
Eating Extra Salt				0.000	
Yes	933 (41)	2713 (45.9)	3646 (44.5)		
No	1442 (59)	3197 (54.1)	4539 (55.5)		
SBP at Risk				0.000	
Yes	1133 (49.8)	1361 (23.0)	2494 (30.5)		
No	1142 (50.2)	4549 (77.0)	5691 (69.5)		
DBP at Risk				0.000	
Yes	1152 (50.6)	1724 (29.2)	2876 (35.1)		
No	1123 (49.4)	4186 (70.8)	5309 (64.9)		
BMI at Risk				0.000	
Yes	252 (11.1)	271 (4.6)	523 (6.4)		
No	2023 (88.9)	5639 (95.4)	(93.6)		
WC at Risk				0.000	
Yes	888 (39.0)	1065 (18.0)	1953 (23.9)		
No	1387 (61.0)	4845 (82.0)	6232 (76.1)		
HC at Risk				0.000	
Yes	1082 (47.6)	1806 (30.6)	2888 (35.3)		
No	1193 (52.4)	4104 (69.4)	5297 (64.7)		
WHtR at Risk				0.000	
Yes	1358 (59.7)	1988 (33.6)	3346 (40.9)		
No	917 (40.3)	3922 (66.4)	4839 (59.1)		
FBG at Risk				0.000	
Yes	635 (27.9)	712 (12.0)	1347 (16.5)		
No	1640 (72.1)	5198 (88.0)	6838 (83.5)		
TGs at Risk				0.000	
Yes	1075 (47.3)	2011 (34.0)	3086 (37.7)		
No	1200 (52.7)	3899 (66.0)	5099 (62.3)		
HDL at Risk				0.036	
Yes	1403 (61.7)	3495 (59.1)	4898 (59.8)		
No	872 (38.3)	2415 (40.9)	3287 (40.2)		
SBP = Systolic Blood Pressure, DBP = Diastolic Blood Pressure, BMI = Body Mass Index, WC = Waist Circumference, HP = Hip Circumference, WHtR = Waist and Height Ratio, FBG = Fasting Blood Glucose, TGs = Triglycerides, HDL = High Density Lipoprotein.

Variable importance and key risk factors of metabolic syndrome

In the initial stage of our analysis, we classified our continuous variables, including WHtR, SBP, DBP, WC, HC, FBG, and TGs, into two categories: Yes and No. The classification of continuous variables is crucial for capturing non-linear relationships in models. Categorizing variables into ranges helps the model capture complex patterns effectively. This process optimizes algorithm performance by simplifying data representation, making it easier for the algorithm to learn and predict. Standardizing input data format and representing all variable types appropriately is essential for dealing with heterogeneous data. Categorizing continuous variables also improves the interpretability of the model’s output, making predictions easier to understand for users [38]. This categorization was based on the ATP III criteria from the original dataset [26]. We made this categorization for our convenience in applying various classification algorithms.

Upon examining Table 3 and Fig 3, we observed that Age, WHtR, SBP, DBP, WC, HC, FBG, Marital Status, TGs, and Residence were identified as key risk factors for Metabolic Syndrome (MetS) through the application of Chi-Square analysis. Among these variables, age had the highest score, followed by WHtR. SBP ranked third in terms of significance. DBP, WC, HC, and FBG held a moderate level of importance. Marital Status, TGs, and Residence were found to have the lowest significance.

10.1371/journal.pone.0309869.g003 Fig 3 Variable importance by using chi square.

10.1371/journal.pone.0309869.t003 Table 3 Top 10 important variables obtained by chi square.

Ranking	Variables	Scores	
01	Age	296.402	
02	WHtR	188.615	
03	SBP	169.351	
04	DBP	116.702	
05	WC	95.258	
06	HC	73.374	
07	FBG	49.488	
08	Marital	47.375	
09	TGs	46.123	
10	Residence	19.108	

After ranking these variables, it became evident that age played a crucial role in the development of Metabolic Syndrome. We observed that individuals above the age of 40 were more susceptible to Metabolic Syndrome diseases. Additionally, WHtR emerged as another prominent risk factor influencing MetS.

In Fig 4, the application of Random Forest on our training dataset reveals that Age, SBP, WHtR, FBG, WC, DBP, marital status, HC, TGs, and smoking are identified as the key risk factors for Metabolic Syndrome (MetS). The Age variable demonstrates the highest score, as confirmed by Chi Square analysis, followed by SBP. WHtR ranks third in importance. FBG, WC, DBP, and marital status hold a medium position in terms of significance. On the other hand, HC, TGs, and smoking are considered the least influential factors. The ranking of these factors, similar to Chi-Square analysis, highlights Age as the most crucial determinant for Metabolic Syndrome, followed by Systolic Blood Pressure. It has been established that individuals with elevated systolic blood pressure face a greater risk of developing Metabolic Syndrome.

10.1371/journal.pone.0309869.g004 Fig 4 Variable importance by using random forest.

Model performance and comparison

The predominant key risk factors of Metabolic Syndrome have been identified as Age, SBP, WHtR, FBG, WC, DBP, HC, and TGs through the Random Forest Variable Importance Measures (VIMs) technique. In order to further investigate these variables, we employed six different ML algorithms, namely Decision Tree, Random Forest, Support Vector Machine, Extreme Gradient Boosting, K-Nearest Neighbors, and Logistic Regression. These algorithms were chosen to explore the potential relationship between the aforementioned eight variables and Metabolic Syndrome.

We evaluate our models through both cross validation and external validation, and we find that our models demonstrate excellent performance. We optimize the parameters for each model through hyperparameter tuning and incorporate them into the model. We mention these parameters in the Table 4.

10.1371/journal.pone.0309869.t004 Table 4 The classification performance results of the models.

Classifier	Precision	Balanced Accuracy	Sensitivity	Specificity	AUC	Recall	F1-Score	
Decision Tree (criterion = ’Gini’, max depth = 4, max leaf nodes = 10, min samples leaf = 0.05, min samples split = 2)	0.75	0.68	0.58	0.77	0.73	0.75	0.71	
Random Forest (max depth = 8, no of estimators = 40)	0.76	0.69	0.61	0.78	0.74	0.76	0.73	
Support Vector Machine(C = 1, gamma = 0.1, kernel = ’rbf’)	0.78	0.70	0.62	0.78	0.68	0.76	0.72	
Extreme Gradient Boosting (learningrate = 0.1,maxdepth = 2,no of estimators = 180	0.74	0.71	0.63	0.79	0.70	0.77	0.73	
K-Nearest Neighbors (no of estimators = 19)	0.75	0.68	0.57	0.79	0.72	0.75	0.73	
Logistic Regression(C = 1, solver = liblinear)	0.77	0.71	0.63	0.78	0.75	0.77	0.74	

Table 4 illustrates that the Support Vector Machine (SVM) achieves the highest precision compared to other algorithms, with an accuracy rate of 78%. The SVM accurately identifies 78% of the relevant items. On the other hand, Logistic regression and XGBoost exhibit the highest balanced accuracy (71%) and sensitivity score (63%) among the other algorithms. However, there is only a 63% probability that the model will correctly detect positive cases of Metabolic Syndrome, which is a significantly lower score. In terms of specificity score, XGBoost and KNN outperform other algorithms, with a rate of 79% probability that these models will correctly reject negative cases. Logistic Regression demonstrates the best results in the AUC score, indicating that 75% of items are correctly classified by this algorithm. In terms of Recall, all classifiers exhibit similarly good performance, capturing 75% to 77% of positive items. The F1 score reveals that all classifiers perform moderately well. Overall, considering all metrics, Logistic Regression emerges as the best classifier among the other algorithms.

Table 5 displays the results of the six models with excluded variables such as BMI, HDL, sex, smoking, alcohol, and using extra salt. The presence of these variables has been found to diminish the overall performance of the models. Consequently, we have opted to remove these variables from our analysis due to their lack of significance in relation to our results.

10.1371/journal.pone.0309869.t005 Table 5 The classification performance results of the models considering excluded variables.

Classifier	Precision	Balanced Accuracy	Sensitivity	Specificity	AUC	Recall	F1-Score	
Decision Tree (criterion = ’Gini’, max depth = 4, max leaf nodes = 10, min samples leaf = 0.05, min samples split = 2)	0.73	0.36	0	0.73	0.36	0	0	
Random Forest (max depth = 8, no of estimators = 40)	0.73	0.86	1	0.73	0.86	1	0.84	
Support Vector Machine (C = 1, gamma = 0.1, kernel = ’rbf’)	0.73	0.36	0	0.73	0.36	0	0	
Extreme Gradient Boosting (learning rate = 0.1, max depth = 2, no of estimators = 180	0.72	0.56	0.39	0.73	0.56	0.39	0.50	
K-Nearest Neighbors (no of estimators = 19)	0.72	0.58	0.43	0.74	0.58	0.43	0.53	
Logistic Regression (C = 1, solver = liblinear)	0.73	0.61	0.48	0.73	0.61	0.48	0.57	

Fig 5 illustrates the identical concept that was previously discussed. Logistic Regression stands out as the most effective classification algorithm when compared to other algorithms in the field. Based on the findings presented in Fig 5, XGBoost can be regarded as the poorest performing classification algorithm, exhibiting the lowest precision in comparison to other algorithms in the same category.

10.1371/journal.pone.0309869.g005 Fig 5 Bar chart of displaying model performance.

Fig 6 illustrates the ROC curve, which depicts the performance of different classification algorithms. According to the results, Logistic Regression emerges as the most effective classification algorithm, while SVM is identified as the least accurate classification algorithm.

10.1371/journal.pone.0309869.g006 Fig 6 ROC curve.

Discussion and recommendation

Our research represents a pioneering effort in predicting potential risk factors of MetS using MLTs. The study encompassed a total of 8185 participants from various regions of Bangladesh and incorporated 16 distinct variables. These variables encompassed anthropometric data, lifestyle-related features gathered through questionnaires, and biochemical test results which were collected from “National STEPS Survey for Non-communicable Diseases Risk Factors in Bangladesh 2018”. Upon conducting the required computations, our findings revealed that 27.8% of the population exhibited a prevalence of MetS [39, 40].

Notably, there were remarkable disparities between sexes, with males comprising 41.7% of the MetS group and females accounting for 58.3% [41, 42]. The higher prevalence of MetS among females can be attributed to various factors, including sociocultural activities, psychosocial behaviors, socioeconomic status, genetic inheritance, and hormonal changes. These factors make females more susceptible to developing MetS compared to males [43].

In this research, a collection of health parameters was prioritized using Chi-Square and then compared to Random Forest to determine the significance of each variable. The findings revealed that the primary predominant variables for MetS in our sample of the Bangladeshi population based on their importance were Age, WHtR, SBP, DBP, WC, FBG, HC, and TGs. Interestingly, six out of these eight variables align with the criteria proposed by ATP III for classifying individuals with MetS [44].

It is worth noting that WHtR was ranked as the second variable in terms of significance, which is a noteworthy discovery, particularly concerning the obesity crisis in our nation and its association with cardiovascular diseases, the leading cause of illness and death globally and in Bangladesh. Abdominal obesity has emerged as an indicator of cardiometabolic risk, prompting considerable endeavors to identify a suitable anthropometric measurement that accurately reflects the accumulation of fat in the abdominal region and can be conveniently obtained without the need for advanced technological equipment [17, 45].

It is also worth noting that anthropometric indexes are significantly impacted by various factors such as age, sex, and ethnicity. Consequently, selecting a suitable index can be a daunting endeavor. When utilizing Chi-Square and Random Forest in our research, age emerges as the most crucial factor. Individuals who are older than 40 years are found to be in the high-risk category for MetS. While BMI has been widely utilized as a measure of body fat content, in our research it fails to accurately reflect abdominal obesity [46, 47].

A simple and effective method for assessing abdominal obesity and associated metabolic risk is WC measurement. It is worth noting that in our research WC is the fifth significant feature as a risk factor for Mets and abdominal obesity significantly influences the emergence of MetS.

The Random Forest Variable Important Measures (VIMs) technique in Fig 4 revealed that BMI, HDL, Marital Status, Residence, Sex, Using Extra Salt, and Smoking had lower scores and were therefore not considered in the modeling. The inclusion of these variables resulted in a decrease in the overall accuracy of the models. Conversely, excluding these variables led to an increase in the overall accuracy of the models [26].

Briefly, our research has revealed a moderately higher prevalence of Metabolic Syndrome (MetS) among the Bangladeshi Population, particularly among females. The primary risk factors identified were Age, Waist and Height Ratio (WHtR), High Systolic and Diastolic Blood Pressure, Waist Circumference (WC), Fasting Blood Glucose (FBG), Hip Circumference (HC), and Triglycerides (TGs), which aligns with the conventional definition of MetS established by various organizations.

Machine learning models have demonstrated potential in clinical settings for early detection and intervention of metabolic syndrome. These models utilize non-invasive factors, making them a cost-effective option for large-scale screening. By taking into account variables such as gender, Age, WHtR, SBP, DBP, WC, FBG, HC, and TGs, these models can pinpoint individuals at high risk. Additionally, they can analyze health records, lifestyle factors, and medical history to provide personalized risk assessments. With a focus on accuracy, these models allow for timely intervention and customized treatment plans for individuals at risk of metabolic syndrome [48, 49]. Future investigation is necessary to further explore the integration of machine learning models into healthcare systems in Bangladesh for early detection and intervention strategies.

Recommendations

It is recommended that expanding on the importance of giving higher priority to older individuals in preventing Metabolic Syndrome (MetS), it is essential to recognize that this population is more susceptible to developing MetS due to age-related physiological changes and lifestyle factors. By prioritizing older individuals, healthcare centers can focus on providing targeted interventions and preventive measures to reduce the risk of MetS. Furthermore, it is crucial to educate females about the risk factors associated with MetS, as they are also more prone to developing this condition. Women often experience hormonal changes throughout their lives, such as during pregnancy and menopause, which can contribute to the development of MetS. By raising awareness among females, healthcare centers can empower them to make informed decisions regarding their health and take necessary steps to prevent MetS [50].

Overall, prioritizing older individuals and providing targeted education on MetS risk factors, with a special emphasis on WHtR, can significantly contribute to the prevention and management of this condition. By promoting awareness and encouraging individuals to be conscious of their weight, WC, and HC, healthcare centers can empower individuals to take control of their health and reduce the burden of MetS in the population [50, 51].

Conclusion

To summarize, our research successfully estimated the prevalence of MetS in Bangladesh, which was found to be 27.8%. Notably, the female population showed a higher prevalence of MetS. Through the implementation of Random Forest and Chi-Square methods, we identified several potential risk factors for MetS, including increased Age, WHtR, SBP, DBP, WC, FBG, HC, and TGs.These models can be utilized to assess the influencing factors of various diseases and contribute to treatment strategies and decision-making processes.

Furthermore, it is crucial to recognize MetSasa global epidemic. In order to limit its further spread and reduce associated morbidity and mortality, it is imperative to implement regional measures that prioritize primary prevention.

Limitations and strength

Limitations

There is a limitation that necessitates attention in our study.

Exclusion of potentially relevant variables

Certain variables were excluded from the analysis due to their insignificant results. Moreover, the performance of our model is not particularly high. Additionally, the significance of BMI, a crucial factor for abdominal obesity, was not observed in our research. Lastly, HDL, an important factor for cardiovascular disease, also exhibited insignificance in our study [52].

Strength

Strengths of our study

Despite these limitations, the primary strength of this study lies in the extensive size of the dataset, which was obtained from the "National STEPS Survey for Non-communicable Diseases Risk Factors in Bangladesh 2018" survey. We employed six significant classification algorithms to predict our results. Notably, our study is the pioneering endeavor of its kind to forecast the prevalence and potential risk factors of MetS in Bangladesh.

Supporting information

S1 Data (ZIP)

We would like to express our sincere appreciation for obtaining the dataset from the "National STEPS Survey for Non-communicable Diseases Risk Factors in Bangladesh 2018".

10.1371/journal.pone.0309869.r001
Decision Letter 0
McGrowder Donovan Anthony Academic Editor
© 2024 Donovan Anthony McGrowder
2024
Donovan Anthony McGrowder
https://creativecommons.org/licenses/by/4.0/ This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.
Submission Version0
8 Mar 2024

PONE-D-24-05709

Metabolic Syndrome Predictive Modeling in Bangladesh applying Machine Learning Approach

PLOS ONE

Dear Dr. Hossain,

Thank you for submitting your manuscript to PLOS ONE. After careful consideration, we feel that it has merit but does not fully meet PLOS ONE’s publication criteria as it currently stands. Therefore, we invite you to submit a revised version of the manuscript that addresses the points raised during the review process.

Please submit your revised manuscript by Apr 22 2024 11:59PM. If you will need more time than this to complete your revisions, please reply to this message or contact the journal office at plosone@plos.org. When you're ready to submit your revision, log on to https://www.editorialmanager.com/pone/ and select the 'Submissions Needing Revision' folder to locate your manuscript file.

Please include the following items when submitting your revised manuscript:

A rebuttal letter that responds to each point raised by the academic editor and reviewer(s). You should upload this letter as a separate file labeled 'Response to Reviewers'.

A marked-up copy of your manuscript that highlights changes made to the original version. You should upload this as a separate file labeled 'Revised Manuscript with Track Changes'.

An unmarked version of your revised paper without tracked changes. You should upload this as a separate file labeled 'Manuscript'.

If you would like to make changes to your financial disclosure, please include your updated statement in your cover letter. Guidelines for resubmitting your figure files are available below the reviewer comments at the end of this letter.

If applicable, we recommend that you deposit your laboratory protocols in protocols.io to enhance the reproducibility of your results. Protocols.io assigns your protocol its own identifier (DOI) so that it can be cited independently in the future. For instructions see: https://journals.plos.org/plosone/s/submission-guidelines#loc-laboratory-protocols. Additionally, PLOS ONE offers an option for publishing peer-reviewed Lab Protocol articles, which describe protocols hosted on protocols.io. Read more information on sharing protocols at https://plos.org/protocols?utm_medium=editorial-email&utm_source=authorletters&utm_campaign=protocols.

We look forward to receiving your revised manuscript.

Kind regards,

Donovan Anthony McGrowder, PhD., MA., MSc

Academic Editor

PLOS ONE

Journal Requirements:

When submitting your revision, we need you to address these additional requirements.

1. Please ensure that your manuscript meets PLOS ONE's style requirements, including those for file naming. The PLOS ONE style templates can be found at 

https://journals.plos.org/plosone/s/file?id=wjVg/PLOSOne_formatting_sample_main_body.pdf and 

https://journals.plos.org/plosone/s/file?id=ba62/PLOSOne_formatting_sample_title_authors_affiliations.pdf

2. Please note that PLOS ONE has specific guidelines on code sharing for submissions in which author-generated code underpins the findings in the manuscript. In these cases, all author-generated code must be made available without restrictions upon publication of the work. Please review our guidelines at https://journals.plos.org/plosone/s/materials-and-software-sharing#loc-sharing-code and ensure that your code is shared in a way that follows best practice and facilitates reproducibility and reuse.

3. We note that you have indicated that there are restrictions to data sharing for this study. PLOS only allows data to be available upon request if there are legal or ethical restrictions on sharing data publicly. For more information on unacceptable data access restrictions, please see http://journals.plos.org/plosone/s/data-availability#loc-unacceptable-data-access-restrictions. 

Before we proceed with your manuscript, please address the following prompts:

a) If there are ethical or legal restrictions on sharing a de-identified data set, please explain them in detail (e.g., data contain potentially identifying or sensitive patient information, data are owned by a third-party organization, etc.) and who has imposed them (e.g., a Research Ethics Committee or Institutional Review Board, etc.). Please also provide contact information for a data access committee, ethics committee, or other institutional body to which data requests may be sent.

b) If there are no restrictions, please upload the minimal anonymized data set necessary to replicate your study findings to a stable, public repository and provide us with the relevant URLs, DOIs, or accession numbers. For a list of recommended repositories, please see

https://journals.plos.org/plosone/s/recommended-repositories. You also have the option of uploading the data as Supporting Information files, but we would recommend depositing data directly to a data repository if possible.

We will update your Data Availability statement on your behalf to reflect the information you provide.

4. When completing the data availability statement of the submission form, you indicated that you will make your data available on acceptance. We strongly recommend all authors decide on a data sharing plan before acceptance, as the process can be lengthy and hold up publication timelines. Please note that, though access restrictions are acceptable now, your entire data will need to be made freely accessible if your manuscript is accepted for publication. This policy applies to all data except where public deposition would breach compliance with the protocol approved by your research ethics board. If you are unable to adhere to our open data policy, please kindly revise your statement to explain your reasoning and we will seek the editor's input on an exemption. Please be assured that, once you have provided your new statement, the assessment of your exemption will not hold up the peer review process.

5. PLOS requires an ORCID iD for the corresponding author in Editorial Manager on papers submitted after December 6th, 2016. Please ensure that you have an ORCID iD and that it is validated in Editorial Manager. To do this, go to ‘Update my Information’ (in the upper left-hand corner of the main menu), and click on the Fetch/Validate link next to the ORCID field. This will take you to the ORCID site and allow you to create a new iD or authenticate a pre-existing iD in Editorial Manager. Please see the following video for instructions on linking an ORCID iD to your Editorial Manager account: https://www.youtube.com/watch?v=_xcclfuvtxQ

6. Your ethics statement should only appear in the Methods section of your manuscript. If your ethics statement is written in any section besides the Methods, please move it to the Methods section and delete it from any other section. Please ensure that your ethics statement is included in your manuscript, as the ethics statement entered into the online submission form will not be published alongside your manuscript. 

7. We note that Figure 1 in your submission contain map/satellite images which may be copyrighted. All PLOS content is published under the Creative Commons Attribution License (CC BY 4.0), which means that the manuscript, images, and Supporting Information files will be freely available online, and any third party is permitted to access, download, copy, distribute, and use these materials in any way, even commercially, with proper attribution. For these reasons, we cannot publish previously copyrighted maps or satellite images created using proprietary data, such as Google software (Google Maps, Street View, and Earth). For more information, see our copyright guidelines: http://journals.plos.org/plosone/s/licenses-and-copyright.

We require you to either (a) present written permission from the copyright holder to publish these figures specifically under the CC BY 4.0 license, or (b) remove the figures from your submission:

a. You may seek permission from the original copyright holder of Figure 1 to publish the content specifically under the CC BY 4.0 license.  

We recommend that you contact the original copyright holder with the Content Permission Form (http://journals.plos.org/plosone/s/file?id=7c09/content-permission-form.pdf) and the following text:

“I request permission for the open-access journal PLOS ONE to publish XXX under the Creative Commons Attribution License (CCAL) CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). Please be aware that this license allows unrestricted use and distribution, even commercially, by third parties. Please reply and provide explicit written permission to publish XXX under a CC BY license and complete the attached form.”

Please upload the completed Content Permission Form or other proof of granted permissions as an "Other" file with your submission.

In the figure caption of the copyrighted figure, please include the following text: “Reprinted from [ref] under a CC BY license, with permission from [name of publisher], original copyright [original copyright year].”

b. If you are unable to obtain permission from the original copyright holder to publish these figures under the CC BY 4.0 license or if the copyright holder’s requirements are incompatible with the CC BY 4.0 license, please either i) remove the figure or ii) supply a replacement figure that complies with the CC BY 4.0 license. Please check copyright information on all replacement figures and update the figure caption with source information. If applicable, please specify in the figure caption text when a figure is similar but not identical to the original image and is therefore for illustrative purposes only.

The following resources for replacing copyrighted map figures may be helpful:

USGS National Map Viewer (public domain): http://viewer.nationalmap.gov/viewer/

The Gateway to Astronaut Photography of Earth (public domain): http://eol.jsc.nasa.gov/sseop/clickmap/

Maps at the CIA (public domain): https://www.cia.gov/library/publications/the-world-factbook/index.html and https://www.cia.gov/library/publications/cia-maps-publications/index.html

NASA Earth Observatory (public domain): http://earthobservatory.nasa.gov/

Landsat: http://landsat.visibleearth.nasa.gov/

USGS EROS (Earth Resources Observatory and Science (EROS) Center) (public domain): http://eros.usgs.gov/#

Natural Earth (public domain): http://www.naturalearthdata.com/

Additional Editor Comments:

Dear Dr. Hossain,

Your manuscript “Metabolic Syndrome Predictive Modeling in Bangladesh applying Machine Learning Approach” has been assessed by our reviewers. They have raised a number of points which we believe would improve the manuscript and may allow a revised version to be published in PLOS ONE. Their reports, together with any other comments given are below.

If you are able to fully address these points, we would encourage you to submit a revised manuscript to PLOS ONE by the date given below.

 Best regards,

Dr. Donovan McGrowder

[Note: HTML markup is below. Please do not edit.]

Reviewers' comments:

Reviewer's Responses to Questions

Comments to the Author

1. Is the manuscript technically sound, and do the data support the conclusions?

The manuscript must describe a technically sound piece of scientific research with data that supports the conclusions. Experiments must have been conducted rigorously, with appropriate controls, replication, and sample sizes. The conclusions must be drawn appropriately based on the data presented.

Reviewer #1: Yes

Reviewer #2: Yes

Reviewer #3: Yes

**********

2. Has the statistical analysis been performed appropriately and rigorously?

Reviewer #1: Yes

Reviewer #2: Yes

Reviewer #3: Yes

**********

3. Have the authors made all data underlying the findings in their manuscript fully available?

The PLOS Data policy requires authors to make all data underlying the findings described in their manuscript fully available without restriction, with rare exception (please refer to the Data Availability Statement in the manuscript PDF file). The data should be provided as part of the manuscript or its supporting information, or deposited to a public repository. For example, in addition to summary statistics, the data points behind means, medians and variance measures should be available. If there are restrictions on publicly sharing data—e.g. participant privacy or use of data from a third party—those must be specified.

Reviewer #1: Yes

Reviewer #2: Yes

Reviewer #3: Yes

**********

4. Is the manuscript presented in an intelligible fashion and written in standard English?

PLOS ONE does not copyedit accepted manuscripts, so the language in submitted articles must be clear, correct, and unambiguous. Any typographical or grammatical errors should be corrected at revision, so please note any specific errors here.

Reviewer #1: Yes

Reviewer #2: Yes

Reviewer #3: Yes

**********

5. Review Comments to the Author

Please use the space provided to explain your answers to the questions above. You may also include additional comments for the author, including concerns about dual publication, research ethics, or publication ethics. (Please upload your review as an attachment if it exceeds 20,000 characters)

Reviewer #1: Abstract:

- The abstract should clearly state the main objectives, methodology, key results, conclusions and limitations of the study. Currently it only summarizes the background.

- Include the key risk factors identified, prevalence of MetS found, and main ML model utilized in the results.

Introduction:

- Expand more on the rationale for identifying risk factors and predicting MetS prevalence in the Bangladeshi population specifically.

- after ''and working in irregular shifts[7].'' discuss the role in particular of obstructive sleep apnea. cite doi:10.3390/life13030702.

- Provide more background on MetS diagnostic criteria and associated health risks to establish significance.

Methods:

- Describe the National STEPS Survey sampling methodology, data collection procedures, and variables obtained in more detail.

- Explain how MetS was defined - which diagnostic criteria used.

- Specify why certain variables were categorized and encoding approaches for ML models.

- State specific ML models, parameters, evaluation metrics, and statistical analysis done.

Results:

- Focus this section only on key results - prevalence, risk factors, model performance. Remove general discussion.

- Include quantitative performance metrics for each ML model - precision, recall, ROC, etc.

Discussion:

- Limitations should note small subset of survey data used, exclusion of potentially relevant variables like lipids. cite doi:10.1016/j.compbiomed.

- Discussion of inflammation markers and future studies seems speculative without any results presented. Would remove.

- Avoid overstating conclusions on predictive capabilities of models until externally validated.

Reviewer #2: Dear authors,

I have now completed the review of the manuscript titled "Metabolic Syndrome Predictive Modeling in Bangladesh applying Machine Learning Approach".

The author aims to predict the prevalence of Metabolic Syndrome (MetS) in Bangladesh using various machine learning algorithms. The study analyzes data from 8185 participants, identifying key risk factors for MetS and assessing the performance of different models in predicting the condition.

The manuscript is interesting and, in general, fairly well-written.

I have some suggestions to further improve the quality of the manuscript.

I would like to suggest that the authors address these limitations in the article, either by discussing them in the limitations section or, where feasible, by making the appropriate revisions:

1. It's crucial to assess how well the models generalize to unseen data. Cross-validation techniques and external validation on datasets from different populations could enhance the robustness of the findings.

2. The study could benefit from a direct comparison with existing predictive models for MetS to highlight its contributions or improvements.

3. While the study employs feature selection techniques, further exploration into dimensionality reduction methods could improve model performance and interpretability.

4. The exclusion of certain variables deemed insignificant could potentially overlook complex interactions between features that contribute to MetS. A more detailed analysis or justification for the exclusion of these variables might provide deeper insights.

5. The practical applicability of the models in clinical settings is not extensively discussed. Future work could focus on how these models can be integrated into healthcare systems for early detection and intervention strategies.

6. Mention breifly deep learning methods article and add in the reference. For example: "An Adaptive Ensemble Deep Learning Framework for Reliable Detection of Pandemic Patients", "Deep Learning Network Selection and Optimized Information Fusion for Enhanced COVID-19 Detection". Deep learning could offer insight into the future research direction.

Thank you for your valuable contributions to our field of research. I look forward to receiving the revised manuscript.

Reviewer #3: In this study, Hossain et al. perform machine learning to identify risk factors for metabolic syndrome. The topic is interesting, and data support the authors’ conclusions. The story is clear; however, the method section is incomplete. The authors write general information of procedures, such as what sensitivity is or how specificity is calculated. This basic and general information is not necessary. The authors should write procedures the authors performed in this study, but no information is provided in this manuscript, which is not appropriate. The authors should provide more detailed information of procedures, such as R and its version, packages, parameters used in analysis, etc.

**********

6. PLOS authors have the option to publish the peer review history of their article (what does this mean?). If published, this will include your full peer review and any attached files.

If you choose “no”, your identity will remain anonymous but your review may still be made public.

Do you want your identity to be public for this peer review? For information about this choice, including consent withdrawal, please see our Privacy Policy.

Reviewer #1: Yes: Antonino

Reviewer #2: No

Reviewer #3: No

**********

[NOTE: If reviewer comments were submitted as an attachment file, they will be attached to this email and accessible via the submission site. Please log into your account, locate the manuscript record, and check for the action link "View Attachments". If this link does not appear, there are no attachment files.]

While revising your submission, please upload your figure files to the Preflight Analysis and Conversion Engine (PACE) digital diagnostic tool, https://pacev2.apexcovantage.com/. PACE helps ensure that figures meet PLOS requirements. To use PACE, you must first register as a user. Registration is free. Then, login and navigate to the UPLOAD tab, where you will find detailed instructions on how to use the tool. If you encounter any issues or have any questions when using PACE, please email PLOS at figures@plos.org. Please note that Supporting Information files do not need this step.

10.1371/journal.pone.0309869.r002
Author response to Decision Letter 0
Submission Version1
26 Apr 2024

Reviewer 1:

Reviewer comments Responses

The abstract should clearly state the main objectives, methodology, key results, conclusions, and limitations of the study. Currently it only summarizes the background. We appreciate your feedback. The primary goals, methodology, important findings, conclusions, and limitations of the study have all been spelled out in detail in the abstract.

Include the key risk factors identified, prevalence of MetS found, and main ML model utilized in the results. Thank you for your recommendations. The major machine learning model used in the results, the prevalence of MetS, and the key risk variables revealed have all been presented in the "abstract".

Expand more on the rationale for identifying risk factors and predicting MetS prevalence in the Bangladeshi population specifically. We appreciate your recommendations. The reasoning behind determining risk variables and forecasting the incidence of MetS in the Bangladeshi population in particular is covered in greater detail in the "Introduction section" of our publication.

after ‘and working in irregular shifts [7].'discuss the role in particular of obstructive sleep apnea. cite doi:10.3390/life13030702. Thanks for your suggestions. According to this research doi:10.3390/life13030702, we have addressed the role of obstructive sleep apnea specifically in the "Introduction section."

Provide more background on MetS diagnostic criteria and associated health risks to establish significance. We appreciate your recommendations. To establish significance, we have included further background information on MetS diagnostic criteria and related health hazards in the "Introduction section" of our publication.

Describe the National STEPS Survey sampling methodology, data collection procedures, and variables obtained in more detail. Thank you for your recommendations. More information on the National STEPS Survey sampling methodology, data collection techniques, and variables gathered may be found in the section titled "Data Source in the Methodology section."

Explain how MetS was defined - which diagnostic criteria used Thanks for your suggestionsThe definition of MetS and the diagnostic criteria that were applied are covered in the section on "diagnostic criteria in the Methodology section."

Specify why certain variables were categorized and encoding approaches for ML models We have explained why specific variables were categorized and encoding methodologies for ML models in the section "Variable importance and key risk factors of metabolic syndrome in the Results and analysis section."

State specific ML models, parameters, evaluation metrics, and statistical analysis done Thanks for your suggestions. Certain machine learning models, parameters, assessment metrics, and statistical analysis have been specified in the "Results and analysis" portion of Tables 2, 3, and 4.

Focus this section only on key results - prevalence, risk factors, model performance. Remove general discussion We appreciate your recommendations. We only included the most important findings—prevalence, risk variables, and model performance—in the "Results and analysis" section.

Include quantitative performance metrics for each ML model - precision, recall, ROC, etc. Thank you for your tremendous recommendations. We incorporated numerical performance indicators, such as precision, recall, ROC, and so on, for every machine learning model in our manuscript's "Results and analysis section in Table 4."

Limitations should note small subset of survey data used, exclusion of potentially relevant variables like lipids. cite doi:10.1016/j.compbiomed We appreciate your recommendations. We have mentioned in the "Limitations Section" that a tiny subset of survey data was utilized, that potentially relevant factors like lipids were excluded, and that we cited this article doi:10.1016/j.compbiomed.

Discussion of inflammation markers and future studies seems speculative without any results presented. Would remove. We appreciate your insightful recommendations. The part about future research and signs of inflammation has been removed.

Avoid overstating conclusions on predictive capabilities of models until externally validated Thanks for your valuable suggestions. This has been deleted from the "Conclusion section."

Reviewer #2

Reviewer comments Responses

It's crucial to assess how well the models generalize to unseen data. Cross-validation techniques and external validation on datasets from different populations could enhance the robustness of the findings. Your insightful recommendations are appreciated. After examining datasets from various populations, we have conducted cross-validation and external validation on our models and found that they perform well.

The study could benefit from a direct comparison with existing predictive models for MetS to highlight its contributions or improvements We appreciate your insightful recommendations. There are no current MetS prediction models available that we can compare our model to from the standpoint of the Bangladeshi population. Our models showed marginally less accuracy when we compared them to the state-of-the-art models from China and Mexico.

While the study employs feature selection techniques, further exploration into dimensionality reduction methods could improve model performance and interpretability. Thanks for your valuable suggestions. Although we have looked into dimensionality reduction techniques, our model's performance remains unchanged.

The exclusion of certain variables deemed insignificant could potentially overlook complex interactions between features that contribute to MetS. A more detailed analysis or justification for the exclusion of these variables might provide deeper insights. Table 5 in the "Results and analysis section" illustrates how these variables lower the models' overall performance. Because of this, we have decided to eliminate these factors from our study because they don't really matter for our findings. We think more research is necessary to ignore the intricate relationships between the characteristics that cause MetS.

The practical applicability of the models in clinical settings is not extensively discussed. Future work could focus on how these models can be integrated into healthcare systems for early detection and intervention strategies We appreciate your recommendations. The "Discussion and recommendations section" of our publication contains a brief discussion of the models' usefulness in clinical settings. To further understand how machine learning models might be incorporated into Bangladeshi healthcare systems for early detection and intervention measures, we think more research is needed.

Mention briefly deep learning methods article and add in the reference. For example: "An Adaptive Ensemble Deep Learning Framework for Reliable Detection of Pandemic Patients", "Deep Learning Network Selection and Optimized Information Fusion for Enhanced COVID-19 Detection". Deep learning could offer insight into the future research direction. Thanks for your important suggestions. In the "Related work section" of our manuscript, we have included a brief reference of an article on deep learning that may provide some light on potential future study directions.

Reviewer #3

Reviewer comments Responses

This basic and general information is not necessary. The authors should write procedures the authors performed in this study, but no information is provided in this manuscript, which is not appropriate. The authors should provide more detailed information of procedures, such as R and its version, packages, parameters used in analysis, etc. Thanks for your suggestions. We removed the general discussion as possible. We included the author’s contribution in this manuscript. We provided detailed information of procedures. We used SPSS Python instead of R. We included its version, libraries and parameter used in analysis.

Attachment Submitted filename: Response to Reviewer.pdf

10.1371/journal.pone.0309869.r003
Decision Letter 1
Klisic Aleksandra Academic Editor
© 2024 Aleksandra Klisic
2024
Aleksandra Klisic
https://creativecommons.org/licenses/by/4.0/ This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.
Submission Version1
13 Aug 2024

Metabolic Syndrome Predictive Modeling in Bangladesh applying Machine Learning Approach

PONE-D-24-05709R1

Dear Dr. Hossain,

We’re pleased to inform you that your manuscript has been judged scientifically suitable for publication and will be formally accepted for publication once it meets all outstanding technical requirements.

Within one week, you’ll receive an e-mail detailing the required amendments. When these have been addressed, you’ll receive a formal acceptance letter and your manuscript will be scheduled for publication.

An invoice will be generated when your article is formally accepted. Please note, if your institution has a publishing partnership with PLOS and your article meets the relevant criteria, all or part of your publication costs will be covered. Please make sure your user information is up-to-date by logging into Editorial Manager at Editorial Manager® and clicking the ‘Update My Information' link at the top of the page. If you have any questions relating to publication charges, please contact our Author Billing department directly at authorbilling@plos.org.

If your institution or institutions have a press office, please notify them about your upcoming paper to help maximize its impact. If they’ll be preparing press materials, please inform our press team as soon as possible -- no later than 48 hours after receiving the formal acceptance. Your manuscript will remain under strict press embargo until 2 pm Eastern Time on the date of publication. For more information, please contact onepress@plos.org.

Kind regards,

Aleksandra Klisic

Academic Editor

PLOS ONE

Additional Editor Comments (optional):

Reviewers' comments:

Reviewer's Responses to Questions

Comments to the Author

1. If the authors have adequately addressed your comments raised in a previous round of review and you feel that this manuscript is now acceptable for publication, you may indicate that here to bypass the “Comments to the Author” section, enter your conflict of interest statement in the “Confidential to Editor” section, and submit your "Accept" recommendation.

Reviewer #1: All comments have been addressed

Reviewer #2: All comments have been addressed

Reviewer #3: All comments have been addressed

**********

2. Is the manuscript technically sound, and do the data support the conclusions?

The manuscript must describe a technically sound piece of scientific research with data that supports the conclusions. Experiments must have been conducted rigorously, with appropriate controls, replication, and sample sizes. The conclusions must be drawn appropriately based on the data presented.

Reviewer #1: Yes

Reviewer #2: Yes

Reviewer #3: (No Response)

**********

3. Has the statistical analysis been performed appropriately and rigorously?

Reviewer #1: Yes

Reviewer #2: Yes

Reviewer #3: (No Response)

**********

4. Have the authors made all data underlying the findings in their manuscript fully available?

The PLOS Data policy requires authors to make all data underlying the findings described in their manuscript fully available without restriction, with rare exception (please refer to the Data Availability Statement in the manuscript PDF file). The data should be provided as part of the manuscript or its supporting information, or deposited to a public repository. For example, in addition to summary statistics, the data points behind means, medians and variance measures should be available. If there are restrictions on publicly sharing data—e.g. participant privacy or use of data from a third party—those must be specified.

Reviewer #1: Yes

Reviewer #2: Yes

Reviewer #3: (No Response)

**********

5. Is the manuscript presented in an intelligible fashion and written in standard English?

PLOS ONE does not copyedit accepted manuscripts, so the language in submitted articles must be clear, correct, and unambiguous. Any typographical or grammatical errors should be corrected at revision, so please note any specific errors here.

Reviewer #1: Yes

Reviewer #2: Yes

Reviewer #3: (No Response)

**********

6. Review Comments to the Author

Please use the space provided to explain your answers to the questions above. You may also include additional comments for the author, including concerns about dual publication, research ethics, or publication ethics. (Please upload your review as an attachment if it exceeds 20,000 characters)

Reviewer #1: The manuscript Is improved and the Authors performed all the revisions required. I can suggest tò accept.

Reviewer #2: All comments were addressed. I declare no further comments to be included. Thank you for considering my opinion.

Reviewer #3: (No Response)

**********

7. PLOS authors have the option to publish the peer review history of their article (what does this mean?). If published, this will include your full peer review and any attached files.

If you choose “no”, your identity will remain anonymous but your review may still be made public.

Do you want your identity to be public for this peer review? For information about this choice, including consent withdrawal, please see our Privacy Policy.

Reviewer #1: Yes: None

Reviewer #2: No

Reviewer #3: No

**********
==== Refs
References

1 Samson SL , Garber AJ . Metabolic Syndrome. Endocrinol Metab Clin North Am. 2014 Mar;43 (1 ):1–23. doi: 10.1016/j.ecl.2013.09.009 24582089
2 World Health Organizations [Internet]. 2023. Diabetes.
3 National Heart, Lung, and Blood Institute [Internet]. 2022. METABOLIC SYNDROME.
4 Wasim Bin Habib , Mohammad Al-Masum Molla . Cardiovascular diseases: A rising concern for young people. The Daily Star. 2022 Sep 29;
5 Chowdhury MZI , Anik AM , Farhana Z , Bristi PD , Abu Al Mamun BM , Uddin MJ , et al . Prevalence of metabolic syndrome in Bangladesh: A systematic review and meta-analysis of the studies. BMC Public Health. 2018 Mar 2;18 (1 ). doi: 10.1186/s12889-018-5209-z 29499672
6 Naghipour M , Joukar F , Nikbakht HA , Hassanipour S , Asgharnezhad M , Arab-Zozani M , et al . High prevalence of metabolic syndrome and its related demographic factors in north of Iran: Results from the PERSIAN guilan cohort study. Int J Endocrinol. 2021;2021 . doi: 10.1155/2021/8862456 33859688
7 Belayneh M , Mekonnen TC , Tadesse SE , Amsalu ET , Tadese F . Sleeping duration, physical activity, alcohol drinking and other risk factors as potential attributes of metabolic syndrome in adults in Ethiopia: A hospital-based cross-sectional study. PLoS One. 2022 Aug 29;17 (8 ):e0271962. doi: 10.1371/journal.pone.0271962 36037175
8 Centers for Disease Control and Prevention [Internet]. 2023. Know Your Risk for High Cholesterol.
9 Seeber B , Morandell E , Lunger F , Wildt L , Dieplinger H . Afamin serum concentrations are associated with insulin resistance and metabolic syndrome in polycystic ovary syndrome. Reproductive Biology and Endocrinology. 2014 Dec 10;12 (1 ):88. doi: 10.1186/1477-7827-12-88 25208973
10 Stubert J , Reister F , Hartmann S , Janni W . The Risks Associated With Obesity in Pregnancy. DtschArztebl Int. 2018 Apr 20; doi: 10.3238/arztebl.2018.0276 29739495
11 Riaz BK , Islam MZ , Islam ANMS , Zaman MM , Hossain MA , Rahman MM , et al . Risk factors for non-communicable diseases in Bangladesh: Findings of the population-based cross-sectional national survey 2018. BMJ Open. 2020 Nov 27;10 (11 ). doi: 10.1136/bmjopen-2020-041334 33247026
12 Indicators K. Bangladesh Demographic and Health Survey 2017–18 [Internet]. 2019. Available from: http://www.niport.gov.bd;
13 Hossain MB , Khan MdN , Oldroyd JC , Rana J , Magliago DJ , Chowdhury EK , et al . Prevalence of, and risk factors for, diabetes and prediabetes in Bangladesh: Evidence from the national survey using a multilevel Poisson regression model with a robust variance. PLOS Global Public Health. 2022 Jun 1;2 (6 ):e0000461. doi: 10.1371/journal.pgph.0000461 36962350
14 Janiesch C , Zschech P , Heinrich K . Machine learning and deep learning. Available from: 10.1007/s12525-021-00475-2
15 Datta S , Schraplau A , Da Cruz HF , Sachs JP , Mayer F , Bottinger E . A machine learning approach for non-invasive diagnosis of metabolic syndrome. In: Proceedings—2019 IEEE 19th International Conference on Bioinformatics and Bioengineering, BIBE 2019. Institute of Electrical and Electronics Engineers Inc.; 2019. p. 933–40.
16 Ivanović D , Kupusinac A , Stokić E , Doroslovački R , Ivetić D . ANN Prediction of Metabolic Syndrome: a Complex Puzzle that will be Completed. J Med Syst. 2016 Dec 1;40 (12 ).
17 Gutiérrez-Esparza GO , Vázquez OI , Vallejo M , Hernández-Torruco J . Prediction of metabolic syndrome in a Mexican population applying machine learning algorithms. Symmetry (Basel). 2020 Apr 1;12 (4 ).
18 Xia SJ , Gao BZ , Wang SH , Guttery DS , Li CD , Zhang YD . Modeling of diagnosis for metabolic syndrome by integrating symptoms into physiochemical indexes. Biomedicine and Pharmacotherapy. 2021 May 1;137 . doi: 10.1016/j.biopha.2021.111367 33588265
19 Zhang H , Chen D , Shao J , Zou P , Cui N , Tang L , et al . Machine learning-based prediction for 4-year risk of metabolic syndrome in adults: A retrospective cohort study. Risk ManagHealthc Policy. 2021;14 :4361–8. doi: 10.2147/RMHP.S328180 34707419
20 Salim Hossain M , ZahedurRahaman M , Banik S , Shahid Sarwar M , Yokota K . PREVALENCE OF THE METABOLIC SYNDROME IN DIABETIC PATIENTS LIVING IN A COASTAL REGION OF BANGLADESH. IJPSR [Internet]. 2012;3 (8 ):8. Available from: www.ijpsr.com
21 Mehata S , Shrestha N , Mehta RK , Bista B , Pandey AR , Mishra SR . Prevalence of the Metabolic Syndrome and its determinants among Nepalese adults: Findings from a nationally representative cross-sectional study. Sci Rep. 2018 Dec 1;8 (1 ). doi: 10.1038/s41598-018-33177-5 30301902
22 Shah HA , Liu J , Yang Z , Zhang X , Feng J . DeepRF: A deep learning method for predicting metabolic pathways in organisms based on annotated genomes. Comput Biol Med. 2022 Aug;147 :105756. doi: 10.1016/j.compbiomed.2022.105756 35759992
23 Janssen I , Katzmarzyk PT , Ross R . Body Mass Index, Waist Circumference, and Health Risk Evidence in Support of Current National Institutes of Health Guidelines [Internet]. Available from: http://archinte.jamanetwork.com/ doi: 10.1001/archinte.162.18.2074 12374515
24 Chobanian Aram V. , Bakris George L. , Black Henry R. , Cushman William C. , Green Lee A. , IzzoJr Joseph L. , et al . Seventh Report of the Joint National Committee on Prevention, Detection, Evaluation, and Treatment of High Blood Pressure. AHA/ASA Journals. 2003 Dec 1;42 (6 ).
25 Executive Summary of the Third Report of the National Cholesterol Education Program (NCEP) Expert Panel on Detection, Evaluation, and Treatment of High Blood Cholesterol in Adults (Adult Treatment Panel III) Expert Panel on Detection, Evaluation, and Treatment of High Blood Cholesterol in Adults T HE THIRD REPORT OF THE EX-pert Panel on Detection, Evaluation, and Treatment of High Blood Cholesterol in Adults (Adult Treatment Panel III, or ATP III) constitutes the National [Internet]. Available from: http://jama.jamanetwork.com/
26 Grundy SM , Brewer HB , Cleeman JI , Smith SC , Lenfant C . Definition of Metabolic Syndrome: Report of the National Heart, Lung, and Blood Institute/American Heart Association Conference on Scientific Issues Related to Definition. In: Circulation. 2004. p. 433–8. doi: 10.1161/01.CIR.0000111245.75752.C6 14744958
27 Java Point [Internet]. Decision Tree Classification Algorithm.
28 Song YY , Lu Y . Decision tree methods: applications for classification and prediction. Shanghai Arch Psychiatry. 2015 Apr 1;27 (2 ):130–5. doi: 10.11919/j.issn.1002-0829.215044 26120265
29 Breiman L. Random Forests. Vol. 45 . 2001.
30 Strobl C , Boulesteix AL , Zeileis A , Hothorn T . Bias in random forest variable importance measures: Illustrations, sources and a solution. BMC Bioinformatics. 2007;8 . doi: 10.1186/1471-2105-8-25 17254353
31 Vijay Kanade . What Is a Support Vector Machine? Working, Types, and Examples. 2002.
32 Ogunleye AA , Qing-Guo W . XGBoost Model for Chronic Kidney Disease Diagnosis.
33 Triguero I , García-Gil D , Maillo J , Luengo J , García S , Herrera F . Transforming big data into smart data: An insight on the use of the k-nearest neighbors algorithm to obtain quality data. Vol. 9 , Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery. Wiley-Blackwell; 2019.
34 Salma Ghoneim. Medium. 2019. Accuracy, Recall, Precision, F-Score & Specificity, which to optimize on?
35 Erickson BJ , Kitamura F . Magician’s corner: 9. performance metrics for machine learning models. Vol. 3 , Radiology: Artificial Intelligence. Radiological Society of North America Inc.; 2021. doi: 10.1148/ryai.2021200126 34136815
36 Namrata Kapoor. Numpy Ninja. 2021. Recall, Specificity, Precision, F1 Scores and Accuracy.
37 Moy FM , Bulgiba A . The modified NCEP ATP III criteria maybe better than the IDF criteria in diagnosing Metabolic Syndrome among Malays in Kuala Lumpur. BMC Public Health. 2010 Dec 6;10 (1 ):678.21054885
38 Riccardo Di Sipio. Medium. 2021. The Definitive Way to Deal With Continuous Variables in Machine Learning.
39 Kim J , Mun S , Lee S , Jeong K , Baek Y . Prediction of metabolic and pre-metabolic syndromes using machine learning models with anthropometric, lifestyle, and biochemical factors from a middle-aged population in Korea. BMC Public Health. 2022 Dec 1;22 (1 ). doi: 10.1186/s12889-022-13131-x 35387629
40 Marbou WJT , Kuete V . Prevalence of metabolic syndrome and its components in Bamboutos division’s adults, West Region of Cameroon. Biomed Res Int. 2019;2019 . doi: 10.1155/2019/9676984 31183378
41 Lin CS , Lee WJ , Lin SY , Lin HP , Chen RC , Lin CH , et al . Subtypes of Premorbid Metabolic Syndrome and Associated Clinical Outcomes in Older Adults. Front Med (Lausanne). 2022 Feb 11;8 . doi: 10.3389/fmed.2021.698728 35223876
42 Jesmin S , Islam S , Akter S , Islam M , Nusrat Sultana S , Yamaguchi N , et al . Metabolic syndrome among pre-and post-menopausal rural women in Bangladesh: result from a population-based study [Internet]. 2013. Available from: http://www.biomedcentral.com/1756-0500/6/157 doi: 10.1186/1756-0500-6-157 23597398
43 Liang X , Or B , Fung Tsoi M , Lung Cheung C , Cheung BM , Bernard Cheung CM . Prevalence of Metabolic Syndrome in the United States National Health and Nutrition Examination Survey (NHANES) 2011–2018. Available from: 10.1101/2021.04.21.21255850
44 Worachartcheewan A , Nantasenamat C , Isarankura-Na-Ayudhya C , Pidetcha P , Prachayasittikul V . Identification of metabolic syndrome using decision tree analysis. Diabetes Res Clin Pract. 2010 Oct;90 (1 ):e15–8. doi: 10.1016/j.diabres.2010.06.009 20619912
45 Chen MS , Chiu CH , Chen SH . Risk assessment of metabolic syndrome prevalence involving sedentary occupations and socioeconomic status. BMJ Open. 2021 Dec 13;11 (12 ):e042802. doi: 10.1136/bmjopen-2020-042802 34903529
46 Tamrakar R , Yang X , Pradhan S , Su X , Luo Z , Li L , et al . Machine learning methods for the prediction of prevalence and potential risk factors of Metabolic Syndrome in Guangxi, China [Internet]. Available from: https://ssrn.com/abstract=4341038
47 Binh TQ , Phuong PT , Nhung BT , Tung DD . Metabolic syndrome among a middle-aged population in the red river delta region of Vietnam. BMC EndocrDisord. 2014 Sep 26;14 (1 ). doi: 10.1186/1472-6823-14-77 25261978
48 Xu W , Zhang Z , Hu K , Fang P , Li R , Kong D , et al . Identifying metabolic syndrome easily and cost effectively using non-invasive methods with machine learning models. Diabetes, Metabolic Syndrome and Obesity. 2023;16 :2141–51. doi: 10.2147/DMSO.S413829 37484515
49 Shin H , Shim S , Oh S . Machine learning-based predictive model for prevention of metabolic syndrome. PLoS One. 2023 Jun 1;18 (6 June). doi: 10.1371/journal.pone.0286635 37267302
50 Lin YH , Chu LL , Kao CC , Chen TB , Lee I , Li HC . The Effects of a Diet and Exercise Program for Older Adults With Metabolic Syndrome. Journal of Nursing Research. 2015 Sep;23 (3 ):197–205. doi: 10.1097/jnr.0000000000000078 25741965
51 Ashwell M , Gunn P , Gibson S . Waist-to-height ratio is a better screening tool than waist circumference and BMI for adult cardiometabolic risk factors: Systematic review and meta-analysis. Vol. 13 , Obesity Reviews. 2012. p. 275–86. doi: 10.1111/j.1467-789X.2011.00952.x 22106927
52 Wu J , Zhou X , Ren J , Zhang Z , Ju H , Diao X , et al . Glycosyltransferase-related prognostic and diagnostic biomarkers of uterine corpus endometrial carcinoma. Comput Biol Med. 2023 Sep;163 :107164. doi: 10.1016/j.compbiomed.2023.107164 37329616
