
==== Front
BMC Pulm Med
BMC Pulm Med
BMC Pulmonary Medicine
1471-2466
BioMed Central London

3252
10.1186/s12890-024-03252-x
Research
Interpretable mortality prediction model for ICU patients with pneumonia: using shapley additive explanation method
Li Jiaxi 1
Zhang Yu 2
He ShengYang 1
Tang Yan 327194072@qq.com

1
1 Department of Clinical Laboratory Medicine, Jinniu Maternity and Child Health Hospital of Chengdu, Chengdu, China
2 grid.13291.38 0000 0001 0807 1581 Information Center, West China Hospital, Sichuan University, Chengdu, China
13 9 2024
13 9 2024
2024
24 44715 12 2023
29 8 2024
© The Author(s) 2024
2024
https://creativecommons.org/licenses/by-nc-nd/4.0/ Open Access This article is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License, which permits any non-commercial use, sharing, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if you modified the licensed material. You do not have permission under this licence to share adapted material derived from this article or parts of it. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by-nc-nd/4.0/.
Background

Pneumonia, a leading cause of morbidity and mortality worldwide, often necessitates Intensive Care Unit (ICU) admission. Accurate prediction of pneumonia mortality is crucial for tailored prevention and treatment plans. However, existing mortality prediction models face limited adoption in clinical practice due to their lack of interpretability.

Objective

This study aimed to develop an interpretable model for predicting pneumonia mortality in ICUs. Leveraging the Shapley Additive Explanation (SHAP) method, we sought to elucidate the Extreme Gradient Boosting (XGBoost) model and identify prognostic factors for pneumonia.

Methods

Conducted as a retrospective cohort study, we utilized electronic health records from the eICU-CRD (2014–2015) for all adult pneumonia patients. The first 24 h of each ICU admission records were considered, with 70% of the dataset allocated for model training and 30% for validation. The XGBoost model was employed, and performance was assessed using the area under the receiver operating characteristic curve (AUC). The SHAP method provided insights into the XGBoost model.

Results

Among 10,962 pneumonia patients, in-hospital mortality was 16.33%. The XGBoost model demonstrated superior predictive performance (AUC: 0.778 ± 0.016)) compared to traditional scoring systems and other machine learning method, which achieved an improvement of 10% points. SHAP analysis identified Aspartate Aminotransferase (AST) as the most crucial predictor.

Conclusions

Interpretable predictive models enhance mortality risk assessment for pneumonia patients in the ICU, fostering transparency. AST emerged as the foremost predictor, followed by patient age, albumin, BMI et al. These insights, rooted in strong correlations with mortality, facilitate improved clinical decision-making and resource allocation.

Keywords

Pneumonia
Interpretable model
Machine learning
ICU
Chengdu Medical Research Project2023195 issue-copyright-statement© BioMed Central Ltd., part of Springer Nature 2024
==== Body
pmcIntroduction

Pneumonia is a common reason for patients to be admitted to the Intensive Care Unit (ICU) [1]. The probability of pneumonia patients experiencing in-hospital mortality is around 17% [2]. Making a complete and accurate prediction and evaluation of in-hospital mortality risk for patients using data from Electronic Health Records (EHR) is an urgent task in medicine, particularly for those admitted to the ICU. However, in real-life situations, it is not feasible to make a quick and effective diagnosis for pneumonia patients, considering factors such as the complexity of the patient’s condition [3]. Therefore, it is meaningful to establish an accurate and efficient predictive model to assist doctors’ decision-making.

Traditional methods for assessing mortality risk, such as the Apache score [4] and Sofa score [5], have their respective advantages and limitations in early diagnosis. These methods rely heavily on doctors’ close attention and long-term clinical experience. In recent years, artificial intelligence has been widely used to explore various disease warning and prediction factors [6–8]. Machine learning algorithms have powerful functions in capturing nonlinear relationships, and more research advocates the use of new predictive models based on machine learning to support clinical treatment of patients [9].

The purpose of this study is to develop an interpretable model based on the free and open eICU Collaborative Research Database (eICU-CRD) [10] to predict the risk of mortality in ICU pneumonia patients. Additionally, the Shapley Additive Explanation (SHAP) [11] algorithm is used to explore the prognostic factors of pneumonia by explaining the XGBoost [12] (extreme gradient boosting) predictive model.

Methods

Data source

The eICU Collaborative Research Database (version 2.0) is a publicly available multicenter database [10], containing deidentified data from over 200,000 ICU admissions across 208 hospitals in the United States from 2014 to 2015. This database provides detailed demographic information, vital sign measurements, diagnostic information, laboratory test data and treatment information for all patients.

Subject

This study is based on the publicly available eICU dataset. In addition to using ICD-9-CM codes (481–487) to filter patients, we simultaneously employed the diagnostic keyword ‘Pneumonia’ as a screening criterion to ensure comprehensive and accurate patient selection [13]. Besides, aged 18 years or older, and excluded if they had a hospital stay of less than 1 day or were missing gender information. Patients with severe immune suppression, such as human immunodeficiency virus (HIV) infection or a neutrophil count of less than (×109/ L) [14], were also excluded. Patients were divided into two groups, the survival group, and the Expired group, based on their discharge outcome.

During the patients’ ICU stay, we collected basic demographic and clinical information, including age, gender, BMI, Apache score, and length of ICU stay, for baseline analysis. We utilized the open-source PostgreSQL management and development tool (PgAdmin) [15] to extract initial laboratory test results, admission vital signs, and other relevant parameters recorded after ICU admission. For preliminary feature selection, we primarily relied on the LASSO (Least Absolute Shrinkage and Selection Operator) regression [16] method to identify important features. Additionally, features were manually selected based on clinical relevance and the judgment of the research team. Specifically, we excluded indicators like Apache score and GCS, which are closely related to the outcome variable, to avoid redundancy and overfitting. We also excluded height and weight, opting instead to use BMI values for a more streamlined analysis. The processed dataset, consisting of 67 features, was then used as input for the machine learning model. To prevent overfitting, the dataset was split into training and test sets in a 7:3 ratio, ensuring the model’s robustness and generalizability.

Missing Data Handling

Missing data variables are common in eICU-CRD. However, ignoring missing data in the analysis may lead to biased results. Therefore, we adopt a random forest regression imputation method [17] to predict missing values by utilizing the correlations between multiple variables. This approach enables us to reduce the impact of missing values on the model as much as possible in a scientific and accurate manner.

Dealing with Imbalanced datasets

ADASYN (Adaptive Synthetic Sampling) [18] is a machine learning algorithm used to address class imbalance in classification problems. In imbalanced datasets, there is a large difference in the number of samples between different classes, which can result in poor classification performance of the model for minority classes. ADASYN balances the dataset by increasing the number of samples of the minority class, thus improving the model’s performance. Considering the imbalance between the survival group and the expired group and the impact on the accuracy of the prediction results, we applied the ADASYN method to oversample the training set, while the test set maintained the original sample ratio.

Additionally, we adopted the SMOTE-ENN (Synthetic Minority Over-sampling Technique combined with Edited Nearest Neighbors) [19] method to perform a comparative analysis of the results. SMOTE-ENN combines two methods: SMOTE, which balances the dataset by generating synthetic samples for the minority class, and Edited Nearest Neighbors, which cleans the dataset by removing misclassified samples. By comparing the performance of the models trained with ADASYN and SMOTE-ENN, we aimed to identify the most effective method for handling class imbalance in our dataset.

Interpretability of machine learning

The prediction model’s explanation is achieved through SHAP, a unified method that accurately calculates the contribution and impact of each feature on the final prediction. The SHAP value can display the degree of contribution of each predictive variable to the target variable, whether it is a positive or negative contribution. Furthermore, SHAP values can explain each observation in the dataset through a specific set of SHAP values. [20].

Numerical analysis and machine learning models

The statistical analysis and calculations in this study were performed using R software and Python version 3.8.0 [21]. Categorical variables are represented by total number and percentage, and intergroup differences are compared using the chi-square test. Continuous variables are represented by mean value and standard deviation, and differences between two groups are compared using the Wilcoxon rank-sum test [22].

Four machine learning models (XGBoost, logistic regression [23], random forest [24], and support vector machine [25]) were used in this study to develop the prediction model. The predictive performance of each model was evaluated using the area under the receiver operating characteristic curve. In addition, considering the recent popularity of neural network models, we have also exploited a multi-layer perceptron model (MLP), aiming to enhance the classification performance of the data [26]. We calculated accuracy, sensitivity, positive predictive value, negative predictive value, and F1 score for all test results.

Results

Characteristics of the data

The eICU-CRD contained 16,146 patients with pneumonia. After applying the inclusion criteria, 10,962 adult patients were eligible for the study. The patient screening process is depicted in Fig. 1.

Fig. 1 Patient inclusion and exclusion process

Table 1 presents a comparison of baseline information between the survival group and the non-survival group, and the mortality rate of severe pneumonia patients was 16.33% (1790/10962). Patients in the Non-survival group were older and had higher Apache scores than those in the survival group, with more significant fluctuations, and ICU length of stay was also increased (P < 0.001).

Table 1 Baseline of patients with pneumonia

Variable	Level	Survival (0)
N = 9172	Non-survival (1)
N = 1790	P
< 0.001	
Age (Mean (SD))	/	65.697(16.007)	71.446(13.684)	< 0.001	
Gender (%)	0	4343(47.35)	799(44.64)	0.0377	
1	4829(52.65)	991(55.36)	
Underlying disease (%)	
CKD	0	7927(86.43)	1488(83.13)	0.0003	
	1	1245(13.57)	302(16.87)	
Diabetes	0	8897(97.00)	1488(83.13)	0.299	
	1	275(3.00)	45(2.51)	
Copd	0	7208(78.59)	1459(81.51)	< 0.001	
	1	1964(21.41)	331(18.49)	
Sepsis	0	6758(73.68)	1227(68.55)	< 0.001	
	1	2414(26.22)	563(31.45)	
Cerebrovascular	0	9042(98.58)	1739(97.15)	< 0.001	
	1	130(1.42)	51(2.85)	
Cardiac	0	5124(55.87)	812(45.36)	< 0.001	
	1	4048(44.13)	978(54.64)	
BMI (Mean (SD))	/	27.41(5.82)	26.43(5.782)	< 0.001	
Apache Score	/	62.517 (23.930)	82.737 (30.688)	< 0.001	
The length of ICU stay(days)	/	5.084 (5.667)	6.869 (7.720)	< 0.001	

SMOTE-ENN and ADASYN

To address the issue of sample imbalance, we applied SMOTE-ENN and ADASYN to the training data, leading to a revised distribution of the training samples. Initially, the training set consisted of 6,433 survival patients and 1,240 non-survival patients. After applying SMOTE-ENN, the ratio of survival to non-survival patients changed to 3,389:5,723. With ADASYN, the ratio was adjusted to 6,433:6,170. we employed a 5-fold cross-validation methodology utilizing the XGBoost algorithm to rigorously assess the model’s performance, benchmarking its performance against a suite of metrics including the Area Under the Curve (AUC), accuracy, mean average precision (mAP), Balanced Accuracy, and the F1 Score., as shown in Table 2. A two-sample t-test was conducted to statistically assess these performance metrics, yielding p-values of less than 0.01, indicating that the differences between ADASYN and SMOTE-ENN were statistically significant. Overall, SMOTE-ENN demonstrated superior performance in handling the imbalanced dataset in our experiments.

Table 2 The outcomes of SMOTE-ENN and ADASYN

Method	AUC	Accuracy	mAP	Balanced-Accuracy	F1-Score	p	
ADASYN	0.732(0.065)	0.759(0.071)	0.257(0.050)	0.657(0.053)	0.355(0.021)	< 0.01	
SMOTE-ENN	0.778(0.016)	0.708(0.036)	0.453(0.207)	0.696(0.077)	0.435(0.081)	

Model comparation

Five models, namely XGBoost, LR, RF, SVM, and MLP, were established using the training dataset and tuned using grid parameter optimization. The models were then tested using respective testing datasets, and AUC values, accuracy, mean average precision (mAP), and F1 Score were calculated to facilitate comparison between the models. The AUC was selected as the primary evaluation metric for the model because it is robust to class imbalance, insensitive to classification thresholds, allows for comparison of performance across different classifiers, and intuitively presents the model’s performance at various thresholds through the ROC curve (as shown in Fig. 2). Among the models, XGBoost achieved the highest predictive performance (AUC = 0.778 ± 0.016) after tuning, representing a notable improvement in scoring results compared to the traditional Apache model (AUC = 0.677). The use of complex neural network models did not lead to an improvement in accuracy; instead, there was a decline in various aspects.

Fig. 2 Receiver operating characteristic curve of different models

we conducted multiple cross-validation tests using k-fold (k = 5) [27] to obtain the average performance results of the model (Table 3).

Table 3 The classification results of different models

Model	AUC	Accuracy	mAP	Balanced-Acc	F1-Score	
xgboost	0.778(0.016)	0.708(0.036)	0.453(0.207)	0.696(0.077)	0.435(0.081)	
Random forest	0.752(0.101)	0.664(0.014)	0.347(0.041)	0.644(0.028)	0.377(0.028)	
LR	0.627(0.034)	0.631(0.027)	0.248(0.022)	0.595(0.031)	0.330(0.031)	
Svm	0.698(0.021)	0.445(0.026)	0.318(0.036)	0.613(0.016)	0.343(0.011)	
MLP	0.697(0.019)	0.518(0.022)	0.320(0.020)	0.627(0.023)	0.354(0.017)	

External validation

To assess the generalizability of the model, we further performed external validation using 549 randomly selected cases from the MIMIC (Medical Information Mart for Intensive Care) database, applying the same inclusion criteria. The confusion matrices for each model are presented in Fig. 3.

Fig. 3 External validation: The confusion matrix of different models

The external validation results are largely consistent with those from the test set, with the XGBoost model outperforming all others. This suggests that the model possesses a certain degree of generalization capability, achieving the highest accuracy rate of 73.9% and a relatively high balanced accuracy of 66.8%. Additionally, its F1 score reached 0.424, indicating a good balance between precision and recall. The final results are presented in Table 4.

Table 4 The classification results of external validation

Model	TN	FP	FN	TP	Accuracy	Balanced accuracy	F1-Score	
xgboost	340	98	40	51	0.739	0.668	0.424	
SVM	235	203	23	88	0.589	0.664	0.436	
LR	186	252	7	104	0.528	0.681	0.445	
RF	334	104	63	48	0.696	0.598	0.363	
MLP	222	216	22	89	0.567	0.655	0.428	
Apache	313	125	88	23	0.612	0.461	0.178	

SHAP visualization

To determine the importance of each predictive variable for the XGBoost model, the SHAP algorithm was utilized. The variable importance plot was generated, which listed the variables in descending order of importance. Among the conventional indicators, AST, which reflects the patient’s metabolic condition, was the strongest predictor for all prediction periods. The average age, albumin, and BMI were also identified as strong predictive features. Figure 4 provides a visualization of these findings.

Fig. 4 The weights of variables importance calculated by SHAP method

To investigate the positive or negative relationship between predictive factors and the target outcome, SHAP values were utilized to identify death risk factors. Figure 5 illustrates the results, with the horizontal position indicating the effect of higher or lower predictions associated with the value, and the color indicating whether the variable is high (red) or low (blue) for the observation value. The findings suggest that an increase in average age has a positive effect on the prediction results and is associated with a higher risk of death, while an increase in AST has a negative effect on the prediction results and is associated with a higher morality, and lower BMI index means a better survival rate.

Fig. 5 SHAP values to show the distribution of the impacts each feature has on the model output. The color represents the feature value (red high, blue low)

To confirm the model’s diagnostic accuracy in individuals, we selected a single patient for visual validation of the prediction model. The patient, who ultimately passed away during their ICU stay, was of advanced age, a natural high-risk factor. Additionally, several other negative predictors for clinical outcomes were present, as indicated in red on the graph. Specifically, the patient had a slightly elevated BMI, which could have increased the risk of complications, and a higher RDW, suggesting a stronger inflammatory response. The slight elevation in AST pointed to potential liver function impairment. These combined factors contributed to the final predicted score, which was greater than 0 (as shown in Fig. 6), forecasting the outcome as expired. This prediction aligned with the actual result, as the patient did not survive their ICU stay. The patient’s unit stay ID was 2,183,008 in the eICU database.

Fig. 6 Visualization of individual classification decision process

Based on the SHAP interaction values analysis, as illustrated in Fig. 7, several key findings were observed: Among patients aged 80 to 90, an increase in AST levels significantly correlates with a heightened risk of mortality. Concurrent increases in BMI and AST are associated with an elevated mortality risk. Similarly, higher levels of AST and total bilirubin are linked to an increased mortality risk. Furthermore, for patients receiving ventilator support, those with higher BMI show a reduced mortality risk, whereas those with lower BMI experience an increased risk of mortality.

Fig. 7 Visualization of SHAP interaction values (e.g.: AST and age, AST and BMI, AST and total bilirubin, BMI and ventilator)

SHAP Robustness

Initial experimental results indicated that SMOTE-ENN outperforms ADASYN in enhancing model performance. Building on this, we selected the top 20 features with the highest importance and performed hyperparameter tuning on the XGBoost model using training sets processed by different data balancing algorithms. Although the optimal parameters (such as max_depth, n_estimators, subsample, etc.) varied, we used 5-fold cross-validation combined with the SHAP algorithm to calculate the mean and variance of feature importance, recording the ranking of feature importance in each iteration. Additionally, we conducted sensitivity analysis to evaluate the SHAP algorithm’s sensitivity to different parameter settings and datasets. Despite significant variations in feature importance values across different models, the feature rankings exhibited a notable degree of stability. Notably, when feature importance was higher, the stability of the rankings increased, which is typically associated with improved model performance (See Table 5).

Table 5 The robustness analyses of SHAP

Feature	Weights (ADASYN)	Ranks_ ADASYN	Avg(std)	Weights (SMOTE-ENN)	Ranks_SMOTE-ENN	Avg(std)	
AST (SGOT)	0.102 (0.012)	1,1,1,1,1	1 ± 0	0.303 (0.021)	1,1,1,1,1	1 ± 0	
age	0.080 (0.007)	2,2,2,2,2	2 ± 0	0.217 (0.018)	2,2,2,2,2	2 ± 0	
BMI	0.061 (0.014)	4,3,3,10,4	4.8 ± 2.64	0.102 (0.016)	3,4,3,6,5	4.2 ± 1.28	
total bilirubin	0.031 (0.006)	3,4,16,20,8	10.2 ± 6.71	0.096 (0.035)	4,6,6,10,7	6.6 ± 1.79	
creatinine	0.020 (0.006)	8,8,6,4,9	7 ± 1.79	0.084 (0.019)	5,3,8, 5,8	5.8 ± 2.17	
TV/kg IBW	0.032 (0.010)	13,13,4,16,3	9.8 ± 5.27	0.073 (0.018)	6,8,10,8,10	8.4 ± 1.57	
paO2	0.033 (0.007)	5,5,7,9,18	8.8 ± 4.83	0.070 (0.019)	7,16,15,19,16	14.6 ± 3.95	
Vent Rate	0.031 (0.007)	7,10,17,7,14	11 ± 3.95	0.060 (0.008)	8,17,16,15,6	12.4 ± 4.64	
vent	0.001 (0.001)	19,20,10,15,6	14 ± 5.33	0.060 (0.010)	9,15,5, 16,15	12.0 ± 4.67	
troponin - I	0.003 (0.003)	17,14,13,13,13	14 ± 1.55	0.059 (0.009)	10,5,11,3,9	7.6 ± 3.10	
albumin	0.006 (0.003)	12,12,8,12,12	11.2 ± 1.60	0.055 (0.008)	11,9,19,9,14	12.4 ± 3.85	
is_copd	0.035 (0.012)	15,17,19,3,10	12.8 ± 5.74	0.054 (0.026)	12,13,4,4,4	9.4 ± 4.25	
intubated	0.030 (0.008)	11,11,11,11,11	11 ± 0	0.053 (0.019)	13,10,7,7,17	11.0 ± 3.74	
RDW	0.035 (0.012)	9,7,5,19,19	11.8 ± 6.01	0.039 (0.009)	14,7,9,11,11	10.8 ± 2.28	
O2 Sat (%)	0.045 (0.010)	6,6,15,5,17	9.8 ± 5.11	0.039 (0.012)	15,14,17,13,3	12.4 ± 5.33	
phosphate	0.021 (0.011)	10,15,14,17,5	12.2 ± 4.26	0.030 (0.002)	16,11,18,12,12	13.0 ± 2.77	
MCHC	0.062 (0.012)	16,16,12,8,16	13.6 ± 3.20	0.025 (0.003)	17,18,12,17,13	15.4 ± 2.53	
PEEP	0.010 (0.003)	20,9,9,6,15	11.8 ± 5.04	0.018 (0.013)	18,20,13,18,20	17.8 ± 1.63	
is_cardiac	0.045 (0.010)	14,19,20,14,7	14.8 ± 4.62	0.014 (0.006)	19,12,14,14,19	15.6 ± 2.70	
PT - INR	0.031 (0.006)	18,18,18,18,20	18.4 ± 0.80	0.010 (0.004)	20, 19,20,20,18	19.4 ± 0.80	

Clinical application

We are actively working on integrating our system with hospital information systems. By interfacing with various hospital systems, we collect real-time patient data and perform data integration and cleaning through application programming interfaces (APIs). The entire AI model is embedded within the intensive care unit (ICU) system. Leveraging the SHAP method, our model not only provides mortality risk predictions but also offers insights into the contribution of each feature to the prediction, enhancing the model’s transparency and credibility. This allows healthcare professionals to better understand the model’s decision-making process, leading to more precise treatment decisions (See Fig. 8).

Fig. 8 The clinical application in ICU Monitoring system

Discussion

Pneumonia has consistently been one of the leading causes of morbidity and mortality worldwide. The mortality rate of pneumonia is intricately associated with age, incidence, and the severity of the disease at the time of admission [28]. In this study, the focus was on developing an interpretable model for predicting mortality risk in ICU patients with pneumonia. The XGBoost model exhibited superior predictive performance, boasting an impressive AUC of 0.778(0.016), surpassing traditional scoring systems and other machine learning methods. External validation using MIMIC data further validated the model’s classification prowess. The SHAP method played a pivotal role in elucidating the predictive factors influencing pneumonia mortality. Notably, AST emerged as the foremost predictor, shedding light on its critical role in prognosis. The study’s findings emphasize the significance of interpretable predictive models in enhancing physicians’ ability to accurately gauge the risk of mortality in ICU patients with pneumonia.

Regarding methods for assessing the severity of pneumonia, such as the Pneumonia Severity Index (PSI) [29], CURB-65 score [30], and clinical scoring systems, such as the Sequential Organ Failure Assessment (SOFA) score [5], the Acute Physiology and Chronic Health Evaluation (APACHE-II) score [4] can assist clinical practitioners in evaluating patients’ overall risk and predict disease progression or formulate treatment plans. However, it’s important to note that numerous scoring systems rely on simplified criteria and parameters, potentially lacking a comprehensive reflection of an individual patient’s condition. The applicability of these scoring systems might be constrained, especially for specific age groups or patients with particular medical conditions. Due to the inherent advantage of learning algorithms in capturing powerful nonlinear characteristics, machine learning predictive models are increasingly and widely applied to explore early warning signs and analyze risk factors in various diseases [31, 32]. This application is aimed at supporting personalized treatment for patients.

In this retrospective cohort study of a large-scale public ICU database, we developed and validated five machine learning algorithms to predict the mortality of patients with pneumonia.

Compared to the traditional Apache scoring system (AUC 0.677), machine learning algorithms such as Random Forest had an AUC of 0.677, Support Vector Machine (SVM) had an AUC of 0.702, and Multi-Layer Perceptron (MLP) had an AUC of 0.683. In our study, the XGBoost model showed a better performance to predict the mortality of pneumonia with an AUC of 0.734 compared with others. In order to validate the efficacy of the model, we conducted external validation using data from MIMIC. Similarly, it is observed that XGBoost achieved the best classification performance. This also indicates that in clinical scenarios, simple machine learning models may be easier to train and less susceptible to overfitting. In our study, the risk analysis for pneumonia patients is just a simple binary classification model. Given the tabular data nature of the task, it is relatively straightforward, requiring no complex feature extraction and pattern recognition. The XGBoost model has proven competent and demonstrated better classification performance. This aligns with relevant literature perspectives and is consistent with our research findings [33]. Moreover, within the medical domain, there is frequently a heightened demand for interpretability in decision-making. Physicians and patients seek to comprehend the rationale behind the model’s predictions. The integration of XGBoost with our SHAP algorithm facilitates a more accessible interpretation of the model’s decision-making process, in contrast to the often-opaque nature of complex neural network models, which are commonly perceived as black-box models.

Using the SHAP method to explain the XGBoost model ensured both its performance and clinical interpretability. This will help doctors better understand the model’s decision-making process and promote the use of predictive results. In our study, we combined the basic clinical information, ventilator parameters, and laboratory characteristics of patients, which provided important vital sign parameters and information related to the severity of pneumonia. We found that age and AST were the most significant variables associated with in-hospital mortality among pneumonia patients. The mortality rate for severe pneumonia patients in the intensive care unit (ICU) was as high as 16.33%, while for patients in the ward, an increase in age was accompanied by an increase in mortality rate, with the average age of the deceased higher than that of the survivors. According to a study by Venceslau Pinto Hespanhol et al., the hospital mortality rate for pneumonia patients over 80 years of age rose sharply after the age of 60, reaching 38.5% after the age of 90 [34]. Previous research on severe pneumonia has found that age (> 65 years), severe leukopenia or leukocytosis, and bacteremia are risk factors for death [35]. The risk of death from pneumonia often increases with patient age and comorbidities, possibly due to systemic inflammation and decreased immunity [36] Excluding the factor of age in the patients’ basic information, AST provides the most important information for predicting patient mortality. Although AST is well known as a marker of liver dysfunction, we found in our study that serum AST concentration is closely related to the mortality rate of severe pneumonia. Enveloped viruses such as SARS-CoV and HCoV-NL63 enter host cells through direct membrane fusion between the host cell surface receptor (ACE2) and the virus (spike protein), leading to the release of viral ssRNA genome into host cells. After the virus enters, ACE2 is down-regulated, leading to excessive ACE/Ang II activity, increased pulmonary vascular permeability, and subsequent lung injury [37–39]. However, the ACE2 receptor is also widely expressed in bile ducts and hepatic epithelial cells [40]. A recent study on the association between liver injury and markers of in-hospital mortality for 2019 coronavirus disease reported that these markers, particularly AST abnormalities and death, were diagnosed in conjunction with liver injury during hospitalization, in comparison to other indicators of liver injury [41]. This suggests that measuring serum AST concentration can serve as a valuable tool for predicting patient outcomes and assessing liver injury in cases of severe pneumonia. Therefore, AST provides critical information that healthcare providers should consider when monitoring patients with severe pneumonia.

To date, none of the severity-of-illness assessment tools commonly used in critical care settings require the evaluation of serum albumin levels. Nevertheless, the relationship between hypoalbuminemia and the incidence, mortality, and length of hospital stay of patients in the ICU has been established [42]. Critically ill patients often exhibit acute-phase reactions, leading to a decrease in albumin levels due to changes in distribution. Concurrently, the synthesis or breakdown metabolism of albumin may also undergo alterations [43]. In previous studies, it was found that decreased serum albumin levels may serve as independent predictors of pneumonia in patients with acute ischemic stroke (AIS), especially in cases of mild stroke [44]. Additionally, the risk of pneumonia may exhibit an inverse correlation with the albumin level. Multiple studies have also identified an overall association between disease severity and obesity, as well as other metabolic risk factors, including diabetes and hypertension [45–47]. Viral pneumonia often necessitates invasive mechanical (IMV), imposing a significant strain on global intensive care resources. In a study conducted by Chetboun et al., it was observed that the need for IMV increased progressively with the body mass index (BMI) among viral pneumonia patients admitted to intensive care units (ICU) [48]. Similarly, recent experiences during the viral pneumonia pandemic suggest that the mortality rate for patients requiring invasive mechanical ventilation falls within the range of 35–50% [49]. Hraiech et al. noted a survival advantage in CAP patients who required mechanical ventilation within 72 h of the onset of Community-Acquired Pneumonia (CAP) compared to those who required mechanical ventilation 4 or more days after the onset of CAP (28% vs. 51%, p = 0.03) [50]. Therefore, any delay in identifying severe illness, recognizing those at risk of mechanical ventilation, or needing ICU-level care, along with associated timely treatment, may have detrimental effects on the prognosis of severe CAP patients [51].

Through interaction analysis, we noticed that an elevated serum level of Aspartate Aminotransferase (AST) is associated with an increased relative risk of mortality when the age is between 80 and 90 years old. Furthermore, elevated Aspartate Aminotransferase (AST) levels were related with increased total bilirubin exacerbate the mortality risk among pneumonia patients, which may be correlated with severe hepatic dysfunction. Previous research indicates that in sepsis induced by severe pneumonia, up to 46% of patients exhibit liver dysfunction, with over half of the affected individuals being 65 years of age or older [52, 53]. As age progresses, there is a gradual decline in hepatic organ function, and the liver’s regenerative capacity is compromised. This leads to a worst prognosis for elderly patients, with higher rates of mortality, disability, prolonged hospitalization, and severe organ dysfunction. Due to systemic inflammation, immune dysregulation, and microcirculatory disturbances, patients with severe pneumonia may experience various types of hepatic dysfunction, such as hypoxia-induced hepatitis, sepsis-associated cholestasis, or hepatic failure [54]. Elevated concentrations of transaminases and total bilirubin, along with jaundice, are observed in the serum of 3–25% of patients with Streptococcus pneumoniae pneumonia [55]. Jaundice associated with pneumonia is largely considered a consequence of hepatocellular damage, with hepatic necrosis commonly found in liver biopsies of pneumonia patients [56]. In patients with Legionella pneumophila causing pneumonia, laboratory serum tests often indicate abnormalities, suggesting impaired liver function, renal insufficiency, thrombocytopenia, and hyponatremia [57].Additionally, AST can serve as one of the biomarkers for myocardial injury induced by pneumonia.Yusuf et al. observed that incidence of myocardial injury is higher in patients with severe Mycoplasma pneumoniae pneumonia, with a significant increase in AST serum levels [58].For patients with a higher Body Mass Index (BMI), the risk of mortality is significantly reduced after timely ventilation measures are implemented. There is a close association between the baseline BMI and the mortality and survival prognosis during the hospital stay in the Intensive Care Unit (ICU). During the ICU stay, compared to patients who lose weight, those who gain weight are associated with higher in-hospital mortality and ICU mortality, especially among patients with more severe conditions; patients with a higher BMI have a longer ICU Length of Stay (LOS) [59]. Among pneumonia patients hospitalized due to COVID-19, those with a higher BMI have a higher mortality rate, more needs for ventilation, and require prolonged use of a ventilator to maintain normal oxygen saturation [60].

Conclusion

In summary, this study successfully developed an interpretable XGBoost model for predicting pneumonia mortality in ICUs, offering improved performance compared to existing models. The SHAP method facilitated a deeper understanding of the model’s decision-making process by revealing the contribution of each feature to the predicted outcomes, particularly in relation to mortality. However, it is important to note that these associations do not imply a direct causal relationship between the features and mortality risk. These findings contribute to advancing individualized care strategies for pneumonia patients in intensive care settings.

Several limitations should be considered when interpreting the results of this study. First, the retrospective nature of the study using electronic health records introduces inherent biases and limitations related to data quality, completeness, and potential confounding variables. Addressing these issues requires ongoing efforts to improve data collection processes and incorporate real-world data. Moreover, the model’s performance, while promising, may vary across different healthcare settings and patient populations. External validation using datasets from diverse sources and geographical locations is imperative to assess the generalizability of the model and ensure its applicability in varied clinical scenarios. In future research, incorporating real-world data and continuous model updates could enable the development of a dynamic prediction tool, responsive to evolving patient conditions and treatment protocols. Collaboration with clinicians is crucial for refining the model and ensuring its seamless integration into clinical workflows [61–63].

Abbreviations

ICU Intensive Care Unit

HER Electronic Health Records

APACHE Acute Physiology and Chronic Health Evaluation

SOFA Sequential Organ Failure Assessment

SHAP Shapley Additive Explanation

XGBoost Extreme Gradient Boosting

HIV Human Immunodeficiency Virus

LASSO Least Absolute Shrinkage and Selection Operator

BMI Body Mass Index

GCS Glasgow Coma Scale

ADASYN Adaptive Synthetic Sampling

SMOTE-ENN Synthetic Minority Over-sampling Technique combined with Edited Nearest Neighbors

SHAP Shapley Additive Explanation

MLP Multi-Layer Perceptron

ROC Receiver Operating Characteristic

AUC Area Under the Curve

mAP Mean Average Precision

LR Logistic Regression

RF Random Forest

SVM Support Vector Machine

MLP Multi-Layer Perceptron

ROC Receiver Operating Characteristic

MIMIC Medical Information Mart for Intensive Care

API Application Programming Interface

AI Artificial Intelligence

AST Aspartate Aminotransferase

PSI Pneumonia Severity Index

SOFA Sequential Organ Failure Assessment

APACHE-II Acute Physiology and Chronic Health Evaluation II

AIS Acute Ischemic Stroke

BMI Body Mass Index

IMV Invasive Mechanical Ventilation

CAP Community-Acquired Pneumonia

LOS Length of Stay

Acknowledgements

Not applicable.

Author contributions

J. L, and Y.Z contributed equally. Y.T conceptualized the study. S.H carried out the collection. J.L and Y.Z analysis of the literature and data and drafted the manuscript.

Funding

This work was supported by Chengdu Medical Research Project 2023195.

Data availability

The data for this study is sourced from a publicly available dataset, Materials used in experiments are available upon request. Contact the corresponding author for further details.

Declarations

Ethics approval and consent to participate

Not applicable.

Consent for publication

Not Applicable.

Competing interests

The authors declare no competing interests.

Publisher’s note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Jiaxi Li and Yu Zhang contributed equally to this work.
==== Refs
References

1. Nair GB Niederman MS Updates on community acquired pneumonia management in the ICU Pharmacol Ther 2021 217 107663 10.1016/j.pharmthera.2020.107663 32805298
Nair GB, Niederman MS. Updates on community acquired pneumonia management in the ICU. Pharmacol Ther. 2021;217:107663. 10.1016/j.pharmthera.2020.107663. Epub 2020 Aug 15. PMID: 32805298; PMCID: PMC7428725.32805298 10.1016/j.pharmthera.2020.107663
2. Ramirez JA Wiemken TL Adults hospitalized with pneumonia in the United States: incidence, epidemiology, and Mortality Clin Infect Dis 2017 65 11 1806 12 10.1093/cid/cix647 29020164
Ramirez JA, Wiemken TL, et al. Adults hospitalized with pneumonia in the United States: incidence, epidemiology, and Mortality. Clin Infect Dis. 2017;65(11):1806–12. 10.1093/cid/cix647.29020164 10.1093/cid/cix647
3. Koenig SM Truwit JD Ventilator-associated pneumonia: diagnosis, treatment, and prevention[J] Clin Microbiol Rev 2006 19 4 637 57 10.1128/CMR.00051-05 17041138
Koenig SM, Truwit JD. Ventilator-associated pneumonia: diagnosis, treatment, and prevention[J]. Clin Microbiol Rev. 2006;19(4):637–57.17041138 10.1128/CMR.00051-05
4. Knaus WA Draper EA Wagner DP APACHE II: a severity of disease classification system[J] Crit Care Med 1985 13 10 818 29 10.1097/00003246-198510000-00009 3928249
Knaus WA, Draper EA, Wagner DP, et al. APACHE II: a severity of disease classification system[J]. Crit Care Med. 1985;13(10):818–29.3928249 10.1097/00003246-198510000-00009
5. Lambden S Laterre PF Levy MM The SOFA score—development, utility and challenges of accurate assessment in clinical trials[J] Crit Care 2019 23 1 1 9 10.1186/s13054-019-2663-7 30606235
Lambden S, Laterre PF, Levy MM, et al. The SOFA score—development, utility and challenges of accurate assessment in clinical trials[J]. Crit Care. 2019;23(1):1–9.30606235 10.1186/s13054-019-2663-7
6. Liu C Wang X Liu C Differentiating novel coronavirus pneumonia from general pneumonia based on machine learning[J] Biomed Eng Online 2020 19 1 1 14 10.1186/s12938-020-00809-9 31915014
Liu C, Wang X, Liu C, et al. Differentiating novel coronavirus pneumonia from general pneumonia based on machine learning[J]. Biomed Eng Online. 2020;19(1):1–14.31915014 10.1186/s12938-020-00809-9
7. Kim SY, Diggans J, Pankratz D, et al. Classification of usual interstitial pneumonia in patients with interstitial lung disease: assessment of a machine learning approach using high-dimensional transcriptional data[J]. Volume 3. The lancet Respiratory medicine; 2015. pp. 473–82. 6.
8. Kuo KM Talley PC Huang CH Predicting hospital-acquired pneumonia among schizophrenic patients: a machine learning approach[J] BMC Med Inf Decis Mak 2019 19 1 1 8
Kuo KM, Talley PC, Huang CH, et al. Predicting hospital-acquired pneumonia among schizophrenic patients: a machine learning approach[J]. BMC Med Inf Decis Mak. 2019;19(1):1–8.
9. Johnson AE, W, Ghassemi MM, Nemati S et al. Machine learning and decision support in critical care[J]. Proceedings of the IEEE, 2016, 104(2): 444–466.
10. Pollard TJ Johnson AEW Raffa JD The eICU Collaborative Research Database, a freely available multi-center database for critical care research[J] Sci data 2018 5 1 1 13 10.1038/sdata.2018.178 30482902
Pollard TJ, Johnson AEW, Raffa JD, et al. The eICU Collaborative Research Database, a freely available multi-center database for critical care research[J]. Sci data. 2018;5(1):1–13.30482902 10.1038/sdata.2018.178
11. Lundberg SM, Lee SI. A unified approach to interpreting model predictions[J]. Adv Neural Inf Process Syst, 2017, 30.
12. Chen T, He T, Benesty M et al. Xgboost: extreme gradient boosting[J]. R package version 0.4-2, 2015, 1(4): 1–4.
13. van de Garde EMW Oosterheert JJ Bonten M International classification of diseases codes showed modest sensitivity for detecting community-acquired pneumonia[J] J Clin Epidemiol 2007 60 8 834 8 10.1016/j.jclinepi.2006.10.018 17606180
van de Garde EMW, Oosterheert JJ, Bonten M, et al. International classification of diseases codes showed modest sensitivity for detecting community-acquired pneumonia[J]. J Clin Epidemiol. 2007;60(8):834–8.17606180 10.1016/j.jclinepi.2006.10.018
14. Van Schyndel SJ Carrier J Bogado Pascottini O The effect of pegbovigrastim on circulating neutrophil count in dairy cattle: a randomized controlled trial[J] PLoS ONE 2018 13 6 e0198701 10.1371/journal.pone.0198701 29953439
Van Schyndel SJ, Carrier J, Bogado Pascottini O, et al. The effect of pegbovigrastim on circulating neutrophil count in dairy cattle: a randomized controlled trial[J]. PLoS ONE. 2018;13(6):e0198701.29953439 10.1371/journal.pone.0198701
15. Robinson C. Basic introduction into pgAdmin III and SQL queries[J]. 2011.
16. Muthukrishnan R, Rohini R. LASSO: A feature selection technique in predictive modeling for machine learning[C]//2016 IEEE international conference on advances in computer applications (ICACA). Ieee, 2016: 18–20.
17. Tang F, Ishwaran H. Sci J. 2017;10(6):363–77. Random forest missing data algorithms[J]. Statistical Analysis and Data Mining: The ASA Data.
18. He H, Bai Y, Garcia EA et al. ADASYN: Adaptive synthetic sampling approach for imbalanced learning[C]//2008 IEEE international joint conference on neural networks (IEEE world congress on computational intelligence). IEEE, 2008: 1322–1328.
19. Muntasir Nishat M Faisal F Jahan Ratul I A Comprehensive Investigation of the performances of different machine learning classifiers with SMOTE-ENN oversampling technique and hyperparameter optimization for Imbalanced Heart failure Dataset[J] Sci Program 2022 2022 1 3649406
Muntasir Nishat M, Faisal F, Jahan Ratul I, et al. A Comprehensive Investigation of the performances of different machine learning classifiers with SMOTE-ENN oversampling technique and hyperparameter optimization for Imbalanced Heart failure Dataset[J]. Sci Program. 2022;2022(1):3649406.
20. Lundberg SM Erion G Chen H From local explanations to global understanding with explainable AI for trees[J] Nat Mach Intell 2020 2 1 56 67 10.1038/s42256-019-0138-9 32607472
Lundberg SM, Erion G, Chen H, et al. From local explanations to global understanding with explainable AI for trees[J]. Nat Mach Intell. 2020;2(1):56–67.32607472 10.1038/s42256-019-0138-9
21. Ihaka R Gentleman R R: a language for data analysis and graphics[J] J Comput Graphical Stat 1996 5 3 299 314 10.1080/10618600.1996.10474713
Ihaka R, Gentleman R. R: a language for data analysis and graphics[J]. J Comput Graphical Stat. 1996;5(3):299–314.10.1080/10618600.1996.10474713
22. Cuzick J A wilcoxon-type test for trend[J] Stat Med 1985 4 1 87 90 10.1002/sim.4780040112 3992076
Cuzick J. A wilcoxon-type test for trend[J]. Stat Med. 1985;4(1):87–90.3992076 10.1002/sim.4780040112
23. LaValley MP Logistic regression[J] Circulation 2008 117 18 2395 9 10.1161/CIRCULATIONAHA.106.682658 18458181
LaValley MP. Logistic regression[J] Circulation. 2008;117(18):2395–9.18458181 10.1161/CIRCULATIONAHA.106.682658
24. Rigatti SJ Random forest[J] J Insur Med 2017 47 1 31 9 10.17849/insm-47-01-31-39.1 28836909
Rigatti SJ. Random forest[J]. J Insur Med. 2017;47(1):31–9.28836909 10.17849/insm-47-01-31-39.1
25. Suthaharan S, Suthaharan S. Support vector machine[J]. Machine learning models and algorithms for big data classification: thinking with examples for effective learning, 2016: 207–235.
26. Taud H, Mas JF. Multilayer perceptron (MLP)[J]. Geomatic approaches for modeling land change scenarios, 2018: 451–455.
27. Rodriguez JD Perez A Lozano JA Sensitivity analysis of k-fold cross validation in prediction error estimation[J] IEEE Trans Pattern Anal Mach Intell 2009 32 3 569 75 10.1109/TPAMI.2009.187
Rodriguez JD, Perez A, Lozano JA. Sensitivity analysis of k-fold cross validation in prediction error estimation[J]. IEEE Trans Pattern Anal Mach Intell. 2009;32(3):569–75.10.1109/TPAMI.2009.187
28. Ito A Ishida T Tokumasu H Prognostic factors in hospitalized community-acquired pneumonia: a retrospective study of a prospective observational cohort BMC Pulm Med 2017 17 1 78 10.1186/s12890-017-0424-4 28464807
Ito A, Ishida T, Tokumasu H, et al. Prognostic factors in hospitalized community-acquired pneumonia: a retrospective study of a prospective observational cohort. BMC Pulm Med. 2017;17(1):78. 10.1186/s12890-017-0424-4. Published 2017 May 2.28464807 10.1186/s12890-017-0424-4
29. Wang D Willis DR Yih Y The pneumonia severity index: Assessment and comparison to popular machine learning classifiers Int J Med Inf 2022 163 104778 10.1016/j.ijmedinf.2022.1047781
Wang D, Willis DR, Yih Y. The pneumonia severity index: Assessment and comparison to popular machine learning classifiers. Int J Med Inf. 2022;163:104778. 10.1016/j.ijmedinf.2022.1047781.10.1016/j.ijmedinf.2022.1047781
30. Patel P S. Calculated decisions: CURB-65 score for pneumonia severity. Emerg Med Pract. 2021;23(Suppl 2):CD1–2. Published 2021 Feb 1. 1.
31. Kang MW Kim J Kim DK Machine learning algorithm to predict mortality in patients undergoing continuous renal replacement therapy Crit Care 2020 24 1 42 10.1186/s13054-020-2752-7 32028984
Kang MW, Kim J, Kim DK, et al. Machine learning algorithm to predict mortality in patients undergoing continuous renal replacement therapy. Crit Care. 2020;24(1):42. 10.1186/s13054-020-2752-7. Published 2020 Feb 6.32028984 10.1186/s13054-020-2752-7
32. Liu C, Liu X, Mao Z, Interpretable Machine Learning Model for Early Prediction of Mortality in ICU Patients with Rhabdomyolysis. Med Sci Sports Exerc., Grinsztajn L, Oyallon E, Varoquaux G et al. Why do tree-based models still outperform deep learning on typical tabular data? [J]. Advances in Neural Information Processing Systems, 2022, 35: 507–520.
33. Grinsztajn L Oyallon E Varoquaux G Why do tree-based models still outperform deep learning on typical tabular data?[J] Adv Neural Inf Process Syst 2022 35 507 20
Grinsztajn L, Oyallon E, Varoquaux G. Why do tree-based models still outperform deep learning on typical tabular data?[J]. Adv Neural Inf Process Syst. 2022;35:507–20.
34. Hespanhol V, Bárbara C. Pneumonia mortality, comorbidities matter? Pulmonology. 2020 May-Jun;26(3):123–129. 10.1016/j.pulmoe.2019.10.003. Epub 2019 Nov 29. PMID: 31787563.
35. Metersky ML, Waterer G, Nsa W et al. Predictors of in-hospital vs postdischarge mortality in pneumonia. Chest. 2012;142(2):476–481. 10.1378/chest.11-2393. PMID: 22383662.
36. Cillo´niz C Polverino E Ewig S Impact of age and comorbidity on cause and outcome in community-acquired pneumonia Chest 2013 144 3 999 1007 10.1378/chest.13-0062 23670047
Cillo´niz C, Polverino E, Ewig S, et al. Impact of age and comorbidity on cause and outcome in community-acquired pneumonia. Chest. 2013;144(3):999–1007.23670047 10.1378/chest.13-0062
37. Malik YA Properties of Coronavirus and SARS-CoV-2 Malays J Pathol 2020 42 1 3 11 32342926
Malik YA. Properties of Coronavirus and SARS-CoV-2. Malays J Pathol. 2020;42(1):3–11. PMID: 32342926.32342926
38. Suryamohan K Diwanji D Stawiski EW Human ACE2 receptor polymorphisms and altered susceptibility to SARS-CoV-2 Commun Biol 2021 4 1 475 10.1038/s42003-021-02030-3 33846513
Suryamohan K, Diwanji D, Stawiski EW, et al. Human ACE2 receptor polymorphisms and altered susceptibility to SARS-CoV-2. Commun Biol. 2021;4(1):475. 10.1038/s42003-021-02030-3. PMID: 33846513; PMCID: PMC8041869.33846513 10.1038/s42003-021-02030-3
39. Bakhshandeh B Sorboni SG Javanmard AR Variants in ACE2; potential influences on virus infection and COVID-19 severity Infect Genet Evol 2021 90 104773 10.1016/j.meegid.2021.104773 33607284
Bakhshandeh B, Sorboni SG, Javanmard AR, et al. Variants in ACE2; potential influences on virus infection and COVID-19 severity. Infect Genet Evol. 2021;90:104773. 10.1016/j.meegid.2021.104773. Epub 2021 Feb 17. PMID: 33607284; PMCID: PMC7886638.33607284 10.1016/j.meegid.2021.104773
40. Chai X, Hu L, Zhang Y et al. Specific ACE2 expression in cholangiocytes may cause liver damage after 2019-nCoV infection. bioRxiv 2020:2020.02.03.931766.
41. Lei F, Liu YM, Zhou F et al. Longitudinal association between markers of liver injury and mortality in COVID-19 in China. Hepatology 2020:1.0.1002/hep.31301.
42. Jellinge ME, Henriksen DP, Hallas P et al. Hypoalbuminemia is a strong predictor of 30-day all-cause mortality in acutely admitted medical patients: a prospective, observational, cohort study. PLoS One. 2014;9(8): e105983. Published 2014 Aug 22. 10.1371/journal.pone.0105983
43. Nicholson JP Wolmarans MR Park GR The role of albumin in critical illness Br J Anaesth 2000 85 4 599 610 10.1093/bja/85.4.599 11064620
Nicholson JP, Wolmarans MR, Park GR. The role of albumin in critical illness. Br J Anaesth. 2000;85(4):599–610. 10.1093/bja/85.4.599.11064620 10.1093/bja/85.4.599
44. Yang X Wang L Zheng L Serum albumin as a potential predictor of Pneumonia after an Acute Ischemic Stroke Curr Neurovasc Res 2020 17 4 385 93 10.2174/1567202617666200514120641 32407279
Yang X, Wang L, Zheng L, et al. Serum albumin as a potential predictor of Pneumonia after an Acute Ischemic Stroke. Curr Neurovasc Res. 2020;17(4):385–93. 10.2174/1567202617666200514120641.32407279 10.2174/1567202617666200514120641
45. Williamson EJ, Walker AJ, Bhaskaran K, et al. Factors associated with COVID-19-related death using OpenSAFELY. Nature. 2020;584(7821):430–6. 10.1038/s41586-020-2521-4.
46. Truog RD Mitchell C Daley GQ The Toughest triage - allocating ventilators in a pandemic N Engl J Med 2020 382 21 1973 5 10.1056/NEJMp2005689 32202721
Truog RD, Mitchell C, Daley GQ. The Toughest triage - allocating ventilators in a pandemic. N Engl J Med. 2020;382(21):1973–5. 10.1056/NEJMp2005689.32202721 10.1056/NEJMp2005689
47. High Prevalence of Obesity in Severe Acute Respiratory Syndrome Coronavirus-2 (SARS-CoV-2) requiring Invasive Mechanical Ventilation Obes (Silver Spring) 2020 28 10 1994 10.1002/oby.23006
High Prevalence of Obesity in Severe Acute Respiratory Syndrome. Coronavirus-2 (SARS-CoV-2) requiring Invasive Mechanical Ventilation. Obes (Silver Spring). 2020;28(10):1994. 10.1002/oby.23006.10.1002/oby.23006
48. Chetboun M Raverdy V Labreuche J BMI and pneumonia outcomes in critically ill covid-19 patients: an international multicenter study Obes (Silver Spring) 2021 29 9 1477 86 10.1002/oby.23223
Chetboun M, Raverdy V, Labreuche J, et al. BMI and pneumonia outcomes in critically ill covid-19 patients: an international multicenter study. Obes (Silver Spring). 2021;29(9):1477–86. 10.1002/oby.23223.10.1002/oby.23223
49. Richardson S, Hirsch JS, Narasimhan M et al. Presenting Characteristics, Comorbidities, and Outcomes Among 5700 Patients Hospitalized With COVID-19 in the New York City Area [published correction appears in JAMA. 2020;323(20):2098]. JAMA. 2020;323(20):2052–2059. 10.1001/jama.2020.6775
50. Hraiech S, Alingrin J, Dizier S et al. Time to intubation is associated with outcome in patients with community-acquired pneumonia. PLoS One. 2013;8(9): e74937. Published 2013 Sep 19. 10.1371/journal.pone.0074937
51. Restrepo MI Mortensen EM Rello J Brody J Anzueto A Late admission to the ICU in patients with community-acquired pneumonia is associated with higher mortality Chest 2010 137 3 552 7 10.1378/chest.09-1547 19880910
Restrepo MI, Mortensen EM, Rello J, Brody J, Anzueto A. Late admission to the ICU in patients with community-acquired pneumonia is associated with higher mortality. Chest. 2010;137(3):552–7. 10.1378/chest.09-1547.19880910 10.1378/chest.09-1547
52. Wester AL, Dunlop O, Melby KK, Dahle UR, Wyller TB. “Age-related differences in symptoms, diagnosis and prognosis of bacteremia.” BMC infectious diseases. 2013;13(346). 10.1186/1471-2334-13-346.
53. Lee CC, Chen SY, Chang IJ, Chen SC, Wu SC. “Comparison of clinical manifestations and outcome of community-acquired bloodstream infections among the oldest old, elderly, and adult patients.” Medicine. 2007;86(3):138–44. 10.1186/1471-2334-13-346.
54. Kasper P Philipp Leberfunktionsstörungen Bei Sepsis Medizinische Klinik Intensivmedizin Und Notfallmedizin 2020 115 7 609 19 10.1007/s00063-020-00707-x 32725325
Kasper P, Philipp, et al. Leberfunktionsstörungen Bei Sepsis. Medizinische Klinik Intensivmedizin Und Notfallmedizin. 2020;115(7):609–19. 10.1007/s00063-020-00707-x. Hepatic dysfunction in sepsis.32725325 10.1007/s00063-020-00707-x
55. Radford AJ Rhodes FA The association of jaundice with lobar pneumonia in the territory of Papua and New Guinea Med J Aust 1967 2 15 678 81 10.5694/j.1326-5377.1967.tb74163.x 6057201
Radford AJ, Rhodes FA. The association of jaundice with lobar pneumonia in the territory of Papua and New Guinea. Med J Aust. 1967;2(15):678–81. 10.5694/j.1326-5377.1967.tb74163.x.6057201 10.5694/j.1326-5377.1967.tb74163.x
56. Jaundice due To bacterial infection Gastroenterology 1979 77 2 362 74 10.1016/0016-5085(79)90293-2 376394
Jaundice due. To bacterial infection. Gastroenterology. 1979;77(2):362–74.376394 10.1016/0016-5085(79)90293-2
57. Kirby BD Snyder KM Meyer RD Finegold SM Legionnaires’ disease: clinical features of 24 cases Ann Intern Med 1978 89 3 297 309 10.7326/0003-4819-89-3-297 686539
Kirby BD, Snyder KM, Meyer RD, Finegold SM. Legionnaires’ disease: clinical features of 24 cases. Ann Intern Med. 1978;89(3):297–309. 10.7326/0003-4819-89-3-297.686539 10.7326/0003-4819-89-3-297
58. Yusuf SO Chen P Clinical characteristics of community-acquired pneumonia in children caused by mycoplasma pneumoniae with or without myocardial damage: a single-center retrospective study World J Clin Pediatr 2023 12 3 115 24 10.5409/wjcp.v12.i3.115 37342450
Yusuf SO, Chen P. Clinical characteristics of community-acquired pneumonia in children caused by mycoplasma pneumoniae with or without myocardial damage: a single-center retrospective study. World J Clin Pediatr. 2023;12(3):115–24. 10.5409/wjcp.v12.i3.115. Published 2023 Jun 9.37342450 10.5409/wjcp.v12.i3.115
59. Zhang J, Du L, Jin X, Ren J, Li R, Liu J, et al. Association between body mass index change and mortality in critically ill patients: a retrospective observational study. Nutrition (Burbank, Los Angeles County, Calif.) 2023;105:111879. 10.1016/j.nut.2022.111879
60. Espiritu AI Reyes NGD Leochico CFD Sy MCC Villanueva Iii EQ Anlacan VMM Jamora RDG Body mass index and its association with COVID-19 clinical outcomes: findings from the Philippine CORONA study Clin Nutr ESPEN 2022 49 402 10 10.1016/j.clnesp.2022.03.013 35623845
Espiritu AI, Reyes NGD, Leochico CFD, Sy MCC, Villanueva Iii EQ, Anlacan VMM, Jamora RDG. Body mass index and its association with COVID-19 clinical outcomes: findings from the Philippine CORONA study. Clin Nutr ESPEN. 2022;49:402–10. 10.1016/j.clnesp.2022.03.013. Epub 2022 Mar 31. PMID: 35623845; PMCID: PMC8968152.35623845 10.1016/j.clnesp.2022.03.013
61. Luo C A machine learning-based risk stratification tool for in-hospital mortality of intensive care unit patients with heart failure J Translational Med 2022 20 1 136 10.1186/s12967-022-03340-8
Luo C, Zhu Y, Zhu Z. A machine learning-based risk stratification tool for in-hospital mortality of intensive care unit patients with heart failure. J Translational Med. 2022;20(1):136.10.1186/s12967-022-03340-8
62. Huang T Machine learning for prediction of in-hospital mortality in lung cancer patients admitted to intensive care unit PLoS ONE 2023 18 1 e0280606 10.1371/journal.pone.0280606 36701342
Huang T, Le D, Yuan L, Xu S, Peng X. Machine learning for prediction of in-hospital mortality in lung cancer patients admitted to intensive care unit. PLoS ONE. 2023;18(1):e0280606.36701342 10.1371/journal.pone.0280606
63. Wang B Novel pneumonia score based on a machine learning model for predicting mortality in pneumonia patients on admission to the intensive care unit Respir Med 2023 217 107363 10.1016/j.rmed.2023.107363 37451647
Wanga B, Li Y, Tian Y, Ju C, Xu X, Pei S. Novel pneumonia score based on a machine learning model for predicting mortality in pneumonia patients on admission to the intensive care unit. Respir Med. 2023;217:107363.37451647 10.1016/j.rmed.2023.107363
