
==== Front
1250250
3132
Comput Biol Med
Comput Biol Med
Computers in biology and medicine
0010-4825
1879-0534

38603899
10.1016/j.compbiomed.2024.108451
nihpa2020726
Article
Early prediction of long hospital stay for Intensive Care units readmission patients using medication information
Zhang Min a
Kuo Tsung-Ting b*
a Applied Statistics, University of Michigan, Ann Arbor, MI, 48109, USA
b UCSD Health Department of Biomedical Informatics, University of California San Diego, La Jolla, CA, 92093, USA
* Corresponding author: tskuo@health.ucsd.edu (T.-T. Kuo).
6 9 2024
5 2024
08 4 2024
10 9 2024
174 108451108451
https://creativecommons.org/licenses/by-nc-nd/4.0/ This is an open access article under the CC BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/).
Objective:

Predicting Intensive Care Unit (ICU) Length of Stay (LOS) accurately can improve patient wellness, hospital operations, and the health system’s financial status. This study focuses on predicting the prolonged ICU LOS (≥3 days) of the 2nd admission, utilizing short historical data (1st admission only) for early-stage prediction, as well as incorporating medication information.

Materials and methods:

We selected 18,572 ICU patients’ records from the MIMIC-IV database for this study. We applied five machine learning classifiers: Logistic regression (LR), Random Forest (RF), Support Vector Machine (SVM), AdaBoost (AB) and XGBoost (XGB). We computed both the sum dose and the average dose for the medication and included them in our model.

Results:

The performance of the RF model demonstrates the highest level of accuracy compared to other models, as indicated by an Area Under the Receiver Operating Characteristic Curve (AUC) of 0.716 and an Expected Calibration Error (ECE) of 0.023.

Discussion:

The calibration improved all five classifiers (LR, RF, SVC, AB, XGB) in terms of ECE. The most important two features for RF are the length of 1st admission and the patient’s age when they visited the hospital. The most important medication features are Phytonadione and Metoprolol Succinate XL. Also, both the sum and the average dose for the medication features contributed to the prediction task.

Conclusion:

Our model showed the capability to predict the prolonged ICU LOS of the 2nd admission by utilizing the demographic, diagnosis, and medication information from the 1st admission. This method can potentially support the prevention of patient complications and enhance resource allocation in hospitals.

Hospital Length of Stay
Intensive Care Unit
Predictive modeling
==== Body
pmc1. Background

Intensive Care Unit (ICU) Length of Stay (LOS) measures the time length between admission and discharge/death for an ICU stay, usually in days, and is an important inspection standard of clinical care efficiency [1]. The average ICU LOS is around 3.3 days [2] in hospitals of the United States. Although the ICU LOS has been decreasing for more than 20 years, inappropriate long-term hospitalizations still exist [3] and have brought serious negative effects on patients and hospitals. Specifically, an improper long ICU LOS leads to potential risks for three parties: (a) Patients. If ICU patients remain in the hospital for an extended period following admission, Healthcare-Associated Infections (HAIs) pose a significant threat [4]. According to the Centers for Disease Control and Prevention (CDC), approximately 1 in 31 hospital patients in the United States has at least one HAI at any given time [5]. (b) Hospitals. When the patients are infected with HAIs, the bed utilization will be reduced as the LOS of the patients increases. The re-infected patients had to stay longer in the hospital for treatment [6]. As a result, the bed utilization in the unit decreased, causing delays in admitting new patients who needed critical care. (c) Healthcare system. In addition, the unnecessary long hospital stay increases the financial costs for both patients and the hospital. In 2017, hospitals incurred over $33 billion in costs due to unnecessary long hospital stays and potentially preventable inpatient complications [7], which brings significant financial pressure to the whole healthcare system.

To address these risks of long ICU LOS, one plausible solution is predictive modeling, such as predicting LOS of patients after cardiac surgery [8], cardiothoracic surgery [9] or patients with severe sepsis [10]; Predicting LOS using general admission features [11], patient vital signs [12] or chart events data [13]; Predicting LOS using individualized single classification algorithm [14], least absolute shrinkage and selection operator [15] or temporal pointwise convolutional networks [16]) By identifying patients who may have extended hospital stays, the patients can have less risk of HAI, the hospital can arrange the beds in an advanced and systematic way, and the overall healthcare system can reduce an amount of avoidable cost. However, there are still several challenges in predicting ICU LOS: First, accurately predicting the LOS in days could be challenging. For example, a study [16] focused on patients who had been in the ICU for at least 5 h and found that the model had a mean absolute deviation of 2.28 days in predicting remaining LOS. Considering the average ICU LOS being 3.3 days [2], such a prediction task is non-trivial. Therefore, a relaxed task such as predicting binarized LOS results (e.g., longer than 3 days or not) that could potentially achieve practical prediction performance [17], can support the use of the learned model. Second, predicting a long ICU LOS based on a longer time of historical data may not be feasible to do early-stage prediction. The patients who have been admitted to the hospital recently may have a limited number of historical clinical data. For example, 81.9 % of patients were readmitted to the hospital equal to or less than twice in a public dataset [18,19] Thus, comparing to studies using multiple times of historical readmission data to predict ICU LOS [13,20], an “early prediction” (e.g., using data from the 1st admission only) can be critical to facilitate clinical decision making. Finally, medication information is yet to be considered in models of some research. Existing studies [12,16] leveraged only basic information and other non-medication features, while recent research in predicting non-ICU LOS showed that drug information can improve the accuracy of prediction [21]. Therefore, incorporating medication information in the predictive model can be desirable.

2. Objective

In this study, we aim to predict long ICU LOS by addressing the above-mentioned three challenges: (1) focusing on the prediction of the days for ICU LOS to binary prediction of the first readmission (i.e., the 2nd admission) (2) using only short-term historical data (i.e., only the 1st admission data); to enable early-stage prediction; and (3) including medication information in the models to improve prediction capability. We utilized demographic information, diagnosis information, and medication information to predict ICU LOS by five classification methods. The novelty of this study is to use only the 1st admission data to early predict the 2nd readmission ICU LOS by including medication information.

3. Materials and method

3.1. Data

Our prediction task was to determine the ICU LOS of patients in their 2nd hospital admission based on their 1st admission data. Therefore, we used the Medical Information Mart for Intensive Care (MIMIC-IV 1.0) database [18,19], which contains 256,878 ICU patients who were admitted to Beth Israel Deaconess Medical Center during 2008–2019 (approved by UCSD IRB# 804237). The data were extracted from three MIMIC-IV relational database tables: admissions, diagnoses frame, and prescript. We selected the patients’ data based on the following inclusion criteria: (i) patients who were admitted to the hospital exactly twice; (i) patients who used at least one with frequently-administered medication (defined by at least 1000 times of total use across patients); and patients with at least one frequent diagnosis (defined by at least 10,000 times of total diagnose across patients). Utilizing frequent features can enhance the generalizability and reduce the risk of overfitting for our predictive model. After filtering patients’ data, the sample size of our dataset is n = 18,572. On the other hand, the prediction outcome is whether the length of the 2nd admission would be long (positive, ICU LOS≥3) or not (negative, ICU LOS <3). Among those 18, 572 patients, 57.7 % of patients had long ICU LOS. The overall patient selection process is shown in Fig. 1.

3.2. Method overview

Our overall system workflow is shown in Fig. 2. First, we initiated our data processing pipeline by meticulously extracting features from three relevant categories (Fig. 2A and Section 3.3). Next, we trained five distinct types of predictive models (Fig. 2B and Section 3.4) and conducted cross validation to tune the hyper-parameters (Fig. 2C and Section 3.5). Finally, we evaluated our models in terms of calibration and discrimination to assess their performance and reliability for the intended application (Fig. 2D and Section 3.6).

3.3. Data preprocessing

We focused on the following five data tables: Admissions, Patients, Diagnoses International Coding Definitions (ICD), Dimension table of ICD (D-ICD) Diagnoses, and Prescriptions. Admissions and Patients tables include the basic and demographic information; Diagnoses ICD and D-ICD Diagnoses tables contain the diagnoses information for each patient and each admission; and the Prescriptions table provides information about prescribed medications during their hospital stays. We first merge the diagnosis and prescriptions tables to the admission table and select relevant data fields as follows. First, to combine Diagnoses ICD and D-ICD Diagnoses tables, we performed a merge based on the ICD code and ICD version. This allows us to match the diagnoses information from Diagnoses ICD with the corresponding detailed diagnoses from D-ICD Diagnoses. By merging the detailed diagnoses with the admission table, we can incorporate the specific diagnoses information into the overall patient admission records. Next, we identified 474 medications with the same name but different units. To ensure data consistency and relevance, we only included patients who were prescribed medications with the same units. Finally, to mitigate the risk of overfitting [22], we removed the low-frequency use (<1000) drugs and low-frequency (<10, 000) diagnoses. After applying these exclusions, we can proceed with merging the remaining relevant medication information from the Prescriptions table with the admission table. After preprocessing the tables, we extracted 220 features in total (as shown in Table 1). The basic information included in our model are Gender, Anchor Age and Length of 1st admission (covariates #1 to #3 in Table 1). For gender, we assigned 0 to female and 1 to male. For the length of 1st admission, we subtracted the admitted time from the discharged time. The length of 2nd admission used the same way. We normalized the Anchor Age and the Length of 1st admission. Also, we kept 81 frequent diagnoses and assigned the diagnoses equal to 1 or 0 if the patients have the diagnoses or not (1 = Yes; 0 = No), listed as covariates #4 to #84 in Table 1. Moreover, we used both the sum dose and the average dose for the medication features (covariates #85 to #220 in Table 1). 68 frequently-used medications were included in our study. Then we have 68 features for the sum dose and 68 features for the average dose of the medication. So we totally kept 136 medication features in the model. For example, suppose a patient took Pravastatin for 10 mg, 15 mg, and 20 mg, then the total dose (covariate #85 in Table 1) will be 45 mg, and the average dose (covariate #153 in Table 1) will be 15 mg, for this patient. The total dose is useful for determining the total amount of medication that a patient has received, which can be important for monitoring the effects of the medication and ensuring that the patient receives the appropriate amount. The average dose provides an estimate of the amount of medication given each dosage.

3.4. Model construction

We built five main predictive models to predict long ICU LOS: Logistic regression (LR), Random Forest (RF), Support Vector Machine (SVM), AdaBoost (AB), and XGBoost (XGB). These machine-learning models are commonly used for biomedical prediction tasks [23]. LR is one of the standard methods of analysis to describe the relationship between the explanatory variables and a discrete dichotomous response variable [24]. RF is known for its high prediction accuracy, especially when dealing with high-dimensional datasets, by combining multiple decision trees to make predictions [25]. SVM learns a classifier by maximizing the margin separating the two classes, and it can handle both linearly and nonlinearly data by utilizing different kernel functions [26]. we utilized the linear kernel in the SVM model. AB builds an additive model by stagewise minimizing the loss function and has demonstrated exceptional performance and robustness across a wide range of tasks [27]. XGB is an implementation of gradient-boosted decision trees designed for speed and performance [28]. We used Pandas [29], NumPy [30], Matplotlib [31], and Scikit-Learn [32] to build our models. We used the same machine (8-core CPU, 8-core GPU and 8 GB memory) and the same programming language (Python) to implement, train, and evaluate all the models.

3.5. Model validation and hyper-parameter tuning

To validate our models, the whole dataset (n = 18,572) was split to 50 % training (9,286), 25 % calibration (4,643), and 25 % evaluation sets (4,643). Since this is a relatively balanced dataset (57.7 % positive), we split the data randomly without using stratified strategy. We adopted 10-fold Cross-Validation (CV) on the 50 % training data to identify the best-performing hyper-parameters. We used Grid Search [33] to find the best combination with the best performance for each model. The search space for each model is as follows. For LR [34], we searched the inverse of regularization strength for the LR model from {100, 10, 1.0, 0.1, 0.01}. The default solver of LR is the Limited-memory Broyden–Fletcher–Goldfarb–Shanno algorithm (LBFGS). For RF [35], we tuned three hyper-parameters: (1) the maximum number of features can be set by taking the square root of the total number of features or by taking the logarithm base 2 of the number of features in each tree { ‘log2’, ‘auto’ }; (2) the minimum number of samples required to split an internal node can choose from {2, 5, 10}; and (3) the number of trees in the forest can choose from {10, 100, 1000}. The base estimator of RF is decision tree. For SVM [36], We searched the inverse of regularization strength for the linear SVM model from {1000, 100, 10, 1.0, 0.1}. Also, we searched the kernel coefficient from {1, 0.1, 0.01, 0.001, 0.0001}. For AB [37], we tuned the weight applied to each classifier at each boosting iteration from {0.0001, 0.001, 0.01, 0.1, 1.0} and the maximum number of estimators from {10, 100, 1000}. The base estimator of AB is decision tree. For XGB [28], we tuned the boosting learning rate from {0.1, 0.2, 0.3}, the number of boosting rounds from {100, 200, 300, 400, 500} and the maximum depth of tree from {2, 3, 4, 5, 6, 7, 8, 9, 10}. The booster of XGB is gradient boosted tree.

3.6. Model calibration and discrimination metrics

We adopted Platt Scaling [38] to calibrate our models on 25 % calibration data. Platt Scaling is a trained sigmoid function to calibrate the output of a probabilistic classifier to improve probability estimates [38]. We used the Pycaleva library [39] to calibrate our models. For discrimination, we used the Area Under the Receiver Operating Characteristic Curve (AUC) as the performance metric. AUC, the average value of the comparison function across all possible positive and negative outcome pairings [40], is a performance metric that can be used to evaluate both balanced and imbalanced datasets. For calibration, we adopted the Expected Calibration Error (ECE) metric [38] to measure how well the predicted probabilities of a classifier align with the true frequencies of the classes. The range of ECE is 0–1 and a low ECE indicates that the classifier is well calibrated.

4. Results

4.1. Demographic analysis

The demographic statistics for the patients included in our study (n = 18,572) are as follows: (i) Among the patients, 48.83 % are female and 51.17 % are male (ii) The statistics data of Anchor Age (i.e., Patient’s age when they visited the hospital). are shown in Table 2. (iii) For the length of the 1st admission, the statistics data are shown in Table 2. 59.52 % of the patients stay in the hospital≥3 days. (iv) For the length of the 2nd admission, 57.70 % of positive patients stay≥3 days.

4.2. Discrimination results and tuned hyper-parameters

The test AUC results of each classifier are shown in Fig. 3. RF exhibited the highest performance, achieving a test AUC of 0.716. And XGB performs well with a 0.712 test AUC. On the other hand, the SVM classifier showcased the lowest performance, with a test AUC of 0.680. The test AUC for the SVM and LR classifiers were relatively close, differing only by 0.001. The detailed 10-fold CV results with the best hyper-parameters are shown in Table 3. The RF classifier performed the best during the validation, which is consistent with the test results. To determine the uncertainty of the average AUC, we also computed the lower bound and the upper bound of 95 % Confidence Intervals (CI) for each classifier. The lower 95 % CI of RF is higher than the upper 95 % CI of LR and SVM, suggesting that RF in general outperformed LR and SVM. Meanwhile, the upper bound of AB and XGB’s 95 % CI is higher than the lower bound of RF’s 95 % CI, implying that the performance of AB and XGB are not far behind that of RF. The p-value of 10-fold CV AUC for five classifiers are shown in Table 4. We used one tail T-test with a 0.05 significance level. The p-values of RF are less than 0.05 compared to LR, AB, and SVM, which indicates that RF is significantly better than LR, AB, and SVM. The p-value of RF and XGB is 0.11539, which is larger than 0.05 and indicates that RF is not significantly better than XGB. Also, AB is significantly better than LR, and LR is significantly better than SVM.

The best combination of hyper-parameters for each classifier are as follows: For LR, the best inverse of regularization strength was 100. For RF, the best maximum number of features was set to the logarithm base 2 of the number of features in each tree, the best minimum number of samples required to split an internal node was 10, and the best number of trees in the forest was 1000. For SVM, the best inverse of regularization strength was 10. The best-performing kernel was linear with a kernel coefficient of 1. For AB, the best boosting iteration weight was 0.1, and the best maximum number of estimators was 1000. For XGB, the best boosting learning rate is 0.1, the best number of boosting rounds is 400 and the best maximum depth of tree is 3.

4.3. Calibration results

The AUC results before and after applying calibration are summarized in Table 5. The results show that all five classifiers kept their discrimination capability after calibration. As shown in the table, the ECE of all five classifiers (LR, RF, SVC, AB, and XGB) are reduced after Platt scaling calibration, particularly the AB classifier. Since RF had the best discrimination performance, with a high AUC both before and after calibration, we further generated a Reliability Diagram [38] to show the relationship between a RF classifier’s predicted probabilities and the true probability of the classes, as shown in Fig. 4. Before the calibration, RF underestimated the actual probability at first and then overestimated the actual probability in the end. After calibration, RF approaches the 45◦ line, indicating better calibration results [38].

4.4. Confusion matrix analysis

For a more comprehensive understanding of our predictive model’s efficacy, we conducted a confusion matrix analysis with a 0.5 decision threshold (Fig. 5). The sensitivity, specificity, precision and F1-Score of five classifiers are summarized in Table 6. RF and AB have the highest sensitivity, which suggests they are better at identifying true positive outcomes. XGB has the highest specificity, which suggests it is more suitable at recognizing true negative outcomes. RF has the highest F1-Score, which indicates RF has the strongest overall performance of the binary classification.

4.5. Feature importance

For our best-performing RF model, we further assessed the influence of covariates by measuring the Mean Decrease in Impurity (MDI) [41]. MDI calculates the average magnitude of how much each feature reduces the impurity of the decision trees within the RF [42]. Features with higher MDI scores are considered more important in prediction. The MDI feature importance for top 10 features in the RF model is plotted in Fig. 6. According to the RF Importance MDI scores, the top three most important features are Length of the 1st Admission (0.055), Anchor Age (0.038), and Phytonadione Average (0.026).

4.6. Running times

The running times used for training 50 % training data for five classifiers are shown in Fig. 7. SVM was the most time-consuming classifier, requiring around 22 s for model training. On the other hand, XGB and LR exhibited a considerably faster training time, completing the training process within 1 s. For RF and AB, their training times were relatively similar, both taking 10–15 s. Training time can be important if a model needs to be retrained frequently due to model drift (i.e., as we acquire new data, the model’s accuracy may decline due to shifts in the data distribution) [43].

5. Discussion

5.1. Findings

Our findings include the following. First, our best predictive model (RF) reached an AUC of 0.716, indicating that the solution to the challenging prediction task of predicting the ICU LOS of the 2nd admission, using only the information collected from the 1st admission, is feasible. For the SVM classifier, we also tried different kernel types (e.g., Polynomial, Radial Basis Function, and Sigmoid), while they did not perform as well as the linear. Second, the test ECE improved all five classifiers (LR, RF, SVC, AB, XGB), and the test AUC of five classifiers either improved slightly or remained at the same level after the calibration. Third, the most important two features are the length of first admission and the patient’s age when they go to the hospital, indicating that the age and previous admission information of the patient have a high correlation with our prediction response. The remaining 8 features from the top 10 important features of RF (Fig. 6) are sum or average dose of the medication, confirming our assumption that the medication information helps to predict ICU LOS. The most important medication features are Phytonadione and Metoprolol Succinate XL, both related to blood control. Phytonadione is used to prevent bleeding [44] and Metoprolol Succinate XL is used to treat high blood pressure [45]. Both total and average dose methods can help extract important features for the prediction task. In addition to employing Grid Search for hyperparameter tuning, we also explored Randomized Search. However, our results indicated that the hyper-parameters identified by the Randomized Search did not outperform Grid Search in terms of model performance. Regarding the binary metrics (Table 6), although the specificity of the models is only moderate, the consequences of failing to identify a long hospital stay when it is present (false negative) can be more serious than the consequences of identifying a long hospital stay when it is not present (false positive), therefore we consider sensitivity as a more critical metric in our study. We also try to adjust decision thresholds from 0.00 to 1.00 at intervals of 0.01 and set the optimal threshold to maximize the F1-Score. The sensitivity slightly increased but the specificity significantly decreased at the optimal threshold. In general, the tuned threshold did not perform as well as the default threshold of 0.5. In summary, the AUC of five classifiers (LR, RF, SVM, AB, XGB) are 0.682, 0.716, 0.681, 0.702, and 0.715. The best predictive model is RF, and the most important feature is the length of first admission.

5.2. Limitations

Below are the limitations of our study. (1) MIMIC-IV 4.0 database [18,19] only contains data from a single medical center, thus we are yet to evaluate our method on multiple datasets sourced from various sites. (2) We only considered basic information, diagnoses, and medication information as our covariates. Therefore, we are yet to extract more related clinical features, such as the clinical test results (e.g., Chartevents or Labevent [18,19]), patients’ information on underlying pathology, or patients’ previous diseases, to potentially improve our prediction. (3) We are yet to consider pandemic-related covariates such as vaccination records or infection history, which could potentially impact the ICU LOS. (4) Including age in our model may introduce the possibility of machine learning bias [46], and we are yet to remove this bias from our model. (5) About 96 % of the features we selected are useful and impact the model performance, yet ~4 % of them (8 features) have a zero MDI score for the RF model. We are yet to remove these features and rebuild our model to investigate their potential impact. (6) The classifiers we applied to the data set are basic and widely used in clinical research studies. Besides, we are yet to consider the time-series information and evaluate more advanced deep learning algorithms, such as Recurrent Neural Network (RNN) [47] or Convolutional Neural Networks (CNN) [48], to further enhance the discrimination results by incorporating information such as the medication functionality in the long term. (7) We are yet to compare our methods empirically with other models and repeat our experiments for multiple trials to ascertain the significance of the differences observed among the comparative model performances. (8) The Grid Search hyper-parameter tuning method is computationally expensive as the number of hyper-parameters and values increases [33]. We are yet to investigate more efficient methods such as Bayesian Optimization [49]. (9) We are yet to expand on the clinical implications of the findings, particularly regarding how predictive models could optimize patient care and resource allocation in hospitals.

6. Conclusion

In this study, we focused on predicting the ICU LOS of the 2nd admission. Our results showed that by using only the information from the 1st admission, our model could potentially predict ICU LOS in the early stage with practical discrimination and calibration performances. The ICU LOS of the 1st admission and the age of patients were identified as the two most significant features for our best performance classifier. Also, we included both sum and average dose of the medication, to consider both cumulative and general dosages of various medications. Our method can help identify patients who may be at risk for prolonged hospital stays, allowing for targeted interactions to prevent complications and improve the performance of patients [50]. Moreover, our effort in model optimization can improve the correctness of prediction and the ability to identify the feature importance accurately. Furthermore, the utilization of LOS predictions aids hospital personnel in effectively handling staffing, bed availability, and achieving cost-efficient care by facilitating improved resource allocation [51]. Finally, the limitations and future work of the study include improving the method by testing it on various datasets, enhancing models with additional relevant clinical features and pandemic-related covariates, rebuilding our models without low-relevance features, exploring alternative machine learning algorithms, and refining the hyper-parameter optimization process.

Supplementary Material

1

Acknowledgement

The authors were funded by the U.S. National Institutes of Health (NIH) (R01EB031030). The content is solely the responsibility of the authors and does not necessarily represent the official views of the NIH. The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.

Data availability statement

The original data used by this article are available in PhysioNet at https://doi.org/10.13026/s6n6-xd98 and https://physionet.org/content/mimiciv/1.0/.

Fig. 1. General data overview. Only patients with frequently-used medications (“frequently” means > 1000 times of total use across patients) and frequent diagnosis (“frequent” means > 10,000 times of total diagnosis across patients) were included.

Fig. 2. Overview of the prediction workflow. AUC: Area Under the Receiver Operating Characteristic Curve, ECE: Expected Calibration Error.

Fig. 3. Test discrimination results before calibration. The test AUC before calibration for our five classifiers. RF has the best performance within the five classifiers.

Fig. 4. Reliability diagram for the Random Forest (RF) model. The red/blue lines represent the performances before/after calibration. The dashed line is a 45-degree “ideal” calibration line.

Fig. 5. Confusion matrices for five classifiers.

Fig. 6. Important features of RF (top 10). MDI: Mean Decrease in Impurity.

Fig. 7. Model training times.

Table 1 Labels and features for the preprocessed MIMIC-IV dataset [18,19]. Three features from each category are enumerated (the full 220 features are listed in Table SA1 in Appendix). The range of all the NUM features is 0–1 after normalization. NOM: Nominal; NUM: Numerical.

Category	#	Name	Description	Data type	Number of possible values (if NOM) or range of values (if NUM)	
Label	–	Prolonged Length of the 2nd Admission	Prolonged length of the second admission (0: <3 days or 1:≥3 days)	NOM	2	
Basic Information	1	Gender	Gender of the patient	NOM	2	
2	Anchor Age	Patient’s age when they go to the hospital	NUM	0–1	
3	Length of the 1st Admission	Length of the first admission	NUM	0–1	
Diagnoses	4	Unspecified Essential Hypertension	Diagnose of the patient (0 = No or 1 = Yes)	NOM	2	
5	Tobacco Use Disorder	Diagnose of the patient (0 = No or 1 = Yes)	NOM	2	
84	Single Liveborn, Born In Hospital, Delivered Without Mention Of Cesarean Section	Diagnose of the patient (0 = No or 1 = Yes)	NOM	2	
Medication (Sum)	85	Pravastatin Sum	Total dose of Pravastatin	NUM	0–1	
86	Enoxaparin Sodium Sum	Total dose of Enoxaparin Sodium	NUM	0–1	
152	Influenza Vaccine Quadrivalent Sum	Total dose of Influenza Vaccine Quadrivalent	NUM	0–1	
Medication (Average)	153	Pravastatin Average	Average dose of Pravastatin	NUM	0–1	
154	Enoxaparin Sodium Average	Average dose of Enoxaparin Sodium	NUM	0–1	
220	Influenza Vaccine Quadrivalent Average	Average dose of Influenza Vaccine Quadrivalent	NUM	0–1	

Table 2 Statistical data for Anchor Age and Length of the 1st admission.

Statistics	Anchor Age (Years Old)	Length of the 1st admission (Days)	
Minimum	0	0.06	
Mean	65.27	5.92	
Median	67.00	3.80	
Maximum	91.00	249.59	
Standard Deviation	16.18	7.81	

Table 3 Discrimination results for five classifiers. AUC: Area Under the Receiver Operating Characteristic Curve; CV: Cross Validation; CI: Confidence Interval.

Classifier		Logistic Regression (LR)	Random Forest (RF)	Support Vector Machine (SVM)	AdaBoost (AB)	XGBoost (XGB)	
10-Fold CV AUC	Average	0.696	0.718	0.694	0.707	0.711	
	AUC 95 % CI Low	0.689	0.708	0.675	0.700	0.705	
	AUC 95 % CI High	0.704	0.729	0.691	0.714	0.717	
Test AUC (Before Calibration)	0.681	0.716	0.680	0.701	0.712	

Table 4 The p-value of 10-Fold CV AUC for five classifiers (before calibration). The significance level is 0.05.

Classifier	Logistic Random Forest (RF)	Regression (LR)	Support Vector Machine (SVM)	AdaBoost (AB)	XGBoost (XGB)	
Logistic Regression (LR)	–	0.00018 (RF is significantly better than LR)	0.02771 (LR is significantly better than SVM)	0.00625 (AB is significantly better than LR)	0.00006 (XGB is significantly better than LR)	
Random Forest (RF)	–	–	0.00012 (RF is significantly better than SVM)	0.00554 (RF is significantly better than AB)	0.11539 (RF is not significantly better than XGB)	
Support Vector Machine (SVM)	–	–	–	0.00311 (AB is significantly better than SVM)	0.00003 (XGB is significantly better than SVM)	
AdaBoost (AB)	–	–	–	–	0.00010 (XGB is significantly better than AB)	
XGBoost (XGB)	–	–	–	–	–	

Table 5 Calibration results for five classifiers. AUC: Area Under the Receiver Operating Characteristic (ROC) Curve; ECE: Expected Calibration Error.

Classifier		Logistic Regression (LR)	Random Forest (RF)	Support Vector Machine (SVM)	AdaBoost (AB)	XGBoost (XGB)	
Test AUC	Before Calibration	0.681	0.716	0.680	0.701	0.712	
	After Calibration	0.682	0.716	0.681	0.702	0.715	
Test ECE	Before Calibration	0.051	0.041	0.017	0.135	0.027	
	After Calibration	0.027	0.023	0.017	0.048	0.021	

Table 6 Sensitivity, Specificity Precision and F1-Score for five classifiers.

Classifier	Logistic Regression (LR)	Random Forest (RF)	Support Vector Machine (SVM)	AdaBoost (AB)	XGBoost (XGB)	
Sensitivity	0.750	0.795	0.752	0.795	0.743	
Specificity	0.486	0.495	0.483	0.463	0.559	
Precision	0.667	0.683	0.667	0.670	0.698	
F1-Score	0.706	0.735	0.707	0.727	0.720	

CRediT authorship contribution statement

Min Zhang: Data curation, Formal analysis, Investigation, Methodology, Software, Validation, Visualization, Writing – original draft. Tsung-Ting Kuo: Formal analysis, Funding acquisition, Investigation, Methodology, Project administration, Resources, Supervision, Validation, Visualization, Writing – review & editing.

Declaration of generative AI and AI-assisted technologies in the Writing process

During the preparation of this work the authors used Grammarly in order to improve readability and language. After using this tool/service, the authors reviewed and edited the content as needed and take full responsibility for the content of the publication.

Declaration of competing interest

None declared

Appendix A. Supplementary data

Supplementary data to this article can be found online at https://doi.org/10.1016/j.compbiomed.2024.108451.
==== Refs
References

[1] Tipton K , Leas BF , Mull NK , , Interventions to Decrease Hospital Length of Stay [Internet], Agency for Healthcare Research and Quality (US), Rockville (MD), 2021 Sep (Technical Brief, No. 40.) Introduction, https://www.ncbi.nlm.nih.gov/books/NBK574438/. (Accessed 6 April 2023).
[2] Hunter A , Johnson L , Coustasse A , Reduction of intensive care unit length of stay: the case of early mobilization, Health Care Manag. 33 (2 ) (2014 Apr-Jun) 128–135, 10.1097/HCM.0000000000000006. https://pubmed.ncbi.nlm.nih.gov/24776831/. (Accessed 19 March 2023).
[3] Li H , Tao H , Li G , Predictors and reasons for inappropriate hospitalization days for surgical patients in a tertiary hospital in Wuhan, China: a retrospective study, BMC Health Serv. Res 21 (1 ) (2021 Sep 1) 900, 10.1186/s12913-021-06845-y.34470637
[4] Monegro AF , Muppidi V , Regunath H , Hospital acquired infections, in: StatPearls [Internet], StatPearls Publishing, Treasure Island (FL), 2023 Jan [Updated 2023 Feb 12], https://www.ncbi.nlm.nih.gov/books/NBK441857/. (Accessed 6 April 2023).
[5] Centers for Disease Control and Prevention. HAI and Antibiotic Use Prevalence Survey. https://www.cdc.gov/hai/eip/antibiotic-use.html (Retrieved on 6 April 2023).
[6] Izadi N , Eshrati B , Mehrabi Y , Etemad K , Hashemi-Nazari SS , The national rate of intensive care units-acquired infections, one-year retrospective study in Iran, BMC Publ. Health 21 (1 ) (2021 Mar 29) 609, 10.1186/s12889-021-10639-6. https://bmcpublichealth.biomedcentral.com/articles/10.1186/s12889-021-10639-6. (Accessed 6 April 2023).
[7] McDermott KW , IBM Watson Health), Jiang HJ (AHRQ). Characteristics and Costs of Potentially Preventable Inpatient Stays, HCUP Statistical Brief #259. June 2020, Agency for Healthcare Research and Quality, Rockville, MD, 2017. www.hcup-us.ahrq.gov/reports/statbriefs/sb259-Potentially-Preventable-Hospitalizations-2017.pdf. (Accessed 6 April 2023).
[8] LaFaro RJ , Pothula S , Kubal KP , Inchiosa ME , Pothula VM , Yuan SC , Inchiosa MA Jr , Neural network prediction of ICU length of stay following cardiac surgery based on pre-incision variables, PLoS One 10 (12 ) (2015) e0145395 (Retrieved on 20 March 2024).26710254
[9] Mollaei N , Londral AR , Cepeda C , Azevedo S , Santos JP , Coelho P , Gamboa H , Length of stay prediction in acute intensive care unit in cardiothoracic surgery patients, in: 2021 Seventh International Conference on Bio Signals, Images, and Instrumentation (ICBSII), IEEE, 2021, March, pp. 1–5 (Retrieved on 20 March 2024).
[10] Chattopadhyay A , Chatterjee S , Predicting ICU length of stay using Apache-IV in persons with severe sepsis–a pilot study, J. Epidemiol. Res 2 (1 ) (2015) 1–8 (Retrieved on. (Accessed 20 March 2024).
[11] Abd-Elrazek MA , Eltahawi AA , Abd Elaziz MH , Abd-Elwhab MN , Predicting length of stay in hospitals intensive care unit using general admission features, Ain Shams Eng. J 12 (4 ) (2021) 3691–3702 (Retrieved on 20 March 2024).
[12] Alghatani K , Ammar N , Rezgui A , Shaban-Nejad A , Predicting intensive care unit length of stay and mortality using patient vital signs: machine learning model development and validation, JMIR Med Inform. 9 (5 ) (2021) e21347, 10.2196/21347. URL: https://medinform.jmir.org/2021/5/e21347. (Accessed 19 March 2023).33949961
[13] Nallabasannagari AR , Reddiboina M , Seltzer R , Zeffiro T , Sharma A , Bhandari M , All data inclusive, deep learning models to predict critical events in the medical information Mart for intensive care III database (MIMIC III). 10.48550/arXiv.2009.01366, 2020. (Accessed 26 June 2023). Retrieved on.
[14] Ma X , Si Y , Wang Z , Wang Y , Length of stay prediction for ICU patients using individualized single classification algorithm, Comput. Methods Progr. Biomed 186 (2020) 105224 (Retrieved on 20 March 2024).
[15] Li C , Chen L , Feng J , Wu D , Wang Z , Liu J , , Prediction of length of stay on the intensive care unit based on least absolute shrinkage and selection operator, IEEE Access 7 (2019) 110710–110721 (Retrieved on 20 March 2024).
[16] Rocheteau Emma , Liò Pietro , Hyland Stephanie , Temporal pointwise convolutional networks for length of stay prediction in the intensive care unit, in: ACM Conference on Health, Inference, and Learning (ACM CHIL ‘21), April 8–10, 2021, Virtual Event, USA. ACM, New York, NY, USA, 2021, p. 19, 10.1145/3450439.3451860. (Accessed 6 April 2023). Retrieved on.
[17] Ruppert MM , Loftus TJ , Small C , Li H , Ozrazgat-Baslanti T , Balch J , Holmes R , Tighe PJ , Upchurch GR Jr. , Efron PA , Rashidi P , Bihorac A , Predictive modeling for readmission to intensive care: a systematic review, Crit Care Explor. 5 (1 ) (2023 Jan 6) e0848, 10.1097/CCE.0000000000000848. https://pubmed.ncbi.nlm.nih.gov/36699252/. (Accessed 19 March 2023).36699252
[18] Goldberger A , Amaral L , Glass L , Hausdorff J , Ivanov PC , Mark R , Stanley HE , PhysioBank, PhysioToolkit, and PhysioNet: components of a new research resource for complex physiologic signals [Online], Circulation 101 (23 ) (2000) e215–e220 (Retrieved on 1 March 2022).10851218
[19] Johnson A , Bulgarelli L , Pollard T , Horng S , Celi LA , Mark R , MIMIC-IV, PhysioNet, 2021, 10.13026/s6n6-xd98, version 1.0.. (Accessed 1 March 2022).
[20] Gupta M , Gallamoza B , Cutrona N , Dhakal P , Poulain R , Beheshti R , An extensive data processing pipeline for MIMIC-IV, Proc Mach Learn Res. 193 (2022 Nov) 311–325.36686986
[21] Bardak B , Tan M , Using clinical drug representations for improving mortality and length of stay predictions, in: 2021 IEEE Conference on Computational Intelligence in Bioinformatics and Computational Biology (CIBCB), 2021, pp. 1–8, 10.1109/CIBCB49929.2021.9562819. Melbourne, Australia, https://ieeexplore.ieee.org/abstract/document/9562819. (Accessed 12 February 2023).
[22] Goyal C , Complete guide to prevent overfitting in neural networks (Part-1). https://www.analyticsvidhya.com/blog/2021/06/complete-guide-to-prevent-overfitting-in-neural-networks-part-1/, June 12, 2021. (Accessed 1 March 2023).
[23] Edelson M , Kuo TT , Generalizable prediction of COVID-19 mortality on worldwide patient data, JAMIA Open 5 (2 ) (2022 May 25) ooac036, 10.1093/jamiaopen/ooac036. Erratum in: JAMIA Open. 2022 Dec 06;5(4): ooac102. (Retrieved on 20 Feb 2024).
[24] Hosmer DW , Lemeshow S , Applied Logistic Regression, second ed., John Wiley & Sons, Inc, New York, NY, 2000 10.1002/0471722146.ch1. (Accessed 1 July 2023).
[25] Qi Y , Random forest for bioinformatics, in: Zhang C , Ma Y (Eds.), Ensemble Machine Learning, Springer, New York, NY, 2012, 10.1007/978-1-4419-9326-7_11. (Accessed 1 July 2023).
[26] Kecman V Support Vector Machines – An Introduction. In: Wang L (eds) Support Vector Machines: Theory and Applications. Studies in Fuzziness and Soft Computing, vol vol. 177 . Springer, Berlin, Heidelberg. 10.1007/10984697_1 (Retrieved on 10 July 2023).
[27] Schapire RE , Explaining AdaBoost, in: Scholköpf B , Luo Z , Vovk V (Eds.), Empirical Inference, Springer, Berlin, Heidelberg, 2013, 10.1007/978-3-642-41136-6_5. (Accessed 11 March 2023). Retrieved on.
[28] Chen T , Guestrin C , XGBoost: a scalable tree boosting system, in: Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, ACM, New York, NY, USA, 2016, pp. 785–794, 10.1145/2939672.2939785. (Accessed 15 March 2024).
[29] McKinney W , others, Data structures for statistical computing in python, in: Proceedings of the 9th Python in Science Conference vol. 445 , 2010, pp. 51–56. Retrieved on. (Accessed 1 March 2022).
[30] Harris CR , Millman KJ , van der Walt SJ , , Array programming with NumPy, Nature 585 (2020) 357–362, 10.1038/s41586-020-2649-2 (Publisher link). (Retrieved on. (Accessed 1 March 2022).32939066
[31] Matplotlib JD , Hunter, Matplotlib: a 2D graphics environment, Comput. Sci. Eng 9 (3 ) (2007) 90–95 (Retrieved on. (Accessed 1 March 2022).
[32] Scikit-learn, Machine learning in Python, Pedregosa , JMLR 12 (2011) 2825–2830.
[33] Shah Rahul , Tune hyperparameters with GridSearchCV, july 20th. https://www.analyticsvidhya.com/blog/2021/06/tune-hyperparameters-with-gridsearchcv/, 2022. (Accessed 4 June 2023).
[34] Ahmed Arafa A , Radad M , Badawy M , El-Fishawy N , Logistic regression hyperparameter optimization for cancer classification, Menoufia J.Electron. Eng. Res 31 (1 ) (2022) 1–8, 10.21608/mjeer.2021.70512.1034 (Retrieved on 06 November 2023).
[35] Probst Philipp , Wright Marvin N. , Boulesteix Anne-Laure , Hyperparameters and tuning strategies for random forest. 10.1002/widm.1301, Jan 28, 2019. (Accessed 6 November 2023).
[36] Gold Carl , Holub Alex , Sollich Peter , Bayesian approach to feature selection and parameter tuning for support vector machine classifiers, Neural Network. 18 (Issues 5–6 ) (2005) 693–701, 10.1016/j.neunet.2005.06.044. ISSN 0893–6080. (Accessed 6 November 2023).
[37] Gao Rongfang , Liu Zhanyu , An improved AdaBoost algorithm for hyperparameter optimization, J. Phys.: Conf. Ser 1631 (2020) 012048 (Retrieved on 06 November 2023).
[38] Huang Yingxiang , Li Wentao , Macheret Fima , Gabriel Rodney A. , Lucila Ohno-Machado, A tutorial on calibration measurements and calibration models for clinical prediction models, J. Am. Med. Inf. Assoc 27 (4 ) (April 2020) 621–633, 10.1093/jamia/ocz228. (Accessed 24 June 2023).
[39] Weigl Martin , pycaleva, GitHub repository (2022) (03/01/, https://github.com/MartinWeigl/pycaleva. (Accessed 24 June 2023).
[40] Lasko TA , Bhagwat JG , Zou KH , Ohno-Machado L , The use of receiver operating characteristic curves in biomedical informatics, J. Biomed. Inf 38 (5 ) (2005) 404–415, 10.1016/j.jbi.2005.02.008. Retrieved on 20 Feb 2024).
[41] Oh Sejong , Predictive case-based feature importance and interaction, Inf. Sci 593 (2022) 155–176, 10.1016/j.ins.2022.02.003. ISSN 0020–0255. (Accessed 20 February 2023).
[42] Louppe Gilles , , Understanding variable importances in forests of randomized trees, Adv. Neural Inf. Process. Syst 26 (2013) (Retrieved on 06 November 2023).
[43] Sahiner Berkman , Chen Weijie , Samala Ravi K. , Petrick Nicholas , Data drift in medical machine learning: implications and potential remedies, Br. J. Radiol 96 (1150 ) (2023) 20220878, 10.1259/bjr.20220878008, 1 October. (Accessed 20 February 2024).36971405
[44] I.B.M. Micromedex, Phytonadione (injection route) side effects, Mayo Foundation for Medical Education and Research (MFMER) (2023), 1998-, https://www.mayoclinic.org/drugs-supplements/phytonadione-injection-route/side-effects/drg-20421713?p=1. (Accessed 20 April 2023).
[45] I.B.M. Micromedex, Metoprolol (oral route) side effects, 1998, Mayo Foundation for Medical Education and Research (MFMER) (2023), https://www.mayoclinic.org/drugs-supplements/metoprolol-oral-route/side-effects/drg-20071141?p=1 (Retrieved on 27 Feb 2023).
[46] Garcia de Alford AS , Hayden SK , Wittlin N , Atwood A , Reducing age bias in machine learning: an algorithmic approach, SMU Data Sci. Rev 3 (2 ) (2020) 11. Retrieved on 20 Feb 2024).
[47] Grossberg Stephen , “Recurrent neural networks, Scholarpedia 8.2 (2013) 1888. Retrieved on 06 November 2023).
[48] Albawi Saad , Mohammed Tareq Abed , Al-Zawi Saad , Understanding of a convolutional neural network. 2017 International Conference on Engineering and Technology (ICET), Ieee, 2017 (Retrieved on 27 Feb 2023).
[49] Tong Yu , Zhu Hong , Hyper-Parameter Optimization, A review of algorithms and applications. 10.48550/arXiv.2003.05689, Mar 12, 2020. Retrieved on 6 June 2023.
[50] Tipton K , Leas BF , Mull NK , , Interventions to Decrease Hospital Length of Stay [Internet], Agency for Healthcare Research and Quality (US), Rockville (MD), 2021 Sep (Technical Brief, Number. 40. Introduction, https://www.ncbi.nlm.nih.gov/books/NBK574438/. (Accessed 26 June 2023).
[51] Viswanadham N , Balaji K , Resource allocation for healthcare organizations, in: 2011 IEEE International Conference on Automation Science and Engineering, 2011, pp. 543–548, 10.1109/CASE.2011.6042473. Trieste, Italy, https://ieeexplore.ieee.org/abstract/document/6042473?casa_token=SLDSgTVSgakAAAAA:9xyqSlxZ8W9ss9h6e3PIQCmLSNmxqWW9kRXltjEgrIvk_cTEY_CIpMAP96bzqzpdSEKdwSV_nz0. (Accessed 26 June 2023).
