
==== Front
BMJ Paediatr Open
BMJ Paediatr Open
bmjpo
bmjpo
BMJ Paediatrics Open
2399-9772
BMJ Publishing Group BMA House, Tavistock Square, London, WC1H 9JR

39038911
10.1136/bmjpo-2023-002365
bmjpo-2023-002365
Original Research
Health Service
1506
Development of machine learning models predicting mortality using routinely collected observational health data from 0-59 months old children admitted to an intensive care unit in Bangladesh: critical role of biochemistry and haematology data
http://orcid.org/0000-0002-7852-6569
Das Subhasish 12subhasish.das.bd@gmail.com

Erdman Lauren 3456larunerdman1@gmail.com

Brals Daniella 7d.brals@aighd.org

Boczek Bartlomiej 7b.boczek@student.vu.nl

Hasan S M Tafsir 1tafsir.hasan@icddrb.org

Massara Paraskevi 38paraskevi.massara@sickkids.ca

Alam Md Ashraful 9mdashraful.alam@uq.net.au

Fahim Shah Mohammad 110mohammad.fahim@icddrb.org

Mahfuz Mustafa 1mustafa@icddrb.org

Hoogendoorn Mark 11m.hoogendoorn@vu.nl

Zuiderent-Jerak Teun 12teun.zuiderent-jerak@vu.nl

Bandsma Robert H J 813robert.bandsma@sickkids.ca

http://orcid.org/0000-0002-4607-7439
Ahmed Tahmeed 1tahmeed@icddrb.org

Voskuijl Wieger 714w.p.voskuijl@amsterdamumc.nl

1 Nutrition Research Division, International Centre for Diarrhoeal Disease Research, Bangladesh, Dhaka, Bangladesh
2 Liggins Institute, University of Auckland, Auckland, New Zealand
3 The Center for Computational Medicine, The Hospital for Sick Children, Toronto, Ontario, Canada
4 Department of Computer Science, University of Toronto, Toronto, Ontario, Canada
5 Translational Medicine Program, The Hospital for Sick Children, Toronto, Ontario, Canada
6 James M. Anderson Center for Health Systems Excellence, Cincinnati Children’s Hospital Medical Center and University of Cincinnati School of Medicine, Cincinnati, Ohio, USA
7 Department of Global Health, Amsterdam Institute for Global Health and Development, Amsterdam, Netherlands
8 Department of Nutritional Sciences, University of Toronto, Toronto, Ontario, Canada
9 The University of Queensland Poche Centre for Indigenous Health, Saint Lucia, Queensland, Australia
10 Division of Nutritional Sciences, Cornell University, Ithaca, New York, USA
11 Faculty of Science, Department of Computer Science, Vrije University, Amsterdam, The Netherlands
12 Athena Institute, Vrije Universiteit Amsterdam, Amsterdam, Netherlands
13 Centre for Global Child Health, Hospital for Sick Children, Toronto, Ontario, Canada
14 Amsterdam UMC, University of Amsterdam, Amsterdam Centre for Global Child Health & Emma Children’s Hospital, Amsterdam, The Netherlands
DrSubhasishDas; subhasish.das.bd@gmail.com
No, there are no competing interests.

Supplemental material: Supplemental material This content has been supplied by the author(s). It has not been vetted by BMJ Publishing Group Limited (BMJ) and may not have been peer-reviewed. Any opinions or recommendations discussed are solely those of the author(s) and are not endorsed by BMJ. BMJ disclaims all liability and responsibility arising from any reliance placed on the content. Where the content includes any translated material, BMJ does not warrant the accuracy and reliability of the translations (including but not limited to local regulations, clinical guidelines, terminology, drug names and drug dosages), and is not responsible for any error and/or omissions arising from translation and adaptation or otherwise.

2024
22 7 2024
8 1 e00236511 11 2023
03 7 2024
Copyright © Author(s) (or their employer(s)) 2024. Re-use permitted under CC BY-NC. No commercial re-use. See rights and permissions. Published by BMJ.
2024
https://creativecommons.org/licenses/by-nc/4.0/ This is an open access article distributed in accordance with the Creative Commons Attribution Non Commercial (CC BY-NC 4.0) license, which permits others to distribute, remix, adapt, build upon this work non-commercially, and license their derivative works on different terms, provided the original work is properly cited, appropriate credit is given, any changes made indicated, and the use is non-commercial. See: http://creativecommons.org/licenses/by-nc/4.0/.

Abstract

Introduction

Treatment in the intensive care unit (ICU) generates complex data where machine learning (ML) modelling could be beneficial. Using routine hospital data, we evaluated the ability of multiple ML models to predict inpatient mortality in a paediatric population in a low/middle-income country.

Method

We retrospectively analysed hospital record data from 0-59 months old children admitted to the ICU of Dhaka hospital of International Centre for Diarrhoeal Disease Research, Bangladesh. Five commonly used ML models- logistic regression, least absolute shrinkage and selection operator, elastic net, gradient boosting trees (GBT) and random forest (RF), were evaluated using the area under the receiver operating characteristic curve (AUROC). Top predictors were selected using RF mean decrease Gini scores as the feature importance values.

Results

Data from 5669 children was used and was reduced to 3505 patients (10% death, 90% survived) following missing data removal. The mean patient age was 10.8 months (SD=10.5). The top performing models based on the validation performance measured by mean 10-fold cross-validation AUROC on the training data set were RF and GBT. Hyperparameters were selected using cross-validation and then tested in an unseen test set. The models developed used demographic, anthropometric, clinical, biochemistry and haematological data for mortality prediction. We found RF consistently outperformed GBT and predicted the mortality with AUROC of ≥0.87 in the test set when three or more laboratory measurements were included. However, after the inclusion of a fourth laboratory measurement, very minor predictive gains (AUROC 0.87 vs 0.88) resulted. The best predictors were the biochemistry and haematological measurements, with the top predictors being total CO2, potassium, creatinine and total calcium.

Conclusions

Mortality in children admitted to ICU can be predicted with high accuracy using RF ML models in a real-life data set using multiple laboratory measurements with the most important features primarily coming from patient biochemistry and haematology.

Health services research
Statistics
==== Body
pmcWHAT IS ALREADY KNOWN ON THE TOPIC

WHAT THIS STUDY ADDS

This study develops a machine learning model showing the outsized importance of biochemistry values for predicting patient death in the International Centre for Diarrhoeal Disease Research, Bangladesh ICU. Predictive power increased substantially when multiple measurements of laboratory tests were added to clinical parameters.

HOW THIS STUDY MIGHT AFFECT RESEARCH, PRACTICE OR POLICY

This work implies that in a setting where basic laboratory testing facilities are available, machine learning could play an additional role in predicting all-cause mortality in sick children admitted to ICU by using the biochemical and haematological features.

Introduction

Children admitted to hospitals for an acute illness are at high risk of clinical deterioration and subsequent death.1 Evidence suggests that clinical warning signs (CWS, eg, obstructed breathing, severe respiratory distress, reduced consciousness) precede life-threatening consequences like clinical deterioration and mortality in these children.2 3 Hence, efforts have been made to identify patients at risk of clinical deterioration and to develop mortality and morbidity prediction models for these vulnerable, sick children.47 These CWS have proven to be predictive for clinical deterioration and systemic illness, and are used in emergency triage assessment and initial treatment for hospitalised children in low and middle-income countries (LMICs).8 9

The risk of adverse clinical outcomes, including mortality, is expected to evolve during hospitalisation, as sick children may deteriorate during hospital stay despite strict adherence to protocolised medical and nutritional treatment.10 11 Existing risk prediction models for sick children in hospital (eg, Paediatric Risk of Mortality, the Paediatric Index of Mortality and the Paediatric Early Warning Score) have been constructed using data from well-resourced settings and are clinically used in many high-resource care settings.5 12 However, many of such early warning scores suffer from methodological issues that might limit their use in LMIC including Bangladesh. Therefore, many physicians in such resource-limited settings do not use these prediction models and only resort to their clinical judgement to manage critically ill children, increasing the risk of not detecting clinical deterioration rapidly.13

The temporal nature, different granularity and irregularity of measurements and the complexity of the relationships between the other factors make clinical data complex. Hence, extracting predictive models from clinical data to detect clinical deterioration can be incredibly challenging. Machine learning (ML) techniques have been shown to be able to cope with such complexity better than conventional statistical methods, as was shown in a data set called Medical Information Mart for Intensive Care database using adult intensive care unit (ICU) data.14

We evaluated multiple ML models to assess their ability to identify the CWS, along with other routinely collected variables predictive of all-cause mortality for critically ill 0–59 months old children admitted to an intensive care unit in Bangladesh.

Methods

The study was conducted at the International Centre for Diarrhoeal Disease Research, Bangladesh (icddr,b). icddr,b is an international health research organisation located in Dhaka, Bangladesh. The ICU of icddr,b Dhaka hospital manages about 1000 critically ill under-5 children each year.15 All children with complications or other associated problems, who required admission (and care) to the ICU (aged 0–59 months) of the Dhaka Hospital of icddr,b from January 2011 to December 2019 were enrolled in this study.

Data source

Dhaka hospital of icddr,b is a paperless hospital. All patient-related data, including information related to registration, triage, follow-up, laboratory investigations and discharge, are routinely recorded in the electronic database (called ‘SHEBA’) of the hospital. All data required for this study was retrieved from this electronic database.

Model predictors

The models developed use demographic, anthropometric, clinical, biochemistry and haematological data for the prediction of mortality (supplementary data, online supplemental tables 1–3). Demographic information was included as potential predictors, including age and sex, as well as anthropometric measurements (eg, weight and height), reduced consciousness, respiratory distress (eg, chest indrawing), fever and respiration rate. Biochemistry values included as potential predictors were: anion gap, chloride, creatinine, potassium, sodium, calcium, magnesium and CO2 and haematological values included as potential predictors were: haemoglobin, haematocrit and haemoglobin cell indices (including red cell distribution width), leucocyte count and white cell differentiation, thrombocyte count, red blood cell count. Biochemistry and haematology measurements were taken for a variable number of days/times per patient.

Data pre-processing

As most baseline clinical features were complete, all variables with missing data, except weight for age z-score (WAZ), were removed. Patients missing WAZ were then removed, as they were systematically different from our data set, leaving a data set of 4060 patients total with complete data (online supplemental figure 1). Additional predictors were added to the data using feature engineering including the number of days admitted in hospital before days to transfer to ICU, whether patients were transferred to ICU the same day they came to hospital and whether systolic and diastolic blood pressure were measured. The data was split to have a 15% test set and 85% training set.

Laboratory measurements were included in the model ordinally (ie, first, second, third measurement of the given laboratory, for example, first creatinine, second creatinine, third creatinine). For missing values in which the laboratory had been measured in the patient at other time points, that patient’s specific average was used. For the remaining patients, with no measurement of that laboratory marker, the weighted overall mean was used. The mean value was weighted to reduce the influence of patients with many measurements and increase the influence of patients with fewer measurements.

proportionk=1nk, where k=patient ID,n=number of repeated measurements for lab value for patient k

weightk=proportionk∑kKproportionk, weight in meancalculation for patient k ∈K

Analyses were tested with unweighted mean imputation as well. Imputation was done in the training and test sets separately. Patients missing all data from any of the clinical or biochemistry or haematology data sets were excluded (14%). Within the remaining (85%) training set, 10-fold cross-validation was used for hyperparameter selection (such as tree depth and class weights) and a second, independent 10-fold cross-validation was used to find the best values for the hyperparameters and done on the training set.

Statistical analysis

Analyses were conducted using R (V.4.0.2).16 Five commonly used ML models for tabular data were evaluated for their ability to predict patient death: Logistic regression (LR), least absolute shrinkage and selection operator (LASSO),17 elastic net (EL),18 gradient boosting trees (GBT)19 and random forest (RF).20 This was done using the glmnet (V.4.1–1),21 gbm (V.2.1.8.1)22 and randomForest (V.4.6–14)23 R packages. Models were compared primarily based on validation performance measured by mean 10-fold cross-validation area under the receiver-operator curve (AUROC), which evaluates the ability of the model to separate the two classes (death vs survival) in terms of the sensitivity-specificity trade-off. An analysis was performed to evaluate the number of laboratory tests that contribute to model performance for each of the models evaluated. This was done by restricting the number of laboratory samples included per person from zero to six laboratory measurements and assessing validation set performance. Top predictors or features were selected using RF mean decrease Gini score (feature importance values).

Results

ICU data from 5669 patients was used for mortality prediction using ML. Following missing data removal, this set was reduced to 3505 patients (10% death, 90% survived). Of these patients, 50% transferred to the ICU the same day they were admitted to the hospital (SD=50%) and the mean patient age was 10.8 months (SD=10.5) for the full data set and 9.6 months (SD=9.2) after removing the patient IDs with missing variables. Of these patients, 42% presented with vomiting, 35% with fever, 74% with diarrhoea, 14% with cough, 8% with oedema and 6% with jaundice. The distribution of the data is comparable with the full data (n=5669) (summary in online supplemental table 1). The complete data (n=3505) were stratified and divided into 2979 training samples (85% of the data set) and 526 test samples (15% of the data set).

Five models were tested for their ability to predict the patient mortality, LR, LASSO, EL, RF and GBT. The top performing models based on the validation performance as measured by mean 10-fold cross-validation AUROC were the tree-based models: RF and GBT (figure 1). In all five models, those predicting mortality using zero laboratory measurements were the worst. LR was least able to take advantage of additional laboratory measurements, whereas RF, LASSO, EL and GBT showed at least marginal improvement in test set prediction with each additional repeated laboratory measurement included. GBT and RF remained the best performing models relative to EL, LASSO and LR, regardless of how many values of the same laboratory measurement were included (0–6). However, GBT validation performance was superior to RF performance when no laboratory values were included whereas RF performance was superior (and model performance overall was superior) when one or more laboratory values were included.

Figure 1 Comparison of models to predict all-cause mortality. Mean performance (AUROC, y-axis) of 10-fold cross validation for each model with 0–6 laboratory values included (x-axis). Training data performance is shown on the left panel and test (held-out, validation) data performance is shown on the right. AUROC, area under the receiver-operator curve; EL, elastic net; GBT, gradient boosting trees; LASSO, least absolute shrinkage and selection operator; LR, logistic regression; RF, random forest.

The tree-based models we selected (RF and GBT) require the selection of a complexity hyperparameter which defines how many decision tree models can be combined for the prediction. When different model complexity (number of trees 500–5000) was considered for the RF and GBT models that used three laboratory values with an additional 10-fold cross-validation, RF consistently outperformed GBT and tree number had minimal impact on model results (figure 2). When selecting for complexity, one usually prefers the highest model performance with the lowest complexity, therefore 1000 trees were found to be optimal for the RF model.

Figure 2 Comparison of random forest versus gradient boosting trees. Boxplots showing the distribution of a second, independent 10-fold cross validation at varying tree counts for the top two performing models in A, gradient-boosted trees (GBT) and random forest (RF). Training data performance is shown on the left panel, validation data performance is shown on the right. AUROC, area under the receiver-operator curve.

The best model was selected as RF with 1000 trees and three laboratory values, with the acknowledgement that inclusion of additional repeated laboratory values would likely lead to marginally improved prediction. The test set AUROC (95% CI) was 0.85 (0.81 to 0.90) for three laboratory values, 0.87 (0.82 to 0.92) for four laboratory values, 0.88 (0.83 to 0.92) for five laboratory values and 0.88 (0.83 to 0.93) for six laboratory values. In contrast, the test set performance was only 0.59 (0.49 to 0.68) for the RF model with no additional laboratory values.

The most important features for the selected best model primarily came from biochemistry and haematology measurements (figure 3). The top predictors were biochemistry measurements of Total CO2, creatinine, potassium and total calcium. Haematocrit and platelet count were the top haematological predictors (figure 3). Age and diarrhoea duration were the only two clinical observations used by the selected model. As repeated, additional laboratory values (eg, fourth, fifth, sixth) were added to the final RF model, age was used when four repeated measurements were included but not when five or six were. Instead, the same laboratory measurements remained as the most important predictors, with the last (most recent) potassium measurement being the most important (figure 4). Indeed, each model shows a feature preference for more recent measurements, with a few exceptions, indicating how predictive utility reveals timing among the features. In contrast to this trend, the first calcium measurement was, indeed all repeated calcium measurements were, always among the top 10 most important features (figures34).

Figure 3 Feature importance of top variables in selected random forest model with 1500 trees and three laboratory measurements. TCO2 1, total CO2 first measurement; T calcium 3: total calcium third measurement; T magnesium 2, total magnesium second measurement.

Figure 4 Feature importance when 4 (panel A), 5 (panel B) or 6 (panel C) laboratory values were included in the model. Feature importance based on mean decreased Gini value in the model, where higher values mean more importance. T, Total; TCO2 1, total CO2 first measurement; PCV, Packed cell volume .

Discussion

We showed that in a real-life data set from critically ill children admitted to ICU in Dhaka, Bangladesh using ML models, mortality can be predicted with an AUROC of up to 0.88 (test set), as well or better than internal held-out testing of comparable ICU risk scores.24 RF was the best performing ML model with the most important features primarily coming from patient biochemistry and haematology. The findings of our paper add to a growing body of investigation on which analytical models are most effective for prediction of mortality in acutely ill children.25

A systematic review done by Mangold et al echoed our findings. They reported that ML models can accurately predict death in neonates.17 We found that logistic regression, least absolute shrinkage and selection operator, EL, GBT were outperformed by the RF ML models. Podda et al, Shulta et al and Kefi et al, applied different ML algorithms to predict mortality in different groups of infants and neonates. The AUC values they reported for the RF models were 0.91,26 0.8227 and 0.828 which are very close to our results. Similarly, our group has recently published on the use of daily clinical data to predict mortality in severely malnourished children, using extended survival models that show results that only have a slightly lower AUC (0.81) than the present ML model.29

Nevertheless, we also need to note that logistic regression has shown strong performance for the prediction of mortality in acutely ill children in a similar setting to our own.15 Moreover, a recent paper from the Childhood Acute Illness and Nutrition (CHAIN) network, using ML to predict mortality in a large international cohort of acutely ill children admitted to a hospital concludes that these complex ML models only offered modest improvements in accuracy compared with simpler tools using a limited subset of (clinical) variables. The authors conclude by suggesting that investing in artificial intelligence algorithms to improve LMIC paediatric management may prove less effective than expanding access to reliable bedside or simple laboratory assessment.30 Although the data from the present study is from a more homogeneous, larger patient cohort, interestingly our work shows additional gains from the use of ML models. Using data from an ICU in Bangladesh more than three biochemistry or haematology tests did provide additional predictive power for mortality with a final AUC of 0.88 when using six additional laboratory tests. Clearly the addition of multiple values of biochemical and haematological parameters to our ML algorithm improved the predictive accuracy as without the laboratory parameters the AUROC was 0.59.

From the features associated with mortality that the RF model identified, biochemistry and haematological parameters were the most influential predictors of mortality during admission to the ICU. Total CO2, serum potassium, creatinine and calcium labels were the top features from the biochemistry group, whereas haematocrit values were the most important feature from the haematology group. Haematocrit can be considered as a proxy for haemoglobin and CO2 as a proxy for body homeostasis, which is disturbed in sepsis. Both low haemoglobin (anaemia) and sepsis are known to be associated with a poor outcome in hospitalised, sick children in Malawi.31 32 Unlike the recent ML paper from the CHAIN cohort,33 anthropometry was not the strongest predictor of death in the present study, rather total CO2, potassium, creatinine and calcium levels are. Here we need to consider the clinical characteristics of the patients. We gathered data from a hospital where all the children were admitted to the ICU with diarrhoeal diseases. Inpatient mortality from diarrhoeal disease is associated with abnormal electrolyte values.34 The association of high creatinine values to the mortality of our patients is also not surprising as acute renal failure due to diarrhoeal diseases is also common.35 36 Additionally, the ML models in our study used most laboratory measurements for prediction, whereas the ML model in the CHAIN paper only selected eight features, which can be explained by the difference between the two data sets, including differing homogeneity and outcomes (clinical deterioration vs death). Indeed, inclusion of so many features limits the utility of methods which seek to enhance explainability of a given prediction, such as local interpretable model-agnostic explanations.37

Finally, this study assessed the number of laboratory measurements needed to obtain such a strong prediction and after the inclusion of a fourth laboratory measurement, we found very minor predictive gains (0.87 vs 0.88) from the inclusion of more laboratory values, thus showing that the benefits from additional testing are not limitless but rather they reach a natural peak.

Strengths of our study are the analytical use of real life, longitudinal data from an ICU in a low-resource setting (icddr,b) which is rare. In addition, this study shows strong test-set prediction of mortality (0.88 AUROC) and emphasises the need for (serial) biochemistry and haematology laboratory measurements to make these predictions. Limitations of this study include the fact that the data is from children who were all admitted to the ICU of a diarrhoeal disease hospital, so they represent a homogeneous group with an a-priori high chance of mortality. Therefore, identifying predictors that can distinguish those at high risk of mortality among this already high-risk group may be difficult and the results from this study might not be generalisable to sick children outside an intensive care setting. The data were collected during an 8-year period and clinical practice might have changed over time. In addition, these models were tested using cross-validation and an internal held-out test set, therefore, they may not generalise to other populations or settings. Finally, high-class imbalance made prediction challenging as relatively very few death examples were present (10%), thus more observations of death would be needed to reach a high and confident level of precision in these predictions.

Conclusion

ML models can predict with high accuracy in critically ill children admitted to ICU in Dhaka, Bangladesh using multiple laboratory measurements. Important added value in the mortality prediction, above and beyond clinical parameters only, comes from biochemistry measurements (CO2, potassium, creatinine and calcium) and haematology measurements (haematocrit and platelet count). In ICUs in an LMIC, like Bangladesh, ML could assist healthcare workers in identification of high-risk patients.

supplementary material

10.1136/bmjpo-2023-002365 online supplemental file 1

10.1136/bmjpo-2023-002365 online supplemental figure 1

Acknowledgements

The authors are gratefully indebted to all parents of the patients who agreed to share their information for research purpose. The authors acknowledge the support of the staff members of Dhaka hospitals of icddr,b. icddr,b is also grateful to the Governments of Bangladesh, Canada, Sweden and the UK for providing core/unrestricted support.

Data availability statement

Data are available upon reasonable request.

Review Process File
22 7 2024

Funding: The authors have not declared a specific grant for this research from any funding agency in the public, commercial or not-for-profit sectors.

Data availability free text: Data are available upon reasonable request. De-identified original transcript data will be shared for academic use only. Please contact the corresponding author for reasonable data requests.

Patient consent for publication: Not applicable.

Ethics approval: This study involves human participants and was approved by Institutional Review Board of the International Centre for Diarrhoeal Disease Research, Bangladesh (PR-21098). Participants gave informed consent to participate in the study before taking part.

Provenance and peer review: Not commissioned; externally peer reviewed.

Patient and public involvement: Patients and/or the public were not involved in the design, or conduct, or reporting, or dissemination plans of this research.
==== Refs
References

1 Hossain M Chisti MJ Hossain MI et al Efficacy of world health organization guideline in Facility‐Based reduction of mortality in severely malnourished children from low and middle income countries: A systematic review and Meta‐Analysis J Paediatr Child Health 2017 53 474 9 10.1111/jpc.13443 28052519
2 Brilli RJ Gibson R Luria JW et al Implementation of a medical emergency team in a large pediatric teaching hospital prevents respiratory and cardiopulmonary arrests outside the intensive care unit Pediatr Crit Care Med 2007 8 236 46 10.1097/01.PCC.0000262947.72442.EA 17417113
3 Tibballs J Kinney S Duke T et al Reduction of Paediatric in-patient cardiac arrest and death with a medical emergency team: preliminary results Arch Dis Child 2005 90 1148 52 10.1136/adc.2004.069401 16243869
4 Kim SY Kim S Cho J et al A deep learning model for real-time mortality prediction in critically ill children Crit Care 2019 23 279 10.1186/s13054-019-2561-z 31412949
5 Pollack MM Holubkov R Funai T et al The pediatric risk of mortality score: update 2015 Pediatr Crit Care Med 2016 17 2 9 10.1097/PCC.0000000000000558 26492059
6 American Medical Informatics Association Interpretable deep models for ICU outcome prediction. AMIA annual symposium proceedings 2016
7 Kennedy CE Turley JP Time series analysis as input for clinical predictive modeling: modeling cardiac arrest in a pediatric ICU Theor Biol Med Model 2011 8 40 10.1186/1742-4682-8-40 22023778
8 World Health Organization Pocket Book of Hospital Care for Children: Guidelines for the Management of Common Childhood Illnesses World Health Organization 2013
9 World Health Organization Updated Guideline: Paediatric Emergency Triage, Assessment and Treatment: Care of Critically-Ill Children World Health Organization 2016
10 Ogero M Sarguta RJ Malla L et al Prognostic models for predicting in-hospital Paediatric mortality in resource-limited countries: a systematic review BMJ Open 2020 10 e035045 10.1136/bmjopen-2019-035045
11 Maitland K Kiguli S Opoka RO et al Mortality after fluid bolus in African children with severe infection N Engl J Med 2011 364 2483 95 10.1056/NEJMoa1101549 21615299
12 Straney L Clements A Parslow RC et al Paediatric index of mortality 3: an updated model for predicting mortality in pediatric intensive care Pediatr Crit Care Med 2013 14 673 81 10.1097/PCC.0b013e31829760cf 23863821
13 Organization WH Pocket Book of Hospital Care for Children: Guidelines for the Management of Common Childhood Illnesses 2005 World Health Organization 2015
14 PMLR Reproducibility in critical care: a mortality prediction case study. machine learning for Healthcare conference 2017
15 Sarmin M Afroze F et al Predictor of death in Diarrheal children under 5 years of age having severe sepsis in an urban critical care ward in Bangladesh Glob Pediatr Health 2019 6 2333794X19862716 10.1177/2333794X19862716
16 RC Team R: A language and environment for statistical computing. R foundation for statistical computing 2013
17 Tibshirani R Regression shrinkage and selection via the lasso J R Stat Soc Series B Stat Methodol 1996 58 267 88 10.1111/j.2517-6161.1996.tb02080.x
18 Zou H Hastie T Regularization and variable selection via the elastic net J R Stat Soc Ser B Stat Methodol 2005 67 301 20 10.1111/j.1467-9868.2005.00503.x
19 Friedman JH Greedy function approximation: a gradient boosting machine Ann Statist 2001 29 1189 232 10.1214/aos/1013203451
20 Breiman L Random forests Mach Learn 2001 45 5 32 10.1023/A:1010933404324
21 Friedman J Hastie T Tibshirani R Regularization paths for generalized linear models via coordinate descent J Stat Softw 2010 33 1 22 1 20808728
22 Ridgeway GG Generalized boosted regression models R package version 2 2013 1
23 Liaw A Wiener M 2 Classification and regression by randomForest R News 2002 18 22
24 van den Brink DA de Vries ISA Datema M et al Predicting clinical deterioration and mortality at differing stages during hospitalization: a systematic review of risk prediction models in children in low-and middle-income countries J Pediatr 2023 260 113448 10.1016/j.jpeds.2023.113448 37121311
25 Diallo AH Sayeem Bin Shahid ASM Khan AF et al Characterising Paediatric mortality during and after acute illness in sub-Saharan Africa and South Asia: a secondary analysis of the CHAIN cohort using a machine learning approach eClin Med 2023 57 101838 10.1016/j.eclinm.2023.101838
26 Podda M Bacciu D Micheli A et al A machine learning approach to estimating Preterm infants survival: development of the Preterm infants survival assessment (PISA) Predictor Sci Rep 2018 8 13743 10.1038/s41598-018-31920-6 30213963
27 Shukla VV Eggleston B Ambalavanan N et al Predictive modeling for perinatal mortality in resource-limited settings JAMA Netw Open 2020 3 e2026750 10.1001/jamanetworkopen.2020.26750 33206194
28 Kefi Z Aloui K Saber M New approach based on machine learning for short-term mortality prediction in neonatal intensive care unit IJACSA 2019 10 10.14569/IJACSA.2019.0100778
29 Wen B Brals D Bourdon C et al Predicting the risk of mortality during hospitalization in sick severely malnourished children using daily evaluation of key clinical warning signs BMC Med 2021 19 222 10.1186/s12916-021-02074-6 34538239
30 Diallo AH Sayeem Bin Shahid ASM Khan AF et al Childhood mortality during and after acute illness in Africa and South Asia: a prospective cohort study Lancet Glob Health 2022 10 e673 84 10.1016/S2214-109X(22)00118-8 35427524
31 Nkosi-Gondwe T Robberstad B Mukaka M et al Adherence to community versus facility-based delivery of monthly malaria Chemoprevention with Dihydroartemisinin-Piperaquine for the post-discharge management of severe anemia in Malawian children: A cluster randomized trial PLoS One 2021 16 e0255769 10.1371/journal.pone.0255769 34506503
32 Rudd KE Johnson SC Agesa KM et al Global, regional, and national sepsis incidence and mortality, 1990–2017: analysis for the global burden of disease study Lancet 2020 395 200 11 10.1016/S0140-6736(19)32989-7 31954465
33 Qiu M Du L Concerns regarding the reliability of subgroup effects eClin Med 2023 56 101794 10.1016/j.eclinm.2022.101794
34 Talbert A Ngari M Bauni E et al Mortality after inpatient treatment for diarrhea in children: a cohort study BMC Med 2019 17 20 10.1186/s12916-019-1258-0 30686268
35 Kondapalli CS Ravideep Yalavarthy RY Laboratory abnormalities in acute diarrhoea in children J Evol Med Dent Sci 2017
36 Kumar SS Paramananthan R Muthusethupathi MA Acute renal failure due to acute Diarrhoeal diseases J Assoc Physicians India 1990 38 164 6 2380138
37 Why should I trust you?" explaining the predictions of any Classifier Proceedings of the 22nd ACM SIGKDD international conference on knowledge discovery and data mining 2016
