
==== Front
NPJ Digit Med
NPJ Digit Med
NPJ Digital Medicine
2398-6352
Nature Publishing Group UK London

1244
10.1038/s41746-024-01244-z
Article
Variational Bayes machine learning for risk adjustment of general outcome indicators with examples in urology
http://orcid.org/0000-0002-4304-2783
Koh Harvey Jia Wei 123
Gašević Dragan 12
Rankin David 23
Heritier Stephane 3
Frydenberg Mark 45
Talic Stella stella.talic@monash.edu

123
1 https://ror.org/02bfwt286 grid.1002.3 0000 0004 1936 7857 Centre for Learning Analytics, Faculty of Information Technology, Monash University, Clayton, VIC Australia
2 https://ror.org/02n8xct56 Digital Health Cooperative Research Centre, Sydney, NSW Australia
3 https://ror.org/02bfwt286 grid.1002.3 0000 0004 1936 7857 School of Public Health and Preventative Medicine, Monash University, Melbourne, VIC Australia
4 Cabrini Healthcare, Malvern, VIC Australia
5 https://ror.org/02bfwt286 grid.1002.3 0000 0004 1936 7857 Department of Surgery, Faculty of Medicine, Nursing and Health Sciences, Monash University, Melbourne, VIC Australia
14 9 2024
14 9 2024
2024
7 2493 11 2023
1 9 2024
© The Author(s) 2024
2024
https://creativecommons.org/licenses/by-nc-nd/4.0/ Open Access This article is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License, which permits any non-commercial use, sharing, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if you modified the licensed material. You do not have permission under this licence to share adapted material derived from this article or parts of it. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by-nc-nd/4.0/.
Risk adjustment is often necessary for outcome quality indicators (QIs) to provide fair and accurate feedback to healthcare professionals. However, traditional risk adjustment models are generally oversimplified and not equipped to disentangle complex factors influencing outcomes that are out of a healthcare professional’s control. We present VIRGO, a novel variational Bayes model trained on routinely collected, large administrative datasets to risk-adjust outcome QIs. VIRGO uses detailed demographics, diagnosis, and procedure codes to provide individualized risk adjustment and explanations on patient factors affecting outcomes. VIRGO achieves state-of-the-art on external datasets and features capabilities of uncertainty expression, explainable features, and counterfactual analysis capabilities. VIRGO facilitates risk adjustment by explaining how patient factors led to adverse outcomes and expresses the uncertainty of each prediction, allowing healthcare professionals to not only explore patient factors with unexplained variance that are associated with worse outcomes but also reflect on the quality of their clinical practice.

Subject terms

Translational research
Quality of life
Urogenital diseases
Risk factors
Digital Health Cooperative Research Centreissue-copyright-statement© Springer Nature Limited 2024
==== Body
pmcIntroduction

Ensuring a high standard of care lies at the heart of every healthcare system’s mission1. To ensure that quality standards are met, quality indicators (QIs) are used to measure healthcare quality2. These measures are predominantly developed by clinical quality registries (CQRs), databases that collect data on patients who have a specific disease or undergo a particular procedure2. Well-developed QIs are important as they provide quantifiable measures on quality of care delivered by healthcare professionals1. These measures can then subsequently be used by healthcare professionals as a benchmark tool to compare against their peers feedback and to reflect on quality of care delivered to patients3.

Well-developed QIs maintained by CQRs remain the gold standard in measuring quality of care, however, due to the high costs, CQRs prioritize diseases with high morbidity and mortality4. Consequently, QIs are underdeveloped in some diseases and specialties’, resulting in lack of coverage and gaps of measurements in quality of care5,6. One such specialty is urology, where outside of uro-oncology, outcome QIs are predominantly used as benchmarks in quality6.

While outcome QIs are useful, outcomes are affected by a myriad of risk factors that are out of healthcare professionals’ control7. Patient variations due to comorbidities, age, disease diagnosis, etc. are all factors that contribute to the outcome for a patient. Without a method of risk-adjusting outcomes to disentangle variations attributable to healthcare professionals, outcome QIs are challenging when used for benchmarking, as it may unfairly target healthcare professionals who treat higher-risk patients7. Currently, the risk adjustment method to detect extended LOS involves the use of 3× the national average of related admissions classified by the Australian Refined-Diagnosis Related Group (AR-DRG)8 code and to detect outliers. While useful, the simpleness of the model may not allow for a more complex and accurate risk adjustment, potentially leading to false positives.

To ameliorate this, the potential secondary use of electronic health data collected by hospitals trained on machine learning models has been explored by multiple studies for risk adjusting/risk stratifying outcome indicators9,10. While coding errors in administrative data exist11, the appeal of using large administrative datasets for quality improvement purposes stems from (1) all hospitals collect administrative data for reporting purposes; (2) the datasets collect data at a population level, allowing for quality improvement cycles to be applied across multiple disciplines and diagnoses; (3) data coding is standardized through the constraints of coding standards. The large amount of rich administrative data could potentially facilitate the building of large complex models for risk adjustment12 by investigating the relationships between covariates and outcomes.

To maximize the utility of a large dataset, complex and flexible models are often required to investigate the relationships between covariates and outcomes. While accurate at predicting outcomes, complex models such as deep artificial neural nets do not necessarily provide information on how they derive the final value, as model complexity often comes at the cost of explainability13. For more complex models to be adopted for risk adjustment, transparency and explainability is paramount to how the model’s predictions are derived.

One way of facilitating a models’ explainability is to quantify the uncertainties of the predicted outcome as well as the factors that affect the outcome. Rather than providing a point estimate (i.e., the model predicts a length of stay of 5 days), a model can express the prediction as a range of expected value (the model predicts that within a 99% probability that the length of stay is within 3 to 7 days). Quantifying uncertainties allows for (1) the detection of meaningful outliers in practice, allowing for healthcare professionals only to be alerted when outliers that are beyond expected ranges occur while minimizing the number of false positives, and (2) provides additional explainability, allowing practitioners to trust in such models and their underlying mechanisms which facilitates a better environment for reflection and quality improvement.

With recent advancements in computational capabilities, as well as new methods for approximation, Bayesian methods have become increasingly popular in quantifying uncertainties when modeling, offering a direct expression of uncertainty by defining the population distribution in which the sample is generated, as compared to frequentist methods14. To that end, we present Variational Inference for Risk adjustment of General Outcome indicators (VIRGO), a high dimensional explainable Bayesian regression model that facilitates quality improvement feedback to be provided to healthcare professionals. VIRGO would adjust for patient factors that affect QIs, explain how interacting patient factors led to the predicted range of values, and express the individualized predicted outcome as a readily interpretable credible interval. By adjusting for complex interacting patient factors, outliers of predicted outcomes can be interpreted as potentially unexplained variance that may be attributable to practice. Thus, allowing healthcare professionals to not only explore patient factors that are associated with worse outcomes but also reflect on the quality of their clinical practice. As part of the experimental setup, we chose urological diseases as evidence has shown that QIs in urological disease are underdeveloped6.

Hence, the objectives of this study were to:Develop a fully interpretable, explainable Bayesian model that is competitive in comparison to other models and can provide accurate expected values as well as deliver uncertainties of outcome QI to healthcare professionals.

Explore the associations of different demographic factors, diagnoses, and procedures on common outcome indicators, e.g., length of stay (LOS) and hospital-acquired complication (HACs).

Demonstrate the clinical applications of VIRGO, featuring full explanation of uncertainties behind individualized predictions of outcomes, as a risk adjustment tool and counterfactual analysis to provide feedback to healthcare professionals on patient outcomes.

Results

VIRGO is defined as a Bayesian hierarchical multiple membership zero-inflated negative binomial (ZINB) regression. Given the large amount of same-day patients and patients with no complications, a mixed distribution that models both the zero-component (logit) and the non-zero components (negative binomial) of the outcomes (LOS and HACs) is required for an accurate representation of the data. This facilitates true to distribution modeling and allows VIRGO to be accurate in both model prediction and inference. For more details, please see Methods.

Cohort characteristics

This study used the Victorian Admitted Episodes Dataset (VAED)15, which contains information on all hospital admissions in Victoria from 2009 to 2019 for model development. We analyzed VAED admissions from 2006 to 2019, which had 3.14 million patients with 12 million admission records. Filtering for urological diseases, there were a total of 186,112 urological patients. After removing patients who were missing gender data (n = 69,723), duplicate admission times and dates (n = 775), missing insurance, gender or postcode information, and non-adults (n = 7161), the total sample size was 108,453 patients with a total of 239,067 admissions (Fig. 1). Overall, urology patients tended to be male, were not culturally and linguistically diverse (CALD) patients, and mostly resided in urban areas (Table 1).Fig. 1 CONSORT Diagram.

CONSORT diagram showing the number of patients and admissions used for training and validating of the model. The Victorian Admitted Episodes Dataset (VAED), which contains all admissions recorded from 2009 to 2019 is used as a training dataset and internal validation dataset (see Methods), while the two external datasets from two separate hospitals from Victoria, Australia, and New South Wales, Australia are used for external validation. The external validation datasets crucially do not include any observations from the VAED.

Table 1 Cohort characteristics

Dataset	VAED (Training Set)	External 1 (VIC)	External 2 (NSW)	
Total number of patients	108,453	2738	2289	
Total number of admissions	239,067	2738	2289	
QIs, Mean ± SD	
 Length of stay per admission	2.7 ± 4.3	2.0 ± 3.6	1.6 ± 4.6	
 HACs rates per admission	0.1 ± 0.3	0.0 ± 0.2	0.0 ± 0.2	
Age group, total number of patients (N, %)	
 79 and above	22,573 (20.8)	0452 (16.5)	614 (26.8)	
 65–78	38,218 (35.2)	1208 (44.1)	918 (40.1)	
 49–64	28,115 (25.9)	790 (28.9)	529 (23.1)	
 34–48	12,944 (11.9)	167 (6.1)	149 (6.5)	
 19–33	6603 (6.1)	121 (4.4)	79 (3.5)	
Gender, total number of patients (N, %)	
 Male	82,789 (76.3)	2352 (85.9)	1861 (81.3)	
 Female	25,664 (23.7)	386 (14.1)	428 (18.7)	
Rurality, total number of patients (N, %)	
 Urban	76,186 (70.2)	2443 (89.2)	2051 (89.6)	
 Rural	21,223 (19.6)	227 (8.3)	161 (7.0)	
 Regional	11,044 (10.2)	68 (2.5)	77 (3.4)	
CALDa, total number of patients (N, %)	
 No	95,555 (88.1)	N/A	N/A	
 Yes	12,898 (11.9)	N/A	N/A	
Hospital type, total number of patients (N, %)	
 Public	63,860 (58.9)	0 (0.0)	0 (0.0)	
 Private	44,593 (41.1)	2738 (100.0)	2289 (100.0)	
Hospital insurance, total number of patients (N, %)	
 No Hospital Insurance	59,104 (54.5)	274 (10.0)	0 (0.0)	
 Hospital Insurance	49,349 (45.5)	2464 (90.0)	2289 (100.0)	
Urological diagnosis count, total number of patients (N, %)	
 One urological diagnosis	78,727 (72.6)	2539 (92.7)	1997 (87.2)	
 Two urological diagnoses	16,196 (14.9)	185 (6.8)	230 (10.0)	
 Three or more urological diagnoses	13,530 (12.5)	14 (0.5)	62 (2.7)	
Prevalence of urological diseasesb, number of patients (prevalence %)	
 Hyperplasia of prostate	27,348 (25.2)	550 (20.1)	318 (13.9)	
 Other disorders of bladder	24,914 (23.0)	257 (9.4)	386 (16.9)	
 Calculus of kidney and ureter	21,626 (19.9)	568 (20.7)	365 (15.9)	
 Malignant neoplasm of prostate	17,314 (16.0)	680 (24.8)	591 (25.8)	
 Other disorders of urinary system	13,479 (12.4)	49 (1.8)	80 (3.5)	
 Urethral stricture	10,513 (9.7)	142 (5.2)	70 (3.1)	
 Malignant neoplasm of bladder	8237 (7.6)	225 (8.2)	310 (13.5)	
 Neuromuscular dysfunction of bladder, not elsewhere classified	3910 (3.6)	71 (2.6)	50 (2.2)	
 Malignant neoplasm of kidney, except renal pelvis	3820 (3.5)	58 (2.1)	61 (2.7)	
 Calculus of lower urinary tract	3820 (3.5)	62 (2.3)	53 (2.3)	
 Redundant prepuce, phimosis, and paraphimosis	3621 (3.3)	68 (2.5)	45 (2.0)	
 Inflammatory diseases of prostate	3328 (3.1)	11 (0.4)	25 (1.1)	
 Other disorders of penis	3311 (3.1)	37 (1.4)	46 (2.0)	
 Other disorders of male genital organs	3302 (3.0)	25 (0.9)	23 (1.0)	
 Other disorders of prostate	2684 (2.5)	28 (1.0)	33 (1.4)	
 Hydrocele and spermatocele	2530 (2.3)	47 (1.7)	33 (1.4)	
 Other disorders of urethra	2006 (1.8)	12 (0.4)	29 (1.3)	
 Orchitis and epididymitis	1786 (1.6)	5 (0.2)	12 (0.5)	
 Inflammatory disorders of male genital organs, not elsewhere classified	1080 (1.0)	2 (0.1)	6 (0.3)	
 Unspecified renal colic	982 (0.9)	9 (0.3)	13 (0.6)	
 Malignant neoplasm of testis	702 (0.6)	16 (0.6)	46 (2.0)	
 Torsion of testis	500 (0.5)	2 (0.1)	-	
 Male infertility	497 (0.5)	1 (0.0)	2 (0.1)	
 Malignant neoplasm of renal pelvis	436 (0.4)	11 (0.4)	38 (1.7)	
 Malignant neoplasm of ureter	377 (0.3)	9 (0.3)	13 (0.6)	
 Urethritis and urethral syndrome	275 (0.3)	1 (0.0)	1 (0.0)	
 Malignant neoplasm of other and unspecified urinary organs	179 (0.2)	2 (0.1)	3 (0.1)	
 Malignant neoplasm of penis	155 (0.1)	2 (0.1)	1 (0.0)	
 Disorders of male genital organs in diseases classified elsewhere	128 (0.1)	1 (0.0)	-	
 Malignant neoplasm of other and unspecified male genital organs	56 (0.1)	-	2 (0.1)	
 Urethral disorders in diseases classified elsewhere	9 (0.0)	-	-	
 Bladder disorders in diseases classified elsewhere	7 (0.0)	-	-	
 Calculus of urinary tract in diseases classified elsewhere	2 (0.0)	-	-	
aCALD: Culturally and linguistically diverse.

bDue to patients being diagnosed with multiple conditions, prevalence total is more than 100%.

For external validation and model performance testing, we used two distinct datasets from private hospitals in two Australian states: Victoria (External 1), encompassing urological admissions between August 2022 and June 2023, and New South Wales (External 2), covering urological admissions from January to December 2023. The external datasets after filtering for urological diagnosis codes had 2738 patients and admissions and 2289 patients and admissions for External 1 and External 2 respectively. External datasets show similar characteristics with the VAED training set.

Model performance and external validation

When testing on the external set for generalizability (Table 2), VIRGO performed the best, substantially outperforming all the other models, including state-of-the art gradient boosting models post-parameter tuning at predicting both LOS and HACs count (Supplementary Table 1). This is despite VIRGO not outperforming state-of-the-art gradient boosting models during internal validation (Supplementary Table 2).Table 2 Model performance on external validation datasets

Length of stay	
Model	External 1 (VIC), RMSE [95% CI]	External 2 (NSW), RMSE [95% CI]	Overall RMSE [95% CI]	
VIRGOa	1.897 [1.319–2.552]	2.544 [2.07–2.976]	2.217 [1.847–2.618]	
XGBoost	2.232 [1.216–3.324]	2.632 [2.166–3.083]	2.413 [1.885–3.042]	
LightGBM	2.26 [1.291–3.282]	2.63 [2.171–3.094]	2.433 [1.89–3.091]	
Elasticnet Regression (Log transformed)	2.246 [1.62 –2.902]	2.869 [2.406–3.382]	2.549 [2.182–2.976]	
Lasso Regression (Log transformed)	2.202 [1.65 –2.711]	3.0 [2.504–3.565]	2.616 [2.222–3.093]	
Random Forest	2.484 [1.479–3.93]	2.868 [2.369–3.373]	2.657 [2.053–3.369]	
SVM Regression (Log transformed)	2.491 [1.7–3.259]	3.033 [2.485–3.602]	2.761 [2.267–3.303]	
Hospital–acquired complications count	
VIRGOa	0.158 [0.122–0.195]	0.211 [0.165–0.258]	0.184 [0.154–0.214]	
LightGBM	0.178 [0.131–0.219]	0.195 [0.156–0.234]	0.187 [0.157–0.217]	
XGBoost	0.181 [0.132–0.226]	0.193 [0.16–0.227]	0.188 [0.16–0.215]	
Random Forest	0.195 [0.146–0.24]	0.198 [0.16–0.234]	0.196 [0.164–0.229]	
LASSO Regression (Log transformed)	0.212 [0.142–0.274]	0.198 [0.148–0.247]	0.205 [0.162–0.249]	
Elasticnet Regression (Log transformed)	0.211 [0.151–0.272]	0.197 [0.15–0.242]	0.205 [0.163–0.247]	
SVM Regression (Log transformed)	0.216 [0.145–0.282]	0.198 [0.152–0.254]	0.208 [0.161–0.25]	
Bold indicates best model performance for the dataset.

aBest performing model overall.

Model inference

Demographics

There were multiple demographical factors that were associated with LOS and HACs. The full regression posterior parameter estimates can be found in Supplementary Table 3.

For LOS, we found that patients who were older were more likely to have longer LOS. When compared to patients who were planned admissions emergency/unplanned admissions were more likely to have longer LOS. Patients in private hospitals were more likely to be predicted with longer LOS when compared to public hospitals. Patients who were from rural areas were likely to have longer LOS when compared to regional or urban areas.

For HACs count, similar to LOS, patients who were older and who were emergency admissions were more likely to have more HACs than planned admissions. Patients who are admitted to public hospitals are more likely to have more HACs than patients admitted to private hospitals.

Diagnosis, comorbidities, and HACs

Patients diagnosed with neuromuscular dysfunction of bladder, other disorders of bladder, other disorders of urinary systems, orchitis and epididymitis, prostate hyperplasia, and inflammatory diseases of the prostate were associated with higher LOS if they were not predicted to have zero LOS when compared to other diagnosis. Compared to other urological diagnosis, patients diagnosed with malignant cancer of kidney, malignant cancer of bladder and malignant cancer of renal pelvis were associated with longer LOS. All comorbidities were associated with higher LOS with patients are not predicted to have zero LOS. All HACs were positively associated with LOS.

For the HACs model, malignant neoplasm of the kidney, malignant neoplasm of bladder, other disorders of urinary system, and other disorders of penis were associated with increased chance of associated with higher HACs count when compared to other urological diagnosis groups. Patients being diagnosed with comorbidities such as acute myocardial infarction, malignant cancer, congestive heart failure, and renal disease also increased the chance of having higher HACs count. The full parameter estimates can be found in Supplementary Table 3.

Procedures

For procedures, after feature selection using horseshoe priors (see Methods for more details), variable selection revealed 293 procedure codes that were predictive of LOS. While HACs variable selection using horseshoe priors revealed 123 procedure codes that were predictive. Overall, procedures that were related to treatment of HACs such as postoperative hemorrhage or control of infection, invasive procedures such as nephrolithotomy with 3 or more stones, nephrectomies were strongly associated with higher LOS and HACs when compared to other procedures. Allied health procedure codes were also associated with higher LOS and HACs, capturing chronic conditions that are not captured by diagnosis codes alone. For full breakdown of all urological procedures which are predictive of LOS and HACs, please see Supplementary Table 3.

Explainable risk adjustment

VIRGO has built-in probabilistic capabilities in detecting outliers due to practice variance in healthcare by adjusting for patient factors such as demographics, procedures, and diagnosis. When testing on external datasets, we found 12 (5 LOS outliers and 7 HACs) cases that were deemed practice outliers post-adjustment. In the case study (Fig. 2), we show that adjusting for patient characteristics and treatment profile, the observed LOS (29 days) is significantly outside of the models’ 99% and 97.5% prediction interval for LOS and HACs respectively. We also show which parameters (logit and/or negative binomial) and features contribute to the outcome.

Further case studies can be found in Supplementary Note 1.Fig. 2 Case study of a patient with longer than expected length of stay (LOS).

This figure illustrates VIRGO’s predictive analysis for a patient with bladder cancer. The top subplot compares the predicted length of stay (LOS) within a 99% probability range against the observed LOS rate, indicating an instance of longer-than-expected LOS. The middle subplot elucidates the contribution of individual model variables, with the count variables demonstrating a more significant influence than the zero inflated (ZI) variables on the predicted outcome. This is visually represented by the overrepresentation of blue bars, signifying the count variables’ contribution to the prediction. The bottom pair of subplots further dissects the distribution of the ZI parameters (bottom left) against the count regression parameters (bottom right), offering a detailed view of each variable’s impact on the prediction. Given the predominance of the count parameters, focus should be placed on the count parameters on the bottom right, since the count parameters provide insight into the factors associated with an increased rate of LOS in this instance when compared to the ZI parameters.

Counterfactual analysis

VIRGO is also endowed with the capabilities of counterfactual analysis. Counterfactual predictions allow for questions like “what if the patient did not have a blood transfusion and surgical related complications for a prostatectomy?” to be asked. VIRGO can run simulated patient profiles, allowing for the addition/removal of comorbidities and other risk factors to understand how expected outcomes may be affected by variables (Fig. 3). We provide a case study (Fig. 3a) on a patient who developed surgical complications and underwent blood transfusion during surgery and simulate the expected outcomes in the absence of aforementioned procedures (Fig. 3b) and how the predicted outcome is affected by the simulation.Fig. 3 Providing feedback to healthcare professionals via counterfactual analysis.

Counterfactual analysis of a patient with surgical complications developed during open prostatectomy for prostate hyperplasia. a Shows the expected range of LOS and the count and zero inflated parameters that contribute to the overall prediction, given the complications and blood transfusions, the expected LOS within a 99% probability range is predicted to be 0–19 days. The feature contribution plot shows that the parameters affecting the overall prediction are mostly count parameters. b Shows the counterfactual analysis of the patient, predicting that the expected range of LOS if the patient did not have complications and blood transfusions. The expected LOS within 99% probability decreases to 0–12 days. The feature contribution plot has now changed, with zero-inflated parameters (highlighted in orange) now marginally contributing to the overall prediction.

More case studies and further information of counterfactual analysis of the model can be found in Supplementary Note 1.

Discussion

In this study, we trained a large Bayesian variational inference model using data from 108,453 patients and 239,067 admissions. Our aim was to create a model that was interpretable, accurate in predicting QIs on an admission-by-admission basis, and provide accurate information regarding the model inherent uncertainty in outcome measurements to healthcare professionals. We found that VIRGO: (1) outperformed all state-of-the-art models in predicting LOS and HACs count when tested on external datasets overall; (2) is able to quantify the uncertainties of predicted outcomes; (3) allowed us to examine both point estimates and estimated distributions of predictors such as demographic, diagnosis and procedures that are strongly associated with LOS and HACs; (4) and able to provide clinical insights into individual admissions by way of counterfactual analysis as well as facilitate individualized risk adjustments.

To our knowledge, while previous studies have used procedure code count as part of predicting outcomes16, this study is the first to use Bayesian machine learning to quantify and estimate the uncertainties of procedure codes that are highly associated with both LOS and HACs. Overall, we found that higher LOS and HACs were positively associated with more invasive surgeries, blood transfusion, and platelet transfusions. It is important to note that while blood transfusions and platelet transfusions are often a proxy for surgical complexity17, the wider body of literature seems to suggest that in urological surgeries, perioperative transfusions are associated with worsening outcomes18,19. From a clinical standpoint, reduction of perioperative blood loss may be warranted to improve patient outcomes20. In contrast to previous findings21, our study also suggests that allied health services are positively associated with HACs and LOS. We hypothesize that patients who receive allied health services tended to have chronic conditions that were not captured by other variables in the model22. Hence, by adjusting for uncaptured variables using procedures, VIRGO is able to provide a better picture on what is much more likely to be clinically associated with higher LOS or HACs.

By providing accurate, readily interpretable uncertainties and point estimate of QI predictions to healthcare professionals, we also demonstrated that VIRGO could be used for benchmarking and quality improvement purposes via risk adjustment. Currently, the risk adjustment method to detect extended LOS involves the use of 1.5x–3x the national average of related admissions classified by the AR-DRG8 code and to detect outliers. While useful, the aforementioned methodology may oversimplify the complex interactions between diagnosis and procedures. In contrast, VIRGO is capable of individualized risk adjustment by adjusting for patient factors and isolating practice variation10. In the context of quality improvement, individualizing the risk adjustment process facilitates a fairer assessment of practice variation and ensures an apples-to-apples comparison, as such healthcare professionals will not be unfairly penalized for having poor outcome QIs by treating worse patients and vice-versa23.

Furthermore, by (1) quantifying the uncertainty of the outcome due to both patient and model variability; (2) facilitating counterfactual analysis of simulated patient conditions to examine how the presence/absence of comorbidities, procedures, and complications may affect the predicted outcome; (3) and allowing for more personalized predictions on a case-by-case basis with explainable parameters. VIRGO allows for quality improvement in a much more clinically meaningful way. In particular, simulations of counterfactuals on real patient cases facilitates for clinicians to understand how risk factors and procedures that are associated with complications may directly affect patient outcomes24. In summary, in a clinical context, the quantification of uncertainty25, counterfactual analysis24 as well as personalized predictions and risk adjustment facilitates a much-improved quality improvement feedback cycle, allowing clinicians to track and review their performance much more effectively and efficiently10,23.

VIRGO also displayed state-of-the-art performance at predicting both LOS and HACs in the external set, outperforming previously state-of-the-art gradient boosting models26 like XGBoost, LightGBM as well as SVM regression and other traditional regression methods during external validation even though VIRGO performed marginally worse during internal validation. This could be explained by gradient boosting models proclivity to overfitting27, resulting in poor generalizability. Outperforming other models on the external test sets implies that VIRGO has a good potential for generalizability in Australia, this is further reinforced by good results on the out of state test set, a hospital that was not included in the training data. Potentially, this means that any hospital in Australia can use VIRGO to provide feedback to healthcare professionals by leveraging the use of the administrative dataset that all healthcare systems routinely collect28. This allows for feedback to be routinely provided for diseases and specialties’ that lack specialized quality monitoring systems (such as the ones provided by CQRs). By doing so, hospitals can potentially fill gaps in quality monitoring, leading to a healthcare system that performs safely, effectively, timely, efficiently, is patient-centered and equitable29.

This study has several limitations. Firstly, it is inherent in the use of large administrative data such as the VAED that coding errors can occur30. Secondly, given the large amount of procedure codes (D = 2254) and diagnosis/comorbidities codes (D = 110) that are attributed to each admissions data, there are potentially DD−12 interactions that can occur with the procedures and diagnosis respectively. VIRGO attempted to mitigate this by; (1) using horseshoe priors and SKIM31 on procedure codes, two state-of-the-art Bayesian variable selection method that seeks to select the most relevant/predictive procedures and discover interactions between procedures; (2) using multiple membership model32 to model the ICD and comorbidities, allowing for random intercepts. Thirdly, Bayesian methods are sensitive to the choice of priors. In our study, sensitivity analysis by using different broad priors did not significantly affect our findings, indicating that the large dataset was informative, and our priors had minimal influence on the results. Fourthly, in our study, we used MFVI to approximate the posterior distribution, which is not asymptotically exact when compared to Monte Carlo Markov Chain (MCMC)33. However, MCMC generally is computationally expensive for large datasets such as ours and may not be feasible for model deployment. Finally, we tested our methods only on a subset of the VAED (urological patients) and validated our findings on external datasets, but the generalizability of this method on other diseases still requires further research.

In summary, we showed that VIRGO achieves state-of-the-art in predicting outcome QIs and when compared to other models, capable of providing uncertainty estimates of predicted outcomes and variables affecting outcomes. We further showed that VIRGO can explain factors that lead to longer LOS and HACs, further facilitate clinical reviews of risk adjusted outliers via counterfactual analysis as well as outlier detection. Future studies should validate these findings testing against known outliers. Overall, we hope that this work will increase trust and utility in Bayesian modeling as a tool to provide valuable feedback to healthcare professionals on expected outcomes of patients.

Methods

Study design and data sources

This study used the VAED15, which contains information on all hospital admissions in Victoria from 2009 to 2019. For this study, urological diseases were chosen as the specialty to be modeled as previous reviews have found gaps in quality monitoring in urology6. The patients were excluded from the analysis if there were missing demographic data.

For external validation and model performance testing, we used two distinct datasets from private hospitals in two Australian states: Victoria (External 1), encompassing urological admissions between August 2022 and June 2023, and New South Wales (External 2), covering urological admissions from January to December 2023.

The dataset and all research have been performed in accordance with the Declaration of Helsinki. All data has been de-identified prior to use for research. The study was approved by the Monash University Human Research Ethics Committee (Project ID: 39871). Site ethics approval was managed through reciprocating Monash Partners ethics and MUHREC data sharing agreements. Patients were asked for informed consent during the data collection process. Finally, all methods adhere to Transparent Reporting of a multivariable prediction model for Individual Prognosis Or Diagnosis reporting standards (Supplementary Table 4).

Outcome variables

Two common outcome QIs, LOS and HACs count were used as the outcome variable6. The selection of these QIs stemmed from their well-documented prevalence in academic literature and their established significance in quality monitoring practices34. The complications were defined by the Australian Commission on Safety and Quality in Health Care (ACSQHC) and codes HACs based on the ICD-10 codes, procedure codes, and diagnosis-related groups (DRG) (Table 3)35. Extreme outliers (outliers that were outside the 99th percentile) were excluded from the study.Table 3 Diagnosis codes for urology

ICD Codes	Description	
C60–C63	Malignant neoplasms of male genital organs and urinary tract	
C64–C68	Malignant neoplasms of urinary tract	
N20–N23	Urolithiasis	
N30–N39	Other diseases of urinary system	
N40–N51	Diseases of male genital organs	

Predictor variables

The variables included as predictors were age, sex, insurance status, CALD, private/public hospital, measures of rurality, procedure codes, diagnoses codes, and comorbidities as recorded in the VAED. Measures of rurality (urban, regional, rural) were calculated and defined using postcodes as defined by the Monash Modified Method36. Urological patients were identified based on the International Classification of Diseases 10th Edition37 (ICD-10-AM 10th Edition) primary diagnosis codes for urology (Table 3). Procedure codes were defined by the Australian Classification of Health Interventions, a modified version of the ICD-10 Procedure Coding System. Comorbidities were calculated based on the Charlson method38, using a previously developed coding algorithm38, and identified by ICD-10-AM 10th diagnosis codes. The comorbidities included as a predictor variable within the study population included AIDS or HIV, malignant cancer, cerebrovascular disease, congestive heart failure, chronic obstructive pulmonary disease, dementia, diabetes with and without complications, hemiplegia or paraplegia, metastatic solid tumor, mild liver disease, moderate or severe liver disease, peptic ulcer disease, peripheral vascular disease, renal disease, and rheumatoid disease.

Model specification

VIRGO is defined as a Bayesian hierarchical multiple membership ZINB regression. Bayesian inference models have the advantage for reporting and providing feedback to clinicians and healthcare professionals as it offers a much more intuitive interpretation of the parameters by expressing it as probabilities. Credible intervals allow for outcome QIs to be expressed as a range of expected values based on n probability, i.e., 95% credible interval of length of stay (LOS) is between 0 to 9 days, it means that there is a 95% probability that the expected LOS is between 0 to 9 days. The model definition is further described below.

Model definition

Let y be the observed count, μ the mean count, and logit the probability of a zero inflation.

Likelihood

The likelihood function py|μ,logit is based on ZINB (Eq. 1):1 py∣μ,logit=logitδy+1−logitΓy+1/ωΓy+1Γ1/ωμμ+1/ωy1/ωμ+1/ω1/ω

where δ(y) is the Dirac delta function and ω is the dispersion parameter.

Linear predictor and multiple membership

μ and logit are modeled as (Eqs. 2 and 3):2 logμ=Xβ+Zu

3 logit=X′β′+Z′u′

Here, u and u′ represent the random effects for ICD groups and comorbidities. Z and Z′ are design matrices that map these random effects onto the data.

Prior Distribution

For the intercepts and coefficients, normal priors centered at 0 and a scale of 1 were used (Eq. 4)4 β∼N0,1,β′∼N0,1

For the random effects (Eq. 5):5 u∼Nα,Σ,u′∼Nα′,Σ′

Hyperpriors for α and Σ (Eqs. 6 and 7):6 α∼N0,1,Σ∼Half−Cauchy25

7 α′∼N0,1,Σ′∼Half−Cauchy25

The choice of Half-Cauchy for Σ and Σ′ follows the rationale that it is a weakly informative prior that allows for heavy tails39.

The multiple membership model allowed for fitting a random intercept for admissions with two or more diagnosis, comorbidities, and complications. We fitted age, sex, insurance status, CALD, private/public hospital, measures of rurality as fixed effects (no hierarchical priors), while diagnosis codes, procedure codes, and comorbidities were modeled as random effects (partial pooling with hyperpriors).

Given the outcomes were both count data with large amounts of zero counts and were overdispersed (variance > mean), ZINB was used to model both outcomes. The zero-inflated component of the model represents excess zeros using a logistic model, while the count is modeled via a negative binomial regression40.

To achieve a parsimonious model, given the substantial number of procedure codes (D = 2254) with potential interactions, we employed two methods, regularized horseshoe prior41 and the sparse kernel interaction model (SKIM)31. The horseshoe method was instrumental in identifying relevant variables, while SKIM was utilized to discern significant two-way interactions between procedures. We considered predictors to be strongly associated with the outcome variables if their credible intervals did not include zero.

Traditionally, Bayesian models require the use of MCMC sampling to compute the posterior distribution. Given the large size of the dataset and parameters, we used an approximation of the posterior distribution via mean-field variational inference (MFVI) to fit and train VIRGO33. We further validated VIRGO by performing Pareto smoothed importance sampling leave one out cross-validation and posterior predictive checks42.

Model performance and external validation

For model selection and internal validation, we conducted a 4-fold cross-validation with hospital randomization, this promotes variances across testing sets and reduces the number of batch effects that may occur in a dataset. The cross-validation is formulated as:

Let C={c1,c2,…,cn} represent the set of all hospitals, where each ci corresponds to a unique hospital. We define S={s1,s2,…,sn} as the set of sizes corresponding to each hospital, where si is the number of samples (or records) associated with hospital ci. The objective is to partition the set C into k disjoint subsets C1,C2,…,Ck (where k is the number of folds in cross-validation), such that each subset Cj contains a random selection of hospitals, and the union of all subsets reconstructs the original set C without overlaps (Eq. 8):8 ⋃j=1kCj=CandCi∩Cj=∅foralli≠j

The randomization process begins with shuffling the hospitals to obtain a random permutation C′. Sequentially, hospitals in C′ are assigned to each fold, ensuring that the cumulative size of the test sets does not exceed a predefined proportion of the total dataset size. Let T be the total number of samples across all hospitals, calculated as (Eq. 9):9 T=∑i=1nsi

The target size for each test set is given by t=α×T, where 0<α<1 represents the proportion of the dataset designated for testing in each fold. For each fold fj, we initialize the cumulative size of the test set tcumj=0. We then iterate over the shuffled hospitals in C′ and assign hospital ci to fold fj if tcumj+si≤t. After assigning ci, we update tcumj=tcumj+si. This process continues until either all hospitals are assigned, or the cumulative size of the test set reaches the target size t. To promote variance and ensure that each fold is representative of the entire dataset, we adopt a stratified approach by considering the size of each hospital during the assignment process. This stratified randomization (Fig. 4, Algorithm 1) helps in achieving a balanced distribution of hospitals across folds, preventing any single fold from being biased towards larger or smaller hospitals.Fig. 4 Stratified K-fold Cross-Validation.

Stratified K-fold cross-validation ensures that during model training, there is a promotion of variance and reduction of batch effects. This cross-validation ensures a balanced distribution of hospitals across folds, preventing any single fold from being biased towards larger or smaller hospitals.

To evaluate the efficacy and generalizability of VIRGO, further to internal validation, we assessed model performance on two external test sets sourced from private hospitals in Victoria (External 1) and New South Wales (External 2), Australia. These datasets were completely unseen by the model during training, allowing for testing for model generalizability.

For the quantification of model performance and uncertainty, we utilized the Root Mean Square Error (RMSE) as our primary metric, accompanied by bootstrap resampling to obtain 95% Confidence Intervals for these estimates. We compared VIRGO against state-of-the-art gradient boosting models (XGBoost, LightGBM, and Random Forest)43, support vector machine (SVM) regression, and log transformed regression (LASSO and elastic net regularization). All parameters for comparator models were tuned using a Bayesian optimization using a tree-structured Parzen Estimator algorithm, tuned using 4-fold cross-validation, with the best model selected based on the performance in the validation set44.

All computational analyses were performed using Python, with scikit-learn facilitating the implementation of LASSO regression, Random Forest, and SVM regression, with additional libraries supporting XGBoost and LightGBM. Numpyro was used for creating and training Bayesian models45.

Model inference, explainable risk adjustment and counterfactual analysis

Given the fully explainable nature of VIRGO, VIRGO is capable of providing individualized predictions on an admissions-by-admissions basis. We provide case studies extracted from the external datasets to demonstrate VIRGO’s ability to provide feedback to healthcare professionals by (1) Explaining the predictions and how demographics, diagnosis, comorbidities, and procedures are used to predict LOS/HACs rate. (2) If there are observed outcomes that are outside the 99% (LOS) and 97.5% (HACs) prediction interval, consider these cases abnormal and flagged for review. (3) Counterfactual predictions, allowing for a patient with the same profile but adding/removing simulated procedures such as blood transfusions, and platelet transfusions as well as comorbidities such as renal disease, or diabetes to demonstrate how that would affect the prediction intervals.

Supplementary information

Supplementary Material

Supplementary information

The online version contains supplementary material available at 10.1038/s41746-024-01244-z.

Acknowledgements

This work was supported by the Digital Health Cooperative Research Centre. Koh, Talic, Gasevic are funded by Digital Health Cooperative Research Centre and Monash University.

Author contributions

H.K. oversaw the study conceptualization and design. H.K., S.T., and D.R. oversaw the acquisition of data. H.K. oversaw the analysis and interpretation of data. H.K., D.G., D.R., M.F., S.H., and S.T. were involved in the drafting of the manuscript and critical revision of the content for important intellectual content. H.K. conducted the statistical analysis. D.G., D.R., M.F, and S.T. oversaw the supervision of the study. All authors have approved the final submitted version of the paper.

Data availability

The data collected is owned by the Department of Health and Human Services Australia and is confidential and hence unavailable to others.

Code availability

Code is available upon request. All code were written in Python 3.12 using packages Numpyro, Pandas, XGBoost, LightGBM and NumPy. Full parameters can be found in Supplementary Table 2.

Competing interests

The authors declare no competing interests.

Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
==== Refs
References

1. Evans SM Bohensky M Cameron PA McNeil J A survey of Australian clinical registries: can quality of care be measured? Intern. Med. J. 2011 41 42 48 10.1111/j.1445-5994.2009.02068.x 19811553
Evans, S. M., Bohensky, M., Cameron, P. A. & McNeil, J. A survey of Australian clinical registries: can quality of care be measured? Intern. Med. J. 41, 42–48 (2011).19811553 10.1111/j.1445-5994.2009.02068.x
2. McNeil JJ Evans SM Johnson NP Cameron PA Clinical-quality registries: their role in quality improvement Med. J. Aust. 2010 192 244 245 10.5694/j.1326-5377.2010.tb03499.x 20201755
McNeil, J. J., Evans, S. M., Johnson, N. P. & Cameron, P. A. Clinical-quality registries: their role in quality improvement. Med. J. Aust. 192, 244–245 (2010).20201755 10.5694/j.1326-5377.2010.tb03499.x
3. Ivers N Audit and feedback: effects on professional practice and healthcare outcomes Cochrane. Database. Syst. Rev. 2012 10.1002/14651858.CD000259.pub3 22696318
Ivers, N. et al. Audit and feedback: effects on professional practice and healthcare outcomes. Cochrane. Database. Syst. Rev.10.1002/14651858.CD000259.pub3 (2012).22696318 10.1002/14651858.CD000259.pub3
4. The Australian Commission on Safety and Quality in Health Care. Economic Evaluation of Clinical Quality Registries: Final Report. https://www.safetyandquality.gov.au/wp-content/uploads/2016/12/Economic-evaluation-of-clinical-quality-registries-Final-report-Nov-2016.pdf (2016).
5. Copnell B Measuring the quality of hospital care: an inventory of indicators Intern. Med. J. 2009 39 352 360 10.1111/j.1445-5994.2009.01961.x 19323697
Copnell, B. et al. Measuring the quality of hospital care: an inventory of indicators. Intern. Med. J. 39, 352–360 (2009).19323697 10.1111/j.1445-5994.2009.01961.x
6. Koh, H. J. W. et al. Quality indicators in the clinical specialty of urology: a systematic review. Eur. Urol. Focus. 9, 435–446 (2022).
7. Lane-Fall MB Neuman MD Outcomes measures and risk adjustment Int. Anesthesiol. Clin. 2013 51 10 21 10.1097/AIA.0b013e3182a70a52 24088885
Lane-Fall, M. B. & Neuman, M. D. Outcomes measures and risk adjustment. Int. Anesthesiol. Clin. 51, 10–21 (2013).24088885 10.1097/AIA.0b013e3182a70a52
8. Ossai CI Rankin D Wickramasinghe N Preadmission assessment of extended length of hospital stay with RFECV-ETC and hospital-specific data Eur. J. Med. Res. 2022 27 128 10.1186/s40001-022-00754-4 35879803
Ossai, C. I., Rankin, D. & Wickramasinghe, N. Preadmission assessment of extended length of hospital stay with RFECV-ETC and hospital-specific data. Eur. J. Med. Res. 27, 128 (2022).35879803 10.1186/s40001-022-00754-4
9. Irvin JA Incorporating machine learning and social determinants of health indicators into prospective risk adjustment for health plan payments BMC Public Health 2020 20 608 10.1186/s12889-020-08735-0 32357871
Irvin, J. A. et al. Incorporating machine learning and social determinants of health indicators into prospective risk adjustment for health plan payments. BMC Public Health 20, 608 (2020).32357871 10.1186/s12889-020-08735-0
10. Kan HJ Exploring the use of machine learning for risk adjustment: a comparison of standard and penalized linear regression models in predicting health care costs in older adults PLOS ONE 2019 14 e0213258 10.1371/journal.pone.0213258 30840682
Kan, H. J. et al. Exploring the use of machine learning for risk adjustment: a comparison of standard and penalized linear regression models in predicting health care costs in older adults. PLOS ONE 14, e0213258 (2019).30840682 10.1371/journal.pone.0213258
11. Assareh H Achat HM Stubbs JM Guevarra VM Hill K Incidence and variation of discrepancies in recording chronic conditions in Australian hospital administrative data PloS ONE 2016 11 e0147087 10.1371/journal.pone.0147087 26808428
Assareh, H., Achat, H. M., Stubbs, J. M., Guevarra, V. M. & Hill, K. Incidence and variation of discrepancies in recording chronic conditions in Australian hospital administrative data. PloS ONE 11, e0147087 (2016).26808428 10.1371/journal.pone.0147087
12. Janssen A Talic S Gasevic D Kay J Shaw T Exploring the Intersection Between Health Professionals’ Learning and eHealth Data: Protocol for a Comprehensive Research Program in Practice Analytics in Health Care JMIR Res. Protoc. 2021 10 e27984 10.2196/27984 34889768
Janssen, A., Talic, S., Gasevic, D., Kay, J. & Shaw, T. Exploring the Intersection Between Health Professionals’ Learning and eHealth Data: Protocol for a Comprehensive Research Program in Practice Analytics in Health Care. JMIR Res. Protoc. 10, e27984 (2021).34889768 10.2196/27984
13. Herm L-V Heinrich K Wanner J Janiesch C Stop ordering machine learning algorithms by their explainability! a user-centered investigation of performance and explainability Int. J. Inf. Manag. 2023 69 102538 10.1016/j.ijinfomgt.2022.102538
Herm, L.-V., Heinrich, K., Wanner, J. & Janiesch, C. Stop ordering machine learning algorithms by their explainability! a user-centered investigation of performance and explainability. Int. J. Inf. Manag. 69, 102538 (2023).10.1016/j.ijinfomgt.2022.102538
14. Kruschke JK Bayesian analysis reporting guidelines Nat. Hum. Behav. 2021 5 1282 1291 10.1038/s41562-021-01177-7 34400814
Kruschke, J. K. Bayesian analysis reporting guidelines. Nat. Hum. Behav. 5, 1282–1291 (2021).34400814 10.1038/s41562-021-01177-7
15. Department of Health. Victoria, Australia. Victorian Admitted Episodes Dataset. https://www.health.vic.gov.au/data-reporting/victorian-admitted-episodes-dataset (2024).
16. Shaaban AN Peleteiro B Martins MRO Statistical models for analyzing count data: predictors of length of stay among HIV patients in Portugal using a multilevel model BMC Health Serv. Res. 2021 21 372 10.1186/s12913-021-06389-1 33882911
Shaaban, A. N., Peleteiro, B. & Martins, M. R. O. Statistical models for analyzing count data: predictors of length of stay among HIV patients in Portugal using a multilevel model. BMC Health Serv. Res. 21, 372 (2021).33882911 10.1186/s12913-021-06389-1
17. Diamantopoulos LN Patterns and timing of perioperative blood transfusion and association with outcomes after radical cystectomy Urol. Oncol. Semin. Orig. Investig. 2021 39 496.e1 496.e8
Diamantopoulos, L. N. et al. Patterns and timing of perioperative blood transfusion and association with outcomes after radical cystectomy. Urol. Oncol. Semin. Orig. Investig. 39, 496.e1–496.e8 (2021).
18. Abu-Ghanem Y Ramon J Impact of perioperative blood transfusions on clinical outcomes in patients undergoing surgery for major urologic malignancies Ther. Adv. Urol. 2019 11 1756287219868054 10.1177/1756287219868054 31447936
Abu-Ghanem, Y. & Ramon, J. Impact of perioperative blood transfusions on clinical outcomes in patients undergoing surgery for major urologic malignancies. Ther. Adv. Urol. 11, 1756287219868054 (2019).31447936 10.1177/1756287219868054
19. Wang Y-L Perioperative blood transfusion promotes worse outcomes of bladder cancer after radical cystectomy: a systematic review and meta-analysis PLoS ONE 2015 10 e0130122 10.1371/journal.pone.0130122 26080092
Wang, Y.-L. et al. Perioperative blood transfusion promotes worse outcomes of bladder cancer after radical cystectomy: a systematic review and meta-analysis. PLoS ONE 10, e0130122 (2015).26080092 10.1371/journal.pone.0130122
20. Callum J Siemens DR We should redouble efforts to minimize transfusions in urological surgery J. Urol. 2023 209 471 473 10.1097/JU.0000000000003146 36629474
Callum, J. & Siemens, D. R. We should redouble efforts to minimize transfusions in urological surgery. J. Urol. 209, 471–473 (2023).36629474 10.1097/JU.0000000000003146
21. Sarkies MN White J Henderson K Haas R Bowles J Additional weekend allied health services reduce length of stay in subacute rehabilitation wards but their effectiveness and cost-effectiveness are unclear in acute general medical and surgical hospital wards: a systematic review J. Physiother. 2018 64 142 158 10.1016/j.jphys.2018.05.004 29929739
Sarkies, M. N., White, J., Henderson, K., Haas, R. & Bowles, J. Additional weekend allied health services reduce length of stay in subacute rehabilitation wards but their effectiveness and cost-effectiveness are unclear in acute general medical and surgical hospital wards: a systematic review. J. Physiother. 64, 142–158 (2018).29929739 10.1016/j.jphys.2018.05.004
22. Barr ML Understanding the use and impact of allied health services for people with chronic health conditions in Central and Eastern Sydney, Australia: a five-year longitudinal analysis Prim. Health Care Res. Dev. 2019 20 e141 10.1017/S146342361900077X 31640837
Barr, M. L. et al. Understanding the use and impact of allied health services for people with chronic health conditions in Central and Eastern Sydney, Australia: a five-year longitudinal analysis. Prim. Health Care Res. Dev. 20, e141 (2019).31640837 10.1017/S146342361900077X
23. Juhnke, C., Bethge, S. & Mühlbacher, A. C. A review on methods of risk adjustment and their use in integrated healthcare systems. Int. J. Integr. Care 16, 4 (2026).
24. Pfohl, S. R., Duan, T., Ding, D. Y. & Shah, N. H. Counterfactual reasoning for fair clinical risk prediction. In Proc. 4th machine learning for healthcare conference 325–358 (PMLR, 2019).
25. Seoni S Application of uncertainty quantification to artificial intelligence in healthcare: a review of last decade (2013–2023) Comput. Biol. Med. 2023 165 107441 10.1016/j.compbiomed.2023.107441 37683529
Seoni, S. et al. Application of uncertainty quantification to artificial intelligence in healthcare: a review of last decade (2013–2023). Comput. Biol. Med. 165, 107441 (2023).37683529 10.1016/j.compbiomed.2023.107441
26. Jain, R., Singh, M., Rao, A. R. & Garg, R. Machine learning models to predict length of stay in hospitals. In Proc. 10th International Conference on Healthcare Informatics (ICHI) 545–546 10.1109/ICHI54592.2022.00105 (IEEE, Rochester, MN, USA, 2022)
27. Dietterich, T. G. An experimental comparison of three methods for constructing ensembles of decision trees: bagging, boosting, and randomization. Mach. Learn. 40, 139–157 (2000).
28. Szakiel J Hospital casemix protocol-medibank private perspective Health Inf. Manag. J. 2010 39 47 49
Szakiel, J. Hospital casemix protocol-medibank private perspective. Health Inf. Manag. J. 39, 47–49 (2010).
29. Institute of Medicine. Crossing the Quality Chasm: A New Health System for the 21st Century. 10.17226/10027 (2001).
30. van Mourik MSM van Duijn PJ Moons KGM Bonten MJM Lee GM Accuracy of administrative data for surveillance of healthcare-associated infections: a systematic review BMJ Open 2015 5 e008424 10.1136/bmjopen-2015-008424 26316651
van Mourik, M. S. M., van Duijn, P. J., Moons, K. G. M., Bonten, M. J. M. & Lee, G. M. Accuracy of administrative data for surveillance of healthcare-associated infections: a systematic review. BMJ Open 5, e008424 (2015).26316651 10.1136/bmjopen-2015-008424
31. Agrawal, R., Trippe, B., Huggins, J. & Broderick, T. The Kernel Interaction Trick: Fast Bayesian Discovery of Pairwise Interactions in High Dimensions. in Proceedings of the 36th International Conference on Machine Learning 141–150 (PMLR, 2019).
32. Leckie, G. Multiple membership multilevel models. 10.48550/arXiv.1907.04148 (2016).
33. Blei DM Kucukelbir A McAuliffe JD Variational inference: a review for statisticians J. Am. Stat. Assoc. 2017 112 859 877 10.1080/01621459.2017.1285773
Blei, D. M., Kucukelbir, A. & McAuliffe, J. D. Variational inference: a review for statisticians. J. Am. Stat. Assoc. 112, 859–877 (2017).10.1080/01621459.2017.1285773
34. Lingsma HF Evaluation of hospital outcomes: the relation between length-of-stay, readmission, and mortality in a large international administrative database BMC Health Serv. Res. 2018 18 116 10.1186/s12913-018-2916-1 29444713
Lingsma, H. F. et al. Evaluation of hospital outcomes: the relation between length-of-stay, readmission, and mortality in a large international administrative database. BMC Health Serv. Res. 18, 116 (2018).29444713 10.1186/s12913-018-2916-1
35. Australian Commission on Safety and Quality in Health Care. Hospital-Acquired Complications (HACs) List. https://www.safetyandquality.gov.au/publications-and-resources/resource-library/hospital-acquired-complications-hacs-list-specifications-version-31-12th-edn (2019).
36. Versace VL Skinner TC Bourke L Harvey P Barnett T National analysis of the modified Monash model, population distribution and a socio-economic index to inform rural health workforce planning Aust. J. Rural Health 2021 29 801 810 10.1111/ajr.12805 34672057
Versace, V. L., Skinner, T. C., Bourke, L., Harvey, P. & Barnett, T. National analysis of the modified Monash model, population distribution and a socio-economic index to inform rural health workforce planning. Aust. J. Rural Health 29, 801–810 (2021).34672057 10.1111/ajr.12805
37. World Health Organization. International Classification of Diseases (ICD). https://www.who.int/standards/classifications/classification-of-diseases (2019).
38. Quan H Coding algorithms for defining comorbidities in ICD-9-CM and ICD-10 administrative data Med. Care 2005 43 1130 1139 10.1097/01.mlr.0000182534.19832.83 16224307
Quan, H. et al. Coding algorithms for defining comorbidities in ICD-9-CM and ICD-10 administrative data. Med. Care 43, 1130–1139 (2005).16224307 10.1097/01.mlr.0000182534.19832.83
39. Gelman A Simpson D Betancourt M The prior can often only be understood in the context of the likelihood Entropy 2017 19 555 10.3390/e19100555
Gelman, A., Simpson, D. & Betancourt, M. The prior can often only be understood in the context of the likelihood. Entropy 19, 555 (2017).10.3390/e19100555
40. Preisser JS Stamm JW Long DL Kincade ME Review and recommendations for zero-inflated count regression modeling of dental caries indices in epidemiological studies Caries Res. 2012 46 413 423 10.1159/000338992 22710271
Preisser, J. S., Stamm, J. W., Long, D. L. & Kincade, M. E. Review and recommendations for zero-inflated count regression modeling of dental caries indices in epidemiological studies. Caries Res. 46, 413–423 (2012).22710271 10.1159/000338992
41. Piironen J Vehtari A Sparsity information and regularization in the horseshoe and other shrinkage priors Electron. J. Stat. 2017 11 5018 5051 10.1214/17-EJS1337SI
Piironen, J. & Vehtari, A. Sparsity information and regularization in the horseshoe and other shrinkage priors. Electron. J. Stat. 11, 5018–5051 (2017).10.1214/17-EJS1337SI
42. Vehtari A Gelman A Gabry J Practical Bayesian model evaluation using leave-one-out cross-validation and WAIC Stat. Comput. 2017 27 1413 1432 10.1007/s11222-016-9696-4
Vehtari, A., Gelman, A. & Gabry, J. Practical Bayesian model evaluation using leave-one-out cross-validation and WAIC. Stat. Comput. 27, 1413–1432 (2017).10.1007/s11222-016-9696-4
43. Florek, P. & Zagdański, A. Benchmarking state-of-the-art gradient boosting algorithms for classification. Preprint at 10.48550/arXiv.2305.17094 (2023).
44. Bergstra, J., Bardenet, R., Bengio, Y. & Kégl, B. Algorithms for Hyper-Parameter Optimization. In Proc. Advances in Neural Information Processing Systems 24 (Curran Associates, Inc., 2011).
45. Phan, D., Pradhan, N. & Jankowiak, M. Composable Effects for Flexible and Accelerated Probabilistic Programming in NumPyro. Preprint at 10.48550/arXiv.1912.11554 (2019).
