
==== Front
Sci Rep
Sci Rep
Scientific Reports
2045-2322
Nature Publishing Group UK London

39261645
72206
10.1038/s41598-024-72206-4
Article
Logistic regression analysis and machine learning for predicting post-stroke gait independence: a retrospective study
Miyazaki Yuta 123
Kawakami Michiyuki michiyukikawakami@hotmail.com

12
Kondo Kunitsugu 12
Hirabe Akiko 12
Kamimoto Takayuki 12
Akimoto Tomonori 12
Hijikata Nanako 12
Tsujikawa Masahiro 12
Honaga Kaoru 14
Suzuki Kanjiro 1
Tsuji Tetsuya 2
1 grid.517658.8 Department of Rehabilitation Medicine, Tokyo Bay Rehabilitation Hospital, Chiba, Japan
2 https://ror.org/02kn6nx58 grid.26091.3c 0000 0004 1936 9959 Department of Rehabilitation Medicine, Keio University School of Medicine, 35 Shinanomachi, Shinjuku-ku, Tokyo, 160-8582 Japan
3 https://ror.org/0254bmq54 grid.419280.6 0000 0004 1763 8916 Department of Physical Rehabilitation, National Center of Neurology and Psychiatry, National Center Hospital, Tokyo, Japan
4 https://ror.org/01692sz90 grid.258269.2 0000 0004 1762 2738 Department of Rehabilitation Medicine, Juntendo University Graduate School of Medicine, Tokyo, Japan
11 9 2024
11 9 2024
2024
14 2127316 11 2023
4 9 2024
© The Author(s) 2024
2024
https://creativecommons.org/licenses/by-nc-nd/4.0/ Open Access This article is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License, which permits any non-commercial use, sharing, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if you modified the licensed material. You do not have permission under this licence to share adapted material derived from this article or parts of it. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by-nc-nd/4.0/.
This study investigated whether machine learning (ML) has better predictive accuracy than logistic regression analysis (LR) for gait independence at discharge in subacute stroke patients (n = 843) who could not walk independently at admission. We developed prediction models using LR and five ML algorithms—specifically, the decision tree (DT), support vector machine, artificial neural network, ensemble learning, and k-nearest neighbor methods. Functional Independence Measure sub-items were used to evaluate the ability to walk independently. Model predictive accuracies were evaluated using areas under receiver operating characteristic curves (AUCs) as well as accuracy, precision, recall, F1 score, and specificity. The AUC for DT (0.812) was significantly lower than those for the other algorithms (p < 0.01); however, the AUC for LR (0.895) did not differ significantly from those for the other models (0.893–0.903). Other performance metrics showed no substantial differences between LR and ML algorithms. In conclusion, the DT algorithm had significantly low predictive accuracy, and LR showed no significant difference in predictive accuracy compared with the other ML algorithms. As its predictive accuracy is similar to that of ML, LR can continue to be used for predicting the prognosis of gait independence, with additional advantages of being easily understandable and manually computable.

Keywords

Machine learning
Logistic regression
Prediction models
Gait independence
Stroke
Subject terms

Cerebrovascular disorders
Stroke
http://dx.doi.org/10.13039/501100001691 Japan Society for the Promotion of Science JP22K21225 Miyazaki Yuta issue-copyright-statement© Springer Nature Limited 2024
==== Body
pmcIntroduction

Stroke is the second most common cause of disability worldwide and is estimated to affect more than 8 million people1. As gait independence is the primary goal of stroke rehabilitation, prognosis in terms of gait independence as compared to that at admission is important2. In this study, following previous research, we defined the acute phase of stroke as 1–7 days post-onset and the subacute phase as 7 days to 6 months post-onset3. Additionally, subacute rehabilitation, in conjunction with acute rehabilitation, improves activities of daily living (ADLs)4. A recent systematic review on post-stroke outcomes identified six studies that predicted post-stroke gait independence: four for acute phase and two for subacute phase patients5. This review also reported that trunk functions2,6 and lower extremity functions, such as those assessed using the Functional Ambulation Categories6, were good prognostic indicators in the acute phase. However, while prognosis after rehabilitation has been investigated, studies in the subacute phase are also necessary, because some patients with acute stroke have disturbances of consciousness, and there is a relative lack of studies on the prognosis of these patients.

Machine learning (ML) has recently gained attention as a new prognostic approach, and recent systematic reviews indicate that ML algorithms had better accuracy and areas under the receiver operating curve (AUCs) than logistic regression analysis (LR)7,8. In addition, an increasing number of studies are using ML for predicting functional outcomes in patients with stroke9, and ML is attracting attention as an alternative to LR. ML has been used to build prediction models for post-stroke outcomes10 based on the modified Rankin scale (mRS)11, Barthel Index (BI)12, and Functional Independence Measure (FIM) assessment13. The decision tree (DT), artificial neural network (ANN), support vector machine (SVM), K-nearest neighbor (KNN), and ensemble learning (EL) algorithms have been adopted as ML methods for classification14. DT15,16 and ANN17 have also been used to predict the prognosis of post-stroke gait independence, with ANN reportedly having a higher AUC than LR, although the statistical significance was not mentioned17.

However, other review articles have reported that the prediction accuracy of ML is comparable to that of LR18,19. Furthermore, ANN and other ML methods can cause overlearning without sufficient cases to develop prediction models. This overlearning increases the prediction accuracy in training data but drastically decreases it in unknown data20. When overlearning occurs, generalization performance deteriorates, and the prediction models become useless. Therefore, prediction models must be established in a way that prevents overlearning by conducting appropriate training with sufficient cases (at least several hundred)21. However, to our knowledge, no previous study has statistically compared the prediction accuracies of ML and LR for gait independence in patients with acute and sub-acute stroke. Therefore, the usefulness of ML for predicting post-stroke functional outcomes remains unclear.

Here, we hypothesized that while there may be differences in prediction accuracies among different ML algorithms, LR may have sufficient accuracy compared with ML for predicting gait independence in patients with stroke. In this study, we aimed to establish prediction models with sufficient cases to prevent overlearning and compare the prediction accuracies for gait independence between ML and LR in patients with subacute stroke.

Methods

Study design and ethical considerations

This single-center, retrospective study was conducted according to the principles of the Declaration of Helsinki22 and approved by the Ethics Committee of Tokyo Bay Rehabilitation Hospital (approval number: 267-7). Informed consent was obtained by using the opt-out method, with the corresponding option provided on the Tokyo Bay Rehabilitation Hospital webpage.

Participants

This study initially included 1552 patients with subacute stroke (approximately 30 days after stroke onset) admitted to Tokyo Bay Rehabilitation Hospital after acute treatment between March 2015 and September 2019. We then excluded patients with (1) a history of stroke, (2) acute care hospital length of stay > 90 days, (3) rehabilitation hospital length of stay < 28 or > 180 days, (4) transfer to an acute care hospital after admission to our hospital, (5) missing data, and (6) gait independence at admission. All participants received standard physical, occupational, and speech therapy for 2–3 h daily.

Measurements

Trained nurses assessed participant ADLs using the Japanese version of the FIM, which consists of 13 motor sub-items and 5 cognitive sub-items23,24. All sub-items were rated on a 7-point scale, with a score of 6 or higher representing “independence,” and a score of 5 or lower representing “requiring assistance”23. In this study, gait independence was defined as a FIM walking sub-items score of 6 or more.

Physicians and physical and occupational therapists also assessed motor function on the basis of bilateral grip strength and using the Stroke Impairment Assessment Scale (SIAS)25—specifically, two upper extremity functions (knee-mouth and finger functions), three lower extremity functions (hip flexion, knee extension, and foot pat), hip flexion on the unaffected side, and trunk function (verticality test) were assessed. All functions were rated on a 6-grade scale (from 0, no movement, to 5, normal movement), except for finger function, which was assessed using an 8-grade scale, and hip flexion on the unaffected side and trunk function, which were evaluated using a 4-grade scale.

The Geriatric Nutritional Risk Index (GNRI)26 was calculated as a measure of nutrition based on albumin levels at admission using the following equation:GNRI=14.89×Alb(g/dL)+41.7×current\;body\;weight / ideal\;body\;weight

Model development

LR as well as DT, SVM, ANN, KNN, and EL prediction models were developed in this study, with sex, age, FIM sub-item score, FIM total score, SIAS sub-item score, bilateral grip strength, and GNRI as features for model development, as these factors have been reported to be predictive of post-stroke functional outcomes27. Participants with a FIM walking sub-items score of 6 or more at discharge were designated to the independent gait group and those with a score of 5 or less to the non-independent gait group; these binary data were used as the objective variable. Further, we developed prediction models to classify whether participants could walk independently at discharge.

Participants were randomly divided into 80% training and 20% test data sets. We developed prediction models based on the six algorithms using the training data set by five-fold cross-validation. In the learning procedure, the hyperparameters were optimized using Bayesian optimization up to 30 times for each prediction model using the Classification Learner application in MATLAB (version 2023a, MathWorks, Natick, MA, USA). Details of the search ranges for the hyperparameters of each algorithm, such as the number of layers and neurons in ANN, are provided in Supplementary Table 1. As this study used cross-validation, training was conducted using a single batch. For the ANN model, the limited-memory Broyden–Fletcher–Goldfarb–Shanno (BFGS) algorithm28 was used to automatically determine the learning rate. Finally, the optimized prediction models were evaluated for overlearning using the test data.

Model evaluation and statistical analysis

The accuracy of the prediction models, as evaluated based on the AUCs, was the primary outcome. AUC is a commonly used measure for binary classification, with values of 0.8–0.9 and ≥ 0.9 considered excellent and outstanding, respectively.29 Overlearning was defined as a statistically significant difference in AUCs between the validation and test data sets for a prediction model. Finally, we also statistical analyzed the AUC data using the pROC package (version 1.18.4)30 in R version 4.3.1 (R Foundation for Statistical Computing, Vienna, Austria). Bootstrapping (2000 iterations) was used to calculate the confidence intervals (CIs)11. All tests for significance were two-tailed, and p < 0.05 was considered statistically significant.

As the secondary outcomes, we calculated the accuracy, precision, recall, F1 score, specificity, and the confusion matrix for each algorithm on both the validation and test data sets. The differences in mean age, days since stroke onset, admission FIM scores, grip strength, and GNRI between the gait-independent and non-independent groups were statistically analyzed using the t-test.

Results

The clinical data of 1552 patients with stroke were retrospectively analyzed. After excluding 289 patients with a history of stroke, 9 patients with acute care hospital length of stay > 90 days, 105 patients with rehabilitation hospital length of stay < 28 days or > 180 days, 103 patients transferred to acute care hospitals, 66 patients with missing data, and 137 patients with gait independence at admission, 843 patients were finally included in this study (Fig. 1). The participant characteristics are summarized in Table 1. Supplementary Table 2 shows the t-test results for the gait-independent and non-independent groups; all items showed statistically significant differences.Fig. 1 Exclusion criteria. The mean age of the participants was 69.6 ± 14.1 years. The mean FIM motor and total scores at admission were 41.5 ± 18.4 and 63.9 ± 24.5 and 70.9 ± 20.4 and 98.2 ± 26.3 at discharge, respectively. A total of 519 participants (61.6%) could independently walk at discharge. FIM, Functional Independence Measure.

Table 1 Participant characteristics.

Age (years)	69.6 ± 14.1(13–98)	
Sex		
 Female	375 (44.5%)	
 Male	467 (55.5%)	
Days since stroke onset	33.6 ± 12.0 (11–86)	
Admission FIM motor score	41.5 ± 18.4 (13–85)	
Discharge FIM motor score	70.9 ± 20.4 (13–91)	
Admission FIM cognitive score	22.4 ± 8.2 (5–35)	
Discharge FIM cognitive score	27.3 ± 7.2 (5–35)	
Admission FIM total score	63.9 ± 24.5 (18–120)	
Discharge FIM total score	98.2 ± 26.3 (18–126)	
Grip strength on the affected side (kg)	10.0 ± 10.6 (0–43.3)	
Grip strength on the unaffected side (kg)	21.1 ± 10.5 (0–50)	
BMI (kg/m2)	22.1 ± 3.7 (13.2–39.3)	
GNRI	97.0 ± 10.5 (64.9–136.9)	
Gait independence at discharge (%)		
 Yes	519 (61.6%)	
 No	323 (38.4%)	
FIM, Functional Independence Measure; BMI, body mass index; GNRI, Geriatric Nutritional Risk Index.

AUCs and overlearning of validation and test data for each prediction model

Figure 2 shows the validation data set ROC curves for the six prediction models. The KNN model AUC was significantly higher than those for the other prediction models. KNN and EL had significantly higher AUCs than LR (p < 0.001), whereas the AUC for DT was significantly lower than those for the other five prediction models (p < 0.001).Fig. 2 ROC curves for the validation and test data sets for each prediction model. A, KNN had the highest AUC, which was significantly higher than those of the other predictive models, followed by EL. The AUC of DT was significantly lower than those of the other five prediction models (p < 0.001). B, The AUC of DT was significantly lower than those of the other five prediction models (p < 0.01). In contrast, the AUC of LR did not differ significantly from those of the ANN, SVM, KNN, and EL models. ANN, artificial neural network; AUC, area under the curve; DT, decision tree; EL, ensemble learning; KNN, K-nearest neighbor; LR, logistic regression analysis; SVM, support vector machine; ROC, receiver operating characteristic.

Figure 2 depicts the test data set ROC curves for the six prediction models. The AUC for DT was the lowest, whereas those for the others were similar to each other. The test data set AUC was 0.812 for DT (95% CI, 0.746–0.872), 0.895 for LR (95% CI, 0.836–0.945), 0.893 for SVM (95% CI, 0.832–0.947), 0.903 for KNN (95% CI, 0.848–0.950), 0.901 (95% CI, 0.847–0.948) for EL, and 0.900 for ANN (95% CI, 0.844–0.948) (Table 2).Table 2 AUC values and other performance metrics for the validation and test data sets.

	Validation data set	Test data set	Differences in AUC	
AUC	95% CI	Accuracy	Precision	Recall	F1 score	Sensitivity	Specificity	AUC	95% CI	Accuracy	Precision	Recall	F1 score	Sensitivity	Specificity	p	
DT	0.835	0.803–0.864	0.826	0.673	0.897	0.769	0.897	0.711	0.812	0.746–0.872	0.810	0.684	0.921	0.785	0.921	0.642	0.516	
LR	0.920	0.895–0.942	0.832	0.661	0.888	0.758	0.888	0.742	0.895	0.836–0.945	0.863	0.634	0.911	0.748	0.911	0.791	0.427	
ANN	0.916	0.891–0.938	0.843	0.657	0.892	0.757	0.892	0.762	0.900	0.844–0.948	0.857	0.639	0.911	0.751	0.911	0.776	0.596	
SVM	0.917	0.892–0.939	0.838	0.658	0.890	0.757	0.890	0.754	0.893	0.832–0.947	0.869	0.630	0.911	0.745	0.911	0.806	0.447	
KNN	1.000	1.000–1.000	0.829	0.680	0.909	0.778	0.909	0.699	0.903	0.848–0.950	0.833	0.664	0.921	0.772	0.921	0.701	0.000	
EL	0.944	0.925–0.960	0.829	0.662	0.885	0.757	0.885	0.738	0.901	0.847–0.948	0.857	0.639	0.911	0.751	0.911	0.776	0.124	
AUC, area under the receiver operating characteristic curve; CI, confidence interval; DT, decision tree; LR, logistic regression analysis; SVM, support vector machine; KNN, K-nearest neighbor; ANN, artificial neural network; EL, ensemble learning.

Statistical comparison of the validation and test data set AUCs to evaluate overlearning revealed a significant difference between the data sets for KNN (p < 0.001), suggesting a risk for overlearning. The other prediction models showed no statistically significant differences.

Comparison of AUCs among the prediction models

Upon investigating significant differences in the validation data set AUCs for the six algorithms using the pROC package, we found that the AUC for DT was significantly lower than those for the other five prediction models (p < 0.001). KNN and EL had significantly higher AUCs than LR (p < 0.001).

The test data set results showed that the AUC for DT was significantly lower than those for the other five prediction models (p < 0.01). Conversely, we observed no significant differences among the AUCs for LR and the other four ML prediction models. Thus, LR did not show a statistically significant decrease compared with ML (Table 3).Table 3 Statistically significant differences in AUCs between DT and LR.

	Validation data	Test data	
p (vs. DT)	p (vs. LR)	p (vs. DT)	p (vs. LR)	
LR	0.000	–	0.003	–	
ANN	0.000	0.111	0.001	0.418	
SVM	0.000	0.150	0.005	0.671	
KNN	0.000	0.000	0.000	0.536	
EL	0.000	0.000	0.001	0.637	
AUC, area under the receiver operating characteristic curve; DT, decision tree; LR, logistic regression analysis; ANN, artificial neural network; SVM, support vector machine; KNN, K-nearest neighbor; EL, ensemble learning.

Other validation and test data set performance metrics for each prediction model

For the validation data, the accuracy, precision, recall, F1 score, and specificity for each algorithm were 0.826–0.843, 0.657–0.680, 0.885–0.909, 0.757–0.778, and 0.699–0.762, respectively. For the test data, the accuracy, precision, recall, F1 score, and specificity for each algorithm were 0.810–0.869, 0.630–0.684, 0.911–0.921, 0.745–0.785, and 0.642–0.806, respectively. The specificity of DT and KNN were lower than those of the other algorithms. Additionally, LR did not show significant differences in these performance metrics compared with the other ML algorithms. The confusion matrices for each algorithm are provided in Supplementary Table 3.

Discussion

In this study, the LR as well as DT, ANN, SVM, KNN, and EL ML methods were used to predict the prognosis of gait independence at discharge in patients with subacute stroke. Comparison of the prediction accuracies of the model showed that DT had significantly lower prediction accuracy than the other models. In contrast, LR showed no decrease in prediction accuracy compared with the other ML models. To our knowledge, this study represents the first to report a comparison of the predictive accuracy of LR and ML algorithms for gait independence prognosis in subacute stroke patients. We had hypothesized that while there may be differences in prediction accuracies among different ML algorithms, LR may have sufficient accuracy compared with ML for predicting gait independence in patients with stroke. There were no significant differences in the AUCs for ML (excluding DT) and LR, with both showing high predictive accuracy (AUC values of approximately 0.9).

As DT had a significantly lower AUC than the LR, SVM, KNN, EL, and ANN algorithms, it was found to be less accurate than them in predicting the prognosis of gait independence in patients with subacute stroke. DT is a powerful tool used to visualize the relationships among explanatory factors in prediction models and has been utilized in previous studies on the prognosis of patients with stroke15,16. However, a systematic review of ML methods reported that DT had lower prediction accuracy of AUC than other ML algorithms18. These findings are similar to the results of our study, in which DT had the worst prediction accuracy. In addition, a systematic review on the use of ML for predicting the prognosis of patients with stroke reported a decrease in the use of DT for this purpose in recent years. In contrast, the number of studies that applied the EL-based random forest (RF) method has remained constant9. Therefore, RF-based prognostic prediction should also be assessed in the future.

In this study, the AUC for LR did not differ significantly from those for SVM, KNN, EL, and ANN. A previous systematic review reported that ML did not outperform LR in terms of AUC accuracy18, which is consistent with our result. However, ML has a “black box” problem, rendering it challenging to understand the prediction model and thus leading to the emerging interest in developing methodologies that can solve and explain the these models31. In contrast, LR has the advantage of handling two or more explanatory variables, allowing manual calculations as well as the understanding of the prediction models in mathematical form32. Therefore, LR is often selected in actual clinical practice, such as with the quick Sepsis Related Organ Failure Assessment tool33. Our results suggest that LR could be adopted for predicting the gait independence of patients with subacute stroke due to its advantages of easy application in actual clinical practice and accuracy comparable to that of ML.

We also evaluated prediction accuracies using a test data set to assess possible overlearning. LR, KNN, EL, and ANN had both high prediction accuracies (AUCs of approximately 0.9) and no overlearning. A review article on the prognosis of gait independence of patients with stroke reported that none of the previous studies had evaluated the internal and external validities of the models5. Another study that did assess external validity reported an AUC of up to 0.8234. Internal validity is important in ML, because generalization performance deteriorates when overlearning occurs. As the present study was conducted at a single center, it was challenging to evaluate external validity; however, we did assess internal validity using a test data set. The results revealed a significant difference in the validation and test data set AUCs for KNN, indicating possible overlearning. Conversely, we observed no such significant differences for the other prediction models, indicating that overlearning did not occur in these models. Hence, the internal validity of the prediction models was considered satisfactory, although future studies on external validity are warranted. In our study, the analysis of a sufficiently high number of participants is likely to be the main reason why overlearning did not occur, except for KNN; moreover, the automatic hyperparameter optimization in MATLAB is also likely to have prevented overlearning.

The AUCs for the test data were approximately 0.9. A review article on the prognosis for patients with stroke extracted data from six studies on gait independence, but none of these studies assessed internal validity5. A recent LR-based study examining gait independence in patients with acute stroke reported an AUC of 0.9112; however, this study also did not examine internal validity, such as by testing for overlearning with test data. The prediction accuracies for the test data set in the present study were expected to be as good as those reported previously. Therefore, we established models with sufficient prediction accuracies.

Furthermore, a previous study that predicted gait independence reported an AUC of 0.958 for DT16, which is higher than that in the present study. However, DT is prone to overlearning without appropriate learning procession, such as setting a stopping criterion14. The evaluation of overlearning in the previous study using a test data set was unclear and may require additional examination. The present study ruled out overlearning by using a test data set that was not used for model development.

Previous studies have also examined the accuracy of predicting mRS using ML. A study that predicted mRS and compared the prediction accuracies of deep neural network (DNN), random forest, LR, and ASTRAL scores reported that DNN showed the best prediction accuracy, with an AUC of 0.88811. Another study that predicted ADL independence by mRS based on imaging findings reported an AUC of 0.85635. The AUCs in the present study were approximately 0.9, higher than those in the previous studies, despite our models predicting gait independence, which is more accurate than mRS. Therefore, our models had sufficient predictive accuracy; however, the differences in findings may be due to differences in timing between the studies. Moreover, previous studies predicted the functional outcomes for patients with acute stroke, whereas the present study predicted outcomes in patients with subacute stroke.

With regard to performance metrics other than AUC, LR did not show significant differences compared with the other ML algorithms. Although these performance metrics cannot determine statistical significance, the lack of substantial differences suggests that LR has comparable performance to other ML algorithms. Therefore, LR may continue to be used alongside ML for the prognostic prediction of gait independence in subacute stroke patients, especially when understanding the predictive model is important.

This study had a few limitations. First, owing to the single-study design, the possibility of overlearning in our hospital cannot be excluded. Additional studies are needed to investigate the external validity of our findings using data from other hospitals. Second, as this was a retrospective study, some cases lacked data and had limited clinical indicators. Third, although imaging findings, such as those from computed tomography and magnetic resonance imaging, can provide additional information, we could not analyze these findings using deep learning in this study. Fourth, the hyperparameters were automatically optimized by MATLAB software; however, further improvement of the algorithms may improve the accuracies. Fifth, this study did not evaluate other algorithms, such as gradient boosting algorithms, which were not implementable in MATLAB. Therefore, future research needs to consider these algorithms as well. Finally, this study was not a predictive study for patients with acute stroke, as the prognosis was predicted based on data from the subacute phase, approximately 1 month after the stroke onset. Therefore, our study cannot be directly compared with previous studies on the acute phase. Nevertheless, the strength of this study is that the findings add important knowledge for predicting the prognosis of patients who cannot walk even 1 month after stroke onset and that it involved a detailed evaluation of gait function at discharge using large-scale data.

Conclusion

Comparisons of the accuracies of the DT, ANN, SVM, KNN, and EL ML algorithms and LR for predicting gait independence at discharge in patients with subacute stroke showed that the DT model had a significantly lower prediction accuracy than the other models. Additionally, the accuracy of LR did not differ significantly from those of the other ML algorithms, all of which showed had high prediction accuracies (AUCs of approximately 0.9) for gait independence at discharge in patients with subacute stroke, even in the data set that was not used for training. Our results indicating that its predictive accuracy was not less than that of the ML approaches suggest that LR can continue to be used to determine the prognosis of gait independence, with the advantage that LR predictive models are understandable and can be computed manually. However, future studies are needed to assess the validity of these findings in other clinical settings and populations, including by incorporating imaging findings, and to further optimize the hyperparameters used in the prediction models.

Supplementary Information

Supplementary Tables.

Supplementary Information

The online version contains supplementary material available at 10.1038/s41598-024-72206-4.

The authors thank Mr. Isao Sato for his assistance in data acquisition.

Author contributions

YM, MK, KK, AH, TK, TA, NH, MT, KH, and KS conceived the idea of the study. YM acquired the data, developed predictive models, and conducted the statistical analysis. YM, MK, KK, AH, TK, TA, MT, KH, and TT contributed to interpreting the results. YM wrote the first draft of the manuscript. MK and KK supervised this study. All authors reviewed and edited the manuscript and approved the final version of the manuscript.

Funding

This study was supported by the Japan Society for the Promotion of Science (JSPS) KAKENHI Grant (Grant No. JP22K21225 to YM) and the Francebed Medical Home Care Research Subsidy Foundation (to YM).

Data availability 

The present research involved human research participant data, which raises ethical issues regarding the protection of personal information. However, the data are available from a reasonable request with permission from the corresponding author and the Ethics Committee in the Tokyo Bay Rehabilitation Hospital.

Code availability 

The code used in this study is available online at https://doi.org/10.5281/zenodo.12685295.

Competing interests

MK is the founder of INTEP, Inc., which is unrelated to this study. The other authors have no conflicts of interest to disclose.

Publisher's note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
==== Refs
References

1. Collaborators GBDS Global, regional, and national burden of stroke, 1990–2016: A systematic analysis for the Global Burden of Disease Study 2016 Lancet Neurol. 2019 18 439 458 10.1016/S1474-4422(19)30034-1 30871944
Collaborators, G. B. D. S. Global, regional, and national burden of stroke, 1990–2016: A systematic analysis for the Global Burden of Disease Study 2016. Lancet Neurol. 18, 439–458. 10.1016/S1474-4422(19)30034-1 (2019).30871944 10.1016/S1474-4422(19)30034-1
2. Ishiwatari M Prediction of gait independence using the Trunk Impairment Scale in patients with acute stroke Ther. Adv. Neurol. Disord. 2022 15 17562864221140180 10.1177/17562864221140180 36506941
Ishiwatari, M. et al. Prediction of gait independence using the Trunk Impairment Scale in patients with acute stroke. Ther. Adv. Neurol. Disord. 15, 17562864221140180. 10.1177/17562864221140180 (2022).36506941 10.1177/17562864221140180
3. Kwakkel G Motor rehabilitation after stroke: European Stroke Organisation (ESO) consensus-based definition and guiding framework Eur. Stroke J. 2023 8 880 894 10.1177/23969873231191304 37548025
Kwakkel, G. et al. Motor rehabilitation after stroke: European Stroke Organisation (ESO) consensus-based definition and guiding framework. Eur. Stroke J. 8, 880–894. 10.1177/23969873231191304 (2023).37548025 10.1177/23969873231191304
4. Deutsch A Poststroke rehabilitation: Outcomes and reimbursement of inpatient rehabilitation facilities and subacute rehabilitation programs Stroke 2006 37 1477 1482 10.1161/01.STR.0000221172.99375.5a 16627797
Deutsch, A. et al. Poststroke rehabilitation: Outcomes and reimbursement of inpatient rehabilitation facilities and subacute rehabilitation programs. Stroke 37, 1477–1482. 10.1161/01.STR.0000221172.99375.5a (2006).16627797 10.1161/01.STR.0000221172.99375.5a
5. Stinear CM Smith MC Byblow WD Prediction tools for stroke rehabilitation Stroke 2019 50 3314 3322 10.1161/strokeaha.119.025696 31610763
Stinear, C. M., Smith, M. C. & Byblow, W. D. Prediction tools for stroke rehabilitation. Stroke 50, 3314–3322. 10.1161/strokeaha.119.025696 (2019).31610763 10.1161/strokeaha.119.025696
6. Veerbeek JM Van Wegen EE Harmeling-Van der Wel BC Kwakkel G Is accurate prediction of gait in nonambulatory stroke patients possible within 72 hours poststroke? The EPOS study Neurorehabil. Neural. Repair 2011 25 268 274 10.1177/1545968310384271 21186329
Veerbeek, J. M., Van Wegen, E. E., Harmeling-Van der Wel, B. C. & Kwakkel, G. Is accurate prediction of gait in nonambulatory stroke patients possible within 72 hours poststroke? The EPOS study. Neurorehabil. Neural. Repair 25, 268–274. 10.1177/1545968310384271 (2011).21186329 10.1177/1545968310384271
7. Senders JT Machine learning and neurosurgical outcome prediction: A systematic review World Neurosurg. 2018 109 476 486.e1 10.1016/j.wneu.2017.09.149 28986230
Senders, J. T. et al. Machine learning and neurosurgical outcome prediction: A systematic review. World Neurosurg. 109, 476-486.e1. 10.1016/j.wneu.2017.09.149 (2018).28986230 10.1016/j.wneu.2017.09.149
8. Benedetto U Machine learning improves mortality risk prediction after cardiac surgery: Systematic review and meta-analysis J. Thorac. Cardiovasc. Surg. 2022 163 2075 2087.e9 10.1016/j.jtcvs.2020.07.105 32900480
Benedetto, U. et al. Machine learning improves mortality risk prediction after cardiac surgery: Systematic review and meta-analysis. J. Thorac. Cardiovasc. Surg. 163, 2075-2087.e9. 10.1016/j.jtcvs.2020.07.105 (2022).32900480 10.1016/j.jtcvs.2020.07.105
9. Wang W A systematic review of machine learning models for predicting outcomes of stroke with structured data PLoS One 2020 15 e0234722 10.1371/journal.pone.0234722 32530947
Wang, W. et al. A systematic review of machine learning models for predicting outcomes of stroke with structured data. PLoS One 15, e0234722. 10.1371/journal.pone.0234722 (2020).32530947 10.1371/journal.pone.0234722
10. Mainali S Darsie ME Smetana KS Machine learning in action: Stroke diagnosis and outcome prediction Front. Neurol. 2021 12 734345 10.3389/fneur.2021.734345 34938254
Mainali, S., Darsie, M. E. & Smetana, K. S. Machine learning in action: Stroke diagnosis and outcome prediction. Front. Neurol. 12, 734345. 10.3389/fneur.2021.734345 (2021).34938254 10.3389/fneur.2021.734345
11. Heo J Machine learning-based model for prediction of outcomes in acute stroke Stroke 2019 50 1263 1265 10.1161/strokeaha.118.024293 30890116
Heo, J. et al. Machine learning-based model for prediction of outcomes in acute stroke. Stroke 50, 1263–1265. 10.1161/strokeaha.118.024293 (2019).30890116 10.1161/strokeaha.118.024293
12. Lin WY Predicting post-stroke activities of daily living through a machine learning-based approach on initiating rehabilitation Int. J. Med. Inform. 2018 111 159 164 10.1016/j.ijmedinf.2018.01.002 29425627
Lin, W. Y. et al. Predicting post-stroke activities of daily living through a machine learning-based approach on initiating rehabilitation. Int. J. Med. Inform. 111, 159–164. 10.1016/j.ijmedinf.2018.01.002 (2018).29425627 10.1016/j.ijmedinf.2018.01.002
13. Miyazaki Y Improvement of predictive accuracies of functional outcomes after subacute stroke inpatient rehabilitation by machine learning models PLoS ONE 2023 18 e0286269 10.1371/journal.pone.0286269 37235575
Miyazaki, Y. et al. Improvement of predictive accuracies of functional outcomes after subacute stroke inpatient rehabilitation by machine learning models. PLoS ONE 18, e0286269. 10.1371/journal.pone.0286269 (2023).37235575 10.1371/journal.pone.0286269
14. Bi QF Goodman KE Kaminsky J Lessler J What is machine learning? A primer for the epidemiologist Am. J. Epidemiol. 2019 188 2222 2239 10.1093/aje/kwz189 31509183
Bi, Q. F., Goodman, K. E., Kaminsky, J. & Lessler, J. What is machine learning? A primer for the epidemiologist. Am. J. Epidemiol. 188, 2222–2239. 10.1093/aje/kwz189 (2019).31509183 10.1093/aje/kwz189
15. Fujita T Functions necessary for gait independence in patients with stroke: A study using decision tree J. Stroke Cerebrovasc. Dis. 2020 29 104998 10.1016/j.jstrokecerebrovasdis.2020.104998 32689598
Fujita, T. et al. Functions necessary for gait independence in patients with stroke: A study using decision tree. J. Stroke Cerebrovasc. Dis. 29, 104998. 10.1016/j.jstrokecerebrovasdis.2020.104998 (2020).32689598 10.1016/j.jstrokecerebrovasdis.2020.104998
16. Inoue Y Imura T Tanaka R Matsuba J Harada K Developing a clinical prediction rule for gait independence at discharge in patients with stroke: A decision-tree algorithm analysis J. Stroke Cerebrovasc. Dis. 2022 31 106441 10.1016/j.jstrokecerebrovasdis.2022.106441 35305537
Inoue, Y., Imura, T., Tanaka, R., Matsuba, J. & Harada, K. Developing a clinical prediction rule for gait independence at discharge in patients with stroke: A decision-tree algorithm analysis. J. Stroke Cerebrovasc. Dis. 31, 106441. 10.1016/j.jstrokecerebrovasdis.2022.106441 (2022).35305537 10.1016/j.jstrokecerebrovasdis.2022.106441
17. Kim JK Choo YJ Chang MC Prediction of motor function in stroke patients using machine learning algorithm: Development of practical models J. Stroke Cerebrovasc. Dis. 2021 30 105856 10.1016/j.jstrokecerebrovasdis.2021.105856 34022582
Kim, J. K., Choo, Y. J. & Chang, M. C. Prediction of motor function in stroke patients using machine learning algorithm: Development of practical models. J. Stroke Cerebrovasc. Dis. 30, 105856. 10.1016/j.jstrokecerebrovasdis.2021.105856 (2021).34022582 10.1016/j.jstrokecerebrovasdis.2021.105856
18. Christodoulou E A systematic review shows no performance benefit of machine learning over logistic regression for clinical prediction models J. Clin. Epidemiol. 2019 110 12 22 10.1016/j.jclinepi.2019.02.004 30763612
Christodoulou, E. et al. A systematic review shows no performance benefit of machine learning over logistic regression for clinical prediction models. J. Clin. Epidemiol. 110, 12–22. 10.1016/j.jclinepi.2019.02.004 (2019).30763612 10.1016/j.jclinepi.2019.02.004
19. Nusinovici S Logistic regression was as good as machine learning for predicting major chronic diseases J. Clin. Epidemiol. 2020 122 56 69 10.1016/j.jclinepi.2020.03.002 32169597
Nusinovici, S. et al. Logistic regression was as good as machine learning for predicting major chronic diseases. J. Clin. Epidemiol. 122, 56–69. 10.1016/j.jclinepi.2020.03.002 (2020).32169597 10.1016/j.jclinepi.2020.03.002
20. Oczkowski WJ Barreca S Neural network modeling accurately predicts the functional outcome of stroke survivors with moderate disabilities Arch. Phys. Med. Rehabil. 1997 78 340 345 10.1016/s0003-9993(97)90222-7 9111450
Oczkowski, W. J. & Barreca, S. Neural network modeling accurately predicts the functional outcome of stroke survivors with moderate disabilities. Arch. Phys. Med. Rehabil. 78, 340–345. 10.1016/s0003-9993(97)90222-7 (1997).9111450 10.1016/s0003-9993(97)90222-7
21. Raudys SJ Jain AK Small sample size effects in statistical pattern recognition: Recommendations for practitioners IEEE Trans. Pattern Anal. Mach. Intell. 1991 13 252 264 10.1109/34.75512
Raudys, S. J. & Jain, A. K. Small sample size effects in statistical pattern recognition: Recommendations for practitioners. IEEE Trans. Pattern Anal. Mach. Intell. 13, 252–264. 10.1109/34.75512 (1991).10.1109/34.75512
22. World Medical, A World Medical Association Declaration of Helsinki: Ethical principles for medical research involving human subjects JAMA 2013 310 2191 2194 10.1001/jama.2013.281053 24141714
World Medical, A. World Medical Association Declaration of Helsinki: Ethical principles for medical research involving human subjects. JAMA 310, 2191–2194. 10.1001/jama.2013.281053 (2013).24141714 10.1001/jama.2013.281053
23. Data management service of the Uniform Data System for Medical, R. & the Center for Functional Assessment, R. Guide for use of the uniform data set for medical rehabilitation. 3rd edn, (State University of New York at Buffalo, 1990).
24. Liu, M., Sonoda, S. & Domen, K. Stroke Impairment Assessment Set (SIAS) and Functional Independence Measure (FIM) and their practical use. In: Chino N, ed. Functional Assessment of Stroke Patients: Practical Aspects of SIAS and FIM. (Splinger, 1997).
25. Tsuji T Liu M Sonoda S Domen K Chino N The stroke impairment assessment set: Its internal consistency and predictive validity Arch. Phys. Med. Rehabil. 2000 81 863 868 10.1053/apmr.2000.6275 10895996
Tsuji, T., Liu, M., Sonoda, S., Domen, K. & Chino, N. The stroke impairment assessment set: Its internal consistency and predictive validity. Arch. Phys. Med. Rehabil. 81, 863–868. 10.1053/apmr.2000.6275 (2000).10895996 10.1053/apmr.2000.6275
26. Bouillanne O Geriatric Nutritional Risk Index: a new index for evaluating at-risk elderly medical patients Am. J. Clin. Nutr. 2005 82 777 783 10.1093/ajcn/82.4.777 16210706
Bouillanne, O. et al. Geriatric Nutritional Risk Index: a new index for evaluating at-risk elderly medical patients. Am. J. Clin. Nutr. 82, 777–783. 10.1093/ajcn/82.4.777 (2005).16210706 10.1093/ajcn/82.4.777
27. Miyazaki Y Comparing the contribution of each clinical indicator in predictive models trained on 980 subacute stroke patients: A retrospective study Sci. Rep. 2023 13 12324 10.1038/s41598-023-39475-x 37516806
Miyazaki, Y. et al. Comparing the contribution of each clinical indicator in predictive models trained on 980 subacute stroke patients: A retrospective study. Sci. Rep. 13, 12324. 10.1038/s41598-023-39475-x (2023).37516806 10.1038/s41598-023-39475-x
28. Liu DC Nocedal J On the limited memory BFGS method for large scale optimization Math. Program. 1989 45 503 528 10.1007/BF01589116
Liu, D. C. & Nocedal, J. On the limited memory BFGS method for large scale optimization. Math. Program. 45, 503–528. 10.1007/BF01589116 (1989).10.1007/BF01589116
29. Mandrekar JN Receiver operating characteristic curve in diagnostic test assessment J. Thorac. Oncol. 2010 5 1315 1316 10.1097/JTO.0b013e3181ec173d 20736804
Mandrekar, J. N. Receiver operating characteristic curve in diagnostic test assessment. J. Thorac. Oncol. 5, 1315–1316. 10.1097/JTO.0b013e3181ec173d (2010).20736804 10.1097/JTO.0b013e3181ec173d
30. Robin X pROC: An open-source package for R and S+ to analyze and compare ROC curves BMC Bioinf. 2011 12 77 10.1186/1471-2105-12-77
Robin, X. et al. pROC: An open-source package for R and S+ to analyze and compare ROC curves. BMC Bioinf. 12, 77. 10.1186/1471-2105-12-77 (2011).10.1186/1471-2105-12-77
31. Rudin C Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead Nat. Mach. Intell. 2019 1 206 215 10.1038/s42256-019-0048-x 35603010
Rudin, C. Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead. Nat. Mach. Intell. 1, 206–215. 10.1038/s42256-019-0048-x (2019).35603010 10.1038/s42256-019-0048-x
32. Sperandei S Understanding logistic regression analysis Biochem. Med. 2014 24 12 18 10.11613/bm.2014.003
Sperandei, S. Understanding logistic regression analysis. Biochem. Med. 24, 12–18. 10.11613/bm.2014.003 (2014).10.11613/bm.2014.003
33. Seymour CW Assessment of clinical criteria for sepsis: for the Third International Consensus Definitions for Sepsis and Septic Shock (Sepsis-3) JAMA 2016 315 762 774 10.1001/jama.2016.0288 26903335
Seymour, C. W. et al. Assessment of clinical criteria for sepsis: for the Third International Consensus Definitions for Sepsis and Septic Shock (Sepsis-3). JAMA 315, 762–774. 10.1001/jama.2016.0288 (2016).26903335 10.1001/jama.2016.0288
34. Langerak AJ Externally validated model predicting gait independence after stroke showed fair performance and improved after updating J. Clin. Epidemiol. 2021 137 73 82 10.1016/j.jclinepi.2021.03.022 33812010
Langerak, A. J. et al. Externally validated model predicting gait independence after stroke showed fair performance and improved after updating. J. Clin. Epidemiol. 137, 73–82. 10.1016/j.jclinepi.2021.03.022 (2021).33812010 10.1016/j.jclinepi.2021.03.022
35. Brugnara G Multimodal predictive modeling of endovascular treatment outcome for acute ischemic stroke using machine-learning Stroke 2020 51 3541 3551 10.1161/strokeaha.120.030287 33040701
Brugnara, G. et al. Multimodal predictive modeling of endovascular treatment outcome for acute ischemic stroke using machine-learning. Stroke 51, 3541–3551. 10.1161/strokeaha.120.030287 (2020).33040701 10.1161/strokeaha.120.030287
