
==== Front
Sci Rep
Sci Rep
Scientific Reports
2045-2322
Nature Publishing Group UK London

39232178
71072
10.1038/s41598-024-71072-4
Article
Multi-modality multi-task model for mRS prediction using diffusion-weighted resonance imaging
Park In-Seo 110
Kim Seongheon 34
Jang Jae-Won 1234
Park Sang-Won 34
Yeo Na-Young 24
Seo Soo Young 56
Jeon Inyeop 6
Shin Seung-Ho 6
Kim Yoon 810
Choi Hyun-Soo choi.hyunsoo@seoultech.ac.kr

910
Kim Chulho gumdol52@hallym.or.kr

7
1 https://ror.org/01mh5ph17 grid.412010.6 0000 0001 0707 9039 Department of Convergence Security, Kangwon National University, Chuncheon, 24253 Korea
2 https://ror.org/01mh5ph17 grid.412010.6 0000 0001 0707 9039 Department of Medical Bigdata Convergence, Kangwon National University, Chuncheon, 24253 Korea
3 https://ror.org/01mh5ph17 grid.412010.6 0000 0001 0707 9039 Department of Medical Informatics, Kangwon National University, Chuncheon, 24253 Korea
4 https://ror.org/01rf1rj96 grid.412011.7 0000 0004 1803 0072 Department of Neurology, Kangwon National University Hospital, Chuncheon, 24253 Korea
5 https://ror.org/03sbhge02 grid.256753.0 0000 0004 0470 5964 Institute of New Frontier Research Team, Hallym University College of Medicine, Chuncheon, 24252 Korea
6 https://ror.org/05hwzrf74 grid.464534.4 0000 0004 0647 1735 Chuncheon Artificial Intelligence Center, Chuncheon Sacred Heart Hospital, Chuncheon, 24253 Korea
7 https://ror.org/05hwzrf74 grid.464534.4 0000 0004 0647 1735 Department of Neurology, Chuncheon Sacred Heart Hospital, Chuncheon, 24253 Korea
8 https://ror.org/01mh5ph17 grid.412010.6 0000 0001 0707 9039 Department of Computer Science and Engineering, Kangwon National University, Chuncheon, 24253 Korea
9 https://ror.org/00chfja07 grid.412485.e 0000 0000 9760 4919 Department of Computer Science and Engineering, Seoul National University of Science and Technology, Seoul, South Korea
10 ZIOVISION, Chuncheon, 24341 Korea
4 9 2024
4 9 2024
2024
14 2057213 12 2023
23 8 2024
© The Author(s) 2024
2024
https://creativecommons.org/licenses/by-nc-nd/4.0/ Open Access This article is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License, which permits any non-commercial use, sharing, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if you modified the licensed material. You do not have permission under this licence to share adapted material derived from this article or parts of it. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by-nc-nd/4.0/.
This study focuses on predicting the prognosis of acute ischemic stroke patients with focal neurologic symptoms using a combination of diffusion-weighted magnetic resonance imaging (DWI) and clinical information. The primary outcome is a poor functional outcome defined by a modified Rankin Scale (mRS) score of 3–6 after 3 months of stroke. Employing nnUnet for DWI lesion segmentation, the study utilizes both multi-task and multi-modality methodologies, integrating DWI and clinical data for prognosis prediction. Integrating the two modalities was shown to improve performance by 0.04 compared to using DWI only. The model achieves notable performance metrics, with a dice score of 0.7375 for lesion segmentation and an area under the curve of 0.8080 for mRS prediction. These results surpass existing scoring systems, showing a 0.16 improvement over the Totaled Health Risks in Vascular Events score. The study further employs grad-class activation maps to identify critical brain regions influencing mRS scores. Analysis of the feature map reveals the efficacy of the multi-tasking nnUnet in predicting poor outcomes, providing insights into the interplay between DWI and clinical data. In conclusion, the integrated approach demonstrates significant advancements in prognosis prediction for cerebral infarction patients, offering a superior alternative to current scoring systems.

Subject terms

Stroke
Stroke
Computer science
Neurology
http://dx.doi.org/10.13039/501100002507 Kangwon National University Korea government (Ministry of Science and ICT)2022R1A5A708390811 2022R1A5A708390811 2022R1A5A708390811 2022R1A5A708390811 2022R1F1A1076454 Kim Seongheon Jang Jae-Won Park Sang-Won Yeo Na-Young Choi Hyun-Soo Ministry of Education (MOE)2023RIS-005 2023RIS-005 2023RIS-005 2023RIS-005 Kim Seongheon Jang Jae-Won Park Sang-Won Yeo Na-Young Ministry of Health & Welfare, Republic of KoreaHR21C0198 HR21C0198 HR21C0198 HR21C0198 Seo Soo Young Jeon Inyeop Shin Seung-Ho Kim Chulho http://dx.doi.org/10.13039/501100002553 Seoul National University of Science and Technology issue-copyright-statement© Springer Nature Limited 2024
==== Body
pmcIntroduction

Stroke is the fifth leading cause of death in the United States, accounting for approximately 795,000 new cases each year1. The symptoms can drastically affect the quality of life for individuals suffering from cerebral infarction, highlighting the importance of tailoring treatment options and rehabilitation strategies to the severity and prognostic status of the stroke2–4. Accurate prediction of stroke outcomes is essential for clinicians to make informed decisions about acute therapies, such as thrombolysis or endovascular therapy, and to plan effective long-term management strategies for survivors. Furthermore, predictive models support the development and assessment of new stroke treatments and interventions.

Recent advances extend beyond traditional decision-making frameworks that use structured data. There have been notable successes in applying deep learning techniques to unstructured data in clinical decision-making across various diseases5,6. Deep learning algorithms that process unstructured data like images and text not only predict disease outcomes effectively but also identify digital phenotypes that correlate with these prognoses7,8.

In this study, we have enhanced the nnUnet architecture to support a multi-task functionality, allowing it to simultaneously predict modified Rankin Scale (mRS) scores and segment affected areas within a single, integrated workflow9. This adaptation represents a significant evolution from its original design, tailored to leverage the detailed imaging data to improve both segmentation and prognosis predictions. We also employed a two-stage methodology: initially using nnUnet for imaging tasks followed by integrating the outputs with clinical data using AutoGluon10. This multi-modality approach aims to synthesize data from diverse sources to enhance the overall predictive accuracy of stroke outcomes.

Additionally, we evaluated whether these deep learning-based prognostic models offer performance benefits over traditional prognostic factors like the Totaled Health Risks in Vascular Events (THRIVE) score or the National Institutes of Health Stroke Scale (NIHSS)11.

Materials and methods

Study population

Consecutive patients who underwent DWI for acute ischemic stroke (AIS) were retrospectively recruited from two hospitals (Hallym University Chuncheon Sacred Heart Hospital (HUCSHH) and Kangwon National University Hospital (KNUH). Only frist-ever ischemic stroke patients with an mRS of 0 were included, as the severity of the previous stroke might influence the prognosis of the current stroke. In addition, we selected the patients were those who presented within 12 hours of symptom onset and had an MRI within 24 hours of presentation. The AIS cases in each hospital were as follows: from January 2011 to December 2019 for HUCSHH, and from July 2014 to October 2019 for KNUH. Eventually, 1925 subjects with DWI images of AIS (for HUCSHH, 657 men and 441 women with a mean ± standard deviation age of 69.8 ± 12.7, and for KNUH, 485 men and 345 women with a mean ± standard deviation age of 72.6 ± 12.0) were included. Additionally, 121 DWI images of control subjects (for HUCSHH, 50 men and 52 women with a mean ± standard deviation age of 59.2 ± 16.4, and for KNUH, 10 men and 11 women with a mean ± standard deviation age of 44.0 ± 17.1) were included in the evaluation of the developed segmentation algorithm. The control group comprised individuals who presented with a headache, and MRI confirmed they were not associated with an acute stroke. This retrospective study was approved by the Institutional Review Boards of HUCSHH and KNUH, which waived the requirement for informed consent (approval no. HUCSHH 2021-06-013 and KNUH-A-2021-021-001) and all methods were performed following applicable guidelines and regulations. In addition, We also utilized data from 3378 subjects provided by Chonnam National University and released through the AI Hub12.

Data

Diffuse-weighted image acquisition

MRI was performed with various machines, including a 1.5T (Magnetom Avanto, Siemens Healthineers, Erlangen, Germany) and a 3T scanner (Ingenia CX, Philips Healthcare; Achieva, Philips Healthcare, Best, The Netherlands). The parameters for the DWI sequences were as follows: repetition time, 3000–8000 ms; echo time, 56–103 ms; flip angle, 90∘; matrix, 256×256–512×512; field of view, 220×220–256×256 mm; the number of excitations, 1–5; the number of slices, 20–50; slice thickness, 3 mm. We used the b-value of 1000 seconds/mm2 for ischemic lesion identification.

Annotation

All lesions on DWI were manually segmented by one neurologist (C.K) in HUCSHH and by two neurologists (S.H.K and J.-W.J) in KNUH using ITK-Snap, an open-source software application used to segment structures in medical images13. All images with masks were saved in Digital Imaging and Communications in Medicine (DICOM) format.

mRS distribusion

In this study, the distribution of mRS scores within the dataset reveals significant insights into patient outcomes post-intervention. The dataset encompasses a total of 5429 cases, segmented by mRS scores from 0 to 6. Notably, scores of 0–2, which are considered indicative of good outcomes, collectively account for 65.93% (3575 cases) of the dataset. This group is characterized by mRS scores of 0 (20.53%, 1111 cases), 1 (26.36%, 1431 cases), and 2 (19.02%, 1035 cases). Conversely, scores ranging from 3 to 6, categorized as bad outcomes, represent 34.06% (1853 cases). This includes mRS scores of 3 (13.36%, 726 cases), 4 (10.75%, 585 cases), 5 (6.82%, 371 cases), and 6 (3.12%, 170 cases). This binary classification of mRS scores into ’good’ (0–2) and ’pool’ (3–6) outcomes serve to streamline our analysis and provides a clear framework for assessing patient recovery trajectories

Data split

In this study, the total dataset of 5429 cases was partitioned into training and test sets with a 9:1 ratio. This resulted in a test set consisting of 520 cases, with the remaining 4909 cases allocated to the training set. The training data was further subdivided into five folds for cross-validation purposes, each split at an 8:2 ratio. Consequently, each fold consisted of 3927 cases for training and 982 cases for validation. This strategic partitioning maintains the ratio of each group across the training, validation, and test datasets, ensuring that the characteristics of each group are consistently represented, thereby facilitating a robust validation of our model’s performance

Clinical features

In this study, we utilized various factors to predict the mRS outcome. These factors included gender, age, body mass index (BMI), NIHSS, Trial of Org 10172 in Acute Stroke Treatment (TOAST) classification14, presence of hypertension, diabetes, hyperlipidemia, smoking status, and atrial fibrillation. Additionally, the study considered the use of thrombolysis, as well as blood test results such as fasting glucose levels and blood pressure measurements at admission. By considering these variables, the researchers aimed to predict the mRS outcome, which is a measure of functional disability or dependence after a stroke.

Methodology

Augmentation and preprocessing

1) Image Preprocessing : DWI data were preprocessed using monai15 and nnUnet. The preprocessing includes three procedures. Firstly, the orientation of all DWIs was adjusted to left-posterior-superior (LPS). Next, non-brain background regions were removed. Finally, voxel intensities were normalized using z-score normalization.

2) Clinical Features Preprocessing : We received confirmation from a specialist regarding the outliers and missing values identified using the 3-sigma and IQR methods. Following the specialist’s recommendations, we corrected these outliers and missing values in our dataset. After making the necessary adjustments, the data was processed through an automated pipeline provided by AutoGlone. The normalization and further preprocessing of the data were managed using AutoGlone’s pipeline, which automatically identifies and appropriately preprocesses various data types, such as numerical and categorical.Especially in cases where NaN values exist, AutoGlone implements imputation techniques for numeric features, often using statistical methods like mean or median substitution depending on the distribution of the data. For categorical data, AutoGlone introduces an ’Unknown’ category for entries with missing values, effectively managing missingness and allowing the model to incorporate the absence of data as an informative feature.

Multi-task mRS estimation

Fig. 1 Multi-task mRS model architecture. Figure 1 illustrates a sophisticated neural network architecture designed to analyze Diffusion-Weighted Imaging (DWI) for predicting the modified Rankin Scale (mRS) scores. In this architecture, DWI inputs are transformed into various-sized image patches and processed through multiple convolutional (Conv2D) and pooling (MaxPool) layers, which are engineered to extract and refine features effectively. As data flows through diverse neural network blocks, it undergoes progressive feature extraction and transformation, culminating in the generation of a segmentation map and mRS pool probability..

Primary outcome was the poor functional outcome at 3 months after the index stroke. In other words, dichotomous mRS of 3–6 was the output variable in this algorithm. We employed the multi-task methodology to predict the mRS using DWI data along with clinical features. The multi-task methodology trains a single model to simultaneously perform multiple tasks. We simultaneously learned two tasks: lesion segmentation and mRS prediction. By training the model to perform both tasks together, the model can focus on identifying lesions during mRS prediction, leading to more accurate predictions.

The model we employed nnUNet, which is an open-source tool designed for the automatic segmentation of medical images. nnUNet utilizes a convolutional neural network (CNN) architecture that is tailored for 3D medical image segmentation. It simplifies the model training process by providing pre-trained models and an automated hyperparameter optimization framework. Within nnUNet, mRS prediction was performed using image feature maps obtained from the encoder component. Simultaneously, lesion segmentation was carried out using nnUNet’s decoder. This approach enabled efficient information extraction from DWI data while leveraging nnUNet’s segmentation capabilities. The model architecture is shown in Fig. 1.

We used the loss function that combines the segmentation loss with the classification loss for performing simultaneous working region segmentation and mRS classification. The segmentation loss is composed of two components: the dice loss and the pixel cross-entropy loss. The predicted pixel class is represented by ypred, and the ground truth pixel class is represented by ypred. The segmentation is defined as follows:1 Ldice=1-2∗∑ytrue∗ypred∑ytrue2+∑ypred2

2 Lpixelcross-entropy=-∑(ytrue∗logypred)

3 Lsegmentation=Ldice+Lpixelcross-entropy

The classification loss is cross-entropy loss. The predicted class for DWI is represented by Ypred, and the ground truth for DWI is represented by Ypred and the ground truth pixel class is represented by Ytrue.4 Lclassification=-∑Ytrue∗logYpred

The total loss function is defined as :5 LMulti-task=12(Lsegmentation+Lclassification)

Multi-modality classification

We employed a multi-modality methodology that leverages both DWI data and clinical information to enhance the accuracy of predicting mRS outcomes. DWI data, obtained through MRI scans, plays a crucial role in detecting and assessing brain tissue abnormalities. The prediction process is divided into two main steps. First, we utilized nnUnet, a deep learning algorithm specifically designed for medical image segmentation, to analyze the DWI data. The nnUnet effectively segmented the brain images and generated predictions for both the lesion area and mRS. This initial step allowed us to identify and quantify the extent of brain tissue damage associated with the stroke. To further refine our predictions and improve the accuracy of mRS outcomes, we integrated the predicted mRS values with relevant clinical information. Clinical data, including patient demographics, medical history, and other pertinent factors, were combined with the nnUnet predictions. By combining the predicted mRS values with the clinical information using autoGluon. The autoGluon is an automated machine learning (auto ml) tool that simplifies the process of building and deploying machine learning models. It automates the model selection process, automatically exploring a wide range of models, including both traditional machine learning algorithms and deep learning architectures, to identify the best-performing model for a given dataset. The autoGluon also automates feature engineering and performs hyperparameter tuning. We achieved a comprehensive prediction of the final mRS outcome. This approach enabled us to consider not only the imaging-based insights provided by DWI data and nnUNet segmentation but also the broader context of patient-specific characteristics captured by the clinical information. Figure 2 visually represents the entire prediction process, illustrating the seamless integration of DWI data, nnUnet segmentation, predicted mRS values, and clinical information. This comprehensive multi-modality approach enhances the accuracy and reliability of predicting mRS outcomes after stroke, thereby assisting clinicians in making more informed decisions regarding patient care and treatment strategies.Fig. 2 Multi-modality mRS prediction process. Figure 2 illustrates the subsequent multi-modality mRS prediction process, building on the initial mRS probability outputs derived from the architecture depicted in Fig.1, which incorporates both imaging and clinical data. This process is segmented into two stages: Stage 1 utilizes the nnU-Net architecture from Fig. 1 to predict the pool probability of mRS from a DWI. The encoder, vector, and decoder structure of the nnU-Net outputs a segmentation map and an mRS pool probability, which serve as the basis for further analysis. Stage 2 enhances the predictive accuracy by combining the mRS pool probability from Stage 1 with clinical information. This combined data is fed into the "Autoglon" framework, which employs a variety of machine learning models, including Random Forest, Neural Network, XG Boost, and Cat Boost. This framework selects the best model dynamically to accurately predict the final mRS..

Evaluation metrics

Consider the imbalance between good outcomes (mRS 0–2) and bad outcomes (mRS 3–6), We evaluated the models based on the F1 score, precision, recall, dice score, and area under the curve (AUC). Each metric is defined as follows,6 Dicescore=2·TruePositive2·TruePositive+FalsePositive+FalseNegative

7 Precision=TruePositiveTruePositive+FalsePositive

8 Recall=TruePositiveTruePositive+FalseNegative

9 F1Score=Recall·PrecisionRecall+Precision

in which the AUC is the area below the receiver operating characteristic (ROC) curve. In the graph of the ROC curve, the x-axis is the false positive rate (FPR), while the y-axis is the true positive rate (TPR). FPR and TPR are defined as follows:10 TPR=TruePositiveTruePositive+FalseNegativeFPR=FalsePositiveFalsePositive+TrueNegative

Results

Implementation details

We used pytorch16 version 1.10 and python version 3.8 to train our deep learning model on an NVIDIA RTX3090 GPU. The training data consisted of DICOM files, which we converted to nib format for compatibility with pytorch. To ensure the robustness of our model, we conducted five-fold cross-validation with a separate validation set for each fold. This allowed us to evaluate the performance of our model on different subsets of the data and reduce the risk of overfitting.

Performance

In this study, we employed five-fold cross-validation and test datasets to evaluate the model’s performance. We evaluated the model in two key areas: lesion segmentation and prediction of the mRS. For lesion segmentation, we utilized dice scores, while for the mRS prediction, we considered metrics such as AUC, F1 score, recall, and precision. The performance results for lesion segmentation, presented in Table 1, yielded a dice score of 0.7391 for validation and 0.7375 for the test dataset.Table 1 Segmentation performance.

	Five-fold validation	Test	
Dice Score	0.7369±0.0116	0.7375	

Table 2 compares the performance of the existing classification model with the model that incorporates our multi-task learning. By applying multi-task learning, we achieved a performance improvement of 0.06 compared to the model without multi-task. This improvement surpassed the performance of other classification models.

Table 3 focuses on performance comparison with the application of multi-modality data. We showed that performance improved when combining DWI with clinical information. Specifically, when utilizing all clinical information, including NIHSS and DWI, we achieved an AUC of 0.8065 on the test data. These results surpassed the existing clinical trials that relied solely on the thrive score17 and NIHSS score, with improvements of 0.16 and 0.06, respectively. Additionally, we plotted a ROC curve to visualize the performance comparison of the models, which can be seen in Fig. 3.Table 2 Comparison with other 3D classification models for DWI.

Model	Dataset	AUC	Accuracy	Precision	Recall	F1-score	
EfficientNet b2	5-fold validation	0.7041 ± 0.0268	0.6670 ± 0.0294	0.6538 ± 0.0220	0.6656 ± 0.0215	0.6529 ± 0.0256	
Test	0.7120	0.6615	0.6324	0.6527	0.6325	
Resnet50	5-fold validation	0.7082 ± 0.0219	0.6682 ± 0.0193	0.6498 ± 0.0181	0.6604 ± 0.0158	0.6489 ± 0.0184	
Test	0.7172	0.6634	0.6386	0.6613	0.6375	
nnUnet encoder	5-fold validation	0.6671 ± 0.0363	0.6521 ± 0.0339	0.6253 ± 0.0303	0.6315 ± 0.0314	0.6260 ± 0.0308	
Test	0.7143	0.6576	0.6224	0.6391	0.6238	
Vit	5-fold validation	0.5724 ± 0.0300	0.5575 ± 0.0700	0.5695 ± 0.0100	0.5679 ± 0.0200	0.5328 ± 0.0700	
Test	0.5722	0.5423	0.5445	0.5528	0.5247	
Ours	5-fold validation	0.7490±0.0269	0.6955±0.0240	0.6787±0.0260	0.6923±0.0260	0.6787±0.0250	
test	0.7681	0.7019	0.6712	0.6914	0.6745	

Table 3 Multi-modality classification performance.

Category	Modality	Dataset	AUC	Accuracy	Precision	Recall	F1-Score	
Assessment score	Thrive	5-fold validation	0.6531±0.0169	0.6741±0.0172	0.6318±0.0167	0.5831±0.0243	0.5755±0.0270	
test	0.6252	0.5307	0.5535	0.5627	0.5211	
NIHSS	5-fold validation	0.7219±0.0179	0.7179±0.0108	0.6889±0.0169	0.6746±0.0127	0.6776±0.0110	
test	0.7350	0.7250	0.6686	0.6530	0.6588	
Single modality	DWI	5-fold validation	0.7490±0.0269	0.6955±0.0240	0.6787±0.0260	0.6923±0.0260	0.6787±0.0250	
test	0.7681	0.7019	0.6712	0.6914	0.6745	
Clinical information	5-fold validation	0.6953±0.006	0.6468±0.0211	0.6376±0.0079	0.6503±0.074	0.6330±0.0159	
test	0.6932	0.6231	0.5982	0.6144	0.5945	
Multi-modality	DWI + Clinical information	5-fold validation	0.7613±0.0248	0.6841±0.0552	0.6892±0.0257	0.7019±0.0258	0.6739±0.0474	
test	0.7754	0.7115	0.6849	0.7139	0.6876	
DWI + NIHSS	5-fold validation	0.7739±0.0225	0.7122±0.0210	0.6951±0.0236	0.7116±0.0235	0.6970±0.0241	
test	0.7910	0.7058	0.6878	0.7206	0.6865	
DWI + NIHSS + Clinical information	5-fold validation	0.7847±0.0220	0.7103±0.0307	0.7028±0.0227	0.7195±0.0205	0.6982±0.0271	
test	0.8080	0.7346	0.7140	0.7503	0.7157	

Fig. 3 ROC Curve for performance of the mRS estimation.

Assessing model components via ablation studies

To rigorously evaluate the intrinsic capabilities of our model configurations, we conducted ablation studies focusing on different segments of the model’s architecture. These investigations are key to delineating how distinct model components–specifically, the U-net based DWI segmentation and the integration of clinical data–affect the prediction of modified Rankin Scale (mRS) outcomes. Initially, the model’s configuration using only U-net for DWI segmentation provided a fundamental assessment of lesion identification and initial prognosis estimation, achieving an AUC of 0.7681, as shown in Table  3. We then assessed the model’s performance with an exclusive focus on clinical data inputs. This configuration, which leverages clinical predictors alone, recorded an AUC of 0.6932, underscoring the value of clinical data in forecasting outcomes independently. Notably, the model’s predictive accuracy witnessed significant enhancement upon the integration of both DWI and clinical data streams. This holistic approach enables the model to leverage a comprehensive dataset, incorporating both patient-specific factors and lesion-specific details, which culminates in more precise and dependable predictions of stroke outcomes. The integration yields a superior AUC of 0.7754, illustrating the synergistic benefits of the combined modalities as detailed in the revised Table  3. This thorough evaluation confirms the robustness of our multi-modal model and underscores the crucial role of integrating diverse data sources to improve the accuracy and reliability of stroke prognosis predictions.

Pool probability correlation

Fig. 4 Correlation for pool probability of DWI.

In order to assess the influence of DWI, we conducted a correlation analysis between the predicted probability of DWI and relevant clinical information, including NIHSS and age, as well as the MRS score. as shown in Fig. 4. Our findings revealed a correlation between the predicted outcome of DWI and NIHSS with a coefficient of 0.51, age with a coefficient of 0.54, and the uncategorized MRS score with a coefficient of 0.47. These results show that the model can identify meaningful features by leveraging the information provided by DWI.

Performance differences based on magnetic field strength

We performed additional experiment to assess if the performance of deep leaning classifier could affect the magnetic field intensity (1.5T vs. 3.0T). For the DWI image data that entered the algorithm, there were slightly more MR images taken at 1.5 T (n= 3647) compared to those with 3.0 T (n= 1367), but we found no statistically significant difference in performance between the two group. We performed a DeLong test for statistically significant difference between the two groups. The p-value for the two groups was 0.2533, and their respective AUC are shown in Table 4Table 4 Results of the performance comparison for magnetic field intensity.

Magnetic field intensity	AUC	p-value	
1.5	0.8153	0.2533	
3.0	0.7832	

Qualitative analysis via class activation and segmentation

Fig. 5 Validation of model’s accurate lesion area perception: (a) Original image (b) Grad-class activation map (CAM) (c) Segmentation results (d) Ground truth.

We compared the Segmentation results, grad-CAM18, and ground truth. As a result, it was confirmed that the model focused on the lesion when predicting mRS. This analysis affirmed the model’s capability to differentiate between lesion and non-lesion regions and accurately capture the crucial aspects of the lesions. In Fig. 5, we show a gradcam of each axis and the segmentation result.

Discussion

In this study, we have proposed a stroke prognostic model using deep learning techniques that incorporate both clinical variables and DWI, which can be obtained during the early stages of hospitalization. The model leverages multi-task methodology to segment lesions in the initial DWI scans and predict patient prognosis. Additionally, it employs multi-modal methodology to integrate clinical variables. Our methodology’s outcomes have showcased a significantly superior performance in prognostic prediction compared to existing deep learning classification models and clinical index.

The cornerstone of our study lies in the use of artificial intelligence for the analysis of diffusion-weighted images related to stroke volumes, locations, and patterns. Previous research established that diffusion-weighted imaging lesion volumes (P=0.047) acted as independent indicators of a positive outcome. Specifically, a DWI lesion volume threshold of less than 16 mL was most efficient in identifying patients who were likely to experience a favorable outcome (mRS between 0 and 1)19. Diffusion location was identified as a significant factor affecting outcomes in prior studies. Voxel-based lesion symptom mapping served to highlight regions that exerted a substantial impact on increased levels on the modified Rankin Scale. These regions accumulated in the corona radiata, internal capsule, and the insular zone. Asymmetrically dispersed impact patterns were also noticed, involving the right inferior temporal gyrus and the left superior temporal gyrus20. Initial DWI lesion patterns can potentially act as predictors for infarct growth or early recurrence. Most patients who exhibited infarct growth showed a territory/lobar or internal border zone-type DWI lesion pattern. Conversely, patients with newly formed lesions displayed small cortical or cortical border zone infarct patterns21.

Additionally, several machine learning studies22–24 have endeavored to predict mRS scores utilizing either clinical or imaging data.Ramos et al. demonstrated that imaging markers, such as hypodensity, hemorrhagic transformation, and leukoaraiosis, enhance the performance of random forests, achieving an AUROC of 0.81 in 1,526 patients undergoing endovascular treatment. Similarly, a convolutional neural network (CNN) employing only diffusion-weighted images measured at admission effectively predicted 90-day mRS in patients receiving endovascular treatment for acute ischemic stroke (AIS). These findings underscore the value of both clinical and imaging data in forecasting patient prognosis post-AIS.Furthermore, Hatami et al.’s research illustrated that integrating clinical variables as weights into an ensemble voting algorithm for image inputs could refine model performance. Consistent with these studies, our results affirm that combining clinical and imaging data within a multi-task framework not only corroborates existing methodologies but also yields incremental enhancements in predictive performance. This alignment with prior research bolsters our multi-task approach, highlighting its potential to improve prognostic predictions in clinical practice.

The findings of this analysis were enlightening. The resultant pool probabilities, generated by the model’s predictions, exhibited substantial correlations with both age and NIHSS scores. These factors have long been established as critical determinants in the field of clinical prognosis prediction. Notably, the correlation coefficient of 0.47 observed with the mRS scores the model’s ability in capturing the relationship between stroke severity and patient functional outcomes. These findings unequivocally validate our model’s effectiveness in discerning and incorporating pivotal features crucial for accurate stroke prognosis prediction. Previous studies predicting stroke outcome have utilized ASTRAL or THRIVE scores for stroke prognosis stratification25,26, of which these scoring systems only include clinical variables of stroke patients. Recently, it has been reported that radiological surrogate markers in addition to these clinical variables may improve the prediction of stroke prognosis27,28, However, these radiologic markers are not readily available for timely decision-making in stroke patients and are obtained after scanning several types of brain images. We used only DWI, not multiple imaging modalities, and tried to predict stroke prognosis using segmented images itself rather than surrogate markers or radiomics parameters. In the future, deep learning-based clinical decision support using images themselves may be more actively utilized than those using radiologic surrogate markers.

As shown in the results in Table 2, nnUnet-based prediction using DWI imaging alone was found to be superior to other tasks using NIHSS or clinical variables in predicting stroke outcome. The most important variables in predicting stroke outcomes are age and NIHSS score29,30, In this study, we showed that nnUnet, which is mainly used for segmentation in the medical imaging field, can also be efficiently used for classification tasks represented by dichotomized mRS score. The nnUnet is an upgraded algorithm for medical image segmentation tasks, and it can provide various useful strategies for preprocessing and data augmentation, but most important is that this is a CNN-based algorithm9,31. The CNN algorithm easily captures local patterns, edges, and texture of image using convolutional layers, and pooling layers enables the algorithm to understand the spatial arrangement or complex patterns32,33. In other words, the nnUnet algorithm’s ability to predict classification outcome as well as segmentation is an indirect evidence that the CNN algorithm can efficiently focus on local lesions in the brain rather than evaluating the entire brain as a whole.

Furthermore, we expanded our analysis by incorporating gradcam and lesion segmentation maps to visualize the prominent brain regions significantly influencing stroke prognosis. These visualizations granted us a more comprehensive understanding of the model’s functioning. Importantly, the outcomes reaffirmed the model’s heavy reliance on lesion-related information when generating prognosis predictions. This dual visualization strategy not only bolsters the model’s reliability but also augments its interpretability.

Our study showed that combining early diffusion MRI and clinical information can predict patient outcomes well. There have been several studies that predicted prognosis based on clinical information. The THRIVE score were well known to stroke outcome prediction. This score is consisted of trichotomized age (≤59, 60–79, and ≥80 years), trichotomized NIHSS (≤10, 11–20, and ≥21), and the chronic disease scale (hypertension, diabetes, and atrial fibrillation). Our study showed more accurate outcome estimation compared with the THRIVE score.

Despite the favorable prognostic effect, our study has several limitations. First, to generalize our findings, a larger and more prospective study cohort is essential. Additionally, external validation on independent datasets is crucial to confirm the model’s performance across various patient demographics and clinical settings. This validation is essential not only for verifying the accuracy of our findings but also for assessing the model’s applicability in real-world scenarios. Second, perfusion and vascular imaging may improve the prognosis of acute ischemic stroke and identify patients with therapeutic targets well beyond the traditional time window for intravenous thrombolysis or embolization. Although initial DWI and NIHSS may reflect perfusion and vascular status to some extent, we believe that the addition of perfusion and vascular imaging analysis may provide a better predictive model. Third, we did not include the revascularization status in our model. We assumed that this clinical information may improve the performance of our algorithm. Fourth, although this study identifies a correlation between GradCAM and variables for model interpretation, it is insufficient to fully elucidate the model’s explanatory power. Therefore, future research should explore new methodologies for model explanation and utilize XAI techniques such as SHAP values to enhance the model’s reliability. Fifth, when modeling networks, we need to address class imbalance to avoid bias towards majority classes (good outcomes) and to ensure better generalization and performance, especially for minority classes (bad outcomes). Strategies such as resampling, class weighting, and appropriate evaluation metrics can effectively mitigate these issues34. Lastly, we did not assess the status of endovascular therapy, which could affect the patient’s outcome. However, patients undergoing EVT represented about 10% of all patients and are not expected to had a significant impact on the performance of this model.

Conclusion

We developed a prognostic prediction model based on deep learning that utilizes multi modal and multi-task. Our model outperformed other clinical indicators such as the THRIVE score and the NIHSS score and existing classification models, which were previously utilized for prognosis prediction. These findings suggest that prognosis prediction using DWI and clinical information can significantly contribute to the development of effective treatment and rehabilitation strategies for patients.

Author contributions

Conceptualization, S.K., C.K., and H.-S.C.; data curation, S.K. J.-W.J., and C.K.; investigation, S.-W.P., N.-Y.Y., S.Y.S., I.J and S.-H.S.; supervision, S.K.,J.-W.J., Y.K.,H.-S.C. and C.K.; validation, I.-S.P.; visualization, I.-S.P.; writing–original draft, I.-S.P. and S.K.; writing–review and editing, I.-S.P., S.K, H.-S.C. and C.K. All authors have read and agreed to the published version of the manuscript.

Funding

This study was supported by 2019 Research Grant from Kangwon National University, supported by a grant of the Korea Health Technology R &D Project through the Korea Health Industry Development Institute (KHIDI), funded by the Ministry of Health & Welfare, Republic of Korea (grant number : HR21C0198), supported by the National Research Foundation of Korea (NRF) grants funded by the Korea government (Ministry of Science and ICT) (2022R1A5A708390811, 2022R1F1A1076454), supported by “Regional Innovation Strategy (RIS)” through the National Research Foundation of Korea (NRF) funded by the Ministry of Education (MOE) (2023RIS-005), and supported by Seoul National University of Science and Technology.

Data availability

Data are available upon request by the corresponding author along with improvement of the data review board.

Code availability

Code is available upon request by the corresponding author.

Competing interests

The authors declare no competing interests.

Publisher's note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

These authors contributed equally: In-Seo Park and Seongheon Kim.
==== Refs
References

1. Powers WJ 2018 guidelines for the early management of patients with acute ischemic stroke: A guideline for healthcare professionals from the American heart association/American stroke association Stroke 2018 49 46 99 10.1161/STR.0000000000000158 29203686
Powers, W. J. et al. 2018 guidelines for the early management of patients with acute ischemic stroke: A guideline for healthcare professionals from the American heart association/American stroke association. Stroke 49, 46–99 (2018).29203686 10.1161/STR.0000000000000158
2. Albers GW Thrombectomy for stroke at 6 to 16 hours with selection by perfusion imaging N. Engl. J. Med. 2018 378 708 718 10.1056/NEJMoa1713973 29364767
Albers, G. W. et al. Thrombectomy for stroke at 6 to 16 hours with selection by perfusion imaging. N. Engl. J. Med. 378, 708–718 (2018).29364767 10.1056/NEJMoa1713973
3. Nogueira RG Thrombectomy 6 to 24 hours after stroke with a mismatch between deficit and infarct N. Engl. J. Med. 2018 378 11 21 10.1056/NEJMoa1706442 29129157
Nogueira, R. G. et al. Thrombectomy 6 to 24 hours after stroke with a mismatch between deficit and infarct. N. Engl. J. Med. 378, 11–21 (2018).29129157 10.1056/NEJMoa1706442
4. Boehme AK Esenwa C Elkind MS Stroke risk factors, genetics, and prevention Circ. Res. 2017 120 472 495 10.1161/CIRCRESAHA.116.308398 28154098
Boehme, A. K., Esenwa, C. & Elkind, M. S. Stroke risk factors, genetics, and prevention. Circ. Res. 120, 472–495 (2017).28154098 10.1161/CIRCRESAHA.116.308398
5. Jeong S Deep learning approach using diffusion-weighted imaging to estimate the severity of aphasia in stroke patients J. Stroke 2022 24 108 117 10.5853/jos.2021.02061 35135064
Jeong, S. et al. Deep learning approach using diffusion-weighted imaging to estimate the severity of aphasia in stroke patients. J. Stroke 24, 108–117 (2022).35135064 10.5853/jos.2021.02061
6. Zekavat SM Deep learning of the retina enables phenome-and genome-wide analyses of the microvasculature Circulation 2022 145 134 150 10.1161/CIRCULATIONAHA.121.057709 34743558
Zekavat, S. M. et al. Deep learning of the retina enables phenome-and genome-wide analyses of the microvasculature. Circulation 145, 134–150 (2022).34743558 10.1161/CIRCULATIONAHA.121.057709
7. Kuntz S Gastrointestinal cancer classification and prognostication from histology using deep learning: Systematic review Eur. J. Cancer 2021 155 200 215 10.1016/j.ejca.2021.07.012 34391053
Kuntz, S. et al. Gastrointestinal cancer classification and prognostication from histology using deep learning: Systematic review. Eur. J. Cancer 155, 200–215 (2021).34391053 10.1016/j.ejca.2021.07.012
8. Zhang R The diagnostic and prognostic value of radiomics and deep learning technologies for patients with solid pulmonary nodules in chest ct images BMC Cancer 2022 22 1118 10.1186/s12885-022-10224-z 36319968
Zhang, R. et al. The diagnostic and prognostic value of radiomics and deep learning technologies for patients with solid pulmonary nodules in chest ct images. BMC Cancer 22, 1118 (2022).36319968 10.1186/s12885-022-10224-z
9. Isensee F Jaeger PF Kohl SA Petersen J Maier-Hein KH nnu-net: A self-configuring method for deep learning-based biomedical image segmentation Nat. Methods 2021 18 203 211 10.1038/s41592-020-01008-z 33288961
Isensee, F., Jaeger, P. F., Kohl, S. A., Petersen, J. & Maier-Hein, K. H. nnu-net: A self-configuring method for deep learning-based biomedical image segmentation. Nat. Methods 18, 203–211 (2021).33288961 10.1038/s41592-020-01008-z
10. Erickson, N. et al. Autogluon-tabular: Robust and accurate automl for structured data. arXiv preprint arXiv:2003.06505 (2020).
11. Banks JL Marotta CA Outcomes validity and reliability of the modified rankin scale: Implications for stroke clinical trials: A literature review and synthesis Stroke 2007 38 1091 1096 10.1161/01.STR.0000258355.23810.c6 17272767
Banks, J. L. & Marotta, C. A. Outcomes validity and reliability of the modified rankin scale: Implications for stroke clinical trials: A literature review and synthesis. Stroke 38, 1091–1096 (2007).17272767 10.1161/01.STR.0000258355.23810.c6
12. Ai hub. available online: https://www.aihub.or.kr. Accessed 1 on May 2023.
13. Itk snap. available online: http://www.itksnap.org. Accessed on 30 December 2020.
14. Adams HP Jr Classification of subtype of acute ischemic stroke definitions for use in a multicenter clinical trial toast trial of org 10172 in acute stroke treatment Stroke 1993 24 35 41 10.1161/01.STR.24.1.35 7678184
Adams, H. P. Jr. et al. Classification of subtype of acute ischemic stroke definitions for use in a multicenter clinical trial toast trial of org 10172 in acute stroke treatment. Stroke 24, 35–41 (1993).7678184 10.1161/01.STR.24.1.35
15. Cardoso, M. J. et al. Monai: An open-source framework for deep learning in healthcare. arXiv preprint arXiv:2211.02701 (2022).
16. Paszke A Pytorch: An imperative style, high-performance deep learning library Adv. Neural Inf. Process. Syst. 2019 32 8024 8035
Paszke, A. et al. Pytorch: An imperative style, high-performance deep learning library. Adv. Neural Inf. Process. Syst. 32, 8024–8035 (2019).
17. Flint AC Thrive score predicts ischemic stroke outcomes and thrombolytic hemorrhage risk in vista Stroke 2013 44 3365 3369 10.1161/STROKEAHA.113.002794 24072004
Flint, A. C. et al. Thrive score predicts ischemic stroke outcomes and thrombolytic hemorrhage risk in vista. Stroke 44, 3365–3369 (2013).24072004 10.1161/STROKEAHA.113.002794
18. Selvaraju, R. R. et al. Grad-cam: Visual explanations from deep networks via gradient-based localization. In Proceedings of the IEEE international conference on computer vision, 618–626 (2017).
19. Kruetzelmann A Pretreatment diffusion-weighted imaging lesion volume predicts favorable outcome after intravenous thrombolysis with tissue-type plasminogen activator in acute ischemic stroke Stroke 2011 42 1251 1254 10.1161/STROKEAHA.110.600148 21415399
Kruetzelmann, A. et al. Pretreatment diffusion-weighted imaging lesion volume predicts favorable outcome after intravenous thrombolysis with tissue-type plasminogen activator in acute ischemic stroke. Stroke 42, 1251–1254 (2011).21415399 10.1161/STROKEAHA.110.600148
20. Cheng B Influence of stroke infarct location on functional outcome measured by the modified rankin scale Stroke 2014 45 1695 1702 10.1161/STROKEAHA.114.005152 24781084
Cheng, B. et al. Influence of stroke infarct location on functional outcome measured by the modified rankin scale. Stroke 45, 1695–1702 (2014).24781084 10.1161/STROKEAHA.114.005152
21. Bang OY Li W Applications of diffusion-weighted imaging in diagnosis, evaluation, and treatment of acute ischemic stroke Precision Future Med. 2019 3 69 76 10.23838/pfm.2019.00037
Bang, O. Y. & Li, W. Applications of diffusion-weighted imaging in diagnosis, evaluation, and treatment of acute ischemic stroke. Precision Future Med. 3, 69–76 (2019).10.23838/pfm.2019.00037
22. Ramos LA Predicting poor outcome before endovascular treatment in patients with acute ischemic stroke Front. Neurol. 2020 11 580957 10.3389/fneur.2020.580957 33178123
Ramos, L. A. et al. Predicting poor outcome before endovascular treatment in patients with acute ischemic stroke. Front. Neurol. 11, 580957 (2020).33178123 10.3389/fneur.2020.580957
23. Hatami, N. et al. CNN-LSTM based multimodal MRI and clinical data fusion for predicting functional outcome in stroke patients. In 2022 44th Annual International Conference of the IEEE Engineering in Medicine & Biology Society (EMBC), 3430–3434 (IEEE, 2022).
24. Nishi H Deep learning-derived high-level neuroimaging features predict clinical outcomes for large vessel occlusion Stroke 2020 51 1484 1492 10.1161/STROKEAHA.119.028101 32248769
Nishi, H. et al. Deep learning-derived high-level neuroimaging features predict clinical outcomes for large vessel occlusion. Stroke 51, 1484–1492 (2020).32248769 10.1161/STROKEAHA.119.028101
25. Ntaios G An integer-based score to predict functional outcome in acute ischemic stroke: The astral score Neurology 2012 78 1916 1922 10.1212/WNL.0b013e318259e221 22649218
Ntaios, G. et al. An integer-based score to predict functional outcome in acute ischemic stroke: The astral score. Neurology 78, 1916–1922 (2012).22649218 10.1212/WNL.0b013e318259e221
26. Flint AC Validation of the totaled health risks in vascular events (thrive) score for outcome prediction in endovascular stroke treatment Int. J. Stroke 2014 9 32 39 10.1111/j.1747-4949.2012.00872.x 22928705
Flint, A. C. et al. Validation of the totaled health risks in vascular events (thrive) score for outcome prediction in endovascular stroke treatment. Int. J. Stroke 9, 32–39 (2014).22928705 10.1111/j.1747-4949.2012.00872.x
27. Xie Y Use of gradient boosting machine learning to predict patient outcome in acute ischemic stroke on the basis of imaging, demographic, and clinical information Am. J. Roentgenol. 2019 212 44 51 10.2214/AJR.18.20260 30354266
Xie, Y. et al. Use of gradient boosting machine learning to predict patient outcome in acute ischemic stroke on the basis of imaging, demographic, and clinical information. Am. J. Roentgenol. 212, 44–51 (2019).30354266 10.2214/AJR.18.20260
28. Kim JK Choo YJ Shin H Choi GS Chang MC Prediction of ambulatory outcome in patients with corona radiata infarction using deep learning Sci. Rep. 2021 11 7989 10.1038/s41598-021-87176-0 33846472
Kim, J. K., Choo, Y. J., Shin, H., Choi, G. S. & Chang, M. C. Prediction of ambulatory outcome in patients with corona radiata infarction using deep learning. Sci. Rep. 11, 7989 (2021).33846472 10.1038/s41598-021-87176-0
29. Saposnik G Guzik AK Reeves M Ovbiagele B Johnston SC Stroke prognostication using age and nih stroke scale: Span-100 Neurology 2013 80 21 28 10.1212/WNL.0b013e31827b1ace 23175723
Saposnik, G., Guzik, A. K., Reeves, M., Ovbiagele, B. & Johnston, S. C. Stroke prognostication using age and nih stroke scale: Span-100. Neurology 80, 21–28 (2013).23175723 10.1212/WNL.0b013e31827b1ace
30. Weimar C Ziegler A König IR Diener H-C Collaborators GSS Predicting functional outcome and survival after acute ischemic stroke J. Neurol. 2002 249 888 895 10.1007/s00415-002-0755-8 12140674
Weimar, C., Ziegler, A., König, I. R., Diener, H.-C. & Collaborators, G. S. S. Predicting functional outcome and survival after acute ischemic stroke. J. Neurol. 249, 888–895 (2002).12140674 10.1007/s00415-002-0755-8
31. Chen X A deep learning-based auto-segmentation system for organs-at-risk on whole-body computed tomography images for radiation therapy Radiother. Oncol. 2021 160 175 184 10.1016/j.radonc.2021.04.019 33961914
Chen, X. et al. A deep learning-based auto-segmentation system for organs-at-risk on whole-body computed tomography images for radiation therapy. Radiother. Oncol. 160, 175–184 (2021).33961914 10.1016/j.radonc.2021.04.019
32. Yang A Yang X Wu W Liu H Zhuansun Y Research on feature extraction of tumor image based on convolutional neural network IEEE Access 2019 7 24204 24213 10.1109/ACCESS.2019.2897131
Yang, A., Yang, X., Wu, W., Liu, H. & Zhuansun, Y. Research on feature extraction of tumor image based on convolutional neural network. IEEE Access 7, 24204–24213 (2019).10.1109/ACCESS.2019.2897131
33. Yuan F Zhang Z Fang Z An effective CNN and transformer complementary network for medical image segmentation Pattern Recogn. 2023 136 109228 10.1016/j.patcog.2022.109228
Yuan, F., Zhang, Z. & Fang, Z. An effective CNN and transformer complementary network for medical image segmentation. Pattern Recogn. 136, 109228 (2023).10.1016/j.patcog.2022.109228
34. Lundberg, S. M. & Lee, S.-I. A unified approach to interpreting model predictions. Adv. Neural Inf. Process. Syst. 30 (2017).
