
==== Front
Sci Rep
Sci Rep
Scientific Reports
2045-2322
Nature Publishing Group UK London

39154039
69870
10.1038/s41598-024-69870-x
Article
Beyond hand-crafted features for pretherapeutic molecular status identification of pediatric low-grade gliomas
Kudus Kareem 12
Wagner Matthias W. 34
Namdar Khashayar 12
Bennett Julie 567
Nobre Liana 89
Tabori Uri 5
Hawkins Cynthia 10
Ertl-Wagner Birgit Betina 12311
Khalvati Farzad farzad.khalvati@utoronto.ca

123111213
1 https://ror.org/057q4rt57 grid.42327.30 0000 0004 0473 9646 Neurosciences & Mental Health Research Program, The Hospital for Sick Children, Toronto, Canada
2 https://ror.org/03dbr7087 grid.17063.33 0000 0001 2157 2938 Institute of Medical Science, University of Toronto, Toronto, Canada
3 https://ror.org/057q4rt57 grid.42327.30 0000 0004 0473 9646 Department of Diagnostic & Interventional Radiology, The Hospital for Sick Children, Toronto, Canada
4 https://ror.org/03b0k9c14 grid.419801.5 0000 0000 9312 0220 Department of Diagnostic and Interventional Neuroradiology, University Hospital Augsburg, Augsburg, Germany
5 https://ror.org/057q4rt57 grid.42327.30 0000 0004 0473 9646 Division of Hematology and Oncology, The Hospital for Sick Children, Toronto, Canada
6 https://ror.org/03zayce58 grid.415224.4 0000 0001 2150 066X Division of Medical Oncology and Hematology, Princess Margaret Cancer Centre, Toronto, Canada
7 https://ror.org/03dbr7087 grid.17063.33 0000 0001 2157 2938 Department of Pediatrics, University of Toronto, Toronto, Canada
8 https://ror.org/0160cpw27 grid.17089.37 Department of Paediatrics, University of Alberta, Edmonton, Canada
9 https://ror.org/05w90nk74 grid.416656.6 0000 0004 0633 3703 Division of Immunology, Hematology/Oncology and Palliative Care, Stollery Children’s Hospital, Edmonton, Canada
10 https://ror.org/057q4rt57 grid.42327.30 0000 0004 0473 9646 Paediatric Laboratory Medicine, Division of Pathology, The Hospital for Sick Children, Toronto, Canada
11 https://ror.org/03dbr7087 grid.17063.33 0000 0001 2157 2938 Department of Medical Imaging, University of Toronto, Toronto, Canada
12 https://ror.org/03dbr7087 grid.17063.33 0000 0001 2157 2938 Department of Computer Science, University of Toronto, Toronto, Canada
13 https://ror.org/03dbr7087 grid.17063.33 0000 0001 2157 2938 Department of Mechanical and Industrial Engineering, University of Toronto, Toronto, Canada
17 8 2024
17 8 2024
2024
14 191028 2 2024
9 8 2024
© The Author(s) 2024
2024
https://creativecommons.org/licenses/by/4.0/ Open Access This article is licensed under a Creative Commons Attribution 4.0 International License, which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made. The images or other third party material in this article are included in the article's Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article's Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by/4.0/.
The use of targeted agents in the treatment of pediatric low-grade gliomas (pLGGs) relies on the determination of molecular status. It has been shown that genetic alterations in pLGG can be identified non-invasively using MRI-based radiomic features or convolutional neural networks (CNNs). We aimed to build and assess a combined radiomics and CNN non-invasive pLGG molecular status identification model. This retrospective study used the tumor regions, manually segmented from T2-FLAIR MR images, of 336 patients treated for pLGG between 1999 and 2018. We designed a CNN and Random Forest radiomics model, along with a model relying on a combination of CNN and radiomic features, to predict the genetic status of pLGG. Additionally, we investigated whether CNNs could predict radiomic feature values from MR images. The combined model (mean AUC: 0.824) outperformed the radiomics model (0.802) and CNN (0.764). The differences in model performance were statistically significant (p-values < 0.05). The CNN was able to learn predictive radiomic features such as surface-to-volume ratio (average correlation: 0.864), and difference matrix dependence non-uniformity normalized (0.924) well but was unable to learn others such as run-length matrix variance (− 0.017) and non-uniformity normalized (− 0.042). Our results show that a model relying on both CNN and radiomic-based features performs better than either approach separately in differentiating the genetic status of pLGGs, and that CNNs are unable to express all handcrafted features.

Subject terms

Paediatric research
Computer science
Cancer imaging
http://dx.doi.org/10.13039/501100000024 Canadian Institutes of Health Research 184015 184015 issue-copyright-statement© Springer Nature Limited 2024
==== Body
pmcIntroduction

Pediatric low-grade gliomas (pLGG) are a heterogeneous set of tumors with various histological features that can occur anywhere in the central nervous system1, and are sometimes referred to as low-grade neuroepithelial tumors. pLGGs account for approximately 40% of central nervous system tumors in children and are the most prevalent brain tumor during childhood2. With 10-year overall survival exceeding 90%3 death from pLGG is relatively rare. Nonetheless, with a progression-free survival rate of 50%, the morbidity of pLGG is high with many patients needing adjuvant therapy3,4.

Recent findings showed that pLGG typically upregulates the mitogen-activated protein kinase pathway, which led to the implementation of targeted therapeutics that can supplement or replace classic cytotoxic treatments3. The use of these targeted therapeutics requires accurate detection of the genetic alteration underlying pLGG, which is typically determined through tumor tissue acquired by surgery5. In some cases, such as in midline pLGG, it may not be feasible to obtain a biopsy due to the risk of neurologic compromise related to the surgery. Brain biopsies are costly, have numerous risks, leave residual tumor, and occasionally fail due to an insufficient sample.

Radiomics is the process of extracting quantitative features, which can be used to predict characteristics of patients, from radiological images6. Over the last decade, radiomics has emerged as a method of decoding tumor phenotypes non-invasively7–9. Radiogenomics, a research area within radiomics that combines “radiology” and “genomics”10, aims to identify relationships between imaging features and genomic data11. Numerous studies have shown the utility of radiomics and radiogenomics in neuro-oncology, including for predicting recurrence, survival, and genetic status12.

The two most common genetic alterations in pLGG are KIAA1549-BRAF fusion (BRAF Fusion) and BRAF p.V600E (BRAF Mutation)13. Wagner et al. showed that it is possible to differentiate between these two alterations using a machine learning (ML) approach relying on radiomic features extracted from T2-weighted fluid-attenuated inversion recovery (FLAIR) MRI sequences14,15. These results demonstrated the feasibility of a non-invasive pLGG genetic alteration identification model. Since then, multiple studies16–18 have shown that radiomics can be used to determine the BRAF status of all pLGGs, rather than just differentiating between the two most common genetic alterations. Tak et al.19, showed that using Convolutional neural networks (CNNs), an alternative approach to radiomics for extracting information from radiological images, it is possible to accurately sort pLGGs into three groups: BRAF Fusion, BRAF Mutation, and non-BRAF altered. Here we evaluated a more advanced approach, relying on both CNNs and radiomics, on the same classification task.

Radiomics relies on handcrafted features, while CNNs are trained to extract differentiative features (Fig. 1)20. The human design of radiomic features limits the amount of useful information that the radiomics approach can extract from medical images21. CNNs do not have this limitation, giving them greater expressive power22, thus they are often thought to be superior to radiomics21. However, CNNs have limitations of their own, namely, they need large datasets to learn from20,23,24. Nevertheless, CNNs have exploded in popularity25, and have become the dominant method to tackle a variety of medical imaging tasks26. The prominence of CNNs for medical imaging tasks in the literature suggests that it is commonly believed that radiomic features are redundant to those discovered by CNNs. According to Orlhac et al., this belief stems from the fact that theoretically, handcrafted features represent only a small subset of the features a neural network can capture22. However, in practice, neural networks may have trouble learning to represent certain handcrafted features due to restrictions such as limited data22.Figure 1 Visualizing the difference between the radiomics and CNN-based approaches for tumor classification. Both take the manually segmented region of interest (ROI) of a pre-processed FLAIR image as inputs. Here the entire brain of a single 2D slice from our dataset is shown, but in practice, a 3D input consisting of just the ROI, which is bounded by the yellow line in the figure, with the rest of the brain zeroed out, was used. The top of the figure shows the radiomics path, which involves extracting features from the image and then using those features to train an ML model to classify the images. The CNN path at the bottom is more direct; the model discovers useful predictive features directly from the ROI.

Zhang et al. showed that radiomic and CNN features extracted from CT images of pancreatic ductal adenocarcinoma are only weakly related, suggesting a complementary relationship27. Other studies have performed more explicit tests and found that CNNs and radiomics work better together for various tasks including the classification of breast tumors28,29, ground glass nodules30, central nervous system tumors31, and adult gliomas32. We hypothesized that a combined approach would result in a more accurate pLGG genetic status classification model, thus we aimed to assess the ability of CNNs to complement handcrafted radiomic features. Additionally, we aimed to evaluate whether CNNs can capture all of the information contained in handcrafted features.

Materials and methods

Patients and data

All methods of this retrospective study were performed in accordance with the guidelines and regulations of the research ethics board of The Hospital for Sick Children (SickKids) (Toronto, Canada), which approved the study and waived informed consent. The electronic health record database at SickKids was screened for patients treated for pLGG between 1999 and 2018 with pre-therapeutic MR brain imaging and molecular characterization available. All MRIs used in this study were acquired prior to any treatment or intervention. We used FLAIR as our primary imaging sequence because it is useful for assessing both tumor volume and cysts, and was available for the most patients. Furthermore, FLAIR depicts the tumor and surrounding area better than contrast-enhanced T1-weighted images4, is known to be sensitive to leptomeningeal spread and non-contrast enhancing lesions5, and is considered to be the most sensitive technique to detect brain tumors overall6. Thus, of the 397 patients found through initial screening, 45 were omitted due to the absence of a FLAIR image, while 16 others were omitted because the FLAIR image was motion-degraded, leaving 336 patients to be included in this study. Many of these patients also had other sequences available: T2-weighted (265), T1-weighted with (270) and without (260) contrast. All available MRI sequences as well as patient molecular status, age, and sex were manually extracted from the electronic health record database. The demographics of the cohort are summarized in Table 1. Some of these patients (11514, 21515, 22033) were used in previous studies which only accounted for the two most common genetic alterations, resulting in less realistic and clinically useful models compared to this current study. Table 1 Demographics of the patient population.

Total number of patients	336	
Patients with BRAF fusion	142 (42%)	
Patients with BRAF mutation	70 (21%)	
Patients confirmed to not have BRAF Fusion or Mutation (non-BRAF altered)	124 (37%)	
   NF1	   7 (2%)	
   FGFR 1	   13 (4%)	
   FGFR 2	   5 (1%)	
   Other	   45 (14%)	
   Unidentified	   54 (16%)	
Histopathology	
 Pilocytic astrocytoma	142 (42%)	
 Ganglioglioma	66 (20%)	
 Low-grade astrocytoma	40 (12%)	
 Dysembryblastic neuroepithelial tumor	25 (7%)	
 Diffuse astrocytoma	21 (6%)	
 Other	42 (13%)	
 Median patient age (Interquartile range)	8.6 (4.8, 12.8)	
 Female	154	
 Male	182	

Molecular subtyping

A stepwise approach was used for the molecular characterization of pLGG, as previously described in14. First immunohistochemistry was used to detect BRAF mutation, then either a nCounter Metabolic Pathways Panel (NanoString Technologies) or fluorescence in situ hybridization was used to identify BRAF Fusion. Samples that were negative for BRAF Mutation and Fusion were analyzed further using other sequencing strategies detailed in13, such as RNA sequencing or panel DNA sequencing. For most patients, formalin-fixed paraffin-embedded tissue obtained during biopsy or resection was used for the molecular analysis, otherwise, frozen tissue was used.

MRI acquisition, segmentation, and preprocessing

Patients underwent brain MRI at 1.5 T or 3 T using either the Achieva, (Philips Healthcare), or Magnetom Skyra (Siemens Healthineers) system. MRI data were deidentified after being extracted from the PACS at SickKids. Each patient had a 2D FLAIR sequence, either axial or coronal.

As described in14, tumor regions were identified through segmentation of the FLAIR images by a pediatric neuroradiologist (MWW). Segmentations were validated by a senior neuroradiologist with more than 20 years of experience (BEW). The level tracing-effect tool of 3D Slicer34,35 (Version 4.10.2) was used to perform semi automated tumor segmentation. The tumor region alone was used as input to our machine-learning models. Images were z-score normalized, bias-corrected, and isotropically resampled to a resolution of 240 × 240 × 155 using 3D Slicer, to help account for differences in slice thickness, field strength, and pixel spacing.

Radiomic feature extraction

The same pipeline as in14 was used to generate radiomic features from the MRIs. The SlicerRadiomics extension of 3D Slicer was used to access PyRadiomics, an open-sourced package for radiomic feature extraction36. Default PyRadiomics settings were used, for example, bin width was set to 25. For each patient, 851 radiomic features, including first-order, shape, and second-order (texture) features, were extracted. The full list of radiomic features can be accessed in the Online Supplemental Data of14.

ML models

Each of our models was trained on a multiclass classification task with three class labels: BRAF Fusion, BRAF Mutation, and non-BRAF altered, a heterogeneous class containing numerous genetic alterations. The radiomics approach relied on a Random Forest (RF) model. We experimented with two different CNN architectures, a well-established deep neural network, the 3D ResNet37, and a custom shallow 3D CNN (Fig. 2) with three convolutional layers and two hidden fully connected layers. For the shallow CNN, the Leaky ReLU activation function and batch normalization were used between each layer, while max-pooling was employed after each convolutional layer. For the combined method we implemented feature-level fusion23 similar to the techniques implemented in27 and32, where CNN features are extracted from the fully-connected hidden layers and then used together with the radiomic features to train an RF (Fig. 3). We trained CNNs on both the FLAIR sequence alone and in conjunction with the other sequences (T2-weighted, T1-weighted with and without contrast), in an attempt to improve the accuracy of the model by including additional information. The only difference between these two configurations was the number of input channels (one versus four), and when training the model on four sequences we employed an input-level dropout approach38 to deal with missing sequences since not all patients had every sequence available.Figure 2 The top of this figure describes the full architecture of our custom shallow 3D CNN. The bottom of the figure describes the contents of the three convolutional blocks (Conv. Block). Convolutional blocks 1, 2, and 3, had 16, 32, and 64 output channels respectively. The first fully connected layer had 64 input and 16 output neurons, while the second fully connected layer had 16 input and 3 output neurons, representing the three classes, BRAF Fusion, BRAF Mutation, and non-BRAF altered.

Figure 3 A visualization of our experiments. A (top left) and B (top right) depict the processes of training the CNN and radiomic models. C (bottom left) describes the combined model which takes the convolutional layers from A and uses them to generate CNN features that are then combined with the radiomic features and fed into a random forest model. D (bottom right) illustrates our attempt to learn radiomic features with CNNs from the MRIs. We first sorted the radiomic features according to their permutation importance in the random forest model from B. These features were then used as labels which we then attempted to learn using the same CNN configuration found as in A.

In addition to testing whether CNNs and radiomics were complementary by combining them, we tested whether the CNNs could learn radiomics features as labels of MR images (Fig. 3). It was not computationally feasible to test the ability of the CNN to learn all 851 radiomic features, so we focused on the features that were the most predictive of pLGG molecular status in our radiomics model according to permutation importance. The rationale for using this subset of radiomic features was that if we found that the CNN could not learn them, we would gain some insight into the limitations of CNNs on this classification task. To the best of our knowledge, our experiment was the first to use real MRI images to test the ability of CNNs to learn radiomic features.

Along with the preprocessing steps outlined above that were performed to prepare the MRIs for radiomic feature extraction, each image was registered to the SRI24 atlas39 prior to being used with the CNN. Registration is commonly included in CNN pipelines, for instance, the Brain Tumor Segmentation challenge40–42 registers images to the SRI24 atlas. Image registration is thought to help CNNs identify positional information by aligning images such that each voxel represents the same anatomical location across all images. The radiomic approach does not account for positional information, which is not captured by typical radiomic features. Thus we did not include registration in the radiomics pipeline, since registration warps and resizes the brain, adding noise to the radiomic features, without any clear benefit.

Experiment configuration

To build the RF radiomics model we split the patients into test (20%) and development (80%) sets, where the development set is the training set plus the validation set, 25 times and performed five-fold cross-validation within each development set to tune hyperparameters (Table 2). Under this nested-cross validation approach, results were unlikely to be influenced by random data splitting. Due to the high computational load of training neural networks, a nested cross-validation scheme was not feasible for experiments using CNNs. We still repeated our experiments to avoid results biased by (un)lucky data splits by resampling test (20%) and development (80%) sets 25 times for the classification tasks and 10 times when attempting to learn radiomics with CNN. However, we only optimized CNN hyperparameters (Table 2) over a single split of the development set (still into five parts, four of which were used for training, while the last one was used for validation), rather than using five-fold cross-validation. Stratified sampling was used for all data splits, to account for class imbalances. When training CNNs the batch size was set to 16 images, and we used a cosine annealing period43 of 50 epochs, with a single warm restart for a total of 100 epochs of training. Kaiming Normal weight initialization44 was implemented to initialize weight parameters in both the convolutional and fully connected layers. Dropout was used during model training and switched off for inference. The final CNN used for evaluation on the test set was taken from the best epoch of the optimal set of hyperparameters, as measured by validation loss on the validation portion of the development set. Table 2 Hyperparameters for radiomics RF and CNN models.

Random Forest hyperparameter	Values explored	
Number of trees	[75, 150, 300]	
Minimum samples per leaf	[1, 2, 4]	
Minimum samples for a split	[2, 4, 8]	
Maximum depth	[4, 8, 12]	
Proportion of samples to use in each tree	[0.5, 0.75, 1.0]	
Proportion of features to use in each tree	[0.25, 0.5, 0.75]	
CNN hyperparameter	Values explored	
Learning rate	[0.1, 0.01, 0.001, 0.0001]	
Dropout (fully connected layers)	[0.3, 0.6]	
Dropout (convolutional layers)	[0.3, 0.6]	
We performed an exhaustive grid search across these hyperparameters to find the optimal model for both our radiomics, CNN, and combined models. For each development/test set split, the optimal model was found through hyperparameter optimization on the development set, and then evaluated on the test set.

Both CNN and RF classification models were trained using cross-entropy loss and evaluated by One-Vs-Rest Area Under the Receiver Operating Characteristic Curve (AUC) on held-out test data. Correlation between predicted and true radiomic feature values was used to evaluate CNNs that were trained, using mean squared error loss, to learn radiomics features as MR image labels. The radiomics, CNN, and combined models were trained and tested on identical data splits. Resampling results in the test set of one trial containing samples from the training set of another trial, which invalidates the independence assumption of the traditional t-test45. To account for this dependence, we use the “corrected resampled t-test”46,47 to test for statistically significant differences in model performance. Python 3.11.0 was used to run all experiments. We relied upon Python’s Scikit-learn package 1.2.048 to execute ML concepts, while the PyTorch 1.13.0 library49 was used to implement deep-learning models.

Results

The performance of each of the models we trained is summarized in Table 3. On average, using the FLAIR sequence alone, the difference between the performance of the custom shallow CNN (mean AUC: 0.764) and the ResNet (0.758) was not statistically significant (p-value: 0.3836). Thus, we proceeded with the more computationally efficient custom shallow CNN, which is 10 × faster than the ResNet. The difference between the performance of the CNN when trained on FLAIR alone and with all four sequences (mean AUC: 0.757) was not significant (p-value: 0.272), so we continued with FLAIR alone for the remaining experiments. Table 3 Listed are the mean AUC and 95% confidence intervals for the mean AUC over 25 different resampled test sets for the custom shallow CNN (FLAIR only, and all sequences), ResNet, radiomics, and combined models.

		95% confidence interval of mean	
Model	Mean AUC	Lower bound	Upper bound	
Shallow CNN (All)	0.757	0.746	0.769	
ResNet (FLAIR)	0.758	0.743	0.772	
Shallow CNN (FLAIR)	0.764	0.751	0.777	
Radiomics (FLAIR)	0.802	0.785	0.819	
Combined (FLAIR)	0.824	0.810	0.838	
Best performing model is in bold.

The performance of the radiomics model (mean AUC: 0.802) eclipsed that of the CNN (0.764), but the combined model performed best (0.824). The differences were significant; p-values when testing whether the combined model was better than the radiomics model or the CNN alone were 0.0344 and 0.0002 respectively, while the p-value when comparing the radiomics model and the CNN was 0.0350.

The most important features in the radiomics model according to permutation importance are listed in Table 4. Figure 4 depicts the ability of the CNN to learn these features as labels from the FLAIR MR images. The CNN was able to learn gray level dependence matrix (GLDM) dependence non-uniformity normalized (average correlation: 0.924) and surface-to-volume ratio (0.864) well. It had more trouble, but some level of success, with gray level size zone matrix (GLSZM) zone percentage (0.735), flatness (0.563), and sphericity (0.532). The CNN had no ability to learn gray level run length matrix (GLRLM) non-uniformity normalized (− 0.042) or variance (− 0.017). Table 4 The seven most predictive radiomic features ranked in order of the ability of the CNN to accurately predict the feature value.

Feature name	Category	Description	Average correlation	
GLDM dependence non-uniformity normalized	Texture	Similarity of dependence	0.924	
Surface-to-volume ratio	Shape	Non-dimensionless measure of sphericity dependent on volume	0.864	
GLSZM zone percentage	Texture	Coarseness of the texture	0.735	
Flatness	Shape	Ratio of the lengths of the two largest principal components	0.563	
Sphericity	Shape	Roundness of the tumor region	0.532	
GLRLM variance	Texture	Variance of the gray-level intensity of the runs	 − 0.017	
GLRLM gray-level non-uniformity normalized	Texture	Similarity of gray-level intensity features	 − 0.042	
The last column contains the average correlation between the actual feature value and the CNNs prediction.

Figure 4 Each boxplot represents the average correlation between actual and predicted radiomic feature values over 10 different splits of the data into development and test sets.

Discussion

This study explored the use of radiomics and CNNs to create a model that can non-invasively identify the underlying genetic alteration of pLGGs, labeling them as either BRAF Fusion, BRAF Mutation, or non-BRAF altered. We found that radiomics (mean AUC: 0.802) outperformed CNNs (0.764) and that a combined model performed better than either approach on its own (0.824). Our experiments also showed that utilizing a deeper model did not improve CNN performance. Furthermore, performance was similar with and without additional MRI sequences beyond FLAIR. We trained CNNs to learn predictive radiomic features from MR images and found that the CNN could estimate the value of certain radiomic features well, however, it was completely incapable of predicting others. These results suggest that the CNN may have performed worse than radiomics because it could not extract all the information contained in the radiomic features.

Handcrafted features have fallen out of favor for medical imaging tasks50, while neural network-based approaches have grown in prominence, in part because, theoretically, they can learn to discover any handcrafted feature from an image; but this does not always play out in practice. We think that it is too early to give up on the traditional ML techniques relying on handcrafted features. It is yet unclear as to whether the drawbacks of radiomics are more problematic than those of CNNs. There have been only a few studies that compare the two approaches directly, and the results have been inconsistent. CNNs have been shown to outperform radiomics for the identification of schizophrenia51, axillary lymph node metastasis21, and malignant breast lesions20. Radiomics surpassed CNNs for the classification of central nervous system tumors31, and the differentiation of malignant and benign ground glass nodules30. Our study adds another data point to this limited body of evidence and suggests that, contrary to popular belief, CNNs alone might not always be the best option for medical imaging tasks.

Klyuzhin et al.24 found that CNNs can learn first-order intensity and size-related radiomic features but are less able to learn shape irregularity and heterogeneity properties on a synthetic PET dataset. Their findings provided preliminary evidence that CNNs used alone in medical imaging are fundamentally limited due to their inability to capture radiomic features associated with clinical outcomes22. Our investigation, performed on MR images from real patients, provides further insights into the relationship between CNNs and radiomic features. In line with the results of24, we found that CNNs are better at learning features influenced by size, like surface-to-volume ratio, than shape-based features, like sphericity and flatness. Furthermore, we explored the ability of the CNN to learn texture features derived from the GLDM, GLRLM, and GLSZM. Klyuzhin et al. concluded that CNNs have a limited ability to learn radiomic texture features generally, whereas our results suggest that CNNs have trouble learning only certain pLGG radiomic texture features. GLDM dependence non-uniformity normalized was learned well, though the CNN performed worse on GLSZM zone percentage and had no ability to learn GLRLM variance or gray-level non-uniformity normalized. Overall, our results provide supporting evidence for the conclusion of Klyuzhin et al.24, that CNNs are not capable of capturing all of the tumor information contained in handcrafted radiomic features.

There are limitations to this study. First, though our dataset is large from a pediatric neuro-oncology perspective (336 patients), from an ML standpoint, it is relatively small and was collected from a single institution. Prospective data and data from other institutions are needed to evaluate the generalizability of our approach and the reliability of our model. Nevertheless, our dataset is diverse, having been collected over two decades using different MRI scanners and field strengths, and our results were found to be statistically significant under a rigorous nested cross-validation approach. Second, our claims about the performance of CNNs, radiomics, and combined models, the (in)ability of CNNs to learn certain radiomic features, and the lack of benefit when including different MRI sequences beyond FLAIR and larger CNNs, are limited to the specific scenarios we explored in this study. There are countless other combinations of radiomics models, neural networks, experiment configurations and preprocessing techniques, that could result in different conclusions; these need to be explored to evaluate the generalizability of our results. However, our results align with related studies24,27–32 that used different configurations, models, and data sources, giving us confidence that our conclusions will be confirmed by future studies. Third, our analysis was based on the entire tumor region. Performance may have been better with more specific labels (edema, necrosis, and enhancing vs non-enhancing structures); this exploration has been left to future work. Finally, it is not clear from our experiments why the CNN underperformed radiomics on the classification task and was unable to learn some radiomic features. Further experiments on larger datasets are required to determine whether CNN performance was limited because of a lack of data, or because there are inherent limitations to the performance of CNNs due to their design.

Conclusion

In this study, we created a model capable of non-invasively classifying pLGGs by BRAF status. We investigated the performance of radiomics and CNNs both separately and combined for this classification task, and found that the combined model relying on both CNNs and radiomic features performed best. Furthermore, we identified radiomic features that CNNs had trouble learning, uncovering limitations of CNNs in terms of the types of information they can extract from medical images. Future studies with large external and prospective datasets are necessary to improve diagnostic accuracy and further validate the robustness of the results.

Acknowledgements

This research has been made possible with the financial support of the Canadian Institutes of Health Research (CIHR) (Funding Reference Number: 184015).

Author contributions

K.K., B.E.W., and F.K. contributed to the design of the concept and study. K.K. contributed to the implementation of machine learning modules and running the experiments. K.K. and K.N. contributed to preprocessing of imaging data. J.B., L.N., U.T., and C.H. contributed to collecting, reviewing, and providing clinical data and genetic markers. M.W. and B.E.W. contributed to reviewing the imaging data and providing tumor segmentation for the data. K.K. wrote the first draft of the manuscript and all authors contributed to the reviewing and editing of the manuscript. All authors read and approved the final manuscript.

Data availability

The datasets generated and/or analyzed during the current study are available from the corresponding author on reasonable request pending the approval of the institution(s) and trial/study investigators who contributed to the dataset.

Competing interests

The authors declare no competing interests.

Publisher's note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

These authors contributed equally: Birgit Betina Ertl-Wagner and Farzad Khalvati.
==== Refs
References

1. Sievert AJ Fisher MJ Pediatric low-grade gliomas J. Child Neurol. 2009 24 1397 1408 10.1177/0883073809342005 19841428
Sievert, A. J. & Fisher, M. J. Pediatric low-grade gliomas. J. Child Neurol. 24, 1397–1408 (2009).19841428
2. Ostrom QT Patil N Cioffi G CBTRUS statistical report: primary brain and other central nervous system tumors diagnosed in the United States in 2013–2017 Neuro Oncol. 2020 22 iv1 iv96 10.1093/neuonc/noaa200 33123732
Ostrom, Q. T. et al. CBTRUS statistical report: primary brain and other central nervous system tumors diagnosed in the United States in 2013–2017. Neuro Oncol. 22, iv1–iv96 (2020).33123732
3. Ryall S Tabori U Hawkins C Pediatric low-grade glioma in the era of molecular diagnostics Acta Neuropathol. Commun. 2020 8 30 10.1186/s40478-020-00902-z 32164789
Ryall, S., Tabori, U. & Hawkins, C. Pediatric low-grade glioma in the era of molecular diagnostics. Acta Neuropathol. Commun. 8, 30 (2020).32164789
4. Armstrong GT Conklin HM Huang S Survival and long-term health and cognitive outcomes after low-grade glioma Neuro Oncol. 2011 13 223 234 10.1093/neuonc/noq178 21177781
Armstrong, G. T. et al. Survival and long-term health and cognitive outcomes after low-grade glioma. Neuro Oncol. 13, 223–234 (2011).21177781
5. Bennett J Erker C Lafay-Cousin L Canadian pediatric neuro-oncology standards of practice Front. Oncol. 2020 10 593192 10.3389/fonc.2020.593192 33415075
Bennett, J. et al. Canadian pediatric neuro-oncology standards of practice. Front. Oncol. 10, 593192 (2020).33415075
6. Lambin P Leijenaar RTH Deist TM Radiomics: The bridge between medical imaging and personalized medicine Nat. Rev. Clin. Oncol 2017 14 749 762 10.1038/nrclinonc.2017.141 28975929
Lambin, P. et al. Radiomics: The bridge between medical imaging and personalized medicine. Nat. Rev. Clin. Oncol 14, 749–762 (2017).28975929
7. Khalvati F Zhang Y Wong A Haider MA Radiomics Encycl. Biomed. Eng. 2019 2 597 603 10.1016/B978-0-12-801238-3.99964-1
Khalvati, F., Zhang, Y., Wong, A. & Haider, M. A. Radiomics. Encycl. Biomed. Eng. 2, 597–603 (2019).
8. Aerts HJW The potential of radiomic-based phenotyping in precision medicine: A review JAMA Oncol. 2016 2 1636 1642 10.1001/jamaoncol.2016.2631 27541161
Aerts, H. J. W. The potential of radiomic-based phenotyping in precision medicine: A review. JAMA Oncol. 2, 1636–1642 (2016).27541161
9. Aerts HJWL Velazquez ER Leijenaar RTH Decoding tumour phenotype by noninvasive imaging using a quantitative radiomics approach Nat. Commun. 2014 5 1 9
Aerts, H. J. W. L. et al. Decoding tumour phenotype by noninvasive imaging using a quantitative radiomics approach. Nat. Commun. 5, 1–9 (2014).
10. Shui L Ren H Yang X The era of radiogenomics in precision medicine: An emerging approach to support diagnosis, treatment decisions, and prognostication in oncology Front. Oncol. 2020 10 570465 10.3389/fonc.2020.570465 33575207
Shui, L. et al. The era of radiogenomics in precision medicine: An emerging approach to support diagnosis, treatment decisions, and prognostication in oncology. Front. Oncol. 10, 570465 (2020).33575207
11. Lo Gullo R Daimiel I Morris EA Pinker K Combining molecular and imaging metrics in cancer: radiOgenomics Insights Imaging 2020 11 1 10.1186/s13244-019-0795-6 31901171
Lo Gullo, R., Daimiel, I., Morris, E. A. & Pinker, K. Combining molecular and imaging metrics in cancer: radiOgenomics. Insights Imaging 11, 1 (2020).31901171
12. Ak M Toll SA Hein KZ Evolving role and translation of radiomics and radiogenomics in adult and pediatric neuro-oncology AJNR Am. J. Neuroradiol. 2021 10.3174/ajnr.A7297 34649914
Ak, M. et al. Evolving role and translation of radiomics and radiogenomics in adult and pediatric neuro-oncology. AJNR Am. J. Neuroradiol.10.3174/ajnr.A7297 (2021).34649914
13. Ryall S Zapotocky M Fukuoka K Integrated molecular and clinical analysis of 1,000 pediatric low-grade gliomas Cancer Cell 2020 37 569 583.e5 10.1016/j.ccell.2020.03.011 32289278
Ryall, S. et al. Integrated molecular and clinical analysis of 1,000 pediatric low-grade gliomas. Cancer Cell 37, 569-583.e5 (2020).32289278
14. Wagner MW Hainc N Khalvati F Radiomics of pediatric low-grade gliomas: Toward a pretherapeutic differentiation of BRAF-mutated and BRAF-fused tumors AJNR Am. J. Neuroradiol. 2021 42 759 765 10.3174/ajnr.A6998 33574103
Wagner, M. W. et al. Radiomics of pediatric low-grade gliomas: Toward a pretherapeutic differentiation of BRAF-mutated and BRAF-fused tumors. AJNR Am. J. Neuroradiol. 42, 759–765 (2021).33574103
15. Kudus K Wagner MW Namdar K Increased confidence of radiomics facilitating pretherapeutic differentiation of BRAF-altered pediatric low-grade glioma Eur. Radiol. 2023 10.1007/s00330-023-10267-1 37803212
Kudus, K. et al. Increased confidence of radiomics facilitating pretherapeutic differentiation of BRAF-altered pediatric low-grade glioma. Eur. Radiol.10.1007/s00330-023-10267-1 (2023).37803212
16. Vafaeikia P Wagner MW Hawkins C MRI-based end-to-end pediatric low-grade glioma segmentation and classification Can. Assoc. Radiol. J. 2024 75 153 160 10.1177/08465371231184780 37401906
Vafaeikia, P. et al. MRI-based end-to-end pediatric low-grade glioma segmentation and classification. Can. Assoc. Radiol. J. 75, 153–160 (2024).37401906
17. Xu J Lai M Li S Radiomics features based on MRI predict BRAF V600E mutation in pediatric low-grade gliomas: A non-invasive method for molecular diagnosis Clin Neurol. Neurosurg. 2022 222 107478 10.1016/j.clineuro.2022.107478 36244075
Xu, J. et al. Radiomics features based on MRI predict BRAF V600E mutation in pediatric low-grade gliomas: A non-invasive method for molecular diagnosis. Clin Neurol. Neurosurg. 222, 107478 (2022).36244075
18. Liu Z Hong X Wang L Radiomic features from multiparametric magnetic resonance imaging predict molecular subgroups of pediatric low-grade gliomas BMC Cancer 2023 23 848 10.1186/s12885-023-11338-8 37697238
Liu, Z. et al. Radiomic features from multiparametric magnetic resonance imaging predict molecular subgroups of pediatric low-grade gliomas. BMC Cancer 23, 848 (2023).37697238
19. Tak, D. et al. Noninvasive molecular subtyping of pediatric low-grade glioma with self-supervised transfer learning. Radiol. Artif. Intell. 6, e230333 (2024).
20. Truhn D Schrading S Haarburger C Radiomic versus convolutional neural networks analysis for classification of contrast-enhancing lesions at multiparametric breast MRI Radiology 2019 290 290 297 10.1148/radiol.2018181352 30422086
Truhn, D. et al. Radiomic versus convolutional neural networks analysis for classification of contrast-enhancing lesions at multiparametric breast MRI. Radiology 290, 290–297 (2019).30422086
21. Sun Q Lin X Zhao Y Deep learning versus radiomics for predicting axillary lymph node metastasis of breast cancer using ultrasound images: Don’t forget the peritumoral region Front. Oncol. 2020 10 53 10.3389/fonc.2020.00053 32083007
Sun, Q. et al. Deep learning versus radiomics for predicting axillary lymph node metastasis of breast cancer using ultrasound images: Don’t forget the peritumoral region. Front. Oncol. 10, 53 (2020).32083007
22. Orlhac F Nioche C Klyuzhin I Radiomics in PET imaging: A practical guide for newcomers PET Clin. 2021 16 597 612 10.1016/j.cpet.2021.06.007 34537132
Orlhac, F. et al. Radiomics in PET imaging: A practical guide for newcomers. PET Clin. 16, 597–612 (2021).34537132
23. Afshar P Mohammadi A Plataniotis KN From handcrafted to deep-learning-based cancer radiomics: Challenges and opportunities IEEE Signal Process. Mag. 2019 36 132 160 10.1109/MSP.2019.2900993
Afshar, P. et al. From handcrafted to deep-learning-based cancer radiomics: Challenges and opportunities. IEEE Signal Process. Mag. 36, 132–160 (2019).
24. Klyuzhin IS Xu Y Ortiz A Testing the ability of convolutional neural networks to learn radiomic features Comput. Methods Progr. Biomed 2022 219 106750 10.1016/j.cmpb.2022.106750
Klyuzhin, I. S. et al. Testing the ability of convolutional neural networks to learn radiomic features. Comput. Methods Progr. Biomed 219, 106750 (2022).
25. Sarvamangala DR Kulkarni RV Convolutional neural networks in medical image understanding: A survey Evol. Intell. 2022 15 1 22 10.1007/s12065-020-00540-3 33425040
Sarvamangala, D. R. & Kulkarni, R. V. Convolutional neural networks in medical image understanding: A survey. Evol. Intell. 15, 1–22 (2022).33425040
26. Yamashita R Nishio M Do RKG Togashi K Convolutional neural networks: An overview and application in radiology Insights Imaging 2018 9 611 629 10.1007/s13244-018-0639-9 29934920
Yamashita, R., Nishio, M., Do, R. K. G. & Togashi, K. Convolutional neural networks: An overview and application in radiology. Insights Imaging 9, 611–629 (2018).29934920
27. Zhang Y Lobo-Mueller EM Karanicolas P Improving prognostic performance in resectable pancreatic ductal adenocarcinoma using radiomics and deep learning features fusion in CT images Sci. Rep. 2021 11 1 11 33414495
Zhang, Y. et al. Improving prognostic performance in resectable pancreatic ductal adenocarcinoma using radiomics and deep learning features fusion in CT images. Sci. Rep. 11, 1–11 (2021).33414495
28. Whitney HM Li H Ji Y Comparison of breast MRI tumor classification using human-engineered radiomics, transfer learning from deep convolutional neural networks, and fusion methods Proc. IEEE 2020 108 163 177 10.1109/JPROC.2019.2950187
Whitney, H. M. et al. Comparison of breast MRI tumor classification using human-engineered radiomics, transfer learning from deep convolutional neural networks, and fusion methods. Proc. IEEE 108, 163–177 (2020).
29. Antropova N Huynh BQ Giger ML A deep feature fusion methodology for breast cancer diagnosis demonstrated on three imaging modality datasets Med. Phys. 2017 44 5162 5171 10.1002/mp.12453 28681390
Antropova, N., Huynh, B. Q. & Giger, M. L. A deep feature fusion methodology for breast cancer diagnosis demonstrated on three imaging modality datasets. Med. Phys. 44, 5162–5171 (2017).28681390
30. Hu X Gong J Zhou W Computer-aided diagnosis of ground glass pulmonary nodule by fusing deep learning and radiomics features Phys. Med. Biol. 2021 66 065015 10.1088/1361-6560/abe735 33596552
Hu, X. et al. Computer-aided diagnosis of ground glass pulmonary nodule by fusing deep learning and radiomics features. Phys. Med. Biol. 66, 065015 (2021).33596552
31. Yun J Park JE Lee H Radiomic features and multilayer perceptron network classifier: A robust MRI classification strategy for distinguishing glioblastoma from primary central nervous system lymphoma Sci. Rep. 2019 9 1 10 10.1038/s41598-019-42276-w 30626917
Yun, J. et al. Radiomic features and multilayer perceptron network classifier: A robust MRI classification strategy for distinguishing glioblastoma from primary central nervous system lymphoma. Sci. Rep. 9, 1–10 (2019).30626917
32. Choi YS Bae S Chang JH Fully automated hybrid approach to predict the IDH mutation status of gliomas via deep learning and radiomics Neuro Oncol. 2021 23 304 313 10.1093/neuonc/noaa177 32706862
Choi, Y. S. et al. Fully automated hybrid approach to predict the IDH mutation status of gliomas via deep learning and radiomics. Neuro Oncol. 23, 304–313 (2021).32706862
33. Wagner MW Namdar K Alqabbani A Dataset size sensitivity analysis of machine learning classifiers to differentiate molecular markers of paediatric low-grade gliomas based on MRI Oncol. Radiother. 2022 16 01 06
Wagner, M. W. et al. Dataset size sensitivity analysis of machine learning classifiers to differentiate molecular markers of paediatric low-grade gliomas based on MRI. Oncol. Radiother. 16, 01–06 (2022).
34. Kikinis R Pieper SD Vosburgh KG Jolesz FA 3D Slicer: A Platform for Subject-Specific Image Analysis, Visualization, and Clinical Support Intraoperative Imaging and Image-Guided Therapy 2014 Springer 277 289
Kikinis, R., Pieper, S. D. & Vosburgh, K. G. 3D Slicer: A Platform for Subject-Specific Image Analysis, Visualization, and Clinical Support. In Intraoperative Imaging and Image-Guided Therapy (ed. Jolesz, F. A.) 277–289 (Springer, 2014).
35. Fedorov A Beichel R Kalpathy-Cramer J 3D Slicer as an image computing platform for the quantitative imaging network Magn. Reson. Imaging 2012 30 1323 1341 10.1016/j.mri.2012.05.001 22770690
Fedorov, A. et al. 3D Slicer as an image computing platform for the quantitative imaging network. Magn. Reson. Imaging 30, 1323–1341 (2012).22770690
36. van Griethuysen JJM Fedorov A Parmar C Computational radiomics system to decode the radiographic phenotype Cancer Res. 2017 77 e104 e107 10.1158/0008-5472.CAN-17-0339 29092951
van Griethuysen, J. J. M. et al. Computational radiomics system to decode the radiographic phenotype. Cancer Res. 77, e104–e107 (2017).29092951
37. Kataoka, H., Wakamiya, T., Hara, K., Satoh, Y. (2020) Would mega-scale datasets further enhance spatiotemporal 3D CNNs? ArXiv
38. Grøvik, E., Yi, D., Iv, M., et al (2021) Handling missing MRI sequences in deep learning segmentation of brain metastases: A multicenter study. npj Digital Medicine 4:1–7
39. Rohlfing T Zahr NM Sullivan EV Pfefferbaum A The SRI24 multichannel atlas of normal adult human brain structure Hum. Brain Map. 2010 31 798 819 10.1002/hbm.20906
Rohlfing, T., Zahr, N. M., Sullivan, E. V. & Pfefferbaum, A. The SRI24 multichannel atlas of normal adult human brain structure. Hum. Brain Map. 31, 798–819 (2010).
40. Bakas S Akbari H Sotiras A Advancing The Cancer Genome Atlas glioma MRI collections with expert segmentation labels and radiomic features Sci. Data 2017 4 1 13 10.1038/sdata.2017.117
Bakas, S. et al. Advancing The Cancer Genome Atlas glioma MRI collections with expert segmentation labels and radiomic features. Sci. Data 4, 1–13 (2017).
41. Baid, U., Ghodasara, S., Bilello, M., et al (2021) The RSNA-ASNR-MICCAI BraTS 2021 Benchmark on brain tumor segmentation and radiogenomic classification. ArXiv
42. Menze BH Jakab A Bauer S The multimodal brain tumor image segmentation benchmark (BRATS) IEEE Trans. Med. Imaging 2015 34 1993 2024 10.1109/TMI.2014.2377694 25494501
Menze, B. H. et al. The multimodal brain tumor image segmentation benchmark (BRATS). IEEE Trans. Med. Imaging 34, 1993–2024 (2015).25494501
43. Loshchilov, I., Hutter, F. (2017) SGDR: Stochastic gradient descent with warm restarts. International Conference on Learning Representations
44. He, K., Zhang, X., Ren, S., Sun, J. (2015) Delving deep into rectifiers: Surpassing human-level performance on ImageNet classification. In: 2015 IEEE International Conference on Computer Vision (ICCV). pp 1026–1034
45. Dietterich TG Approximate statistical tests for comparing supervised classification learning algorithms Neural Comput. 1998 10 1895 1923 10.1162/089976698300017197 9744903
Dietterich, T. G. Approximate statistical tests for comparing supervised classification learning algorithms. Neural Comput. 10, 1895–1923 (1998).9744903
46. Bouckaert RR Frank E Evaluating the Replicability of Significance Tests for Comparing Learning Algorithms. Advances in Knowledge Discovery and Data Mining 2004 Springer
Bouckaert, R. R. & Frank, E. Evaluating the Replicability of Significance Tests for Comparing Learning Algorithms. Advances in Knowledge Discovery and Data Mining (Springer, 2004).
47. Nadeau C Bengio Y Inference for the generalization error Mach. Learn. 2003 52 239 281 10.1023/A:1024068626366
Nadeau, C. & Bengio, Y. Inference for the generalization error. Mach. Learn. 52, 239–281 (2003).
48. Pedregosa F Varoquaux G Gramfort A Scikit-learn: Machine learning in python J. Mach. Learn. Res. 2011 12 2825 2830
Pedregosa, F. et al. Scikit-learn: Machine learning in python. J. Mach. Learn. Res. 12, 2825–2830 (2011).
49. Paszke, A., Gross, S., Massa, F., et al (2019) PyTorch: an imperative style, high-performance deep learning library. In: Proceedings of the 33rd International Conference on Neural Information Processing Systems. Curran Associates Inc., Red Hook, NY, USA, pp 8026–8037
50. Dutta, P., Upadhyay, P., De, M., Khalkar, RG. (2020) Medical image analysis using deep convolutional neural networks: CNN architectures and transfer learning. In: 2020 International Conference on Inventive Computation Technologies (ICICT). pp 175–180
51. Hu, M., Sim, K., Zhou, JH., et al (2020) Brain MRI-based 3D convolutional neural networks for classification of schizophrenia and controls. arXiv [cs.CV]
