
==== Front
Sci Rep
Sci Rep
Scientific Reports
2045-2322
Nature Publishing Group UK London

73140
10.1038/s41598-024-73140-1
Article
Machine learning based predictive analysis of DNA cleavage induced by diverse nanomaterials
Niu Jie 12
Wang Xufeng 3
Chen Jiangling 3
Zhao Yingcan yingcanzhao@uic.edu.cn

4
Chen Xiaohui 5
Yang Baoling 6
Liu Na 6
Wu Pan pwu@gzu.edu.cn

12
1 https://ror.org/02wmsc916 grid.443382.a 0000 0004 1804 268X College of Resources and Environmental Engineering, Guizhou University, Guiyang, 550025 China
2 https://ror.org/03m01yf64 grid.454828.7 0000 0004 0638 8050 Key Laboratory of Karst Georesources and Environment, Ministry of Education, Guiyang, 550025 China
3 https://ror.org/02xe5ns62 grid.258164.c 0000 0004 1790 3548 School of Environment, Jinan University, Guangzhou, 510632 China
4 https://ror.org/0145fw131 grid.221309.b 0000 0004 1764 5980 Environmental Science Program, Department of Life Sciences, Beijing Normal University-Hong Kong Baptist University United International College, No. 2000 Jintong Road, Tangjiawan, Zhuhai, 519087 Guangdong China
5 grid.258164.c 0000 0004 1790 3548 College of Chemistry and Materials Science, Jinan University, Guangzhou, 510632 China
6 https://ror.org/02xe5ns62 grid.258164.c 0000 0004 1790 3548 College of Life Science and Technology, Jinan University, Guangzhou, 510632 China
20 9 2024
20 9 2024
2024
14 219666 4 2024
13 9 2024
© The Author(s) 2024
2024
https://creativecommons.org/licenses/by-nc-nd/4.0/ Open Access This article is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License, which permits any non-commercial use, sharing, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if you modified the licensed material. You do not have permission under this licence to share adapted material derived from this article or parts of it. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by-nc-nd/4.0/.
DNA cleavage by nanomaterials has the potential to be utilized as an innovative tool for gene editing. Numerous nanomaterials exhibiting DNA cleavage properties have been identified and cataloged. Yet, the exploitation of property data through data-driven machine-learning approaches remains unexplored. A database was developed, compiling thirty distinctive characteristics, encompassing physical and chemical properties, as well as experimental conditions of nanomaterials that have demonstrated DNA cleavage capability such as in articles published over the past two decades. The DNA cleavage effect and efficiency of nanomaterials were predicted using machine learning algorithms such as support vector machines, deep neural networks, and random forest, and a classification accuracy of 0.93 for the cleavage effect was achieved. Moreover, the potential of utilizing larger datasets to enhance the predictive capacity of models was discussed. The findings indicate the feasibility of predicting nanomaterial properties based on experimental data. Evaluating the performance and effectiveness of the machine learning models trained using the existing data can furnish valuable insights for future materials research endeavors, especially for the design of DNA cleavage with specific sites.

Subject terms

Nanotoxicology
Nanoscale materials
issue-copyright-statement© Springer Nature Limited 2024
==== Body
pmcIntroduction

In the field of biomedicine, an intensified focus is placed on the exploration of nanomaterials, particularly in terms of their DNA cleavage capabilities. Notably, nanomaterials such as fullerenes1, carbon nanotubes2, and graphene3,4 have demonstrated efficient DNA cleavage, attributed to their exceptional characteristics, including vast specific surface area, excellent electrical conductivity, and mechanical properties5. Nanomaterials exhibit significant application potential in gene editing owing to their exceptional biocompatibility and customizable surface chemistry6. The utilization of nanomaterial-based DNA cleavage technology is anticipated to catalyze revolutionary advancements in gene therapy, synthetic biology, and innovative biosensors. Hence, this novel area of using nanomaterials to serve as DNA endonuclease is extremely important since it can solve a series of problems induced by the traditional DNA endonuclease, such as instability, cost, etc. However, the vast number of potential modifying materials and the varied effects of exposure conditions render the comprehensive evaluation of all nanomaterials through traditional laboratory testing methods alone as impractical. Consequently, leveraging predictive models for nanomaterial properties offers substantial benefits in efficiency and resource allocation, overcoming the limitations of conventional laboratory testing methods.

Experimental methodologies, computational simulations, and machine-learning techniques are widely utilized for the prediction of nanomaterial properties. Experimental techniques encompass both conventional and high-throughput methods, facilitating the quantification of diverse attributes including structural, optical, and electrical properties. Indeed, the experimental assessment of nanomaterials’ DNA cleavage properties is often a laborious and expensive process, with the resulting data emerging as intricate and challenging to interpret. Computational simulations, in contrast, employ methods like molecular dynamics (MD)7 and density functional theory (DFT)8 to predict properties such as structure, thermodynamics, and mechanical characteristics. Meanwhile, machine learning algorithms, including SVM, neural network (NN), and RF, provide a data-driven methodology for to the prediction of nanomaterial properties through the process of model training9. Data-driven machine learning models have emerged as a robust method predicting the physical and chemical properties of nanomaterials to make references up to date10,11.

For instance, Fernandez et al. utilized structural features of graphene nanosheets to configure a machine-learning model, facilitating the prediction of discrepancies in electronic properties at two varied levels of approximation12. Wang et al. developed a comprehensive set of 36 machine-learning models by amalgamating nano descriptors with six distinct machine-learning algorithms. Notably, the random forest-catalase (CAT) model and the k-nearest neighbor classifier (KNN)-glutathione peroxidase (GPx) model demonstrated exceptional performance, both in terms of training set accuracy and external validation13. Mirzaei et al. curated a comprehensive dataset of nanomaterial properties and investigated the efficacy of machine learning algorithms in predicting the antioxidant efficiency of these materials14. They demonstrated that the data-driven process could provide the possibility to predict the properties of nanomaterials, enabling the screening and optimization of materials prior to synthesis, the design of experiments with enhanced precision, and the simplification of interpretation and comprehension of experimental outcomes. Consequently, the deployment of machine learning in this arena embodies significant potential for propelling the field forward and yielding valuable insights into the DNA cleavage capabilities of a broad spectrum of nanomaterials.

Given the considerable time and cost constraints associated with experimental testing of nanomaterials15, the development of a computational method for predicting DNA cleavage properties would be highly advantageous. To the best of our knowledge, this is the inaugural effort utilizing machine learning tools to predict the DNA cleavage properties of nanomaterials. As shown in Fig. 1, in this study, a data-driven method for predicting DNA cleavage properties of nanomaterials has been established. An exhaustive compilation of experimental data from DNA cleavage experiments involving nanomaterials, as documented in scientific articles over the past two decades, was conducted. Machine learning models were utilized to predict the DNA cleavage capabilities of nanomaterials, and a comparative analysis was performed to assess the prediction performance of various models. Additionally, the study investigated the key factors influencing the prediction results. Despite the challenges posed by limited data availability, the study effectively validated the potential of machine learning tools in predicting the DNA cleavage effects of nanomaterials and identified significant properties that exerted a notable impact on the predictions.Fig. 1 Overview of the workflow of the model training and prediction process. Model training and validation flowchart showing the different tools in the training process. The entire process begins with data collection and ends with verification of predictive efficiency. Each link in the process is divided into boxes, with solid arrows indicating the direction of the process and dotted arrows indicating specific data flow (DNN: Deep Neural Networks, SVR: Support Vector Machines, RF: Random Forest, BR: Bayesian Ridge Regression, GBR: Gradient Boosting Regression, SVC: Support Vector Classification, EV: Explained Variance, MAE: Mean Absolute Error, MSE: Mean Square Error).

In addition, introducing nanomaterials into the fields of gene editing or medicine requires considerations of safety, ethical concerns, and regulatory requirements6. Therefore, when promoting the application of this technology, it is essential to assess its safety and establish detailed safety protocols rigorously. Given the unique nature of this technology, governments worldwide should enact corresponding laws and regulations for oversight. This includes strict control over laboratory operations, product applications, and human trials, and enhancing international cooperation to establish unified regulatory standards. Therefore, integrating machine learning into predicting and screening the properties of nanomaterials, especially in predicting DNA shearing efficiency, can support the safety assessment of nanomaterials.

Results

Classification models for DNA cleavage nanomaterials

The objective of the classification model was to ascertain the capability of nanomaterials to cleave DNA, predicated on the values of input features. A bespoke dataset, annotated with DNA cleavage effects, served for both training and testing purposes. This dataset was apportioned into a training set, accounting for 80% of the data, and a testing set, encompassing 20% of the data. Subsequent to the model’s training phase, its performance was appraised using the testing set.

Figure 2a showcases the accuracy, precision, recall, and F1 scores of the DNN, SVM, and RF models. As indicated in the figure, all three models achieved an accuracy of 0.9. Additionally, a precision rate of 0.92 signifies a reduced rate of false positive (FP). A more elevated recall rate of 0.93 indicates a diminished rate of false negative (FN), demonstrating that the models are highly effective in identifying positive instances. As depicted in Fig. 2b, the ROC curve was utilized to evaluate the models’ performance. This curve maps the model’s efficacy across various thresholds, positioning the true positive (TP) rate on the vertical axis and the FP rate on the horizontal axis. A ROC curve nearing the upper left corner signifies a greater likelihood of correctly identifying the TPs and a reduced likelihood of falsely recognizing the FPs, heralding enhanced model performance. Given the lack of distinct differences in the shape of the ROC curves among the three models, a comparative analysis of their respective Area Under the Curve (AUC) values was performed. The AUC, a widely recognized metric for quantifying a model’s classification performance, spans from 0 to 1, where a value closer to 1 denotes superior classification efficacy and model robustness. Among the three models, the DNN registered an AUC of 0.94, marginally outperforming the other two models. Additionally, the DNN model demonstrates a higher TP rate at a lower FP rate. As the FP rate exceeds 0.2, the DNN model maintains a higher TP rate compared to the SVC and RF models. As illustrated in Fig. 2c, d, and e, confusion matrices were generated to gain further insights into the predictive performance of the three models for both label types. The confusion matrix outlines classification statistics aligning the model’s predictions with the actual labels. From the analysis of the confusion matrix, it is evident that the three models exhibit proficient abilities in correctly identifying positive samples. Nonetheless, there is a tendency for these models to misclassify negative samples as positive. This phenomenon may be attributed to the predominance of positive samples over negative ones within the dataset, notwithstanding efforts to augment the count of negative samples.Fig. 2 Evaluation of DNN, SVC, and RF classification models using the test dataset. (a) The efficacy of DNN, SVC, and RF in discerning the presence or absence of DNA cleavage effects with a fivefold cross-test. (b) The Receiver Operating Characteristic (ROC) curves of DNN, SVC, and RFC. A ROC curve approaching the upper left corner signifies superior model performance. (c–e) Confusion matrices for DNN, SVC, and RF, respectively, illustrate each model’s predictive accuracy.

The learning curves of the three models are presented, showcasing the performance trends on both the training and validation datasets as the number of training samples increases. The curves for the training set and validation set of the three models exhibit convergence and proximity, accompanied by low error values. This indicates that the models are effectively fitting the training data and possess robust generalization capabilities. Among them, the RF model stands out with the highest score, indicating superior predictive performance. Specifically, as shown in Fig. 3a, the DNN achieves stable prediction results after approximately 75 epochs, with the training sample size at 400, marking an accuracy of approximately 0.85. Moreover, as illustrated in Fig. 3b, once the SVC model’s score attains 0.85, the scores for both the validation and training sets begin to show a parallel fluctuation pattern, signifying that enhancements in the training set performance are mirrored in the validation set performance. As highlighted in Fig. 3c, while the RF model does not exhibit similar fluctuations—potentially due to the dataset’s size constraints—it nevertheless demonstrates the capability for effective classification within the domain, underscoring the model’s potential for tasks in this field.Fig. 3 Learning curves of three classification models. (a) The learning curve of the DNN model showcases how the model’s prediction accuracy stabilizes after approximately 75 epochs with a training sample size of 400. (b) The learning curve of the SVC model highlights the parallel fluctuation patterns between the validation and training set scores once the model score reaches 0.85, indicating synchronized improvements across both sets. (c) The learning curve of the RFC model with a fivefold cross-test, with accuracy as the metric, demonstrates the models’ superior predictive performance despite the absence of fluctuation patterns observed in the SVC model, suggesting strong potential for classification tasks within the field.

Regression models for DNA cleavage nanomaterials

The objective of the regression model was to predict the DNA cleavage efficiency of nanomaterials based on the values of input characteristics. A custom-constructed dataset, annotated with DNA cleavage efficiency, was utilized for both the training and testing phases, maintaining a training-to-testing set ratio of 0.8–0.2, respectively. After the completion of the model training phase, the efficacy of the model was assessed using the test set.

As shown in Fig. 4a, while the EV and R2 values of all models do not exceed 0.6, RF and GBR stand out with higher EV and R2 scores, along with lower MSE and MAE values. This indicates that RF and GBR excel in terms of model fit compared to other models. In terms of stability, DNN demonstrates smaller standard deviations in EV and R2 metrics, suggesting a notable advantage in prediction consistency. Nonetheless, given its lower EV and R2 scores, DNN may not be the optimal choice for a prediction model. Considering both the regression efficacy and stability, RF emerges as the most proficient model for regression tasks.Fig. 4 Performance and prediction error distribution of regression models using test dataset. (a) The efficacy of DNN, SVR, RF, BR, LR, and GBR in predicting the efficiency of DNA cleavage. Each performance metric of every model with a fivefold cross-test to ensure reliability. (b–g) The error distributions of each model’s prediction. A concentration of the error distribution close to 0 indicates superior model performance.

As shown in Fig. 4b–g, the prediction errors of all six models are predominantly clustered around zero, with RF and GBR exhibiting a denser concentration of errors near this point. This observation implies that RF and GBR hold the potential for more accurate predictions of DNA cleavage efficiency compared to the other four models. Additionally, the area where the prediction error is less than zero surpasses that where the prediction error is greater than zero, revealing a propensity for the models to yield conservative predictions. Consequently, such predictions are typically lower than the actual values, a pattern possibly linked to the predominance of lower efficiency labels within the training dataset. RF, in particular, demonstrates superior regression performance, which could be ascribed to the characteristics of the dataset features. The dataset contains a substantial mix of binary and numeric features, domains where RF excels due to its proficiency in managing such data types and its capability to navigate high-dimensional spaces efficiently. This advantage translates into enhanced regression results with minimal requirements for data preprocessing, underscoring RF’s suitability for this application based on the dataset’s characteristics.

As shown in Fig. 5, RF and GBR models exhibit higher R2 scores, underscoring their superior capability to explain the variations in DNA cleavage efficiency based on the input features. Furthermore, analyzing the distribution of predicted values across the six models reveals a notable concentration within the range of 0.25–0.5. This pattern suggests a propensity for overestimation in instances where the actual label falls below 0.25, and underestimation when it exceeds 0.5, reflecting a bias towards the mean of the distribution. This bias likely stems from the true labels’ distribution, which predominantly spans the 0.25–0.5 interval. In an effort to optimize the MSE during the training process, the models appear to adopt a conservative prediction strategy, aiming to minimize extreme errors by gravitating towards the distribution’s central tendency.Fig. 5 Comparisons of predicted values by models and true values using test dataset. Scatter points proximal to the diagonal dashed line signify a high degree of accuracy in the model’s prediction. (a) DNN, (b) SVR, (c) RF, (d) BR, (e) LR, and (f) GBR. The alignment of scatter points with the diagonal line across these subplots provides a visual representation of each model’s ability to accurately predict DNA cleavage efficiency, with a closer alignment indicating superior predictive accuracy.

Among the regression prediction models evaluated, GBR and RF demonstrate notably superior performance over the other four models. As depicted in Fig. 6a, the GBR model’s performance trajectory with increasing training sample sizes is evident. Initially, the model scores highly at 0.9 on the training set, but this score diminishes to 0.6 as the dataset expands. Simultaneously, the validation set score sees a gradual ascent from 0 to 0.4, signaling the model’s adaptation to mitigate overfitting and bolster its generalization capacity. Figure 6b showcases the RF model’s performance, where a slight decline in the training set score from 0.8 to 0.75 is observed, against a validation set score increment from 0 to 0.4. Despite this process, the persisting disparity between training and validation scores hints at potential overfitting, indicating an overly specialized learning of the training data that compromises its performance on unseen data. Adjustments to the RF model’s hyperparameters have yet to rectify the overfitting dilemma. Additionally, an increase in the test score corresponding to a rise in sample numbers suggests that the dataset’s limited scope could primarily be fueling the overfitting issue. Expanding the dataset is posited to counteract overfitting challenges, thereby enhancing the model’s predictive accuracy.Fig. 6 Learning curves of two regression models. This figure delineates the training progression and validation performance of GBR and RF models, with R2 as the evaluation metric for both. (a) The learning curve of GBR illustrates a notable trend where the model’s performance on the training set starts at a high score of 0.9, gradually decreasing to 0.6 as the dataset size increases. Simultaneously, the performance on the validation set improves from 0 to 0.4, indicating the model’s adaptation to enhance its generalization ability and mitigate overfitting. (b) The learning curve of the RF model shows a slight decline in the training set score from 0.8 to 0.75, while the validation set score increases from 0 to 0.4. However, the gap between the training and validation scores points to potential overfitting issues, suggesting that the model may be too closely fitted to the training data, affecting its performance on new, unseen data.

Feature analysis in prediction

In both the regression and classification prediction tasks, the RF model emerges as the standout performer, prompting a focused analysis of the feature sensitivity within the RF model for both types of tasks.

As depicted in Fig. 7a, the distinct influence of the lighting feature within the RF model’s input features, where it stands out with a relatively high main effect (0.5438). This contrasts sharply with the main effects of other characteristics, which all fall below 0.15. Additionally, the lighting feature also boasts the highest total effect (0.6467), underscoring its paramount importance in influencing the model’s performance. Figure 7b further highlights the main effects of each input feature in the RF model are predominantly low, with most values ranging between 0 and 0.15. Notably, illumination characteristics exhibit the most significant total effect (0.3775), followed by imino groups (0.2265), and benzene ring groups (0.2265), indicating these features’ substantial impact on model predictions.Fig. 7 Sensitivity of RF models to input features. Illustrating the sensitivity of Random Forest (RF) models to various input features in both classification and regression contexts. (a) Sensitivity of the RF model in tasks, showcasing the impact of input features on the model’s precision in predicting DNA cleavage efficiency. (b) Sensitivity of the RF model in regression classification tasks, highlighting how different features influence the model’s ability to accurately categorize DNA cleavage effects.

From the analysis of the RF model in predicting DNA cleavage effects, it can be deduced that no single feature exerts a substantial impact on the prediction outcomes by itself. Rather, it is the synergistic effect of illumination features alongside other attributes that significantly influences the variations in prediction results. This observation suggests a complex interplay of factors, where the presence of light and its interaction with various nanomaterial properties become critical in determining DNA cleavage capabilities. Furthermore, when utilizing the RF model to predict DNA cleavage efficiency, the significance of lighting characteristics is doubly underscored. They not only significantly affect the prediction results on their own but also enhance the model’s predictive accuracy in conjunction with other features. This multifaceted impact of lighting features reflects the nuanced role that environmental conditions play in the interaction between nanomaterials and DNA. This insight corroborates the research findings of Zhang et al.3, who highlighted that nanomaterials demonstrate DNA cleavage effects under specific light conditions. Such an alignment between model-based predictions and empirical research underscores the importance of considering environmental factors, particularly illumination, in studies related to nanomaterial-induced DNA cleavage. This also emphasizes the RF model’s capability to capture and reflect the complex dynamics of DNA-nanomaterial interactions, providing a robust framework for predicting DNA cleavage effects and efficiency with high accuracy.

The importance of each feature within an RF model can be quantified by calculating the average Gini coefficient or the average variance attributed to each feature across all the decision trees in the ensemble. This calculation of feature importance offers a metric for accessing the extent to which each feature contributes to the prediction results, thereby facilitating a deeper understanding and explanation of the prediction process. Moreover, conducting a feature importance analysis serves a dual purpose: it not only elucidates the contribution of individual features to the model’s predictive accuracy but also aids in identifying anomalies within the features. An outlier scenario may arise if a feature, which is theoretically significant, exhibits very low importance in the analysis. Such a discrepancy could signal potential issues within the data collection or processing phases, indicating that the data may not accurately represent the feature’s true influence on the prediction results. Hence, feature importance analysis is not merely a tool for model interpretation; it also functions as a diagnostic mechanism to ensure the integrity and efficacy of the model’s input data. By identifying and rectifying any anomalies detected through this analysis, researchers can enhance the reliability and accuracy of their predictive models, ensuring that they reflect the true dynamics of the phenomena being studied.

As depicted in Fig. 8, the top five factors that significantly influence the prediction outcomes of both the classification and regression models are Light, NPC, PC, RT, and pH value. This highlights that the models heavily weigh the characteristics of experimental conditions in their predictive processes. Notably, the importance assigned to temperature is zero, suggesting its negligible impact on model predictions. This observation could be attributed to the approach taken during the data processing phase, where missing temperature values were imputed using the mode value, reflective of the most common experimental conditions for DNA cleavage observed in the dataset. In the context of the classification model, features related to light are deemed most critical. This prominence of light-related features can be linked to the predominance of photocatalysis-driven DNA cleavage instances within the training dataset. Such a finding underscores the significance of experimental conditions, particularly illumination, in influencing DNA cleavage outcomes, and by extension, the predictive accuracy of models focused on this biological phenomenon.Fig. 8 Feature importance in RF classification and regression models. (a) Classification. (b) Regression. The x-axis represents the feature importance scores assigned to each feature, while the y-axis lists the features of the samples, organized in descending order of importance. The graph highlights the key features contributing to the model’s predictions, with the most influential features at the top and the least at the bottom. NPC: NPs Concentration; PC: Plasmid Concentration; RT: Reaction Time.

Discussion

This study set forth with several objectives aimed at enhancing our understanding of DNA cleavage induced by nanomaterials through a data-driven approach. These objectives were to compile data from various sources related to DNA cleavage properties of nanomaterials, evaluate the distribution of this data and identify informational gaps, assess the viability of employing self-constructed datasets for machine learning predictions, and ascertain the influence of specific features on prediction outcomes. Here, we reflect on the achievements and challenges encountered in pursuit of these goals.

A notable inconsistency was observed in the reporting standards of nanomaterial properties across different studies, complicating efforts to aggregate a coherent dataset. In predictive modeling, the classification models demonstrated superior performance over regression models, with the Deep Neural Network (DNN) model achieving marginally higher accuracy (0.914) compared to the Support Vector Machine (SVM) (0.900) and Random Forest (RF) (0.906) models. Conversely, regression models underperformed, evidenced by an R-squared value below 0.6, likely due to incomplete data. Feature importance and sensitivity analysis underscored the critical roles of illumination conditions and nanomaterial concentration for classification tasks, whereas nanomaterial concentration and reaction time were pivotal for regression tasks. This suggests that while machine learning models are capable of classifying DNA cleavage properties of nanomaterials effectively, regression tasks present considerable challenges under current data conditions.

Among several methods for predicting the DNA cleavage properties of nanomaterials, molecular dynamics simulations offer the advantage of delving into the microscopic mechanisms of material-DNA interactions to predict cleavage sites and efficiency. However, they entail high computational demands, requiring precise force field parameters and complex simulation settings, making them unsuitable for large-scale screening8. Experimental testing provides the benefit of directly assessing the actual cleavage performance of materials, yielding authentic and reliable results. Yet, it involves extensive experimental work and is not conducive to rapid screening of new materials7. Computational chemistry methods have the advantage of predicting material-DNA interactions and cleavage mechanisms from a quantum mechanics perspective. Nevertheless, they are computationally intensive, demanding substantial computational resources, and the correlation with experimental results requires further validation. Data-driven models offer the advantage of swiftly and cost-effectively predicting the cleavage properties of new materials once a reliable predictive model is established. However, they necessitate a substantial amount of reliable experimental data as training samples, making the modeling process complex.

The study also highlighted several systemic issues at the data level, including a scarcity of comprehensive reports on newly discovered nanomaterials and the absence of standardized reporting protocols. These deficiencies underline the urgent need for concerted efforts to develop and adopt uniform data reporting standards to bolster the utility of databases in advancing nanomaterial research.

At the dataset level, limitations include the small size and incomplete nature of the dataset, which hinder the training of highly accurate machine learning models. Moreover, the reliance on gel electrophoresis strip images for analyzing cleavage efficiency introduces potential measurement inaccuracies. Additionally, the complexity of quantifying DNA cleavage mechanisms poses a significant obstacle, making it difficult to incorporate such mechanisms as input features for predictive models. The limited size and diversity of the manually collected database further exacerbate the challenges in predicting DNA cleavage properties of novel nanomaterials.

While this study has made strides in employing machine learning tools for predicting DNA cleavage effects of nanomaterials, it also sheds light on the substantial hurdles that remain. Addressing these challenges, particularly those related to data standardization and completeness, is crucial for advancing the field and enhancing the predictability of nanomaterial interactions with DNA.

Materials and methods

Data compilation

Given the multifaceted nature of nanomaterial types, their challenging synthesis processes, and the intricate mechanisms underlying DNA cleavage, it is impractical to solely rely on experimental data for generating comprehensive datasets15. DNA cleavage is modulated by a wide array of factors, including the nanomaterial’s structure, surface functional groups, particle size, surface properties, crystal structure, lattice defects, composition, surface modification, functionalization, and environmental conditions. In response to this challenge, a meticulously curated dataset was compiled by extracting information from published articles.

The dataset utilized in this study was procured through an exhaustive literature review of studies investigating the DNA cleavage effects of various classes of nanomaterials. This dataset incorporates 18 input features that detail various nanomaterial properties, including molecular weight, shape, particle size, water solubility, zeta potential, and functional groups. Additionally, it comprises 8 input features that characterize exposure conditions, such as nanoparticle concentration, light exposure, temperature, and pH value. Lastly, the dataset contains 2 output tags that categorize the DNA cleavage effect and efficiency of nanomaterials’ DNA cleavage properties. The input features that characterize the properties of nanomaterials encompass shape, particle size, water solubility, zeta potential, surface area, and functional groups. These features were explicitly outlined in the articles, whereas the molecular weight was either directly furnished or inferred using structural formula and spectral data, including FT-IR, 1HNMR, mass spectra, and elemental analysis, as documented in the publication16. Regarding the input features pertaining to exposure conditions, the articles supplied information on nanoparticle concentration, plasmid type, plasmid concentration, light exposure, temperature, pH, exposure time, and the presence of additional substrates. The output labels, which elucidate the DNA cleavage properties of nanomaterials, were ascertained through qualitative and quantitative analysis of gel electrophoresis band images presented in the publications. The cleavage effect was assessed by analyzing the size, number, and distribution of DNA fragments visible in the gel electrophoresis bands17. Similarly, the cleavage efficiency was quantified by measuring the band intensities using ImageJ software18. The URL of ImageJ software used in the study is https://imagej.net/software/fiji/, and the version is 2.35.

This dataset stands to serve other researchers by enabling the training models, augmentation of the dataset, and bridging any data gaps through sophisticated grouping methods. Such initiatives are anticipated to significantly improve the predictive efficiency of models concerning the DNA cleavage effects of nanomaterials.

Data processing

To guarantee the data’s appropriateness for model training, minimize the influence of biases and errors, and enhance the accuracy and reliability of the models, a thorough data cleaning procedure was implemented.

Table 1 delineates a comparison between the original and the processed datasets. Specific attributes, such as molecular weight, shape, particle size, and water solubility, were omitted from the dataset due to the absence of values. Missing information on illumination conditions was substituted with “no illumination” conditions. For materials exhibiting DNA shearing properties under light conditions, based on researchers’ insights, an assumption is made that these properties are absent under no-light conditions, and the data is supplemented accordingly. This approach aims to reduce model bias resulting from imbalanced sample classes. Furthermore, absent data for temperature, pH, and reaction time were supplemented with the mean value for numeric features and the mode value for categorical features, enhancing data integrity19. To facilitate the inclusion of categorical features in regression models, a one-hot encoding strategy20 was applied. Moreover, the normalization technique21 was utilized to scale the values, effectively minimizing magnitude bias while maintaining the proportional distinctions within the value ranges. After data processing, the data set partitioning was performed.Table 1 Summary of input and target data and processing methods.

Category	Variables	Type	Original data	Processed data	Application	
Min–max/labels	Missing (%)	Data processing	Min–max/labels	
Feature of NPs	Benzene	Nominal	1, 0	0	No processing	1, 0	Input	
R–OH	1, 0	0	1, 0	
R–O–R	1, 0	0	1, 0	
R=O	1, 0	0	1, 0	
R–COOH	1, 0	0	1, 0	
R–NO2	1, 0	0	1, 0	
R–SH	1, 0	0	1, 0	
R=N–R	1, 0	0	1, 0	
R–NH–R	1, 0	0	1, 0	
R–NH2	1, 0	0	1, 0	
Exposure condition	NPs concentration	Numeric	3.13*10–5–50, NaN	0.7	Drop NaN	3.13*10–5–50	Input	
Plasmid concentration	0.008–60, NaN	4.0	0.008–60	
Temperature	20–37 (°C), NaN	56.5	Majority filling	20–37 (°C)	
pH	6–9, NaN	34.6	6–9	
Reaction time	0.08–24(h)	23.0	0.08–24(h)	
Light	Nominal	1, 0, NaN	40.2	Augmentation	1, 0	
Plasmid type	pUC19, pBR322, pSP72, etc., NaN	0	One-hot encoding	pUC19, pBR322, pSP72, etc.,	
Effect	Cutting effect	Nominal	1, 0	0	No processing	Select	Target	
Cutting efficiency	Numeric	0–1	0	Select	
The binary encoding of “1” and “0” denotes the presence and absence, respectively, of specific functional groups in nanoparticles. “Drop NaN” refers to the exclusion of samples containing null values from the database. “Majority Filling” describes the imputation of missing values with the most frequently occurring value within the dataset. “Augmentation” is the process of enriching the dataset based on the principle that nanomaterials do not induce DNA cleavage in the absence of light exposure. “One-hot Encoding” is the transformation of categorical variables into a binary vector representation, facilitating their integration into machine learning models.

Figure 9 illustrates the distribution of original data labels used for model training. From Fig. 9, it can be observed that out of 496 samples, approximately 200 samples have a cleavage efficiency lower than 0.2, while around 300 samples have an efficiency higher than 0.2. This indicates that after data augmentation, the imbalance in the data is not prominent.Fig. 9 Probability density distribution of cutting efficiency in raw data. The horizontal axis represents cutting efficiency, while the histogram and curve represent sample count (vertical axis on the left) and Kernel Density Estimation (KDE) density (vertical axis on the right), respectively.

The processed dataset was partitioned into a training set, accounting for 80% of the dataset, designated for model training and performance evaluation, and a test set, comprising the remaining 20%, utilized for assessing the models’ predictive accuracy22.

Model selection

Regression models and classification models constitute two fundamental types of predictive models in machine learning. A regression model is designed to predict the values of one or more continuous target variables based on input features, making it ideal for predicting quantitative outcomes. Conversely, a classification model is engineered to assign instances into distinct classes based on input features, yielding discrete variables as outputs that signify the respective categories, suitable for categorical outcomes. Within the scope of this study, a classification model was developed to determine the cleavage effect, whereas a regression model was deployed to estimate the cleavage efficiency.

Various machine learning models are applicable for both classification and regression tasks, including DNN23, SVM24, and RF25, among others. DNN stands as a robust model that emulates the human nervous system, featuring a complex network of interconnected neurons. These neurons can be meticulously trained to modify weights and biases for either classification or regression tasks, making DNN exceptionally adept at managing nonlinear problems. SVM, conversely, is a supervised learning algorithm utilized for both classification and regression analyses. It identifies an optimal hyperplane by projecting the dataset into a high-dimensional space to distinguish between different categories of data. SVM excels in scenarios involving high-dimensional data and is capable of achieving high accuracy even with limited sample sizes. RF is an ensemble learning algorithm that constructs multiple decision trees and aggregates their predictions through random feature selection and sampling. RF has demonstrated strong performance in both classification and regression tasks, providing numerous advantages including the capacity to manage high-dimensional data with diverse features, resilience to outliers and missing values, and the capability to assess feature importance. Additionally, other models are noteworthy for their efficacy in regression tasks, such as the Bayesian Ridge Regression Model (BR)26, Linear Regression Model (LR)27, and Gradient Boosted Regression Model (GBR)28, which exhibit excellent performance in regression analysis.

In this study, classification models were developed using DNN, SVM, and RF algorithms, whereas regression models were constructed using DNN, SVR, RF, BR, LR, and GBR algorithms.

Model evaluation

In order to reveal the performance of the models, evaluation metrics for the prediction capability and evaluation methods for the learning process were adopted29. Four evaluation metrics were utilized to gauge the performance of the classification models: accuracy, precision, recall, and F1 score. Accuracy represents a commonly adopted metric quantifying the proportion of samples correctly classified by the model. Precision measures the fraction of samples identified as the positive class that are True Positive (TP) instances, whereas recall determines the proportion of TP instances correctly predicted as the positive class. The F1 score, which is the harmonic mean of precision and recall, provides a holistic evaluation of the model by incorporating both precision and recall. The aforementioned metrics could be used to evaluate classification models but are not suitable for assessing regression models. Different metrics are required for evaluating regression models.

In addition, another four metrics were utilized to evaluate the performance of the regression model: Explained Variance (EV), Mean Absolute Error (MAE), Root Mean Squared Error (RMSE), and Coefficient of Determination (R2). EV quantifies the proportion of variance in the dependent variable that is predictable from the independent variables. A higher EV signifies a model’s enhanced capability to elucidate the variance in the predicted values. MSE, computed as the average of the squared differences between the predicted and actual values, offers an indication of the prediction accuracy, with a lower MSE denoting higher model precision. MAE, the average of the absolute differences between predicted and actual values, provides insight into the average error magnitude. R2, a statistical measure of the fit quality, indicates the proportion of the dependent variable’s variance that is explained by the independent variables in the model. A greater R2 value suggests a better model fit and more substantial explanatory power. The analysis of model performance could be conducted not only based on the prediction results but also by examining the prediction process itself.

For visualizing the training performance of the model across varying dataset sizes, we also adopted a crucial tool, the learning curve, assessed through the fivefold cross-validation. This method facilitates an in-depth assessment of the model’s efficacy in data-driven scenarios. Throughout the evaluation phase, the same training cessation strategy implemented during the model’s training is applied. For instance, training for the GBR model ceases after 100 iterations, whereas the RF model relies on the default stopping strategy inherent to random forests.

Model sensitivity

To reveal key factors within feature data that significantly influence both classification and regression tasks30, an in-depth analysis of the sensitivity of the random forest model to input features was conducted, alongside the computation of feature importance. Sensitivity analysis serves as a crucial instrument for assessing parameter importance, exploring parameter interactions, and evaluating model reliability and robustness. The widely utilized Sobol’s method31, a global sensitivity analysis approach, was employed in this study. Sobol’s method is a variance-based Monte Carlo approach that involves decomposing the model into a function consisting of single and multiple input variables. The sensitivity coefficients are then calculated to determine the influence of the variance of individual or multiple input variables on the overall output variance. To investigate the sensitivity of the random forest model towards input features, we engaged in both single-factor and multi-factor sensitivity analyses. The main effect is utilized as an indicator for single-factor sensitivity, aiming to gauge the impact of individual features independently. Conversely, the total effect is leveraged for multi-factor sensitivity analysis, designed to evaluate the combined effects of multiple features on the model’s performance. For this analysis, the processed dataset is utilized as a sample, with the variation range for parameter sampling defined by the maximum and minimum values observed in the features.

Author contributions

Conceptualization: J.N., Y.Z. Methodology: J.N., X.W., Y.Z., X.C. Investigation: X.W., B.Y., J.C. Visualization: X.W. Supervision: J.N., N.L., P.W. Writing—original draft: X.W., Writing—review and editing: J.N., Y.Z.

Funding

National Natural Science Foundation of China [41972244]; The Project of Science and Technology Department of Guizhou Province (the technical system of prevention and control to mine groundwater pollution in karst areas, 2022); Hong Kong HKSTP & HKBU Joint Innovation Fund [IRF24-115]; Guangdong University Key Research Project [2024ZDZX2095]; Zhuhai Basic and Applied Basic Research Foundation [2320004002479]; UIC Start-up Research Fund (No. UICR0700084-24); The High-Level Talent Training Program in Guizhou Province (GCC[2023]045).

Data availability

The DNA cleavage data and machine learning code supporting the findings of this study have been deposited for future reference and reproducibility. The code for this study is stored in figshare, 10.6084/m9.figshare.25556961. The dataset for this study is also stored in figshare, 10.6084/m9.figshare.25556955.

Declarations

Competing interests

The authors declare no competing interests.

Publisher’s note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

These authors contributed equally: Jie Niu and Xufeng Wang.
==== Refs
References

1. Chen Y Zhao D Liu Y Polysaccharide-porphyrin-fullerene supramolecular conjugates as photo-driven DNA cleavage reagents Chem. Commun. 2015 51 12266 12269 10.1039/C5CC04625D
Chen, Y., Zhao, D. & Liu, Y. Polysaccharide-porphyrin-fullerene supramolecular conjugates as photo-driven DNA cleavage reagents. Chem. Commun. 51, 12266–12269 (2015).
2. Rozhina E Comparative cytotoxicity of kaolinite, halloysite, multiwalled carbon nanotubes and graphene oxide Appl. Clay Sci. 2021 205 106041 10.1016/j.clay.2021.106041
Rozhina, E. et al. Comparative cytotoxicity of kaolinite, halloysite, multiwalled carbon nanotubes and graphene oxide. Appl. Clay Sci. 205, 106041 (2021).
3. Zhang J Wu S Ma L Wu P Liu J Graphene oxide as a photocatalytic nuclease mimicking nanozyme for DNA cleavage Nano Res. 2020 13 455 460 10.1007/s12274-020-2629-8
Zhang, J., Wu, S., Ma, L., Wu, P. & Liu, J. Graphene oxide as a photocatalytic nuclease mimicking nanozyme for DNA cleavage. Nano Res. 13, 455–460 (2020).
4. Wang X DNA damage caused by light-driven graphene oxide: A new mechanism Environ. Sci. Nano 2023 10 519 527 10.1039/D2EN00948J
Wang, X. et al. DNA damage caused by light-driven graphene oxide: A new mechanism. Environ. Sci. Nano 10, 519–527 (2023).
5. Champa-Bujaico E Garcia-Diaz P Diez-Pascual AM Machine learning for property prediction and optimization of polymeric nanocomposites: A state-of-the-art Int. J. Mol. Sci. 2022 23 10712 10.3390/ijms231810712 36142623
Champa-Bujaico, E., Garcia-Diaz, P. & Diez-Pascual, A. M. Machine learning for property prediction and optimization of polymeric nanocomposites: A state-of-the-art. Int. J. Mol. Sci. 23, 10712 (2022).36142623
6. Singh AV Navigating regulatory challenges in molecularly tailored nanomedicine Explor. BioMat X 2024 1 124 134 10.37349/ebmx.2024.00009
Singh, A. V. et al. Navigating regulatory challenges in molecularly tailored nanomedicine. Explor. BioMat X 1, 124–134 (2024).
7. Dao M Lu L Asaro RJ De Hosson JTM Ma E Toward a quantitative understanding of mechanical behavior of nanocrystalline metals Acta Mater. 2007 55 4041 4065 10.1016/j.actamat.2007.01.038
Dao, M., Lu, L., Asaro, R. J., De Hosson, J. T. M. & Ma, E. Toward a quantitative understanding of mechanical behavior of nanocrystalline metals. Acta Mater. 55, 4041–4065 (2007).
8. Mathew K Sundararaman R Letchworth-Weaver K Arias TA Hennig RG Implicit solvation model for density-functional study of nanocrystal surfaces and reaction pathways J. Chem. Phys. 2014 140 084106 10.1063/1.4865107 24588147
Mathew, K., Sundararaman, R., Letchworth-Weaver, K., Arias, T. A. & Hennig, R. G. Implicit solvation model for density-functional study of nanocrystal surfaces and reaction pathways. J. Chem. Phys. 140, 084106 (2014).24588147
9. Prasad KRKV Machine learning algorithms are applied in nanomaterial properties for nanosecurity J. Nanomater. 2022 2022 1 14 10.1155/2022/5450826
Prasad, K. R. K. V. et al. Machine learning algorithms are applied in nanomaterial properties for nanosecurity. J. Nanomater. 2022, 1–14 (2022).
10. Singh AV Artificial intelligence and machine learning empower advanced biomedical material design to toxicity prediction Adv. Intell. Syst. 2020 2 2000084 10.1002/aisy.202000084
Singh, A. V. et al. Artificial intelligence and machine learning empower advanced biomedical material design to toxicity prediction. Adv. Intell. Syst. 2, 2000084 (2020).
11. Singh AV Artificial intelligence and machine learning disciplines with the potential to improve the nanotoxicology and nanomedicine fields: A comprehensive review Arch. Toxicol. 2023 97 963 979 10.1007/s00204-023-03471-x 36878992
Singh, A. V. et al. Artificial intelligence and machine learning disciplines with the potential to improve the nanotoxicology and nanomedicine fields: A comprehensive review. Arch. Toxicol. 97, 963–979 (2023).36878992
12. Fernandez M Bilic A Barnard AS Machine learning and genetic algorithm prediction of energy differences between electronic calculations of graphene nanoflakes Nanotechnology 2017 28 38LT03 10.1088/1361-6528/aa82e5 28752822
Fernandez, M., Bilic, A. & Barnard, A. S. Machine learning and genetic algorithm prediction of energy differences between electronic calculations of graphene nanoflakes. Nanotechnology 28, 38LT03 (2017).28752822
13. Wang X Li F Teng Y Ji C Wu H Characterization of oxidative damage induced by nanoparticles via mechanism-driven machine learning approaches Sci. Total Environ. 2023 871 162103 10.1016/j.scitotenv.2023.162103 36764549
Wang, X., Li, F., Teng, Y., Ji, C. & Wu, H. Characterization of oxidative damage induced by nanoparticles via mechanism-driven machine learning approaches. Sci. Total Environ. 871, 162103 (2023).36764549
14. Mirzaei M Furxhi I Murphy F Mullins M Employing supervised algorithms for the prediction of nanomaterial’s antioxidant efficiency Int. J. Mol. Sci. 2023 24 2792 10.3390/ijms24032792 36769135
Mirzaei, M., Furxhi, I., Murphy, F. & Mullins, M. Employing supervised algorithms for the prediction of nanomaterial’s antioxidant efficiency. Int. J. Mol. Sci. 24, 2792 (2023).36769135
15. Murugadoss S Identifying nanodescriptors to predict the toxicity of nanomaterials: A case study on titanium dioxide Environ. Sci. Nano 2021 8 580 590 10.1039/D0EN01031F
Murugadoss, S. et al. Identifying nanodescriptors to predict the toxicity of nanomaterials: A case study on titanium dioxide. Environ. Sci. Nano 8, 580–590 (2021).
16. Patel MB Novel cationic fullerene derivatized s-triazine scaffolds as photoinduced DNA cleavage agents: Design, synthesis, biological evaluation and computational investigation RSC Adv. 2013 3 8734 8746 10.1039/c3ra40950c
Patel, M. B. et al. Novel cationic fullerene derivatized s-triazine scaffolds as photoinduced DNA cleavage agents: Design, synthesis, biological evaluation and computational investigation. RSC Adv. 3, 8734–8746 (2013).
17. Lebedová J Hedberg YS Odnevall Wallinder I Karlsson HL Size-dependent genotoxicity of silver, gold and platinum nanoparticles studied using the mini-gel comet assay and micronucleus scoring with flow cytometry Mutagenesis 2018 33 77 85 10.1093/mutage/gex027 29529313
Lebedová, J., Hedberg, Y. S., Odnevall Wallinder, I. & Karlsson, H. L. Size-dependent genotoxicity of silver, gold and platinum nanoparticles studied using the mini-gel comet assay and micronucleus scoring with flow cytometry. Mutagenesis 33, 77–85 (2018).29529313
18. Rueden CT Image J2: ImageJ for the next generation of scientific image data BMC Bioinform. 2017 18 529 10.1186/s12859-017-1934-z
Rueden, C. T. et al. Image J2: ImageJ for the next generation of scientific image data. BMC Bioinform. 18, 529 (2017).
19. Rafsunjani S Safa RS Imran AA Rahim S Nandi D An empirical comparison of missing value imputation techniques on APS failure prediction Int. J. Inf. Technol. Comput. Sci. 2019 11 21 29
Rafsunjani, S., Safa, R. S., Imran, A. A., Rahim, S. & Nandi, D. An empirical comparison of missing value imputation techniques on APS failure prediction. Int. J. Inf. Technol. Comput. Sci. 11, 21–29 (2019).
20. Yu L Zhou R Chen R Lai KK Missing Data preprocessing in credit classification: One-hot encoding or imputation? Emerg. Mark. Finance Trade 2022 58 472 482 10.1080/1540496X.2020.1825935
Yu, L., Zhou, R., Chen, R. & Lai, K. K. Missing Data preprocessing in credit classification: One-hot encoding or imputation?. Emerg. Mark. Finance Trade 58, 472–482 (2022).
21. Jo J-M Effectiveness of normalization pre-processing of big data to the machine learning performance J. Korea Inst. Electron. Commun. Sci. 2019 14 547 552
Jo, J.-M. Effectiveness of normalization pre-processing of big data to the machine learning performance. J. Korea Inst. Electron. Commun. Sci. 14, 547–552 (2019).
22. Yousef WA Kundu S Learning algorithms may perform worse with increasing training set size: Algorithm–data incompatibility Comput. Stat. Data Anal. 2014 74 181 197 10.1016/j.csda.2013.05.021
Yousef, W. A. & Kundu, S. Learning algorithms may perform worse with increasing training set size: Algorithm–data incompatibility. Comput. Stat. Data Anal. 74, 181–197 (2014).
23. Schmidhuber J Deep learning in neural networks: An overview Neural Netw. 2015 61 85 117 10.1016/j.neunet.2014.09.003 25462637
Schmidhuber, J. Deep learning in neural networks: An overview. Neural Netw. 61, 85–117 (2015).25462637
24. Chang C-C Lin C-J LIBSVM: A library for support vector machines ACM Trans. Intell. Syst. Technol. 2011 2 1 27 10.1145/1961189.1961199
Chang, C.-C. & Lin, C.-J. LIBSVM: A library for support vector machines. ACM Trans. Intell. Syst. Technol. 2, 1–27 (2011).
25. Strobl C Malley J Tutz G An introduction to recursive partitioning: Rationale, application, and characteristics of classification and regression trees, bagging, and random forests Psychol. Methods 2009 14 323 348 10.1037/a0016973 19968396
Strobl, C., Malley, J. & Tutz, G. An introduction to recursive partitioning: Rationale, application, and characteristics of classification and regression trees, bagging, and random forests. Psychol. Methods 14, 323–348 (2009).19968396
26. Firinguetti-Limone L Pereira-Barahona M Bayesian estimation of the shrinkage parameter in ridge regression Commun. Stat. Simul. Comput. 2020 49 3314 3327 10.1080/03610918.2018.1547395
Firinguetti-Limone, L. & Pereira-Barahona, M. Bayesian estimation of the shrinkage parameter in ridge regression. Commun. Stat. Simul. Comput. 49, 3314–3327 (2020).
27. Su X Yan X Tsai C Linear regression WIREs Comput. Stat. 2012 4 275 294 10.1002/wics.1198
Su, X., Yan, X. & Tsai, C. Linear regression. WIREs Comput. Stat. 4, 275–294 (2012).
28. Konstantinov AV Utkin LV Interpretable machine learning with an ensemble of gradient boosting machines Knowl. Based Syst. 2021 222 106993 10.1016/j.knosys.2021.106993
Konstantinov, A. V. & Utkin, L. V. Interpretable machine learning with an ensemble of gradient boosting machines. Knowl. Based Syst. 222, 106993 (2021).
29. Zhou J Gandomi AH Chen F Holzinger A Evaluating the quality of machine learning explanations: A survey on methods and metrics Electronics 2021 10 593 10.3390/electronics10050593
Zhou, J., Gandomi, A. H., Chen, F. & Holzinger, A. Evaluating the quality of machine learning explanations: A survey on methods and metrics. Electronics 10, 593 (2021).
30. Novello P Poëtte G Lugato D Congedo PM Goal-oriented sensitivity analysis of hyperparameters in deep learning J. Sci. Comput. 2023 94 45 10.1007/s10915-022-02083-4
Novello, P., Poëtte, G., Lugato, D. & Congedo, P. M. Goal-oriented sensitivity analysis of hyperparameters in deep learning. J. Sci. Comput. 94, 45 (2023).
31. Dimov I Georgieva R Ostromsky TZ Monte Carlo sensitivity analysis of an Eulerian large-scale air pollution model Reliab. Eng. Syst. Saf. 2012 107 23 28 10.1016/j.ress.2011.06.007
Dimov, I., Georgieva, R. & Ostromsky, T. Z. Monte Carlo sensitivity analysis of an Eulerian large-scale air pollution model. Reliab. Eng. Syst. Saf. 107, 23–28 (2012).
