
==== Front
Hum Genomics
Hum Genomics
Human Genomics
1473-9542
1479-7364
BioMed Central London

39256852
667
10.1186/s40246-024-00667-9
Research
AI-derived comparative assessment of the performance of pathogenicity prediction tools on missense variants of breast cancer genes
https://orcid.org/0000-0002-7531-5264
Ahmad Rahaf M. 1
https://orcid.org/0000-0003-1306-6618
Ali Bassam R. 2
https://orcid.org/0000-0002-0792-9598
Al-Jasmi Fatma 13
https://orcid.org/0000-0002-1384-2915
Al Dhaheri Noura 13
https://orcid.org/0000-0001-7017-336X
Al Turki Saeed 1
https://orcid.org/0000-0003-2687-6476
Kizhakkedath Praseetha 2
https://orcid.org/0000-0002-1079-4559
Mohamad Mohd Saberi saberi@uaeu.ac.ae

14
1 https://ror.org/01km6p862 grid.43519.3a 0000 0001 2193 6666 Health Data Science Lab, Department of Genetics and Genomics, College of Medical and Health Sciences, United Arab Emirates University, Tawam road, Al Maqam district, Al Ain, Abu Dhabi United Arab Emirates
2 https://ror.org/01km6p862 grid.43519.3a 0000 0001 2193 6666 Department of Genetics and Genomics, College of Medical and Health Sciences, United Arab Emirates University, Tawam road, Al Maqam district, Al Ain, Abu Dhabi United Arab Emirates
3 https://ror.org/007a5h107 grid.416924.c 0000 0004 1771 6937 Division of Metabolic Genetics, Department of Pediatrics, Tawam Hospital, Al Ain, United Arab Emirates
4 https://ror.org/04zrbnc33 grid.411865.f 0000 0000 8610 6308 Center for Engineering Computational Intelligence, Faculty of Engineering and Technology, Multimedia University, Melaka, Malaysia
11 9 2024
11 9 2024
2024
18 9918 3 2024
22 8 2024
© The Author(s) 2024
2024
https://creativecommons.org/licenses/by-nc-nd/4.0/ Open Access This article is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License, which permits any non-commercial use, sharing, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if you modified the licensed material. You do not have permission under this licence to share adapted material derived from this article or parts of it. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by-nc-nd/4.0/.
Single nucleotide variants (SNVs) can exert substantial and extremely variable impacts on various cellular functions, making accurate predictions of their consequences challenging, albeit crucial especially in clinical settings such as in oncology. Laboratory-based experimental methods for assessing these effects are time-consuming and often impractical, highlighting the importance of in-silico tools for variant impact prediction. However, the performance metrics of currently available tools on breast cancer missense variants from benchmarking databases have not been thoroughly investigated, creating a knowledge gap in the accurate prediction of pathogenicity. In this study, the benchmarking datasets ClinVar and HGMD were used to evaluate 21 Artificial Intelligence (AI)-derived in-silico tools. Missense variants in breast cancer genes were extracted from ClinVar and HGMD professional v2023.1. The HGMD dataset focused on pathogenic variants only, to ensure balance, benign variants for the same genes were included from the ClinVar database. Interestingly, our analysis of both datasets revealed variants across genes with varying penetrance levels like low and moderate in addition to high, reinforcing the value of disease-specific tools. The top-performing tools on ClinVar dataset identified were MutPred (Accuracy = 0.73), Meta-RNN (Accuracy = 0.72), ClinPred (Accuracy = 0.71), Meta-SVM, REVEL, and Fathmm-XF (Accuracy = 0.70). While on HGMD dataset they were ClinPred (Accuracy = 0.72), MetaRNN (Accuracy = 0.71), CADD (Accuracy = 0.69), Fathmm-MKL (Accuracy = 0.68), and Fathmm-XF (Accuracy = 0.67). These findings offer clinicians and researchers valuable insights for selecting, improving, and developing effective in-silico tools for breast cancer pathogenicity prediction. Bridging this knowledge gap contributes to advancing precision medicine and enhancing diagnostic and therapeutic approaches for breast cancer patients with potential implications for other conditions.

Keywords

Single nucleotide variants (SNVs)
Breast cancer
Pathogenicity prediction tools
Health data science
Artificial intelligence (AI)
United Arab Emirates UniversityStrategic Research Program (#12R111) issue-copyright-statement© BioMed Central Ltd., part of Springer Nature 2024
==== Body
pmcIntroduction

Genomic medicine implementation and research have recently been increasingly relying on comprehensive sequencing, making a shift from targeted or single gene testing approaches to whole exome or whole genome. The advent of next generation DNA sequencing allowed for the identification of hundreds of thousands of rare variants within the genomes of individuals. However, interpreting the functional effects of these variants, especially novel missense variants, presents challenges that can lead to inconclusive genomic clinical reports and uncertainties for clinicians as well as patients and their families. Consequently, it is currently impractical for researchers to study the effects of every possible variant within the approximately 20,000 genes of the human genome [1]. Therefore, there is a pressing need for novel clinical-grade approaches to assist in determining the pathogenicity of variants including missense variants.

One promising approach is the use of machine learning, which has led to the development of various tools for predicting pathogenicity including REVEL [2], Polyphen [3] and many other well-known tools [4]. These tools utilize variant features and pre-assigned labels to predict whether a variant is pathogenic or benign. Databases such as the Human Gene Mutation Database (HGMD), the Leiden Open Variation Database, ClinVar, and the Genome Aggregation Database (gnomAD) offer annotated variants for training and testing pathogenicity predictors and classifiers.

The increasing numbers of identified single nucleotide variants (SNVs), has been steadily rising due to the increasing use of next-generation sequencing (NGS) and SNVs array technologies [5]. Despite genome-wide association studies (GWAS) and candidate gene association studies presents data about SNVs [6], difficulties persist in accurately characterizing and predicting the functional consequences of these genetic variants [7]. Experimental methods for functional classification such as enzyme activity assays, protein-protein interaction assays, reporter gene assays, and cell-based functional assays are time-consuming and very costly [8]. These limitations led to a shift towards computational predictor tools and algorithms that are based on biochemical and biological data [9]. The growing submission of disease-related variations to publicly accessible databases warrant the development of accurate functional prediction tools [10]. Understanding the functional consequences of molecular pathway changes and disease etiology is critical for understanding the causes and prognosis of numerous disorders, including cancer [11]. A recent review by Ahmad et al. [4] covered the biological and computational aspects of predicting the pathogenicity of breast cancer. Despite the use of NGS for genome-wide or panel gene sequencing of cancer-related genes, the relationship between cancer phenotypes and their underlying molecular causes remains unknown [12].

One of the most common genomic variations in human genomes are SNVs, many of which have influence biological functions including gene expression, susceptibility or causation of disease, and protein interactions [13]. SNVs can have positive, neutral, or negative (pathogenic) impacts on phenotypic outcomes [14]. Benign and neutral variations typically result in mild to manageable or even advantageous outcomes, whereas most pathogenic variations have detrimental effects on individuals, leading or contributing to deteriorated health status and decreased state of well-being [15]. Therefore, benign variants are often found at high frequencies among populations and are transmitted through generations [16].

For the interpretation and categorization of variations, several country-specific consensus criteria have been established such as the American College of Medical Genetics (ACMG) recommendations [17]. The Association for Molecular Pathology (AMP), American Society of Clinical Oncology (ASCO), and College of American Pathologists (CAP) have all developed clinical interpretation guidelines for cancer-related variations [18]. As a major categorization, reflecting their occurrence timeline, cancer variations are divided into somatic and germline variants, which are determined primarily whether they are acquired or are present at birth as well as their prevalence in the body cells [19]. Somatic variations (that are acquired) have garnered more attention in clinical cancer than germline variants, owing to their greater influence on medication responses and treatment outcomes [20]. Whereas germline mutations are more relevant to detecting susceptibility to cancer. However, according to recent studies, both somatic and germline variations and their combinations play important roles in medication sensitivity, toxicity, and selection, impacting cancer development, progression and prognosis. These mutations are rapidly being acknowledged as important inputs to our understanding of cancer development [21]. Another method of categorizing cancer variations is to categorize them as “driver” or “passenger” mutations. Driver variations are those that provide tumor cells a biological advantage, increasing tumor growth. Passenger variations, on the other hand, do not directly contribute to tumor development but are a consequence of increased mutation rates and loss of control over the tumor such as loss of DNA repair [20].

To predict the impact and categorize variations in clinical settings, ACMG, AMP, ASCO and CAP encouraged the use of in-silico techniques and artificial intelligence (AI) algorithms. These in-silico prediction algorithms and tools use a variety of techniques and approaches [22]. With the availability of large amounts of data generated by NGS, the need for curated variant databases has been highlighted [23]. As a result, various variation databases have been built to assist variant collection, annotation, curation, and categorization and are used by bioinformaticians, physicians, and experimentalists [24]. Open variation databases, such as ClinVar [25] and gnomAD [26], have gathered a large number of disease-causing variations, including those related to cancer. ClinVar [25] stands out as one of the most extensive variation databases.

The data collection approach of ClinVar involves gathering variant submissions supported by clinical or experimental observations. It then employs a cumulative interpretation methodology to assess the level of agreement or disagreement among submitters [27]. It is important to note, however, that while many cancer variations have been deposited to this database, a large number of them need final interpretation and categorization [28]. Additionally, HGMD provides a comprehensive set of published germline variants in genes that are thought to underlie or are closely associated with human inherited disease including cancer. In clinical contexts, misinterpretation of cancer variations leads to various obstacles, including inaccurate clinical diagnosis, a faulty understanding of variant harmful consequences, and obsolete submissions [29]. Furthermore, human variation databases contain a large number of variations whose clinical consequences, interpretation, or categorization are uncertain, leading to their classification as “variants of unknown significance” (VUS) [30].

The performance of in-silico prediction tools has been evaluated using a variety of variant datasets [15, 27, 30–33]. However, due to the large number and continuous expansion of variations and prediction methods involved, these evaluations create significant obstacles. Overfitting of variant interpretation data is a key concern, as AI algorithms in prediction tools are frequently trained using duplicate data [27]. Furthermore, the performance of a tool might differ greatly based on the type of dataset used for training [15]. Researchers have recommended using datasets collected from online variant databases such as gnomAD [26], dbSNP [34], the 1000 Genomes Project [35] HGMD [36] and ClinVar [25] to address both over-fitting and variability in tool performance. These datasets have also been categorized based on their intended uses, such as effect-specific datasets, molecule-specific datasets, and disease-specific datasets [37]. Disease-specific datasets concentrating on cancer-related variants have been curated for cancer-related studies [38]. However, cancer-specific datasets may not contain a high number of genes and variants due to a lack of experimental verification of the functional impact of cancer variations [39, 40].

Several studies have been conducted on cancer variants [41, 42]. However, to the best of our knowledge, none of these studies have specifically focused on assessing the functional consequence of comprehensive breast cancer variants extracted from ClinVar [25] and HGMD [36] by employing various in-silico pathogenicity tools. Therefore, our aim was to evaluate the performance of AI-derived in-silico pathogenicity prediction tools to investigate the functional consequences of breast cancer-related variants from ClinVar [25] and HGMD [36] that are known to have clinical significance. We used a total of 21 different in-silico tools to evaluate the performance on the datasets to gain insights into their effectiveness in accurately discriminating between pathogenic and benign variants using several metrics.

Methods

Extraction of variants and dataset generation

The dataset in this study was created in July 2023 by extraction from the ClinVar database [25] and HGMD professional v2023.1 [36]. Primarily, a comprehensive list of breast cancer related genes was collected through a literature review, those genes were used as keywords to look for variants related to breast cancer. For the ClinVar dataset, the keywords “breast” and “cancer” were used as disease/phenotype feature in the advanced search engine in addition to “gene name” in which each gene was individually used. Variants with the classification of Pathogenic or Likely Pathogenic were included in Pathogenic category of the dataset, and variants with the classification of Benign or Likely Benign were included in the Benign category of the dataset. All other classifications were excluded from the dataset. For the HGMD dataset, the phenotype field was filtered to include the words “breast” and “cancer” and similarly, the genes were further filtered based on the relatedness of the gene on breast cancer and vice versa. For this study, only missense single nucleotide variants were included. Through the literature review, the predisposition status of the genes was collected. According to the data presented along with each gene’s information, genes associated with low and moderate penetrance predisposition exhibit a significant number of variants. This implies that if a tool’s training data encompasses a comprehensive gene list specifically linked to a particular disease, the tool will possess a higher level of specificity and quality than generalized or overly specific tools. The schematic pathways representation covered by the top genes in the evaluated dataset is shown if Fig. 1. Additionally, the genes that were found to have breast cancer variants, the number of variants in each gene and the predisposition of each gene are shown in Fig. 2.

Fig. 1 The schematic pathways representation covered by the top genes in the evaluated dataset

Furthermore, variations from the ClinVar database were filtered depending on their ClinVar review status, with “No assertion criteria” variants being eliminated. In addition, duplicate variations were removed from each of the final datasets. The dataset from ClinVar was divided into two categories: variations with known clinical impact and variants with unknown clinical significance. VUSs represented the highest percentage of all classes of variants, this is due to the increasing number of sequenced genomes and the development of new techniques for variants identification. The ultimate goal has been and still is the determination of the functional impact of those variants and their clinical significance. Although VUSs had higher representation in the primary dataset, with the presence of 14,920 variants as shown in Fig. 3A, they were not used to evaluate the performance of the tools as this was not possible due to their unknown clinical significance status. The final dataset derived from ClinVar used in this study included 1,710 variants, all known clinical impact. For HGMD dataset, it only included variants of known clinical impact, with the final dataset containing 589 pathogenic variants and 4 benign variants. To avoid any bias by the tools, 570 benign variants from ClinVar were added to the HGMD dataset to make the final dataset balanced, including 1,163 variants.

Fig. 2 The distribution of variants among the genes in the A) ClinVar and B) HGMD datasets. Note Red dotted box highlights the top 10 genes

Predictions and scores thresholds

In this study, 21 in-silico pathogenicity prediction tools were used to predict the functional effects of the variants in the evaluated dataset. The tools used were CADD [43], DANN [44], Fathmm [45], Fathmm-MKL [46], Fathmm-XF [47], Genocanyon [48], Clinpred [49, 50], LIST-S2 [51], LRT [52], M-CAP [53], Meta-LR [54], Meta-RNN [55], Meta-SVM [54], Mutpred [56], MVP [57], Polyphen-2 [3], Primate-AI [58], Provean [59], Revel [2], SIFT [60], and SIFT4G [61] for missense variants. Each tool was annotated using Ensemble’s Variant Effect Predictor (VEP) web-platform [50]. The SIFT and PolyPhen-2 [3] scores were created within the VEP platform using the tools’ native modules. The scores of the other tools were annotated in VEP utilizing dbNSFP [62] modules. Variants that received no ratings from the tools were marked with a “-” in the dataset. Following that, all the scores were translated into binary functional impact predictions, which were classified as “Benign” or “Pathogenic” based on the pathogenicity prediction findings or score thresholds of each tool. Some tools can generate scores only, including CADD, DANN, GenoCanyon, Mutpred, MVP and REVEL. For these tools, the threshold values by the original authors were used to classify the scores into Pathogenic or Benign. The overall thresholds used for each tool are summarized in Table 1.

Table 1 The threshold used for each tool used

Tool	Score threshold / Prediction	Reference	
CADD	>=20 → Pathogenic

<=10 → Benign

	[63]	
ClinPred	Damaging → Pathogenic

Tolerated → Benign

	[49, 50]	
DANN	>= 0.5 → Pathogenic

< 0.5→ Benign

	[44, 64]	
FATHMM-MKL	Damaging → Pathogenic

Neutral → Benign

	[50]	
FATHMM-XF	Damaging → Pathogenic

Neutral → Benign

	[50]	
GenoCanyon	>=0.5 → Pathogenic

<=0.5 → Benign

	[50]	
LIST S2	Damaging → Pathogenic

Tolerated → Benign

	[50]	
LRT	Damaging → Pathogenic

Neutral → Benign

	[50]	
M-CAP	Damaging → Pathogenic

Tolerated → Benign

	[50]	
META-LR	Damaging → Pathogenic

Tolerated → Benign

	[50]	
META-RNN	Damaging → Pathogenic

Tolerated → Benign

	[50]	
META-SVM	Damaging → Pathogenic

Tolerated → Benign

	[50]	
MutPred	>=0.5 → Pathogenic

< 0.5 → Benign

	[65]	
MVP	>= 0.75 → Pathogenic

<=0.75 → Benign

	[57]	
PolyPhen	Possibly damaging→ Pathogenic

Probably damaging → Pathogenic

Benign → Benign

	[66]	
PrimateAI	Damaging → Pathogenic

Tolerated → Benign

	[66]	
Provean	Damaging → Pathogenic

Neutral → Benign

	[66]	
REVEL	>=0.75 → Pathogenic

< 0.75 → Benign

	[2]	
SIFT	Deleterious → Pathogenic

Tolerated → Benign

	[66]	
SIFT4G	Deleterious → Pathogenic

Tolerated → Benign

	[66]	
Fathmm	Damaging → Pathogenic

Neutral → Benign

	[66]	

Performance assessment of the prediction tools

The canonical transcript was used for all variations in the assessment of in-silico tools and categorization of variants in the breast cancer datasets using VEP. The clinical outcome data was translated into binary classifications of “Benign” and “Pathogenic”. Several measures were used to assess the efficacy of the in-silico tools, including accuracy, precision, specificity, sensitivity, negative predictive value (NPV), Matthew’s correlation coefficient (MCC), and false positive rate. These indicators were derived using a confusion (contingency) matrix, which involves comparing the predictions from each instrument against the datasets. The equations used are shown in Table 2, where true positives, true negatives, false positives, and false negatives are represented by TP, TN, FP, and FN respectively. The proportion of missing values for each tool’s predictions was also calculated. The confusion matrix was used to evaluate the discrimination power of each tool for the pathogenicity of variants using the confusion matrixfunction available in Scikit-learn [67], a Python package for machine learning.

Table 2 The metrics used to evaluate the performance of the tools, and the equations used to calculate them

Metrics	Equations	
\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\:\text{A}\text{c}\text{c}\text{u}\text{r}\text{a}\text{c}\text{y}$$\end{document}	\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\:=\frac{TP+TN}{TP+TN+FP+FN}$$\end{document}	
\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\:\text{P}\text{r}\text{e}\text{c}\text{i}\text{s}\text{i}\text{o}\text{n}$$\end{document}	\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\:=\frac{TP}{TP+FP}$$\end{document}	
\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\:\text{S}\text{p}\text{e}\text{c}\text{i}\text{f}\text{i}\text{c}\text{t}\text{y}$$\end{document}	\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\:=\frac{TN}{TN+FP}$$\end{document}	
\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\:\text{S}\text{e}\text{n}\text{s}\text{i}\text{t}\text{i}\text{v}\text{t}\text{y}$$\end{document}	\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\:=\frac{TP}{TP+FN}$$\end{document}	
\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\:\text{T}\text{r}\text{u}\text{e}\:\text{n}\text{e}\text{g}\text{a}\text{t}\text{i}\text{v}\text{e}\:\text{r}\text{a}\text{t}\text{e}\:\left(\text{T}\text{N}\text{R}\right)$$\end{document}	\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\:=\frac{TN}{TN+FN}$$\end{document}	
\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\:\text{M}\text{C}\text{C}$$\end{document}	\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\:=\frac{TP\:x\:TN-FN\:x\:FP}{\sqrt{(TP+FN)(TP+FP)(TN+FN)(TN+FP)}}$$\end{document}	
\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\:\text{F}\text{a}\text{l}\text{s}\text{e}\:\text{p}\text{o}\text{s}\text{i}\text{t}\text{i}\text{v}\text{e}\:\text{r}\text{a}\text{t}\text{e}\:\left(\text{F}\text{P}\text{R}\right)$$\end{document}	\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\:=1-\text{S}\text{p}\text{e}\text{c}\text{i}\text{f}\text{i}\text{c}\text{i}\text{t}\text{y}$$\end{document}	

Results and discussion

The dataset structure

The dataset utilized in the current study originates from the ClinVar database and includes a total of 16,630 variations, focusing on breast cancer variants in 27 cancer genes. This dataset has been split into two subsets: known importance variations (1,710 variants) and unknown significance variants (14,920 variants). The dataset from the HGMD database includes a total of 593 variations in 35 cancer genes plus 570 benign variations from ClinVar for balancing. When the gene distribution data (Fig. 2A and B; Table 3) was examined, it was observed that the number variants in tumor suppressor genes was greater than the number of variants in oncogenes among the top 10 genes in both datasets. The same genes were also identified in the Cancer Gene Census (CGC) [68], which is a meticulously curated catalogue of genes encompassing mutations implicated in cancer development in humans. Additionally, the study by Yazar and Ozbek has shown similar results, confirming the use of an open-source dataset for the research. In the evaluated dataset, most of the high susceptibility genes had higher number of variants, but interestingly, the number of variants per gene are not directly corelated as some of the high susceptibility genes like PTEN which had only 7 variants compared to PALB2 which had 61 variants in this dataset as shown in Table 1. The top genes were the same for both datasets, except for Rad51D which was in the top 10 genes among ClinVar dataset genes and PTEN which was among the top 10 genes in the HGMD dataset. The top genes identified in the evaluated dataset are critically involved in breast cancer pathways through their roles in DNA repair, cell cycle regulation, and maintenance of genomic stability. BRCA1 and BRCA2 are key players in homologous recombination repair of DNA double-strand breaks. TP53, known as the “guardian of the genome,” is pivotal for cell cycle control and apoptosis. CHEK2 and ATM are essential for DNA damage response signaling. BRIP1, PALB2, BARD1, RAD51D, and RAD51C facilitate the recruitment and function of the BRCA1/2 complexes. CDH1 and PTEN act as tumor suppressors, with CDH1 maintaining cell adhesion and PTEN regulating cell growth and survival pathways. Mutations or dysfunctions in these genes disrupt their normal protective roles, leading to increased susceptibility to breast cancer by allowing the accumulation of genetic mutations and uncontrolled cell proliferation [69, 70]. The schematic representation of the pathways discussed above are shown in Fig. 1 below.

Table 3 Information about the top genes with the most variants in breast cancer dataset

Gene name	Protein/enzyme	Biological function	Tumor Suppressor Gene / Oncogene	
BRCA1	Breast Cancer Type 1 Susceptibility Protein	Maintains the genome stability	Tumor Suppressor Gene	
BRCA2	Breast Cancer Type 2 Susceptibility Protein	Maintains the genome stability	Tumor Suppressor Gene	
TP53	Tumor Protein P53	Induce cell cycle arrest, apoptosis, senescence, DNA repair, or metabolism changes	Oncogene/ Tumor Suppressor Gene	
CHEK2	Checkpoint Kinase 2	Regulates cell cycle checkpoint	Tumor Suppressor Gene	
BRIP1	BRCA1 Interacting Protein C-Terminal Helicase 1	Maintains the chromosomal stability	Tumor Suppressor Gene	
ATM	ATM Serine/Threonine Kinase	A cell cycle checkpoint kinase	Tumor Suppressor Gene	
PALB2	Partner And Localizer of BRCA2	Homologous recombination repair	Tumor Suppressor Gene	
BARD1	BRCA1 Associated RING Domain 1	Maintains the genome stability	Tumor Suppressor Gene	
RAD51D	RAD51 Paralog D	Homologous recombination and repair of DNA.	Tumor Suppressor Gene	
RAD51C	RAD51 Paralog C	Homologous recombination and repair of DNA.	Tumor Suppressor Gene	
CDH1	Cadherin 1	Calcium-dependent cell adhesion	Tumor Suppressor Gene	
PTEN	Phosphatase and TENsin homolog	Maintains the genome stability and regulate cell growth and proliferation	Tumor Suppressor Gene	

In the datasets utilized for this study, variants with the classification of “Pathogenic”, “Likely Pathogenic”, or “risk_factor” / “drug-response” in addition to “Pathogenic” or “Likely Pathogenic” were included in Pathogenic category of the dataset. While variants with the classification of “Benign”, “Likely Benign” or “risk_factor” / “drug-response” in addition to “Benign” or “Likely Benign” were included in the Benign category of the dataset. All other classifications were excluded from the dataset. The distribution of clinical importance among the variations indicated that the majority (14,920) belonged to VUSs and (1,710) belonged to variants with known clinical impact, as shown in Fig. 3A. Following that, the pathogenic variations (1,140) were the second biggest category, and the Benign variants (570) were the smallest, as illustrated in Fig. 3B. The distribution of the HGMD dataset variants and the balancing of it using ClinVar benign variants are shown in Fig. 3C and D, respectively.

Fig. 3 The distribution of variants of known clinical impact on VUSs. A) Number of variants in each classification, B) The distribution of the variants in each classification. And the distribution of variants of known clinical impact (Pathogenic to Benign). C) This distribution of variants in the HGMD dataset D) The balancing of the HGMD dataset using benign variants from the ClinVar dataset

Pathogenicity predictions obtained from the tools

The frequency distributions of pathogenicity predictions among all tools are shown in Tables 4 and 5. It was not possible to generate a perfectly balanced dataset as the distribution of the benign variants in ClinVar is limited, this can be due to the limited reporting of the benign variants specifically. Furthermore, as previously demonstrated in Fig. 2, the distribution of variants is different from one gene to another, which makes it difficult to balance the number of variants represented by each gene. From a general point of view, and while considering the ClinVar dataset as the benchmark in which it is represented by 33.33% Benign variants and 66.67% Pathogenic variants, the prediction distribution frequencies of the variants obtained from in-silico tools revealed that the benign variant frequencies of REVEL (32.69%), Fathmm-XF (29.24%), Meta-RNN (28.42%) and Meta-SVM (27.95%) have the most similar four results to ClinVar (33.33%) in variants with known clinical impact. These tools take evolutionary conservation data into account as a feature. Evolutionary conservation measures how well a particular sequence or genetic variant is preserved across different species during evolution. In other words, it is the occurrence of homologous genes, gene segments, or chromosome regions in different species, indicating a shared evolutionary origin and underscoring a critical functional significance of the conserved elements. The conservation information provides insights into the functional importance of a genetic variant and can aid in predicting its pathogenicity. Moreover, the prediction distribution frequencies of the variants obtained from in-silico tools revealed that the pathogenic variant frequencies of DANN (47.89%), M-CAP (41.81%), MVP (37.49%) and GenoCanyon (35.20%) have shown the top four closest results to ClinVar (66.67%) in variants with known clinical impact.

On the other hand, for the HGMD dataset, as the database is specialized in pathogenic variants mainly, there is no reference of average distribution of significance like for the ClinVar dataset, therefore we compared the results with the ClinVar results as well just to give an idea of the performance on different datasets. The most interesting finding was that the percentage of missing values predicted using the HGMD dataset was much lower than that on the ClinVar dataset. We assume that this is due to the high disease specificity variants that are included in the HGMD database. Furthermore, the top four similar in-silico tools prediction frequency distributions were SIFT (32.76%), PolyPhen (36.63%), CADD (34.14%) and Fathmm-MKL (32.76%) for benign variants in addition to GenoCanyon (64.92%), M-CAP (73.60%), MVP (66.21%) and Fathmm-MKL (56.58%) for pathogenic variants.

Table 4 The frequency distributions of pathogenicity predictions among all 21 tools on ClinVar dataset

Tool	Pathogenic predicted	Benign predicted	Missing values	
SIFT	32.34%	19.36%	48.30%	
PolyPhen-2	30.64%	21.17%	48.19%	
CADD	31.46%	20.58%	47.95%	
ClinPred	25.03%	27.02%	47.95%	
DANN	47.89%	4.15%	47.95%	
Fathmm	26.43%	23.80%	49.77%	
GenoCanyon	35.20%	16.84%	47.95%	
LIST-S2	22.46%	14.62%	62.92%	
LRT	24.27%	27.02%	48.71%	
M-CAP	41.81%	2.87%	55.32%	
MVP	37.49%	11.23%	51.29%	
MetaLR	26.84%	24.74%	48.42%	
MetaRNN	23.16%	28.42%	48.42%	
MetaSVM	23.63%	27.95%	48.42%	
MutPred	15.50%	6.61%	77.89%	
PROVEAN	22.16%	27.60%	50.23%	
PrimateAI	1.35%	48.65%	50.00%	
REVEL	16.49%	32.69%	50.82%	
SIFT4G	27.54%	23.98%	48.48%	
Fathmm-MKL	32.51%	19.53%	47.95%	
Fathmm-XF	22.69%	29.24%	48.07%	
Benchmarking dataset percentage	66.67%	33.33%	-	
Note The bold values are for the best performing tools

Table 5 The frequency distributions of pathogenicity predictions among all 21 tools on HGMD dataset

Tool	Pathogenic predicted	Benign predicted	Missing values	
SIFT	56.49%	32.76%	10.75%	
PolyPhen	52.71%	36.63%	10.66%	
CADD	55.89%	34.14%	9.97%	
ClinPred	46.09%	43.25%	10.66%	
DANN	82.98%	6.36%	10.66%	
Fathmm	40.15%	46.17%	13.67%	
GenoCanyon	64.92%	24.42%	10.66%	
LIST-S2	34.91%	25.02%	40.07%	
LRT	41.10%	46.69%	12.21%	
M-CAP	73.60%	6.53%	19.86%	
MVP	66.21%	20.64%	13.16%	
MetaLR	41.79%	47.12%	11.09%	
MetaRNN	41.44%	47.46%	11.09%	
MetaSVM	36.46%	52.45%	11.09%	
MutPred	28.37%	18.31%	53.31%	
PROVEAN	38.44%	46.78%	14.79%	
PrimateAI	3.96%	83.75%	12.30%	
REVEL	25.37%	61.48%	13.16%	
SIFT4G	46.26%	42.48%	11.26%	
Fathmm-MKL	56.58%	32.76%	10.66%	
Fathmm-XF	39.47%	49.70%	10.83%	
Benchmarking dataset percentage	66.67%	33.33%	-	
Note The bold values are for the best performing tools

Performance assessment

The pathogenicity thresholds defined for each tool as given in literature were used to study the performance metrics of in-silico tools and the statistical classification of variants. The confusion (contingency) matrix approach was used to create performance assessment measures of variants with known clinical impact group and the results are shown in Table 6 for ClinVar dataset and in Table 7 for HGMD dataset. All the performance metrics calculated are summarized in Table 8 for ClinVar dataset and in Table 9 for HGMD dataset. Accordingly, for the ClinVar dataset, M-CAP (0.961), DANN (0.956), MutPred (0.900) and MVP (0.887) were the top four tools with the highest sensitivity values. However, PrimateAI (0.993), REVEL (0.859), Meta-RNN (0.770) and Fathmm-XF (0.759) were the top four tools with highest specificity values. Moreover, the top four tools in terms of accuracy were MutPred (0.749), Meta-RNN (0.724), ClinPred (0.712) and Fathmm-XF (0.700). Among the tools used in this study, ClinPred, LRT, Meta-RNN, Meta-SVM and Fathmm-XF displayed a high, balanced values of sensitivity and specificity in addition to high accuracy values.

Table 6 The results of the confusion matrices including TP, TN, FP and FN for the ClinVar dataset

Tool	Actual Positives	Actual Negatives	Total #	Missing Values	True Positives Predicted	False Positives Predicted	True Negatives Predicted	False Negatives Predicted	
Pathogenic	Benign	Total	
SIFT	1140	570	1710	707	119	826	331	222	229	102	
PolyPhen	1140	570	1710	706	118	824	322	202	250	112	
CADD	1140	570	1710	704	116	820	347	191	263	89	
ClinPred	1140	570	1710	704	116	820	304	124	330	132	
DANN	1140	570	1710	704	116	820	417	402	52	19	
Fathmm	1140	570	1710	715	136	851	275	177	257	150	
GenoCanyon	1140	570	1710	704	116	820	279	323	131	157	
LIST-S2	1140	570	1710	802	274	1076	252	132	164	86	
LRT	1140	570	1710	707	126	833	286	129	315	147	
M-CAP	1140	570	1710	752	194	946	373	342	34	15	
MVP	1140	570	1710	733	144	877	361	280	146	46	
MetaLR	1140	570	1710	707	121	828	299	160	289	134	
MetaRNN	1140	570	1710	707	121	828	293	103	346	140	
MetaSVM	1140	570	1710	707	121	828	287	117	332	146	
MutPred	1140	570	1710	928	404	1332	191	74	92	21	
PROVEAN	1140	570	1710	719	140	859	229	150	280	192	
PrimateAI	1140	570	1710	721	134	855	20	3	433	399	
REVEL	1140	570	1710	725	144	869	222	60	366	193	
SIFT4G	1140	570	1710	707	122	829	305	166	282	128	
Fathmm-MKL	1140	570	1710	704	116	820	353	203	251	83	
Fathmm-XF	1140	570	1710	704	118	802	279	109	343	157	
Note The bold values are for the best performing tools, Mathew’s correlation coefficient (MCC)

Table 7 The results of the confusion matrices including TP, TN, FP and FN for the HGMD dataset

Tool	Actual Positives	Actual Negatives	Total #	Missing Values	True Positives Predicted	False Positives Predicted	True Negatives Predicted	False Negatives Predicted	
Pathogenic	Benign	Total	
SIFT	589	574	1163	6	119	125	435	222	233	148	
PolyPhen	589	574	1163	6	118	124	411	202	254	172	
CADD	589	574	1163	0	116	116	459	191	267	130	
ClinPred	589	574	1163	6	116	124	412	124	334	169	
DANN	589	574	1163	8	116	124	560	405	53	21	
Fathmm	589	574	1163	23	126	159	289	178	260	277	
GenoCanyon	589	574	1163	8	116	124	431	324	134	150	
LIST-S2	589	574	1163	191	275	466	274	132	167	124	
LRT	589	574	1163	16	126	142	349	129	319	224	
M-CAP	589	574	1163	33	198	231	514	342	34	42	
MVP	589	574	1163	9	144	153	486	284	146	94	
MetaLR	589	574	1163	8	121	129	326	160	293	255	
MetaRNN	589	574	1163	8	121	129	379	103	350	202	
MetaSVM	589	574	1163	8	121	129	307	117	336	274	
MutPred	589	574	1163	212	408	620	256	74	92	121	
PROVEAN	589	574	1163	31	141	172	297	150	283	261	
PrimateAI	589	574	1163	8	135	143	43	3	436	538	
REVEL	589	574	1163	9	144	153	234	61	369	346	
SIFT4G	589	574	1163	9	122	131	372	166	286	208	
Fathmm-MKL	589	574	1163	8	116	124	455	203	255	126	
Fathmm-XF	589	574	1163	8	118	126	350	109	347	231	
Note The bold values are for the best performing tools, Mathew’s correlation coefficient (MCC)

Table 8 The calculated performance metrics for all 21 tools on the ClinVar dataset

Tool	Sensitivity (TPR)	Specificity (TNR)	False Positive rate	False Negative rate	Accuracy	Precision	F1-Score	MCC	Misclassification	
SIFT	0.76443418	0.507760532	0.492239468	0.2355658	0.633484163	0.598553345	0.671399594	0.28114822	0.36651584	
PolyPhen	0.741935484	0.553097345	0.446902655	0.2580645	0.645598194	0.614503817	0.67223382	0.30002987	0.35440181	
CADD	0.79587156	0.579295154	0.420704846	0.2041284	0.685393258	0.644981413	0.712525667	0.38355973	0.31460674	
ClinPred	0.697247706	0.726872247	0.273127753	0.3027523	0.712359551	0.710280374	0.703703704	0.42434296	0.28764045	
DANN	0.956422018	0.114537445	0.885462555	0.043578	0.526966292	0.509157509	0.664541833	0.13092132	0.47303371	
Fathmm	0.647058824	0.592165899	0.407834101	0.3529412	0.619324796	0.60840708	0.62713797	0.23954051	0.3806752	
GenoCanyon	0.639908257	0.288546256	0.711453744	0.3600917	0.460674157	0.46345515	0.537572254	-0.0764467	0.53932584	
LIST-S2	0.74556213	0.554054054	0.445945946	0.2544379	0.65615142	0.65625	0.698060942	0.30586787	0.34384858	
LRT	0.660508083	0.709459459	0.290540541	0.3394919	0.685290764	0.689156627	0.674528302	0.37047083	0.31470924	
M-CAP	0.961340206	0.090425532	0.909574468	0.0386598	0.532722513	0.521678322	0.676337262	0.10563337	0.46727749	
MVP	0.886977887	0.342723005	0.657276995	0.1130221	0.608643457	0.563182527	0.688931298	0.27263716	0.39135654	
MetaLR	0.690531178	0.643652561	0.356347439	0.3094688	0.666666667	0.651416122	0.670403587	0.33440742	0.33333333	
MetaRNN	0.676674365	0.770601336	0.229398664	0.3233256	0.724489796	0.73989899	0.706875754	0.44954865	0.2755102	
MetaSVM	0.662817552	0.739420935	0.260579065	0.3371824	0.701814059	0.71039604	0.685782557	0.40359531	0.29818594	
MutPred	0.900943396	0.554216867	0.445783133	0.0990566	0.748677249	0.720754717	0.800838574	0.49342842	0.25132275	
PROVEAN	0.543942993	0.651162791	0.348837209	0.456057	0.598119859	0.604221636	0.5725	0.1962704	0.40188014	
PrimateAI	0.047732697	0.993119266	0.006880734	0.9522673	0.529824561	0.869565217	0.090497738	0.12622274	0.47017544	
REVEL	0.534939759	0.85915493	0.14084507	0.4650602	0.699167658	0.787234043	0.637015782	0.41734861	0.30083234	
SIFT4G	0.704387991	0.629464286	0.370535714	0.295612	0.666288309	0.647558386	0.674778761	0.33460692	0.33371169	
Fathmm-MKL	0.809633028	0.552863436	0.447136564	0.190367	0.678651685	0.634892086	0.711693548	0.37425216	0.32134831	
Fathmm-XF	0.639908257	0.758849558	0.241150442	0.3600917	0.70045045	0.719072165	0.677184466	0.40190259	0.29954955	
Note The bold values are for the best performing tools, Mathew’s correlation coefficient (MCC)

Table 9 The calculated performance metrics for all 21 tools on the HGMD dataset

Tool	Sensitivity (TPR)	Specificity (TNR)	False Positive rate	False Negative rate	Accuracy	Precision	F1-Score	MCC	Misclassification	
SIFT	0.746140652	0.5120879	0.487912088	0.253859348	0.643545279	0.66210046	0.7016129	0.265826996	0.356454721	
PolyPhen	0.704974271	0.5570175	0.442982456	0.295025729	0.640038499	0.67047308	0.68729097	0.264343956	0.359961501	
CADD	0.779286927	0.5829694	0.417030568	0.220713073	0.693409742	0.70615385	0.7409201	0.370385923	0.306590258	
ClinPred	0.709122203	0.7292576	0.270742358	0.290877797	0.717998075	0.76865672	0.73769024	0.435516884	0.282001925	
DANN	0.963855422	0.1157205	0.884279476	0.036144578	0.589990375	0.58031088	0.72445019	0.153611276	0.410009625	
Fathmm	0.510600707	0.5936073	0.406392694	0.489399293	0.546812749	0.61884368	0.55953533	0.103609792	0.453187251	
GenoCanyon	0.741824441	0.2925764	0.707423581	0.258175559	0.543792108	0.57086093	0.64520958	0.03832282	0.456207892	
LIST-S2	0.688442211	0.5585284	0.441471572	0.311557789	0.632711621	0.67487685	0.68159204	0.247863709	0.367288379	
LRT	0.609075044	0.7120536	0.287946429	0.390924956	0.654260529	0.73012552	0.6641294	0.319360692	0.345739471	
M-CAP	0.924460432	0.0904255	0.909574468	0.075539568	0.587982833	0.60046729	0.72804533	0.026684839	0.412017167	
MVP	0.837931034	0.3395349	0.660465116	0.162068966	0.625742574	0.63116883	0.72	0.206163701	0.374257426	
MetaLR	0.561101549	0.6467991	0.353200883	0.438898451	0.598646035	0.67078189	0.61105904	0.206673424	0.401353965	
MetaRNN	0.65232358	0.7726269	0.227373068	0.34767642	0.705029014	0.78630705	0.7130762	0.42265155	0.294970986	
MetaSVM	0.528399312	0.7417219	0.258278146	0.471600688	0.621856867	0.7240566	0.61094527	0.272488349	0.378143133	
MutPred	0.679045093	0.5542169	0.445783133	0.320954907	0.640883978	0.77575758	0.7241867	0.220100925	0.359116022	
PROVEAN	0.532258065	0.6535797	0.346420323	0.467741935	0.585267407	0.66442953	0.59104478	0.185242978	0.414732593	
PrimateAI	0.074010327	0.9931663	0.006833713	0.925989673	0.469607843	0.93478261	0.13716108	0.160280261	0.530392157	
REVEL	0.403448276	0.8581395	0.141860465	0.596551724	0.597029703	0.79322034	0.53485714	0.284447223	0.402970297	
SIFT4G	0.64137931	0.6327434	0.367256637	0.35862069	0.637596899	0.69144981	0.66547406	0.272253556	0.362403101	
Fathmm-MKL	0.78313253	0.5567686	0.443231441	0.21686747	0.683349374	0.69148936	0.73446328	0.350185312	0.316650626	
Fathmm-XF	0.602409639	0.7609649	0.239035088	0.397590361	0.672131148	0.76252723	0.67307692	0.363123816	0.327868852	
Note The bold values are for the best performing tools, Mathew’s correlation coefficient (MCC)

Moreover, for the HGMD dataset, CADD (0.779), DANN (0.963), M-CAP (0.924) and MVP (0.837) were the top four tools with the highest sensitivity values. Nonetheless, Meta-RNN (0.772), PrimateAI (0.993), REVEL (0.858) and Fathmm-MKL (0.760) were the top four tools with highest specificity values. Moreover, the top four tools in terms of accuracy were CADD (0.693), ClinPred (0.717), Meta-RNN (0.705) and Fathmm-MKL (0.683). Tools with distinct performance should have higher specificity and sensitivity values at same time [71, 72]. Furthermore, these values in distinct tools should also be balanced, suggesting that their ratio to each other is close to 1. Among the tools used in this study, on both datasets, ClinPred, LRT, Meta-RNN, Meta-SVM and Fathmm-XF displayed a high, balanced values of sensitivity and specificity in addition to high accuracy values.

As a result, the thresholds provided by the original authors of most of the tools would not be appropriate for all datasets due to the relatively low specificity and sensitivity values; therefore, more precise thresholds are necessary for this type of variation dataset [30, 31]. MCC is another performance statistic that reflects the degree of correlation between actual and anticipated binary classification [73]. Consequently, for the ClinVar dataset, MutPred (0.493), Meta-RNN (0.450), ClinPred (0.424), and REVEL (0.417) had the greatest MCC values, indicating that these approaches had greater positive correlations between observed and anticipated binary predictions. Only GenoCanyon (-0.076) demonstrated a somewhat negative correlation between the actual and predicted binary categorization. For the HGMD dataset, CADD (0.370), ClinPred (0.435), Meta-RNN (0.422) and Fathmm-XF (0.363) had the greatest MCC values.

Furthermore, a bar plot of classification accuracies to examine the instruments’ discriminating ability in variations with established clinical importance was plotted. According to Fig. 4A, the top tools with the highest discriminatory power among all in-silico prediction tools for the ClinVar dataset were MutPred (Accuracy = 0.73), Meta-RNN (Accuracy = 0.72), ClinPred (Accuracy = 0.71), Meta-SVM, REVEL, and Fathmm-XF (Accuracy = 0.70). Furthermore, and as shown in Fig. 4B, for the HGMD dataset the top tools were ClinPred (Accuracy = 0.72), MetaRNN (Accuracy = 0.71), CADDd (Accuracy = 0.69), Fathmm-MKL (Accuracy = 0.68), and Fathmm-XF (Accuracy = 0.68). These tools are shown to have the highest overall performance when performance metrics, and prediction distribution frequency data are combined.

Fig. 4 The classification accuracy bar plots of known clinically significant variations from A) ClinVar database and B) HGMD database using 21 different in-silico tools. Scikit-learn [67], a Python machine learning package, was used to perform the analysis

Comparison and discussion

Based on the findings of this paper, in addition to a recent publication by Ahmad et al. [4], tools that are trained on breast cancer data performed better then tools only tested but not trained on breast cancer data as shown in Table 10 below. In addition to the type of dataset a tool is trained on, the volume of the data used for training is also an important factor. For example, CADD, is one of the well-known pathogenicity prediction tools, and it is trained on a large volume and wide range of genes and diseases, it has also performed well on a breast cancer dataset as shown in Table 10. This made its ability for prediction better than general tools that have been trained on limited datasets and therefore its performance on breast cancer datasets is not surprising. Additionally, Revel, which is an ensemble meta-prediction method based on several other individual pathogenicity prediction tools, has also demonstrated a good performance on the breast cancer dataset as shown in Table 10. This is due to the inclusion of a wide range of single pathogenicity prediction tools scores in the development of Revel, as each tool is trained on a different type and volume of data, the prediction is based on the collection of them. Furthermore, the tools evaluated in this paper, has performed better on their training datasets or on the open-source datasets than on the evaluated breast cancer dataset in terms of accuracy. The accuracy values used in Table 11 were calculated in Tables 8 and 9 for the evaluated datasets and rounded to three decimal places, while for the open-source dataset, the accuracy values were collected from the literature, as shown in Table 11 below. From Table 10, we can say that tools that are trained on breast cancer data, or that are trained on very large volumes of data, or that are build based on several individual tools are shown to have the best performance when tested on breast cancer data as per the literature. While in Table 11, comparing the accuracy results of the evaluated specific dataset to the open-source general dataset, the general dataset performance has almost always outperformed the specific dataset. This is because the tools are not trained on breast cancer data. This leads us to conclude that disease specific tools for breast cancer can be more clinically significant and that the thresholds of the tools evaluated in this paper are suggested to be revised or the tool to be retrained, when possible, for specific diseases such as breast cancer to achieve better accuracy.

Table 10 Comparison of the performance of different pathogenicity prediction tools on BC data [4]

Tools	Relevance	Type (ML/Non-ML)	Accuracy	Reference	
Trained on conservation data	Trained on BC data	Tested on BC data	
CADD	✓	-	✓	ML	0.99	[78]	
Polyphen	✓	-	✓	ML	0.88	
Gene specific model	✓	✓	✓	ML	0.999	
SIFT	✓	-	✓	Non-ML	0.11	
Polyphen	✓	-	✓	ML	0.64	[79]	
Lyrus	✓	✓	✓	ML	0.89	
SIFT	✓	-	✓	Non-ML	0.66	
Polyphen	✓	-	✓	ML	0.628	[80]	
CADD	✓	-	✓	ML	0.621	
Revel	✓	-	✓	ML	0.63	
SIFT	✓	-	✓	Non-ML	0.40	
Polyphen	✓	-	✓	ML	0.77	[81]	
SIFT	✓	-	✓	Non-ML	0.55	
Polyphen	✓	-	✓	ML	0.87	[81]	
Revel	✓	-	✓	ML	0.97	
SIFT	✓	-	✓	Non-ML	0.85	

Table 11 Comparison of the accuracy of the evaluated tools on the evaluated datasets and on general datasets

Tool	ClinVar Dataset	HGMD Dataset	Open-source general datasets	Reference	
Accuracy	
SIFT	0.633	0.644	0.853	[74]	
PolyPhen	0.646	0.640	0.873	[74]	
CADD	0.685	0.693	0.860	[75]	
ClinPred	0.712	0.718	0.993	[74]	
DANN	0.527	0.590	0.775	[76]	
Fathmm	0.619	0.547	0.770	[77]	
GenoCanyon	0.461	0.544	0.620	[77]	
LIST-S2	0.656	0.633	0.860	[76]	
LRT	0.685	0.654	0.745	[76]	
M-CAP	0.533	0.588	0.925	[76]	
MVP	0.609	0.626	0.943	[76]	
MetaLR	0.667	0.599	0.917	[76]	
MetaRNN	0.724	0.705	0.977	[76]	
MetaSVM	0.702	0.622	0.921	[76]	
MutPred	0.749	0.641	0.919	[76]	
PROVEAN	0.598	0.585	0.930	[75]	
PrimateAI	0.530	0.470	0.848	[76]	
REVEL	0.699	0.597	0.970	[74]	
SIFT4G	0.666	0.638	0.873	[76]	
Fathmm-MKL	0.679	0.683	0.800	[76]	
Fathmm-XF	0.700	0.672	0.684	[76]	

Conclusions

Finding the functional impact of cancer-related variations is the primary objective of precision and predictive cancer medicine. While several research have examined cancer-related variations and prediction methods, none have specifically examined the functional importance of a complete gene list of breast cancer missense variants obtained from ClinVar and HGMD databases. Additionally, there hasn’t been a thorough computational analysis of such breast cancer-related variants using in-silico prediction methods. Therefore, using the breast cancer variant datasets from ClinVar and HGMD, our work evaluated the functional relevance and performance metrics of 21 functional prediction algorithms in an effort to meet this demand.

Our findings from Fig. 2 demonstrate that genes associated with low and moderate penetrance predisposition display a notable abundance of variants. Consequently, utilizing a disease-specific tool trained on a comprehensive gene list linked to the specific disease leads to enhanced specificity and quality, surpassing generalized or overly specific tools. Moreover, the findings indicates that MutPred exhibited the highest discriminatory power when applied to the breast cancer-related variant dataset obtained from ClinVar after evaluating the statistical analysis of performance metrics, prediction distribution frequency data, while REVEL and CADD out performs MutPred on the HGMD dataset. Furthermore, our results suggest that adjusting the thresholds of each prediction tool according to the specific dataset could enhance the specificity and sensitivity values. In conclusion, we can say that the top tools with the highest discriminatory power on the ClinVar dataset were MutPred, Meta-RNN, ClinPred, Meta-SVM, REVEL and Fathmm-XF and on the HGMD dataset were REVEL, CADD, MutPred, ClinPred, and Meta-RNN and they are good candidates to be applied on breast cancer variants.

In conclusion, this paper is addressing an existing gap in the literature by specifically focusing on the assessment of functional relevance in comprehensive breast cancer variants sourced from ClinVar and HGMD, using various in-silico pathogenicity tools. The primary objective of our research was to evaluate the performance of AI-derived in-silico pathogenicity prediction tools in investigating the functional effects of breast cancer-related variants with known clinical impact. To achieve this, we employed a total of 21 different in-silico tools and analyzed their performance metrics on the dataset, providing valuable insights into their effectiveness. Our study successfully demonstrated the applicability of in-silico pathogenicity prediction tools for assessing the functional effects of breast cancer variants. Finally, we recommend changing the thresholds of the available tools if there is no specific tool available for the targeted disease, this can be done by using several thresholds of an arithmetic sequence, then choosing the best values that results in the highest numbers of true positives and true negatives for the specific dataset. If the pre-built model can be retrained, then using a specific dataset and training the model on it would be an option. Another recommendation can be designing a specific tool based on the limitations of the available tools, this includes the training dataset, the features used, or others. By filling the research gap and providing comprehensive insights into the performance of various tools, our findings contribute to the understanding of breast cancer genetics and pave the way for future research and clinical applications in this field.

Acknowledgements

None.

Author contributions

R.M.A. and M.S.M. conceptualized the research, R.M.A., P.K. S.T. and N.D. designed the methodology, R.M.A., P.K. and F.J. validated the results, R.M.A. and M.S.M. curated the data, R.M.A. wrote the original draft, R.M.A., B.R.A., S.T., M.S.M. revised and edited the draft, M.S.M., B.R.A., F.J., and N.D. supervised the research. All authors have read and agreed to the published version of the manuscript.

Funding

This work was supported by the United Arab Emirates University through Strategic Research Program (#12R111) and Research Start-up Program (#12M109). Rahaf M. Ahmad is supported by a PhD fellowship from the United Arab Emirates University.

Data availability

No datasets were generated or analysed during the current study.

Declarations

Competing interests

The authors declare no competing interests.

Conflict of interest

In addition to being an author of this manuscript, Professor Bassam R. Ali serves as a Deputy Editor-in-Chief for the Journal of Human Genomics. The authors declare no conflict of interest.

Publisher’s note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
==== Refs
References

1. Collins FS Patrinos A Jordan E Chakravarti A Gesteland R Walters LR 1998 New goals for the U.S. Human Genome Project: 1998–2003 10.1126/science.282.5389.682
Collins FS, Patrinos A, Jordan E, Chakravarti A, Gesteland R, Walters LR. Science. 1998;282(23):682–9. 10.1126/science.282.5389.682. New goals for the U.S. Human Genome Project: 1998–2003.10.1126/science.282.5389.682
2. Ioannidis NM, et al. REVEL: an Ensemble Method for Predicting the pathogenicity of rare missense variants. Am J Hum Genet. Oct. 2016;99(4):877–85. 10.1016/j.ajhg.2016.08.016.
3. Adzhubei I Jordan DM Sunyaev SR Predicting functional effect of human missense mutations using PolyPhen-2 Curr Protoc Hum Genet no SUPPL 76 2013 10.1002/0471142905.hg0720s76
Adzhubei I, Jordan DM, Sunyaev SR. Predicting functional effect of human missense mutations using PolyPhen-2. Curr Protoc Hum Genet no SUPPL 76. 2013. 10.1002/0471142905.hg0720s76.10.1002/0471142905.hg0720s76
4. Ahmad RM, Ali BR, Al-Jasmi F, Sinnott RO, Dhaheri NA, Mohamad MS. A review of genetic variant databases and machine learning tools for predicting the pathogenicity of breast cancer. Brief Bioinform. Nov. 2023;25(1). 10.1093/bib/bbad479.
5. Rich KA, Roggenbuck J, Kolb SJ. Searching Far and Genome-Wide: The Relevance of Association Studies in Amyotrophic Lateral Sclerosis, Front Neurosci, vol. 14, no. January, pp. 1–11, 2021, 10.3389/fnins.2020.603023
6. Gyulkhandanyan A Analysis of protein missense alterations by combining sequence- and structure-based methods Mol Genet Genomic Med 2020 no November 2019 1 28 10.1002/mgg3.1166
Gyulkhandanyan A, et al. Analysis of protein missense alterations by combining sequence- and structure-based methods. Mol Genet Genomic Med. 2020;no November 2019:1–28. 10.1002/mgg3.1166.10.1002/mgg3.1166
7. Tam V Patel N Turcotte M Bossé Y Paré G Meyre D Benefits and limitations of genome-wide association studies Nat Rev Genet 2019 20 8 467 84 10.1038/s41576-019-0127-1 31068683
Tam V, Patel N, Turcotte M, Bossé Y, Paré G, Meyre D. Benefits and limitations of genome-wide association studies. Nat Rev Genet. 2019;20(8):467–84. 10.1038/s41576-019-0127-1.31068683 10.1038/s41576-019-0127-1
8. Kircher M Witten DM Jain P O’roak BJ Cooper GM Shendure J A general framework for estimating the relative pathogenicity of human genetic variants Nat Genet 2014 46 3 310 5 10.1038/ng.2892 24487276
Kircher M, Witten DM, Jain P, O’roak BJ, Cooper GM, Shendure J. A general framework for estimating the relative pathogenicity of human genetic variants. Nat Genet. 2014;46(3):310–5. 10.1038/ng.2892.24487276 10.1038/ng.2892
9. Kucukkal TG, Petukh M, Li L, Alexov E. Structural and physico-chemical effects of disease and non-disease nsSNPs on proteins, Curr Opin Struct Biol, vol. 32, no. 3, pp. 18–24, Jun. 2015, 10.1016/j.sbi.2015.01.003
10. Li MX Predicting mendelian disease-causing non-synonymous single nucleotide variants in Exome sequencing studies PLoS Genet 2013 9 1 1 11 10.1371/journal.pgen.1003143
Li MX, et al. Predicting mendelian disease-causing non-synonymous single nucleotide variants in Exome sequencing studies. PLoS Genet. 2013;9(1):1–11. 10.1371/journal.pgen.1003143.10.1371/journal.pgen.1003143
11. Ponzoni L Bahar I Structural dynamics is a determinant of the functional significance of missense variants Proc Natl Acad Sci U S A 2018 115 16 4164 9 10.1073/pnas.1715896115 29610305
Ponzoni L, Bahar I. Structural dynamics is a determinant of the functional significance of missense variants. Proc Natl Acad Sci U S A. 2018;115(16):4164–9. 10.1073/pnas.1715896115.29610305 10.1073/pnas.1715896115
12. Chen H Comprehensive assessment of computational algorithms in predicting cancer driver mutations Genome Biol 2020 21 1 1 17 10.1186/s13059-020-01954-z
Chen H, et al. Comprehensive assessment of computational algorithms in predicting cancer driver mutations. Genome Biol. 2020;21(1):1–17. 10.1186/s13059-020-01954-z.10.1186/s13059-020-01954-z
13. Marian AJ Clinical interpretation and management of genetic variants JACC Basic Transl Sci 2020 5 10 1029 42 10.1016/j.jacbts.2020.05.013 33145465
Marian AJ. Clinical interpretation and management of genetic variants. JACC Basic Transl Sci. 2020;5(10):1029–42. 10.1016/j.jacbts.2020.05.013.33145465 10.1016/j.jacbts.2020.05.013
14. Petukh M, Kucukkal TG, Alexov E. On human disease-causing amino acid variants: statistical study of sequence and structural patterns. Hum Mutat. May 2015;36(5):524–34. 10.1002/humu.22770.
15. Niroula A, Vihinen M. How good are pathogenicity predictors in detecting benign variants? bioRxiv. 2018;1–17. 10.1101/408153.
16. Telenti A Deep sequencing of 10,000 human genomes Proc Natl Acad Sci U S A 2016 113 42 11901 6 10.1073/pnas.1613365113 27702888
Telenti A, et al. Deep sequencing of 10,000 human genomes. Proc Natl Acad Sci U S A. 2016;113(42):11901–6. 10.1073/pnas.1613365113.27702888 10.1073/pnas.1613365113
17. Richards S Standards and guidelines for the interpretation of sequence variants: a joint consensus recommendation of the American College of Medical Genetics and Genomics and the Association for Molecular Pathology Genet Sci 2015 17 5 405 24 10.1038/gim.2015.30
Richards S, et al. Standards and guidelines for the interpretation of sequence variants: a joint consensus recommendation of the American College of Medical Genetics and Genomics and the Association for Molecular Pathology. Genet Sci. 2015;17(5):405–24. 10.1038/gim.2015.30.10.1038/gim.2015.30
18. Li MM Standards and guidelines for the interpretation and reporting of sequence variants in Cancer: a Joint Consensus Recommendation of the Association for Molecular Pathology, American Society of Clinical Oncology, and College of American Pathologists J Mol Diagn 2017 19 1 4 23 10.1016/j.jmoldx.2016.10.002 27993330
Li MM, et al. Standards and guidelines for the interpretation and reporting of sequence variants in Cancer: a Joint Consensus Recommendation of the Association for Molecular Pathology, American Society of Clinical Oncology, and College of American Pathologists. J Mol Diagn. 2017;19(1):4–23. 10.1016/j.jmoldx.2016.10.002.27993330 10.1016/j.jmoldx.2016.10.002
19. Chatrath A, et al. The pan-cancer landscape of prognostic germline variants in 10,582 patients. medRxiv. 2019;1–18. 10.1101/19010264.
20. Bailey MH, et al. Comprehensive Characterization of Cancer Driver Genes and Mutations. Cell. Apr. 2018;173(2):371–85. 10.1016/j.cell.2018.02.060. .e18.
21. Menden MP The germline genetic component of drug sensitivity in cancer cell lines Nat Commun 2018 9 1 1 8 10.1038/s41467-018-05811-3 29317637
Menden MP, et al. The germline genetic component of drug sensitivity in cancer cell lines. Nat Commun. 2018;9(1):1–8. 10.1038/s41467-018-05811-3.29317637 10.1038/s41467-018-05811-3
22. Kucukkal TG Yang Y Chapman SC Cao W Alexov E Computational and experimental approaches to reveal the effects of single nucleotide polymorphisms with respect to disease diagnostics 15 6 2014 10.3390/ijms15069670
Kucukkal TG, Yang Y, Chapman SC, Cao W, Alexov E. Computational and experimental approaches to reveal the effects of single nucleotide polymorphisms with respect to disease diagnostics. 15 6. 2014. 10.3390/ijms15069670.10.3390/ijms15069670
23. Brown DK, Tastan Bishop Ö. The role of structural bioinformatics in drug discovery via computational SNP analysis – a proposed protocol for analyzing variation at the protein level. Glob Heart. Jun. 2017;12(2):151–61. 10.1016/j.gheart.2017.01.009.
24. Ganesan K Kulandaisamy A Binny Priya S Gromiha MM HuVarbase: a human variant database with comprehensive information at gene and protein levels PLoS ONE 2019 14 1 1 7 10.1371/journal.pone.0210475
Ganesan K, Kulandaisamy A, Binny Priya S, Gromiha MM. HuVarbase: a human variant database with comprehensive information at gene and protein levels. PLoS ONE. 2019;14(1):1–7. 10.1371/journal.pone.0210475.10.1371/journal.pone.0210475
25. Landrum MJ, et al. ClinVar: improving access to variant interpretations and supporting evidence. Nucleic Acids Res. 2018;46. 10.1093/nar/gkx1153. D1, pp. D1062–D1067.
26. Karczewski KJ, et al. The mutational constraint spectrum quantified from variation in 141,456 humans. Nature. May 2020;581(7809):434–43. 10.1038/s41586-020-2308-7.
27. Gunning AC, et al. Assessing performance of pathogenicity predictors using clinically relevant variant datasets. J Med Genet. 2020;p. jmedgenet-2020-10700310.1136/jmedgenet-2020-107003.
28. Lek M Analysis of protein-coding genetic variation in 60,706 humans Nature 2016 536 7616 285 91 10.1038/nature19057 27535533
Lek M, et al. Analysis of protein-coding genetic variation in 60,706 humans. Nature. 2016;536(7616):285–91. 10.1038/nature19057.27535533 10.1038/nature19057
29. Stella A, et al. Accurate classification of NF1 gene variants in 84 Italian patients with neurofibromatosis type 1. Genes (Basel). Apr. 2018;9(4):216. 10.3390/genes9040216.
30. Li J Performance evaluation of pathogenicity-computation methods for missense variants Nucleic Acids Res 2018 46 15 7793 804 10.1093/nar/gky678 30060008
Li J, et al. Performance evaluation of pathogenicity-computation methods for missense variants. Nucleic Acids Res. 2018;46(15):7793–804. 10.1093/nar/gky678.30060008 10.1093/nar/gky678
31. Thusberg J, Olatubosun A, Vihinen M. Performance of mutation pathogenicity prediction methods on missense variants, Hum Mutat, vol. 32, no. 4, pp. 358–368, Apr. 2011, 10.1002/humu.21445
32. Grimm DG The evaluation of tools used to predict the impact of missense variants is hindered by two types of circularity Hum Mutat 2015 36 5 513 23 10.1002/humu.22768 25684150
Grimm DG, et al. The evaluation of tools used to predict the impact of missense variants is hindered by two types of circularity. Hum Mutat. 2015;36(5):513–23. 10.1002/humu.22768.25684150 10.1002/humu.22768
33. Riera C Padilla N de la Cruz X The Complementarity between protein-specific and general pathogenicity predictors for amino acid substitutions Hum Mutat 2016 37 10 1013 24 10.1002/humu.23048 27397615
Riera C, Padilla N, de la Cruz X. The Complementarity between protein-specific and general pathogenicity predictors for amino acid substitutions. Hum Mutat. 2016;37(10):1013–24. 10.1002/humu.23048.27397615 10.1002/humu.23048
34. Sherry ST DbSNP: the NCBI database of genetic variation Nucleic Acids Res 2001 29 1 308 11 10.1093/nar/29.1.308 11125122
Sherry ST, et al. DbSNP: the NCBI database of genetic variation. Nucleic Acids Res. 2001;29(1):308–11. 10.1093/nar/29.1.308.11125122 10.1093/nar/29.1.308
35. 1000 T, Consortium GP. A global reference for human genetic variation, Nature, vol. 526, no. 7571, pp. 68–74, Oct. 2015, 10.1038/nature15393
36. Stenson PD et al. The Human Gene Mutation Database: towards a comprehensive repository of inherited mutation data for medical research, genetic diagnosis and next-generation sequencing studies, Human Genetics, vol. 136, no. 6. Springer Verlag, pp. 665–677, Jun. 01, 2017. 10.1007/s00439-017-1779-6
37. Sarkar A Yang Y Vihinen M Variation benchmark datasets: update, criteria, quality and applications Database 2020 2020 1 16 10.1093/database/baz117 33002137
Sarkar A, Yang Y, Vihinen M. Variation benchmark datasets: update, criteria, quality and applications. Database. 2020;2020:1–16. 10.1093/database/baz117.33002137 10.1093/database/baz117
38. Niroula A Vihinen M Harmful somatic amino acid substitutions affect key pathways in cancers BMC Med Genomics 2015 8 1 1 12 10.1186/s12920-015-0125-x 25582225
Niroula A, Vihinen M. Harmful somatic amino acid substitutions affect key pathways in cancers. BMC Med Genomics. 2015;8(1):1–12. 10.1186/s12920-015-0125-x.25582225 10.1186/s12920-015-0125-x
39. Goncearenco A Rager SL Li M Sang QX Rogozin IB Panchenko AR Exploring background mutational processes to decipher cancer genetic heterogeneity Nucleic Acids Res 2017 45 W514 22 10.1093/nar/gkx367 28472504
Goncearenco A, Rager SL, Li M, Sang QX, Rogozin IB, Panchenko AR. Exploring background mutational processes to decipher cancer genetic heterogeneity. Nucleic Acids Res. 2017;45:W514–22. 10.1093/nar/gkx367. no. W1.28472504 10.1093/nar/gkx367
40. Yue Z Zhao L Xia J DbCPM: a manually curated database for exploring the cancer passenger mutations Brief Bioinform 2018 21 1 309 17 10.1093/bib/bby105
Yue Z, Zhao L, Xia J. DbCPM: a manually curated database for exploring the cancer passenger mutations. Brief Bioinform. 2018;21(1):309–17. 10.1093/bib/bby105.10.1093/bib/bby105
41. Sengupta D, Bhattacharya G, Ganguli S, Sengupta M. Structural insights and evaluation of the potential impact of missense variants on the interactions of SLIT2 with ROBO1/4 in cancer progression. Sci Rep. Dec. 2020;10(1):21909. 10.1038/s41598-020-78882-2.
42. Raimondi D Passemiers A Fariselli P Moreau Y Current cancer driver variant predictors learn to recognize driver genes instead of functional variants BMC Biol 2021 19 1 1 13 10.1186/s12915-020-00930-0 33407428
Raimondi D, Passemiers A, Fariselli P, Moreau Y. Current cancer driver variant predictors learn to recognize driver genes instead of functional variants. BMC Biol. 2021;19(1):1–13. 10.1186/s12915-020-00930-0.33407428 10.1186/s12915-020-00930-0
43. Rentzsch P, Witten D, Cooper GM, Shendure J, Kircher M, CADD. Predicting the deleteriousness of variants throughout the human genome. Nucleic Acids Res. Jan. 2019;47:D886–94. 10.1093/nar/gky1016.
44. Quang D, Chen Y, Xie X. DANN: A deep learning approach for annotating the pathogenicity of genetic variants, Bioinformatics, vol. 31, no. 5, pp. 761–763, Mar. 2015, 10.1093/bioinformatics/btu703
45. Shihab HA, et al. An integrative approach to predicting the functional effects of non-coding and coding sequence variation. Bioinformatics. May 2015;31(10):1536–43. 10.1093/bioinformatics/btv009.
46. Shihab HA et al. Jan., Predicting the Functional, Molecular, and Phenotypic Consequences of Amino Acid Substitutions using Hidden Markov Models, Hum Mutat, vol. 34, no. 1, pp. 57–65, 2013, 10.1002/humu.22225
47. Rogers MF, Shihab HA, Mort M, Cooper DN, Gaunt TR, Campbell C. FATHMM-XF: Accurate prediction of pathogenic point mutations via extended features, Bioinformatics, vol. 34, no. 3, pp. 511–513, Feb. 2018, 10.1093/bioinformatics/btx536
48. Lu Q, Hu Y, Sun J, Cheng Y, Cheung KH, Zhao H. A statistical framework to predict functional non-coding regions in the human genome through integrated analysis of annotation data. Sci Rep. May 2015;5. 10.1038/srep10576.
49. Alirezaie N, Kernohan KD, Hartley T, Majewski J, Hocking TD. Am J Hum Genet. Oct. 2018;103(4):474–83. 10.1016/j.ajhg.2018.08.005. ClinPred: Prediction Tool to Identify Disease-Relevant Nonsynonymous Single-Nucleotide Variants.
50. McLaren W, et al. The Ensembl variant effect predictor. Genome Biol. Jun. 2016;17(1). 10.1186/s13059-016-0974-4.
51. Malhis N Jacobson M Jones SJM Gsponer J LIST-S2 Taxonomy based sorting of deleterious missense mutations across species Nucleic Acids Res 2020 48 W154 61 10.1093/NAR/GKAA288 32352516
Malhis N, Jacobson M, Jones SJM, Gsponer J, LIST-S2. Taxonomy based sorting of deleterious missense mutations across species. Nucleic Acids Res. 2020;48:W154–61. 10.1093/NAR/GKAA288. no. W1.32352516 10.1093/NAR/GKAA288
52. Chun S, Fay JC. Identification of deleterious mutations within three human genomes, Genome Res, vol. 19, no. 9, pp. 1553–1561, Sep. 2009, 10.1101/gr.092619.109
53. Jagadeesh KA et al. Dec., M-CAP eliminates a majority of variants of uncertain significance in clinical exomes at high sensitivity, Nat Genet, vol. 48, no. 12, pp. 1581–1586, 2016, 10.1038/ng.3703
54. Dong C et al. Apr., Comparison and integration of deleteriousness prediction methods for nonsynonymous SNVs in whole exome sequencing studies, Hum Mol Genet, vol. 24, no. 8, pp. 2125–2137, 2015, 10.1093/hmg/ddu733
55. Li C, Zhi D, Wang K, Liu X. MetaRNN: differentiating rare pathogenic and rare benign missense SNVs and InDels using deep learning. Genome Med. Dec. 2022;14(1). 10.1186/s13073-022-01120-z.
56. Pejaver V, et al. Inferring the molecular and phenotypic impact of amino acid variants with MutPred2. Nat Commun. Dec. 2020;11(1). 10.1038/s41467-020-19669-x.
57. Qi H, et al. MVP predicts the pathogenicity of missense variants by deep learning. Nat Commun. Dec. 2021;12(1). 10.1038/s41467-020-20847-0.
58. Sundaram L et al. Aug., Predicting the clinical impact of human mutation with deep neural networks, Nat Genet, vol. 50, no. 8, pp. 1161–1170, 2018, 10.1038/s41588-018-0167-z
59. Choi Y, Chan AP. PROVEAN web server: a tool to predict the functional effect of amino acid substitutions and indels. Bioinformatics. Jan. 2015;31:2745–7. 10.1093/bioinformatics/btv195.
60. Ng PC, Henikoff S, SIFT. Jul., : Predicting amino acid changes that affect protein function, Nucleic Acids Res, vol. 31, no. 13, pp. 3812–3814, 2003, 10.1093/nar/gkg509
61. Vaser R, Adusumalli S, Leng SN, Sikic M, Ng PC. SIFT missense predictions for genomes. Nat Protoc. Jan. 2016;11(1):1–9. 10.1038/nprot.2015.123.
62. Liu X, Li C, Mou C, Dong Y, Tu Y. dbNSFP v4: a comprehensive database of transcript-specific functional predictions and annotations for human nonsynonymous and splice-site SNVs. Genome Med. Dec. 2020;12(1). 10.1186/s13073-020-00803-9.
63. Gu F A suite of automated sequence analyses reduces the number of candidate deleterious variants and reveals a difference between probands and unaffected siblings Genet Sci 2018 10.1038/s41436
Gu F, et al. A suite of automated sequence analyses reduces the number of candidate deleterious variants and reveals a difference between probands and unaffected siblings. Genet Sci. 2018. 10.1038/s41436.10.1038/s41436
64. Sun H, Yu G. New insights into the pathogenicity of non-synonymous variants through multi-level analysis. Sci Rep. Dec. 2019;9(1). 10.1038/s41598-018-38189-9.
65. Pejaver V, Mooney SD, Radivojac P. Missense variant pathogenicity predictors generalize well across a range of function-specific prediction challenges, Hum Mutat, vol. 38, no. 9, pp. 1092–1108, Sep. 2017, 10.1002/humu.23258
66. Ensembl. Variant Effect Predictor. Accessed: Apr. 05, 2023. [Online]. Available: https://grch37.ensembl.org/Tools/VEP
67. Pedregosa F Scikit-learn: machine learning in Python J Mach Learn Res 2011 12 2825 30 10.1289/EHP4713
Pedregosa F, et al. Scikit-learn: machine learning in Python. J Mach Learn Res. no. 2011;12:2825–30. 10.1289/EHP4713.10.1289/EHP4713
68. Sondka Z Bamford S Cole CG Ward SA Dunham I Forbes SA The COSMIC Cancer Gene Census: describing genetic dysfunction across all human cancers Nat Rev Cancer 2018 18 11 696 705 10.1038/s41568-018-0060-1 30293088
Sondka Z, Bamford S, Cole CG, Ward SA, Dunham I, Forbes SA. The COSMIC Cancer Gene Census: describing genetic dysfunction across all human cancers. Nat Rev Cancer. 2018;18(11):696–705. 10.1038/s41568-018-0060-1.30293088 10.1038/s41568-018-0060-1
69. Öfverholm A, et al. Extended genetic analysis and tumor characteristics in over 4600 women with suspected hereditary breast and ovarian cancer. BMC Cancer. Dec. 2023;23(1). 10.1186/s12885-023-11229-y.
70. Breast Cancer Risk Genes — Association Analysis in More than 113,000 Women, New England Journal of Medicine, vol. 384, no. 5, pp. 428–439. Feb. 2021, 10.1056/NEJMoa1913948
71. McNamara LA, Martin SW. Principles of Epidemiology and Public Health, Fifth edit. Elsevier Inc.; 2018. 10.1016/B978-0-323-40181-4.00001-3.
72. Sahin IE The sensitivity and specificity of the balance evaluation systems test-BESTest in determining risk of fall in stroke patients NeuroRehabilitation 2019 44 1 67 77 10.3233/NRE-182558 30814369
Sahin IE, et al. The sensitivity and specificity of the balance evaluation systems test-BESTest in determining risk of fall in stroke patients. NeuroRehabilitation. 2019;44(1):67–77. 10.3233/NRE-182558.30814369 10.3233/NRE-182558
73. Vihinen M. How to evaluate performance of prediction methods? Measures and their interpretation in variation effect analysis. BMC Genomics. 2012;13 Suppl 4(no Suppl 4). 10.1186/1471-2164-13-S4-S2.
74. Gunning AC et al. Aug., Assessing performance of pathogenicity predictors using clinically relevant variant datasets, J Med Genet, vol. 58, no. 8, pp. 547–555, 2021, 10.1136/jmedgenet-2020-107003
75. Cannon S, Williams M, Gunning AC, Wright CF. Evaluation of in silico pathogenicity prediction tools for the classification of small in-frame indels. BMC Med Genomics. Dec. 2023;16(1). 10.1186/s12920-023-01454-6.
76. Sayeed MA, Aldarmaki H, Ben Amor B. Gene pathogenicity prediction using genomic Foundation models, 2024. [Online]. Available: www.aaai.org.
77. Tarnovskaya SI, Kostareva AA, Zhorov BS. In silico analysis of TRPM4 variants of unknown clinical significance, PLoS One, vol. 18, no. 12 DECEMBER, Dec. 2023, 10.1371/journal.pone.0295974
78. Khandakji MN, Mifsud B. Gene-specific machine learning model to predict the pathogenicity of BRCA2 variants. Front Genet. Sep. 2022;13. 10.3389/fgene.2022.982930.
79. Lai J, Yang J, Gamsiz Uzun ED, Rubenstein BM, Sarkar IN. LYRUS: a machine learning model for predicting the pathogenicity of missense variants. Bioinf Adv. Jan. 2022;2(1). 10.1093/bioadv/vbab045.
80. Yazar M, Ozbek P. Assessment of 13 in silico pathogenicity methods on cancer-related variants. Comput Biol Med. Jun. 2022;145. 10.1016/j.compbiomed.2022.105434.
81. Poon KS. In silico analysis of BRCA1 and BRCA2 missense variants and the relevance in molecular genetic testing. Sci Rep. Dec. 2021;11(1). 10.1038/s41598-021-88586-w.
