
==== Front
Brief Bioinform
Brief Bioinform
bib
Briefings in Bioinformatics
1467-5463
1477-4054
Oxford University Press

10.1093/bib/bbae097
bbae097
Problem Solving Protocol
AcademicSubjects/SCI01060
FusionNW, a potential clinical impact assessment of kinases in pan-cancer fusion gene network
Yang Chengyuan School of Public Health, The University of Texas Health Science Center at Houston, Houston, TX 77030, USA

https://orcid.org/0000-0002-4335-4517
Kumar Himansu McWilliams School of Biomedical Informatics, The University of Texas Health Science Center at Houston, Houston, TX 77030, USA

https://orcid.org/0000-0002-8321-6864
Kim Pora McWilliams School of Biomedical Informatics, The University of Texas Health Science Center at Houston, Houston, TX 77030, USA

Corresponding author. Pora Kim, Department of Bioinformatics and Systems Medicine, McWilliams School of Biomedical Informatics, The University of Texas Health Science Center at Houston, 7000 Fannin St., Houston, TX 77030, USA. Tel.: +713-500-3636; E-mail: Pora.Kim@uth.tmc.edu
3 2024
16 3 2024
16 3 2024
25 2 bbae0976 12 2023
16 2 2024
22 2 2024
© The Author(s) 2024. Published by Oxford University Press.
2024
https://creativecommons.org/licenses/by-nc/4.0/ This is an Open Access article distributed under the terms of the Creative Commons Attribution Non-Commercial License (https://creativecommons.org/licenses/by-nc/4.0/), which permits non-commercial re-use, distribution, and reproduction in any medium, provided the original work is properly cited. For commercial re-use, please contact journals.permissions@oup.com

Abstract

Kinase fusion genes are the most active fusion gene group in human cancer fusion genes. To help choose the clinically significant kinase so that the cancer patients that have fusion genes can be better diagnosed, we need a metric to infer the assessment of kinases in pan-cancer fusion genes rather than relying on the sample frequency expressed fusion genes. Most of all, multiple studies assessed human kinases as the drug targets using multiple types of genomic and clinical information, but none used the kinase fusion genes in their study. The assessment studies of kinase without kinase fusion gene events can miss the effect of one of the mechanisms that enhance the kinase function in cancer. To fill this gap, in this study, we suggest a novel way of assessing genes using a network propagation approach to infer how likely individual kinases influence the kinase fusion gene network composed of ~5K kinase fusion gene pairs. To select a better seed of propagation, we chose the top genes via dimensionality reduction like a principal component or latent layer information of six features of individual genes in pan-cancer fusion genes. Our approach may provide a novel way to assess of human kinases in cancer.

kinase
fusion gene
gene assessment
network propagation
variational autoencoder
feature reduction
National Institutes of Health 10.13039/100000002 R35GM138184 University of Texas Health Science Center at Houston 10.13039/100012615
==== Body
pmcINTRODUCTION

The fusion gene, formed by diverse types of structural variants induced by the double-strand DNA breakages, is one of the hallmarks of the cancer genome. So far, there are more than 103K fusion genes reported from transcriptomics data analysis of cancer cells, patients and cancer cohorts [1]. Out of this big number of existing fusion genes, it is necessary to assess the clinical importance of individual genes. One of the major criteria of high-impact mutation genes was having a high sample frequency. However, gene fusion, which has a unique pair of two different loci genes, has a unique context in terms of multiple combinations of two genes, breakpoints and expression in multiple cancers of partner genes. Previously, we suggested multiple ways of ranking the impact of individual human genes that are involved in fusion genes. For example, individual genes can involve in the formation of fusion genes from multiple breakpoints with multiple partner genes in multiple cancer types. Based on these numerical features, we introduced a concept of the degree of frequency (DoF) score, which is the product of three values above each gene in our previous study. For example, in the application of this DoF to ~500 human kinomes, we identified multiple features of driver kinase fusion genes [2]. We also introduced another metric, called the Major Active Isofusion Index (MAII), which is calculated by dividing the number of observed samples with a particular fusion gene by its DoF score [3]. Here, the isofusion refers to one particular gene fusion combination, with one particular partner gene and one particular breakpoint, in one particular cancer type. This new score (MAII) can give us the average frequency of each gene for each possible isofusion. By applying the MAII scoring to the transcription factor fusion genes, we found PLAU, which encodes plasminogen activator urokinase and serves as a biomarker for tumor invasion, was found to be consistently activated in the samples with the highest MAII scores. However, these trials are not utilizing all numerical features of individual genes in pan-cancer fusion genes. To better rank the importance of individual genes in pan-cancer fusion genes, we can utilize all existing features.

Kinase fusion genes are the driver fusion genes with therapeutics of kinase inhibitors. In this study, we made a novel gene assessment scoring scheme in a pan-cancer kinase fusion gene network using network propagation and dimensionality reduction approaches. To use six numerical features and present one numerical value of individual genes in pan-cancer fusion genes, we performed multiple dimensionality reduction approaches [i.e. principal component analysis (PCA) and variational autoencoder (VAE) model]. Then, we selected the top 5% genes as the seed genes in the kinase fusion network and performed the network propagation by random walk process. From this process, we can infer how the seed genes influence to form of a kinase gene fusion. The output value per gene can be interpreted as the kinase fusion relevance score. Comparing with the previous metrics (such as fusion-positive sample frequency, DoF score and MAII score) of assessment of fusion panel genes and ClinVar genes, we found that our new method, FusionNW, showed the largest overlap with the fusion panel target genes and ClinVar genes. We hope our study can help in the assessment of human kinases in pan-cancer fusion genes.

Human kinase fusion gene network

In humans, there are 7273 reported kinase fusion genes, which were fused among 5578 genes including 501 kinases (Supplementary Table S1). This fusion gene information was downloaded from our human fusion gene knowledgebase (FusionGDB2.0). For each gene, we have six features that were calculated based on the identified fusion genes such as the number of cancer types, number of partner genes, number of breakpoints, number of expressed samples, DoF and MAII scores (Supplementary Table S2). Figure 1 shows the sorted information of the first four values per kinase group (number of cancer types, number of partner genes, number of breakpoints, number of expressed samples). This figure provides a landscape of the activity of involvement in gene fusion of individual kinases across nine different kinase groups. Based on the number of samples, the TK (tyrosine kinase) and TKL (tyrosine kinase–like) groups were the most actively involved kinase groups in human fusion genes. The light-yellow background kinases in the x-axis labels of these plots are the ones that were expressed fusion genes in more than 20 samples. For example, in TK group, ABL1, RET, ALK, TYK2, ERBB2, ROS1, PTK2, NTRK1, NTRK3, EGFR, FYN, MET, ERBB3, JAK2, TYRO3, FGFR1, INSR, JAK1, ERBB4 and FGFR2 were expressed in more than 20 samples. In the TKL group, there are BRAF, RAF1, MAP3K13 and KSR1. In the Atypical group, there are BCR, BRD4, SMG1, ABR, MTOR PRKDC and BRD2. In the CMGC [cyclin-dependent kinases (CDK), mitogen-activated protein kinases (MAPK), glycogen synthase kinase (GSK3) and CDC-like kinase (CLK) group of protein kinases], there are CDK12, CDK14 and CDK19. In the Other group, there are EIF2AK2, EIF2AK1 and ULK4. In the AGC group, there are PRKCA, PRKCH, PRKCB, GRK5 and AKT3. In the STE group, there are PAK1, MAP2K5, TAOK3, STK10, MAP3K5, STK24 and PAK2. In the CK1 group, there are CSNK1D and CSNK1A1. Figure 2A shows the human kinase fusion gene network composed of 5578 genes including 501 kinases. The network diameter is 22, and there are 1485 neighbor genes on average. In Figure 2B, we visualized the same kinase fusion gene network with different weights on the node color of individual genes reflecting our FusionNW output scores. The red nodes are the high-scored genes from FusionNW analysis.

Figure 1 Kinases involved in fusion genes expressed in more than five samples per kinase group. The shaded kinases were expressed in more than 20 samples.

Figure 2 Network of kinase fusion genes in humans. (A) Human kinase fusion gene–gene network without prioritization. (B) Human kinase fusion gene–gene network with our suggested FusionNW assessment scores in the node color.

Kinase impact assessment in pan-cancer kinase fusion genes with network propagation Currently, we have 7273 kinase fusion genes in ~102K all known fusion genes from FusionGDB2.0. To highlight the actively involved kinases in pan-cancer fusion genes, we introduced the DoF and the MAII scores that were in our previous study on investigating the affected downstream genes by the TF fusion genes [3]. However, these scores are the transformed values through the basic calculations of the number of involved cancer types, number of partner genes and number of breakpoints of individual kinases, which do not have any learning processes on the activeness of individual kinases in the pan-cancer. In this study, we tried to highlight the kinases that were actively involved in the formation of fusion genes in pan-cancer by using a network propagation algorithm in the kinase fusion gene–gene network map. A recent study used network propagation (i.e., random walk with restart [RWR] [4, 5]) to map causal variants to their relevant cellular context at single-cell resolution. In our study, we tried application of this concept to assess the kinase impact across the pan-cancer fusion genes. Figure 3A shows the overall process of FusionNW, and Figure 3B shows the example six features of multiple kinases.

Figure 3 Overview of FusionNW. This overview shows the kinase fusion gene network propagation with dimension reduction of the six features of genes in pan-cancer. (A) Overall process of FusionNW. (B) Example six feature information of the top 10 most frequent kinases in pan-cancer fusion genes. (C) PCA visualization of the FusionNW output scores based on the PCA feature reduction-based ranking of genes. We performed PCA analysis for six features of individual genes. (D) Overview of making a ranking of genes involved in human fusion genes using variational autoencoder trained based on the six gene features. (E) Comparison of the error between PCA or variational autoencoder for the feature reduction.

First, we constructed the kinase fusion gene (gene–gene) network using 7273 kinases fusion gene pairs composed of 5579 unique genes. To initiate the propagation, we needed to choose the seed genes. To choose the seed genes, we performed two different feature dimensionality reduction analyses (i.e. PCA and VAE) based on the six features (i.e. the number of cancer types, gene partners, breakpoints, sample frequencies, DoF scores and MAII scores) of all 20K individual genes (Supplementary Table S3) that are involved in the formation of ~102K human fusion genes and ranked the genes from the PCA transformed values and latent information, using PCA and VAE, respectively. Figure 3C shows the scatter plot of the reduced feature values in the three major principal components of 3D.

Figure 3D shows the overview of the VAE algorithm. First, we made a symmetric matrix by multiplication of standardized N genes by six features and six features by N genes (20K genes). This matrix carries information about pairwise relations of the features so that the VAE model could focus on capturing patterns and structures in these relationships. The latent distribution represents how each gene is related to the other. This relationship is more useful for selecting seed genes since we want to use the most ‘relevant’ genes as the initial nodes of network propagation. If using gene*features (20 k*6) only, the VAE learns a lower-dimensional latent distribution of each gene focusing on the relationships based on features. After optimization of the VAE model, we took the latent information of N genes. We regarded this reduced feature space as the ranking of individual genes that were involved in human cancer fusion genes. After overlapping with 5579 genes making kinase fusion genes, we chose the top 5% ranked genes as the seed of the kinase fusion gene network propagation and got the final KFG relevance scores. We further compare this VAE model with the PCA. In our implementation of the PCA algorithm, we utilized the standardized dataset comprising N genes across six features. The PCA function from Scikit-learn [6] was applied on this dataset. The dimensionality of the latent vector was subsequently reduced from six features. To evaluate the performance of this reduction, we reconstructed the latent vectors and computed the mean square error.

Based on these two different seeds from PCA and VAE, we performed the network propagation until the absolute difference between the current and next network structure was less than \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $1\ast{10}^{-5}$\end{document}. The final stable network’s stationary probabilities are the list of the network propagation scores. Each model shows how the seed genes influence a kinase fusion gene. Therefore, a kinase gene with a higher network propagation score indicates it has a greater relevance with the fusion genes. In this way, we calculated the kinase fusion gene relevance scores per gene. We compared the performance of this fusion gene profile feature reduction approach to other approaches like PCA by checking the mean squared error (MSE) across different numbers of the principal components. As shown in Figure 3E, the MSE of PCA with one component is 1.75, which is less than VAE’s MSE. However, since VAE extract the non-linear latent variables and PCA extracts the linear latent variables, it is worthy to compare their assessment score after the FusionNW network propagation step.

FusionNW interpretation in the clinical assessment

The PCA-based seed gene selection for the KFG network propagation outperformed the one using the seed genes based on the VAE latent information in the kinase fusion assessment as shown in the difference of the output scores between kinases versus non-kinases, fusion panel genes versus non-fusion panel genes and ClinVar genes versus non-ClinVar genes (Figure 4A, Supplementary Table S4). Next, we checked the binomial linearity using the logistic regression whether the gene is in the fusion gene panel or not based on multiple metrics of fusion gene assessments such as the number of samples, DoF, MAII and FusionNW output values in fusion genes of 501 human kinases. As shown in Figure 4B, The FusionNW based on the PCA seed gene selection method showed a similar pattern with the logistic regression line using the number of samples and showed the most statistical significance. We also checked the relationship between our FusionNW output score versus six individual features (Figure 4C). There was a highly significant correlation. Last, for the top half of 500 kinases chosen by multiple fusion gene assessment metrics (number of samples, DoF, MAII and FusionNW output values), we counted the number ClinGenes and TruSight RNA fusion panel target genes. The output values of the FusionNW were specific to the fusion panel target gene selection compared to the ClinGene selection (Figure 4D). All these statistical analysis results show the relevance of our FusionNW output scores to assess the impact of individual genes in pan-cancer fusion genes.

Figure 4 Interpretation of the KFG relevance scores. (A) Difference of KFG relevance scores in the clinical gene groups (between kinase and non-kinase, fusion gene target panel or not, ClinGene or not). (B) Logistic regression between fusion gene target panel and four assessment criteria (# samples, DoF, MAII, VAE-based and PCA-based FusionNW KFG relevance scores). (C) Correlation between six features and final KFG relevance scores of all kinase fusion network genes. (D) Comparison of selection of ClinGene or TruSight RNA fusion panel genes from multiple gene assessment metrics.

Potential fusion panel target genes from FusionNW analysis

Table 1 shows the FusionNW output score-based top 100 kinases related information to check the potential assessment in the kinase fusion genes. Overall, almost of the kinases were targeted by drugs as shown in the column, named ‘International Union of Basic and Clinical Pharmacology (IUPHAR)’ (96/100). Nineteen of them were ClinGenes, and 16 were included in the TruSeq fusion panel target gene list. The fusion genes with these kinases were expressed in 45 cell lines or patient samples on averages. Among these, we checked the top 10 kinases that were not listed in the fusion panel target genes such as PTK2, ERBB2, PAK1, IGF1R, MELK, BMPR2, STK39, CDK12, CAMK1D and CDK19. For these kinases, we checked what fusion genes are existing with how many frequencies in Figure 5. The center is the kinase and the other nodes are the partner genes in fusion genes. The width of the edges reflects the fusion gene expressed sample frequencies. It seems PTK2 and ERBB2 forms diverse fusion genes with high frequency. Protein tyrosine kinase 2 (PTK2) is a focal adhesion-associated protein kinase (FAK) that plays an important role in regulation of cellular responses such as cell adhesion, migration and survival [7]. PTK2 fusion genes are reported in the most common alterations in PTK2, such as PTK2 mutation (1.37%), PTK2 amplification (0.46%), PTK2 fusion (0.05%), PTK2-AGO2 fusion (0.01%) and PTK2 E789K (0.02%) from AACR Project GENIE [8]. Amplification or overexpression of HER2 (ERBB2) occurs in approximately 15–30% of breast cancers and 10–30% of gastric/gastroesophageal cancers and serves as a prognostic and predictive biomarker. ERBB2 Amplification is present in 3.30% of AACR GENIE cases, with breast-invasive ductal carcinoma, invasive breast carcinoma, esophageal adenocarcinoma, lung adenocarcinoma and colon adenocarcinoma having the greatest prevalence [8]. Among the patients with ERBB2 fusions in 32 131 Chinese patients with solid tumors, most patients with ERBB2 fusion also had ERBB2 amplification (75.28%; 67/89), which was similar to the data in the The Cancer Genome Atlas Program (TCGA) database (88.00%; 44/50) [9]. From these reports, we can see the importance of these fusions. With the accumulated genomic sequencing data of cancer patients, these candidate kinases may have potential to be identified more frequently. We may consider these kinases also in the fusion gene panel list in the future.

Table 1 Related information of top 100 KFG relevance scored kinases in pan-cancer fusion genes including gene group, six features and assessment metrics

Gene	Kinase	IUPHAR	ClinGene	ChimerKB4	Panel	Curated fusion protein	n_ctypes	n_partners	n_bps	DoF	n_samples	MAII	log2(output+1)	
PTK2	1	1	0	1	0	0	18	60	57	64,800	99	−9.35	1.14E-04	
ERBB2	1	1	0	0	0	0	17	80	85	######	129	−9.72	1.11E-04	
PAK1	1	1	0	1	0	0	13	53	36	36,517	68	−9.07	9.91E-05	
IGF1R	1	1	0	0	0	0	8	41	31	13,448	41	−8.36	9.10E-05	
ABL1	1	1	0	1	1	1	17	42	203	29,988	260	−6.85	9.03E-05	
MELK	1	1	0	0	0	0	7	87	31	52,983	102	−9.02	7.94E-05	
BRAF	1	1	1	1	1	1	21	70	82	######	106	−9.92	7.59E-05	
BCR	1	1	0	1	1	1	20	44	206	38,720	266	−7.19	7.51E-05	
BMPR2	1	1	1	0	0	0	9	33	26	9801	42	−7.87	7.25E-05	
SIK3	1	1	0	0	1	0	11	27	23	8019	34	−7.88	6.86E-05	
STK39	1	1	0	0	0	0	16	28	20	12,544	35	−8.49	6.62E-05	
CDK12	1	1	0	1	0	0	19	67	45	85,291	90	−9.89	6.61E-05	
CAMK1D	1	1	0	0	0	0	10	43	42	18,490	48	−8.59	6.57E-05	
CDK19	1	1	0	0	0	0	9	27	25	6561	36	−7.51	6.40E-05	
ROR1	1	1	1	0	0	0	7	21	15	3087	23	−7.07	6.36E-05	
FGFR2	1	1	1	0	1	1	14	33	30	15,246	55	−8.11	6.26E-05	
ERBB3	1	1	0	0	1	0	21	46	31	44,436	55	−9.66	6.08E-05	
MAPK10	1	1	0	0	0	0	10	23	21	5290	38	−7.12	6.07E-05	
CDK14	1	1	0	0	0	0	15	41	39	25,215	55	−8.84	6.00E-05	
EGFR	1	1	1	1	1	0	15	64	49	61,440	117	−9.04	5.90E-05	
TAOK1	1	1	1	0	1	0	20	56	35	62,720	70	−9.81	5.87E-05	
SMG1	1	1	0	0	0	0	11	48	42	25,344	60	−8.72	5.80E-05	
CSNK1D	1	1	0	0	0	0	17	36	35	22,032	51	−8.75	5.77E-05	
RPS6KB1	1	1	0	1	0	0	17	23	23	8993	91	−6.63	5.71E-05	
TAOK3	1	1	0	0	0	0	8	32	31	8192	42	−7.61	5.53E-05	
CDK6	1	1	0	1	1	1	11	24	23	6336	27	−7.87	5.49E-05	
PRKCA	1	1	0	1	0	0	17	70	45	83,300	75	−10.12	5.37E-05	
TRIO	1	1	1	0	0	0	13	43	37	24,037	50	−8.91	5.33E-05	
FGFR1	1	1	1	1	1	1	12	41	59	20,172	43	−8.87	5.22E-05	
STK25	1	1	0	0	0	0	14	17	8	4046	64	−5.98	5.22E-05	
SRPK2	1	1	0	0	0	0	13	35	30	15,925	41	−8.6	5.03E-05	
PRKG1	1	1	1	0	0	0	7	27	26	5103	29	−7.46	4.90E-05	
TTN	1	1	1	0	0	0	8	23	27	4232	26	−7.35	4.76E-05	
PRKAA1	1	1	0	1	0	0	6	15	14	1350	18	−6.23	4.74E-05	
MAST2	1	1	0	1	0	1	18	40	32	28,800	47	−9.26	4.68E-05	
DYRK1A	1	1	1	0	0	0	12	46	35	25,392	59	−8.75	4.62E-05	
ADRBK2	1	0	0	0	0	0	8	12	6	1152	16	−6.17	4.62E-05	
JAK2	1	1	0	1	1	1	15	33	38	16,335	36	−8.83	4.61E-05	
VRK1	1	1	0	0	0	0	9	20	12	3600	27	−7.06	4.52E-05	
AKT3	1	1	1	1	0	0	12	38	34	17,328	45	−8.59	4.32E-05	
TTBK2	1	1	0	0	0	0	9	18	17	2916	20	−7.19	4.27E-05	
HIPK2	1	1	0	0	0	0	14	26	23	9464	38	−7.96	4.24E-05	
MAP3K5	1	1	0	0	0	0	13	30	29	11,700	40	−8.19	4.16E-05	
JAK1	1	1	0	0	1	0	8	32	30	8192	39	−7.71	4.07E-05	
MAP4K3	1	1	0	1	0	0	12	19	15	4332	27	−7.33	4.00E-05	
WNK1	1	1	0	0	0	0	16	35	35	19,600	58	−8.4	3.98E-05	
GSK3B	1	1	0	0	0	0	13	35	25	15,925	43	−8.53	3.94E-05	
STK32A	1	1	0	0	0	0	4	13	12	676	12	−5.82	3.91E-05	
NTRK2	1	1	0	1	1	1	10	19	20	3610	20	−7.5	3.88E-05	
SGK1	1	1	0	0	0	0	7	15	19	1575	23	−6.1	3.81E-05	
MLKL	1	1	0	0	0	0	2	18	13	648	20	−5.02	3.80E-05	
TLK2	1	1	0	0	0	0	12	33	28	13,068	48	−8.09	3.80E-05	
MAST4	1	1	0	0	0	0	9	22	23	4356	27	−7.33	3.79E-05	
TYRO3	1	1	0	0	0	0	6	9	8	486	33	−3.88	3.78E-05	
ICK	1	0	0	0	0	0	5	6	7	180	7	−4.68	3.78E-05	
TGFBR2	1	1	1	0	0	0	8	14	12	1568	17	−6.53	3.78E-05	
ALK	1	1	1	1	1	1	22	68	86	######	109	−9.87	3.69E-05	
INSR	1	1	0	0	0	0	9	32	24	9216	36	−8	3.67E-05	
CSNK1E	1	1	0	0	0	0	12	25	16	7500	30	−7.97	3.66E-05	
EIF2AK2	1	1	0	0	0	0	17	48	15	39,168	53	−9.53	3.51E-05	
UHMK1	1	1	0	0	0	0	4	12	12	576	13	−5.47	3.50E-05	
SRC	1	1	1	0	0	0	7	19	14	2527	23	−6.78	3.49E-05	
EEF2K	1	1	0	0	0	0	10	13	12	1690	19	−6.47	3.49E-05	
PAN3	1	0	0	0	0	0	14	35	22	17,150	40	−8.74	3.46E-05	
STK24	1	1	0	1	0	0	16	35	17	19,600	50	−8.61	3.41E-05	
BAZ1A	1	1	0	0	0	0	7	24	19	4032	26	−7.28	3.40E-05	
FYN	1	1	0	0	0	0	14	30	24	12,600	41	−8.26	3.39E-05	
PRKACB	1	1	0	0	0	0	7	19	15	2527	20	−6.98	3.35E-05	
ULK4	1	1	0	0	0	0	9	25	31	5625	34	−7.37	3.31E-05	
WEE1	1	1	0	0	0	0	9	12	14	1296	18	−6.17	3.23E-05	
DAPK2	1	1	0	0	0	0	10	17	20	2890	28	−6.69	3.22E-05	
ABR	1	0	0	0	0	0	19	40	31	30,400	72	−8.72	3.18E-05	
MAP4K4	1	1	0	0	0	0	13	28	26	10,192	33	−8.27	3.15E-05	
AKT2	1	1	0	1	0	0	14	24	22	8064	27	−8.22	3.11E-05	
ALPK1	1	1	0	0	0	0	6	9	13	486	13	−5.22	3.11E-05	
MAP4K5	1	1	0	0	0	0	10	26	27	6760	36	−7.55	3.08E-05	
CDC42BPA	1	1	0	0	0	0	10	23	20	5290	23	−7.85	3.07E-05	
STK17A	1	1	0	0	0	0	6	8	11	384	13	−4.88	3.04E-05	
MARK2	1	1	0	0	0	0	12	19	15	4332	32	−7.08	3.03E-05	
CAMK2D	1	1	0	0	0	0	10	17	18	2890	21	−7.1	3.01E-05	
RET	1	1	1	1	1	1	12	39	39	18,252	98	−7.54	3.00E-05	
MAP2K1	1	1	1	0	0	0	7	14	14	1372	15	−6.52	2.97E-05	
SCYL1	1	1	0	0	0	0	9	13	16	1521	24	−5.99	2.93E-05	
PTK7	1	1	0	0	0	0	13	25	17	8125	32	−7.99	2.90E-05	
RIPK4	1	1	0	0	0	0	3	6	6	108	7	−3.95	2.89E-05	
PKN1	1	1	0	0	0	0	9	18	18	2916	26	−6.81	2.85E-05	
DCLK1	1	1	0	0	0	0	8	24	22	4608	25	−7.53	2.83E-05	
CAMKK2	1	1	0	0	0	0	10	17	22	2890	30	−6.59	2.83E-05	
EIF2AK4	1	1	1	0	0	0	13	19	30	4693	36	−7.03	2.81E-05	
PRKCB	1	1	0	0	0	0	13	29	24	10,933	36	−8.25	2.80E-05	
MAP3K14	1	1	0	0	0	0	8	14	15	1568	18	−6.44	2.79E-05	
SCYL2	1	1	0	0	0	0	10	16	18	2560	29	−6.46	2.76E-05	
PDGFRA	1	1	1	1	1	1	10	26	34	6760	28	−7.92	2.76E-05	
EPHA7	1	1	0	0	0	0	5	9	8	405	10	−5.34	2.72E-05	
ATR	1	1	0	0	0	0	8	16	15	2048	22	−6.54	2.68E-05	
NRBP1	1	1	0	0	0	0	11	16	17	2816	26	−6.76	2.67E-05	
TRPM7	1	1	0	0	0	0	8	26	26	5408	31	−7.45	2.67E-05	
MAP2K5	1	1	0	0	0	0	14	24	29	8064	37	−7.77	2.66E-05	
MOK	1	1	0	0	0	0	7	18	18	2268	20	−6.83	2.66E-05	
CSNK1A1	1	1	0	0	0	0	12	32	36	12,288	45	−8.09	2.60E-05	

Figure 5 TOP 11 FusionNW relevance scored kinases not listed in the fusion gene target panel. The node represents the gene involved in each kinase fusion gene and the edge is the fusion gene pair. The width of the edges represents the number of expressed samples of fusion genes.

From Table 1, we checked any kinases that were not listed in the fusion panel target genes but manually curated fusion proteins. Out of these top 100 kinases, there was one kinase, MAST2. As shown in Figure 5, MAST2 has 37 known fusion genes with 37 different partner genes. Out of these, 11 are annotated as potentially being translated and make fusion proteins. Those are ARID1A-MAST2, MAST2-COQ6, MAST2-NDC1, MAST2-ST3GAL3, FOS-MAST2, HOMER1-MAST2, MAST2-CARD16, MAST2-HP1BP3, MAST2-SERBP1, MAST2-TM4SF1 and RAD54L-MAST2. Three fusion genes (MAST2-COQ6, MAST2-NDC1, MAST2-ST3GAL3) were expressed in more than two samples, and ARID1A-MAST2 was the one that were manually curated from ChimerKB4 in ChimerDB4.

New kinase assessment map through the fusions

Next, we tried to make a new landscape of the assessment of individual human kinases in pan-cancer fusion genes to provide intuitive visualization of our result. In Figure 6, we made a circus plot of TK group kinases using their four-assessment metrics in pan-cancer. From this assessment landscape, one can pick up relatively highly implacable kinases in pan-cancer fusion genes. From the outer layer, it shows the log2 (FusionNW kinase fusion relevance score + 1), number of samples (sample frequency), DoF (degree of freedom in fusion gene combination) score and MAII (major active iso-fusion gene index) score, respectively. For other kinase groups, the landscape images are shown in Figure 7. We hope this map of kinase impact assessment metrics would be helpful to select driver fusion genes and to design a drug development to treat fusion gene patients in cancer.

Figure 6 Circos plot of the assessment metrics of TK group kinases. In this circus plot, we compare four available assessment metrics of genes in pan-cancer fusion genes.

Figure 7 Circos plots of the assessment metrics of all other kinase groups.

DISCUSSION

Protein kinases play a crucial role in signal transduction and regulate multiple cellular processes such as metabolism, membrane transport, motility and the cell cycle. Despite their significant involvement in various diseases, comprehensive information regarding their interactions is currently limited [10]. Tyrosine kinase fusion genes constitutes a significant class of oncogenes related with leukemia and solid tumors. These genes arise from translocations and various chromosomal rearrangements involving specific tyrosine kinase genes like ABL, PDGFRA, PDGFRB, FGFR1, SYK, RET, JAK2 and ALK [11].

In this study, we developed a novel computational approach to assess how actively each kinase gene is involved in the pan-cancer fusion genes using network propagation. To initiate the network propagation, we chose the seed genes from their six-feature information with dimension reduction. We tried two representative dimension reduction methods (PCA and VAE) and we identified that the PCA-based gene ranking information was more helpful to suggest the seed genes in human kinase fusion gene–gene network. Dimensionality reduction method reduces the number of random variables under different purpose-driven considerations by feature extraction. PCA has been widely used in many biological datasets and is the first of its type proposed for such purpose. PCA digests the high-dimensional cellular data into principal components in a linear fashion in such a way that the first and second principal components contain segregated major populations that resemble their relatedness [12]. Many recent preprints accomplish this using VAEs, generative models that learn underlying structure of data by compress it into a constrained, low-dimensional space. The low-dimensional spaces generated by VAEs have revealed complex patterns and novel biological signals from large-scale gene expression data and drug response predictions [13].

Multiple studies suggested clinical assessment of human kinases [14–16], but none considered the gene fusion aspect of human kinases. The assessment studies of kinase without kinase fusion gene events cannot get the effect of one of the mechanisms that enhance the kinase function in cancer, which is a critical mechanism in almost all major cancers. In this study, we suggest a novel way of assessing human kinases from six different fusion gene–related numerical values using a network propagation approach to infer how likely individual kinases influence the kinase fusion gene network composed of ~5K kinase fusion gene pairs. Our result suggests that the linear transformation of fusion genes across the six-feature information was more suitable for capturing the underlying structure compared to the non-linear transformations. To support this conclusion, we examined whether there are linear relationships among the six features’ scores of all human genes involved in fusion genes. As shown in Supplementary Figure S1, there were strong linear relationships among the six features. Furthermore, we assumed a Gaussian distribution for the latent space of VAE, while this choice may be a limitation for our model even though it is a simplifying assumption made to enable tractable computations for training process. This assumption could lead to challenges if there are heavy tails for the true distribution. If there is no clear sign of linear relationships between selected genes, we may try some semi-supervised [17] or unsupervised [18] Gaussian mixture VAE. With the selected genes, we further used propagation to study the gene–gene pair networks and downstream analysis. Our study results suggest potentially active genes to update the fusion gene panel. Based on our suggested new kinase assessment maps, it may be required to check the interaction between all kinome-involved fusion proteins and all approved kinase inhibitors. These days, we are also working for the in silico prediction of the potential small molecules of kinase fusion proteins with Bioinformatics. We hope we can identify new insights in kinase inhibitor therapeutics in the future.

MATERIALS AND METHODS

Fusion gene and other gene group information

We downloaded fusion gene information of ~102K human fusion genes from our resource, FusionGDB2 [1]. The fusion genes in FusionGDB 2.0 were downloaded from ChiTaRS 5.0 and ChimerDB 4.0, respectively. Of these, 50 360 and 50 931 fusion genes were from Entrez Sanger sequences and TCGA samples. The fusion genes from the Entrez database were identified by running BLAT for the expressed sequence tags (ESTs). More details of individual samples that have expressed these ESTs can be obtained from the EST libraries. The fusion genes from the TCGA were identified by running STAR-fusion and FusionScan. TCGA data have been sequenced in 33 different cancer types. The sample ID information of individual fusion genes in TCGA is available in Supplementary Table S1. More detailed clinical information can be downloaded from the GDC data portal. For individual genes that were involved to form fusion genes, we made a profile of number of cancer types that formed fusion genes, number of partner genes to form fusion genes, number of breakpoints inside of gene and number of sample sizes that expressed fusion genes. Based on the kinase fusion gene–gene pair information, we drew the kinase fusion gene network in Figure 1 using the Cytoscape 3.9.1. We downloaded kinase information from Kinase mutation database (KinaseMD, https://bioinfo.uth.edu/kmd/) [19]. We downloaded ClinGene information from the website of ClinGen (https://clinicalgenome.or) [20]. We downloaded IUPHAR drug target gene information from the website of IUPHAR (https://www.guidetopharmacology.org/) [21]. We downloaded the TruSight RNA Fusion Panel gene information from the website of the illumine (https://www.illumina.com/products/by-type/clinical-research-products/trusight-rna-fusion.html).

Highlighting kinases that actively involving in the formation of fusion genes in cancers with kinase fusion gene network propagation

FusionNW, learning from the SCAVENGE [4], is based on a random walk with restart (RWR) [5] to propagate the set of seed genes for a trait of interest to discover the FG-relevance association hidden in the gene–gene interaction (GGI). The initial state of the GGI graph is defined by selected seed nodes (genes). The initial seed genes in our RWR are based on six features: (1) number of cancer types per gene, (2) number of partner genes per gene, (3) number of breakpoint (BPs) per gene, (4) sample frequency per fusion gene, (5) MAII score per gene and (6) DoF score per gene. From each feature, we sorted the values and chose the highest 1337 genes (5%) as the seed genes and use these seed genes to run separate models. We used igraph in R 4.2.0 for the GGI networks. Out of 12 769 pairs of fusion genes, 8756 genes were included in this study since they have connections with other fusion genes. More formally, suppose there are selected seed genes \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $I\in$\end{document}  V defining the starting stage of a directed GGI network graph, \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $G=\left(V,E\right)$\end{document}. Two sparse matrices are shown to represent network structure.

\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $$ {A}_{i,j}=\left\{\!\begin{array}{c}1\ if\ {e}_{i,j}\in E\\{}0\ otherwise\end{array}\right. $$\end{document}

\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $$ {M}_{i,j}=\frac{A_{i,j}}{\Sigma_j{A}_{i,j}}, $$\end{document}

where \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} ${A}_{i,j}$\end{document} denotes the adjacency matrix of G and \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} ${M}_{i,j}$\end{document} is the transition probability matrix that is the column normalization of \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} ${A}_{i,j}$\end{document}. We set the equation of stationary probability for each node as the same as the SCAVENGE’s:

\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $$ {v}_{s+1}^T=\left(1-\gamma \right)M{v}_s^T+\gamma{v}_0^T $$\end{document}

where \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $\gamma$\end{document} is the restart probability (\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $\gamma =0.01$\end{document}) and s is the number of iterations for \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $\forall s\in \mathbb{N}$\end{document}. The propagation process keeps walking until the absolute difference between \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} ${v}_{s+1}$\end{document} and \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} ${v}_s$\end{document} is less than \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $1\ast{10}^{-5}$\end{document}. Then we say the random walk is stable. The final \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} ${v}_s$\end{document}is a list of the network propagation scores, which indicates how a kinase fusion gene is influenced by the seed genes. Therefore, a kinase gene with a higher network propagation score indicates it has a greater relevance with the fusion genes. After we got the kinase FusionNW relevance scores for each model based on each feature out of six fusion gene features, we took the average output scores per gene as our final TF FG relevance scores.

Select the seed genes for the network propagation using a PCA feature reduction approach

We executed a PCA on a standardized dataset of N genes by six features. The PCA was done using the Scikit-learn package in Python, which was used to compute the covariance matrix for six scaled variables. Following this, we calculated the eigenvalues associated with the covariance matrix. Ultimately, we reduced the six-dimensional feature space into a single dimension by selecting the first principal component, which has the highest eigenvalue. We constructed a reconstruction matrix to compute the MSE using the same Scikit-learn package in Python. Then, we manually calculated the MSE by subtracting the reconstructed output from the input and dividing the results by the total ~20 k genes. Lastly, we feed the top 5% of the kinase fusion genes from the sorted first principal component of PCA as the seed genes for the kinase fusion gene network propagation.

Select the seed genes for the network propagation using a VAE feature reduction approach

We used a variational autoencoder r(VAE) model to learn a continuous non-linear latent space \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $Z={R}^1$\end{document} given ~26K gene samples (26 735 genes involved in fusion genes). Our VAE model consists of the generative model \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $p\left(y|z\right)$\end{document} given a prior \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $p(z)$\end{document} and the recognition model (without noise) \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $q\left(z|y\right)$\end{document}. This model was trained by concurrently minimizing the reconstruction error and Kullback–Leibler (KL) divergence loss [22]:

\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $$ \mathrm{KL}\left(q(z)|p\left(z|y\right)\right)={\mathbb{E}}_{q(z)}\left[\ln \left(\frac{q(z)}{p\left(z|y\right)}\right)\right] $$\end{document}

We assumed that both the conditional probability distribution q(z|x) and the prior distribution p(z) adhered to Gaussian distributions. This is a typical assumption in VAE models as it aligns with the common assumption in machine learning that data and error often follow a normal distribution. Both reconstruction error and KL divergence loss error were saved for VAE model for the comparison with PCA. This VAE model produced latent values for ~26 k genes, forming a compact representation of the genes in the lower-dimensional latent space. The process of model training and testing were facilitated using the PyTorch library [23] in Python. Post-training, these latent values were sorted and mapping with the Kinase fusions genes. The top 5% were selected for the FusionNW analysis.

Key Points

Our study provides new useful features in kinase fusion gene assessment.

Our study provides new methods assessing the kinase genes in pan-cancer fusion genes.

Our study provides new ranking of the kinase genes in pan-cancer fusion genes.

Supplementary Material

Supplementary_Figure_S1_bbae097

Supplementary_Table_S1_Human_kinase_fusion_gene_information_bbae097

Supplementary_Table_S2_Six_features_of_5K_genes_in_kinase_fusion_genes_bbae097

Supplementary_Table_S3_Six_features_of_all_genes_involved_in_human_fusion_genes_bbae097

Supplementary_Table_S4_all_information_of_human_kinome_in_pan-cancer_fusion_genes_bbae097

ACKNOWLEDGEMENTS

We thank the members of the Center for Computational Systems Medicine at SBMI at UTHealth for valuable questions, discussions and suggestions.

FUNDING

This work was partially supported by the National Institutes of Health grants [R35GM138184] to P.K. The funders had no role in study design, data collection and analysis, decision to publish or preparation of the manuscript. Funding for open access charge: Startup Fund to Dr. Kim from the University of Texas Health Science Center at Houston.

DATA AVAILABILITY

All fusion gene information is available from the FusionGDB website (https://compbio.uth.edu/FusionGDB2). All related analyses information is available from the supplementary files. Further information and requests should be directed to Dr Pora Kim (Pora.kim@uth.tmc.edu).

CODE AVAILABILITY

FusionNW model and code are available from our website (https://compbio.uth.edu/FusionPDB/FusionNW/). Further information and requests should be directed to Dr Pora Kim (Pora.kim@uth.tmc.edu).

Himansu Kumar is a Research Scientist at the McWilliams School of Biomedical Informatics, The University of Texas Health Science Center at Houston. His research interests are bioinformatics, cancer genomics, and cheminformatics.
==== Refs
References

1. Kim  P, Tan  H, Liu  J, et al.  FusionGDB 2.0: fusion gene annotation updates aided by deep learning. Nucleic Acids Res  2022;50 :D1221–30.34755868
2. Kim  P, Jia  P, Zhao  Z. Kinase impact assessment in the landscape of fusion genes that retain kinase domains: a pan-cancer study. Brief Bioinform  2018;19 :450–60.28013235
3. Kim  P, Ballester  LY, Zhao  Z. Domain retention in transcription factor fusion genes and its biological and clinical implications: a pan-cancer study. Oncotarget  2017;8 :110103–17.29299133
4. Yu  F, Cato  LD, Weng  C, et al.  Variant to function mapping at single-cell resolution through network propagation. Nat Biotechnol  2022;40 :1644–53.35668323
5. Tong H, Faloutsos C Pan Jy. Fast Random Walk with Restart and Its Applications, Sixth International Conference on Data Mining (ICDM'06), Hong Kong, China, 2006, pp. 613–622. 10.1109/ICDM.2006.70.
6. Buitinck, L, Louppe, G, Blondel, M, et al.  2013.
7. Lietha  D, Cai  X, Ceccarelli  DF, et al.  Structural basis for the autoinhibition of focal adhesion kinase. Cell  2007;129 :1177–87.17574028
8. Consortium  APG . AACR project GENIE: powering precision medicine through an international Consortium. Cancer Discov  2017;7 :818–31.28572459
9. Guan  Y, Wang  Y, Li  H, et al.  Molecular and clinicopathological characteristics of ERBB2 gene fusions in 32,131 Chinese patients with solid tumors. Front Oncol  2022;12 :986674.36276102
10. Buljan  M, Ciuffa  R, van  Drogen  A, et al.  Kinase interaction network expands functional and disease roles of human kinases. Mol Cell  2020;79 :504–520.e9.32707033
11. Medves  S, Demoulin  JB. Tyrosine kinase gene fusions in cancer: translating mechanisms into targeted therapies. J Cell Mol Med  2012;16 :237–48.21854543
12. Cheng  Y, Newell  EW. Deep profiling human T cell heterogeneity by mass cytometry. Adv Immunol  2016;131 :101–34.27235682
13. Hu  Q, Greene  CS. Parameter tuning is a key part of dimensionality reduction via deep variational autoencoders for single cell RNA transcriptomics. Pac Symp Biocomput  2019;24 :362–73.30963075
14. Southekal  S, Mishra  NK, Guda  C. Pan-cancer analysis of human Kinome gene expression and promoter DNA methylation identifies dark kinase biomarkers in multiple cancers. Cancers (Basel)  2021;13 :1–17.
15. Essegian  D, Khurana  R, Stathias  V, Schurer  SC. The clinical kinase index: a method to prioritize understudied kinases as drug targets for the treatment of cancer. Cell Rep Med  2020;1 :100128.33205077
16. Gerdes  H, Casado  P, Dokal  A, et al.  Drug ranking using machine learning systematically predicts the efficacy of anti-cancer drugs. Nat Commun  2021;12 :1850.33767176
17. Kingma  DP, Mohamed  S, Jimenez Rezende  D, Welling  M. Semi-supervised learning with deep generative models. Advances in Neural Information Processing Systems  2014;27.
18. Dilokthanakul  N, Mediano  PA, Garnelo  M, et al.  Deep unsupervised clustering with gaussian mixture variational autoencoders. 2016; arXiv preprint arXiv:1611.02648.
19. Hu  R, Xu  H, Jia  P, Zhao  Z. KinaseMD: kinase mutations and drug response database. Nucleic Acids Res  2021;49 :D552–61.33137204
20. Rehm  HL, Berg  JS, Brooks  LD, et al.  ClinGen — the clinical genome resource. N Engl J Med  2015;372 :2235–42.26014595
21. Harding  SD, Armstrong  JF, Faccenda  E, et al.  The IUPHAR/BPS guide to PHARMACOLOGY in 2022: curating pharmacology for COVID-19, malaria and antibacterials. Nucleic Acids Res  2022;50 :D1282–94.34718737
22. Csiszar  I . $I$-divergence geometry of probability distributions and minimization problems. The Annals of Probability  1975;3 :146, 113–58.
23. Paszke  A, Gross  S, Massa  F, et al. In: Wallach  H, Larochelle  H, Beygelzimer  A  et al. (eds). Advances in Neural Information Processing Systems 32. Curran Associates, Inc., 2019, 8024–35.
