
==== Front
Brief Bioinform
Brief Bioinform
bib
Briefings in Bioinformatics
1467-5463
1477-4054
Oxford University Press

10.1093/bib/bbae452
bbae452
Problem Solving Protocol
AcademicSubjects/SCI01060
GeneRaMeN enables integration, comparison, and meta-analysis of multiple ranked gene lists to identify consensus, unique, and correlated genes
https://orcid.org/0000-0002-7346-803X
Yousefi Meisam Program in Emerging Infectious Diseases, Duke-NUS Medical School, 8 College Road, Singapore 169857, Singapore
Centre for Computational Biology, Duke-NUS Medical School, 8 College Road, Singapore 169857, Singapore

See Wayne Ren Program in Emerging Infectious Diseases, Duke-NUS Medical School, 8 College Road, Singapore 169857, Singapore

Aw-Yong Kam Leng Program in Emerging Infectious Diseases, Duke-NUS Medical School, 8 College Road, Singapore 169857, Singapore

Lee Wai Suet Program in Emerging Infectious Diseases, Duke-NUS Medical School, 8 College Road, Singapore 169857, Singapore

Yong Cythia Lingli Program in Emerging Infectious Diseases, Duke-NUS Medical School, 8 College Road, Singapore 169857, Singapore

Fanusi Felic Program in Emerging Infectious Diseases, Duke-NUS Medical School, 8 College Road, Singapore 169857, Singapore

Smith Gavin J D Program in Emerging Infectious Diseases, Duke-NUS Medical School, 8 College Road, Singapore 169857, Singapore

Ooi Eng Eong Program in Emerging Infectious Diseases, Duke-NUS Medical School, 8 College Road, Singapore 169857, Singapore

Li Shang Program in Cancer and Stem Cell Biology, Duke-NUS Medical School, 8 College Road, Singapore 169857, Singapore

Ghosh Sujoy Centre for Computational Biology, Duke-NUS Medical School, 8 College Road, Singapore 169857, Singapore
Laboratory of Computational Biology, Pennington Biomedical Research Center, 6400 Perkins Road, Baton Rouge, LA 70808, United States

https://orcid.org/0000-0001-9014-1365
Ooi Yaw Shin Program in Emerging Infectious Diseases, Duke-NUS Medical School, 8 College Road, Singapore 169857, Singapore
A*STAR Infectious Diseases Laboratories (ID Labs), Agency for Science and Technology Research (A*STAR), 8A Biomedical Grove, #05-13 Immunos, Singapore 138648, Singapore

Corresponding authors. Yaw Shin Ooi, Program in Emerging Infectious Diseases, Duke-NUS Medical School, 8 College Road, Singapore 169857, Singapore. E-mail: yawshin.ooi@duke-nus.edu.sg; Meisam Yousefi, Program in Emerging Infectious Diseases, Duke-NUS Medical School, 8 College Road, Singapore 169857, Singapore. E-mail: meisam.yousefi@u.duke.nus.edu
9 2024
18 9 2024
18 9 2024
25 5 bbae45211 6 2024
15 7 2024
30 8 2024
© The Author(s) 2024. Published by Oxford University Press.
2024
https://creativecommons.org/licenses/by-nc/4.0/ This is an Open Access article distributed under the terms of the Creative Commons Attribution Non-Commercial License (https://creativecommons.org/licenses/by-nc/4.0/), which permits non-commercial re-use, distribution, and reproduction in any medium, provided the original work is properly cited. For commercial re-use, please contact journals.permissions@oup.com

Abstract

High-throughput experiments often produce ranked gene outputs, with forward genetic screening being a notable example. While there are various tools for analyzing individual datasets, those that perform comparative and meta-analytical examination of such ranked gene lists remain scarce. Here, we introduce Gene Rank Meta Analyzer (GeneRaMeN), an R Shiny tool utilizing rank statistics to facilitate the identification of consensus, unique, and correlated genes across multiple hit lists. We focused on two key topics to showcase GeneRaMeN: virus host factors and cancer dependencies. Using GeneRaMeN ‘Rank Aggregation’, we integrated 24 published and new flavivirus genetic screening datasets, including dengue, Japanese encephalitis, and Zika viruses. This meta-analysis yielded a consensus list of flavivirus host factors, elucidating the significant influence of cell line selection on screening outcomes. Similar analysis on 13 SARS-CoV-2 CRISPR screening datasets highlighted the pivotal role of meta-analysis in revealing redundant biological pathways exploited by the virus to enter human cells. Such redundancy was further underscored using GeneRaMeN’s ‘Rank Correlation’, where a strong negative correlation was observed for host factors implicated in one entry pathway versus the alternate route. Utilizing GeneRaMeN’s ‘Rank Uniqueness’, we analyzed human coronaviruses 229E, OC43, and SARS-CoV-2 datasets, identifying host factors uniquely associated with a defined subset of the screening datasets. Similar analyses were performed on over 1000 Cancer Dependency Map (DepMap) datasets spanning 19 human cancer types to reveal unique cancer vulnerabilities for each organ/tissue. GeneRaMeN, an efficient tool to integrate and maximize the usability of genetic screening datasets, is freely accessible via https://ysolab.shinyapps.io/GeneRaMeN.

genome-scale CRISPR screening
meta-analysis
rank aggregation
virus-host interactions
flaviviruses
human coronaviruses
virus host factors
cancer dependencies
Khoo Bridge Fund, Singapore KBrFA/2022/0060 Louisiana Clinical and Translational Science Center 10.13039/100018106 NIGMS 2U54GM104940 Ministry of Education of Singapore MOE-T2EP30123-0008 MOE-000095-01 National Research Foundation of Singapore 10.13039/501100001381 NRF-MOST Joint Grant MOH-000928 Duke-NUS Medical School 10.13039/100016017 Duke/Duke-NUS/RECA(Pilot)/2019/0047
==== Body
pmcIntroduction

With the upsurge in the popularity and accessibility of high throughput genetic screening techniques, the quantity of data generated through such methods are growing exponentially. This highlights the need to develop meta-analysis tools capable of effectively integrating and comparing datasets derived from multiple independent studies addressing similar biological inquiries. Among such genetic screening approaches, clustered regularly interspaced short palindromic repeats/CRISPR associated protein 9 (CRISPR/Cas9) genetic screenings have gained substantial prominence in various fields of biology [1]. These screenings have been proven invaluable for the identification of drug targets [2–4], pathogen host factors discovery [5–8], cancer dependencies [9, 10], and numerous other pertinent biological phenomena. Given the escalating trend of having multiple independent CRISPR/Cas9 screens studying the same biological context, the requirement for a standardized and comprehensive method to integrate and compare such screening outcomes has become evident. One of the biggest challenges associated with meta-analyzing these datasets lies in the multitude of factors that can significantly impact the enrichment scores of genes, often resulting in scores being non-comparable across studies. Various factors, including the choice of the CRISPR single guide ribonucleic acid (sgRNA) libraries and cell lines, experimental conditions, batch effects, and bioinformatics pipelines for next-generation sequencing (NGS) data analysis, can exert substantial influence on the final score obtained for each hit. For instance, following the Coronavirus disease 2019 (COVID-19) pandemic, various groups performed genetic screens aiming to uncover novel host factors and gain a better understanding of the biology of severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) infection, which may potentially innovate the development of new antiviral strategies [11–19]. While each of those studies discovered important host factors, the extent of overlap observed across the studies has been minimal, primarily due to the abovementioned technical challenges. Some studies have tried to compare their screening results with those of the previous ones, typically employing Venn diagrams as tools to highlight mutual genes [11, 20]. While using Venn diagrams for such purpose is straightforward and thus gained popularity, there are several pitfalls accompanying this approach. Firstly, this method is not systematic as the researchers must rely on setting arbitrary thresholds for comparisons (e.g. top 50 hits) which can significantly impact the combination of shared hits recovered. Additionally, Venn diagrams are agnostic to the enrichment scores of the genes; therefore, they treat all entities equally as long as they rank above the threshold. Importantly, Venn diagrams exhibit limitations in scalability, in particular, it is troublesome to visualize and interpret results involving more than six groups. In addition to these challenges regarding the identification of shared hits among datasets, there is also a lack of established methodologies and easily accessible tools for discerning distinct high-ranking hits unique to a subset of the dataset but absent in others, as well as tools to examine how genes correlate with each other across multiple ranked datasets. These challenges have called for the development of novel tools to perform meta-analyses on ranked gene lists.

Here, we have developed Gene Rank Meta Analyzer (GeneRaMeN), an R Shiny tool that allows the user to integrate multiple genetic screening results based on their ‘ranks’, bypassing the need to compare and normalize the enrichment scores. This provides the user with great flexibility to include data analyzed via different methods or from disparate sources. To demonstrate the utility of GeneRaMeN in analyzing ranked gene lists, we applied it to analyze several published and unpublished in-house genetic screening datasets, spanning topics such as host factors for medically important viruses and essential genes of human cancer. We also conducted new genome-wide CRISPR/Cas9 screens involving flaviviruses and coronaviruses and integrated these results with those of previously available datasets for analysis in GeneRaMeN. We applied the same rank-based methodology on the dependency scores from the Cancer Dependency Map (DepMap) initiative to discern gene dependencies for specific cancer types, exhibiting robust alignment with previously validated markers, and proposing candidate genes for future research.

Materials and methods

Preprocessing of screening data

GeneRaMeN imports sorted gene names, subsequently re-ranking them for downstream analysis. To ensure the compatibility of datasets before integrating their ranks, the program initially maps all gene name aliases to their most recent official symbol, as provided by the Bioconductor ‘annotationDbi’ package (v1.64.1) [21]. This measure avoids the presentation of the same gene under multiple names, a circumstance that may arise from the utilization of different versions of genome annotations across original datasets.

Rank aggregation of multiple ranked gene lists

GeneRaMeN offers users various methods for aggregating ranked lists, including robust rank aggregation (RRA), geometric means, and arithmetic means. The RRA relies on the ‘RobustRankAggreg’ package (v1.2.1) in R, calculating the probability of genes having certain ranks compared to a null model where all input ranks are random [22]. Geometric means for gene ranks are determined using the formula:

\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $$ {\bar{x}}_g={e}^{\frac{\sum_{i=1}^n\log \left({x}_{g,i}\right)}{n}} $$\end{document}

in which \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} ${x}_{g,i}$\end{document} denotes the rank of gene g in the ith dataset. The mean is computed using the base R mean function. An important problem in the final gene rank outputs obtained from widely used CRISPR screen analysis tools, such as MAGeCK [23] is related to encompassing all input genes in the output. This can compromise the reliability of rankings, particularly toward the lower (less significant) end of the ranked lists, which are often arranged arbitrarily/alphabetically. To avoid this issue, GeneRaMeN allows for customization, permitting users to refine the inclusion criteria by setting a threshold for gene rank values. By specifying a threshold value, the application effectively treats genes ranked below this threshold as having the specified threshold rank for all subsequent computations, such that lower threshold values result in more stringent selection of hits by the application.

Comparison of multiple ranked gene lists to identify unique genes

GeneRaMeN provides two primary methods to identify genes exhibiting significantly distinct ranks between two study subsets. The application can conduct one-way two-sample t-tests to identify significant differentially ranked genes, arranged by the magnitudes of their effect size. Alternatively, it can execute rank aggregation independently on each subset, then compute a rank subtraction statistic for each gene which can be used for ‘Uniqueness’ ranking. For these parallel rank aggregations, users again have the option to select between RRA, rank geometric means, and arithmetic rank means as the method of choice. These latter alternative methods resemble the identification of up/downregulated genes in multiple gene expression datasets proposed by Rank Product and Rank Sum algorithms [24]. Similar to the Rank Aggregation section above, users can set a threshold concerning maximum ranks. For Rank Uniqueness analysis of DepMap, the dataset is not included in the web app instance due to size constraints.

Identifying correlated and anti-correlated set of genes

To identify genes whose ranks exhibit similar or opposing trends to a gene of interest across multiple ranked datasets, GeneRaMeN offers correlation tests. The correlation coefficients are determined using Pearson’s method, and the resulting values are subsequently sorted in descending and ascending order to identify genes demonstrating the highest degrees of positive and negative correlation, respectively. The calculation of correlations is performed using the WGCNA R package. Similar to the rank aggregation and uniqueness functions described above, it is possible to define maximum rank thresholds.

Implementation

GeneRaMeN is developed in R and powered by Shiny (https://shiny.rstudio.com/). The application consists of two main components: a user interface and a server function for executing computations. The user interface design prioritized simplicity without compromising functionality, primarily utilizing shiny (v1.8.0), bslib (v0.6.1), shinyWidgets (0.8.0), and shinycssloaders (v1.0.0) packages. The application incorporates predefined datasets, available either for standalone demonstration purposes or for amalgamation with user inputs for further analysis. Gene ontology and pathway enrichment over-representation analysis is mediated using gprofiler2 (v0.2.2) package [25]. Data cleaning and manipulations on the server side mostly employ the tidyverse (v2.0.0) packages and data.table (v1.14.8), while visualizations are facilitated by ggplot2 (v3.4.4) and pheatmap (v1.0.12). Rank aggregations are performed via RobustRankAggreg (v1.2.1) and multiple t-tests via Rfast (v2.1.0) package with modifications [26]. A comprehensive tutorial on using GeneRaMeN can be found under the ‘tutorial’ tab from within the web app, or in the supplementary attached to this paper (Supplementary Data 1).

Cells and viruses

African green monkey kidney Vero clone E6 (VeroE6) cells (CRL-1586, ATCC), human hepatocellular carcinoma Huh7.5.1 cells (provided by Jan E Carette), and human embryonic kidney 293FT subclones (R70007, Thermo Fisher) were cultured in Dulbecco’s Modified Eagle Medium (DMEM) supplemented with 1× penicillin–streptomycin, 1× L-glutamine, and 10% heat-inactivated fetal bovine serum (FBS). Aedes albopictus clone C6/36 cells (CRL-1660, ATCC, provided by EEO) were cultured in RPMI-1640 supplemented with 10% FBS, 1× penicillin–streptomycin, and 1× L-glutamine. All cell lines underwent mycoplasma testing and were confirmed negative.

Dengue virus serotype 1 (DENV1; NR-3782; BEI Resources) was propagated on C6/36 cells. Human coronavirus 229E (HCoV-229E; VR-740; ATCC, provided by GJDS) was propagated and titered on Huh7.5.1 cells [27]. Japanese encephalitis virus (JEV, SA14–14-2; vaccine strain provided by EEO) was propagated and titered on VeroE6 cells.

Genome-scale CRISPR/Cas9 screening

Genome-wide knockout (KO) libraries were generated as previously described [19]. Briefly, for DENV1 and JEV screens, 293FT cells were first transduced with the lentivirus harboring lentiCas9-Blast (Addgene #52962), and subsequently treated with 15 μg/ml of blasticidin to select for cells with stable expression of the Cas9 protein (293FT-Cas9). These 293FT-Cas9 cells were then transduced with lentiviruses carrying human Brunello CRISPR KO sgRNA library (Addgene #73178) at the multiplicity of infection (MOI) of 0.3. At 2 days post-transduction, 2 μg/ml of puromycin was introduced to select cells that are stably expressing the sgRNA constructs. A pool of ~40 million mutagenized cells—about 500-fold coverage for each unique sgRNA from the library—were used for each of the screens. The seeded mutagenized cells were subsequently infected with DENV or JEV at an MOI of 0.2 and 0.05, respectively, the next day. An additional round of infection was carried out at 18 days post-infection. Resistant cells were harvested in two batches for each screen, one at 21 days post-infection and the other at 27 days post-infection. Cells harvested in the second batch were also re-infected at 21 days post-initial infection. A total of 44 million non-infected mutagenized cells were simultaneously harvested to serve as the unselected control. For HCoV-229E screen, a similar strategy was employed using Huh7.5.1 cells. Cas9 expressing Huh7.5.1 cells were transduced with the Brunello library to generate a pool of mutagenized cells. A pool of ~60 million mutagenized cells was subsequently seeded and infected with HCoV-229E at a MOI of 0.003. Resistant cells were harvested at 15 days post-infection. A total number of 40 million non-infected populations of mutagenized cells were harvested in parallel as the unselected control. Genomic DNA was extracted from each harvested cell pellet using QIAamp DNA Mini Kit (Qiagen). The sgRNA containing sequences were PCR amplified and barcoded with Illumina NGS indexes. The amplicons were gel purified, quantified, and sequenced using Illumina HiSeq or MiSeq platform. The resulting NGS data files were quality controlled with FastQC, and subsequently subjected to mapping and sgRNA enrichment analysis using MAGeCK (v0.5.9) [23].

Results

The overview of GeneRaMeN

The GeneRaMeN platform empowers integration and meta-analysis of ranked gene lists addressing the same biological question in both online and offline modes. In its current version, the platform is pre-loaded with multiple screening datasets—13 published SARS-CoV-2 datasets, 20 published with four previously unpublished flavivirus datasets, and 9 published plus 1 new seasonal human coronavirus dataset. Users have the option to import their datasets alone or integrate their datasets into the pre-loaded datasets for executing GeneRaMeN analyses. GeneRaMeN’s ‘Rank Aggregation’ function enables integration of datasets to identify a list of aggregated rank hits, while ‘Rank Uniqueness’ allows comparison of two groups of user-defined datasets for determining unique hits, and ‘Rank Correlation’ performs correlation analysis between genes (Fig. 1). While RRA is the default algorithm that drives GeneRaMeN’s ‘Rank Aggregation’ analyses, alternative choices such as rank geometric mean and rank mean are also available. The resulting aggregated rank list can be exported as a table and visually represented as a scatter plot. Using the heatmap feature, users can visualize an overview of aggregated hits across individual datasets. Additionally, GeneRaMeN enables the gene ontology and pathway enrichment analyses of selected aggregated hits. GeneRaMeN’s ‘Rank Uniqueness’ analysis adopted multiple one-way t-tests (by default) to uncover hits that are unique to a user-defined subset of datasets. Alternatively, users have the option to use other methods such as RRA, geometric mean, or mean. The heatmap feature allows visualization of unique hits from both subsets of data. Furthermore, GeneRaMeN’s ‘Rank Correlation’ allows the users to determine groups of genes which show similar or opposite trends to a gene of interest, by performing correlation analyses.

Figure 1 Schematic diagram of GeneRaMeN workflow. The analysis input consists of ranked gene lists, which can undergo ‘Rank Aggregation’, ‘Rank Uniqueness’, or ‘Rank Correlation’ processes. In the ‘Rank Aggregation’ function, all ranked lists are integrated to identify commonly shared hits across the datasets. The ‘Rank Uniqueness’ feature enables the assignment of datasets into two distinct subsets for comparative analysis, elucidating hits specific to each subset. ‘Rank Correlation’ identifies genes showing similar and opposing trends to a specific gene of interest across multiple datasets. All results can be downloaded from GeneRaMeN in various publication-ready outputs.

Integrating CRISPR screens can depict the redundant entry pathways of SARS-CoV-2

SARS-CoV-2, an enveloped RNA virus belonging to the Coronaviridae family, is the etiologic agent responsible for the COVID-19 pandemic that has caused over 700 million confirmed infections, with close to 7 million deaths worldwide [28]. We collated 13 individually published genome-scale CRISPR screening datasets aiming to uncover SARS-CoV-2 host factors [11–20, 29, 30]. While most of these screens successfully identified ACE2, the viral receptor, as one of the top 10 hits, minimal overlapping top hits were observed (Fig. 2a). To establish a consensus list of SARS-CoV-2 host factors, we employed GeneRaMeN’s ‘Rank Aggregation’ to integrate and analyze all these screening datasets using three different maximum rank thresholds (top 1000, 2000, and 5000 hits). As a result, regardless of the threshold settings, ACE2, TMPRSS2, and CTSL, the most critical host factors required for SARS-CoV-2 to enter host cells were consistently ranked by GeneRaMeN as the top aggregate hits (Fig. 2b, Supplementary Table 1). Upon binding to the ACE2 receptor on the plasma membrane via its Spike (S) protein, SARS-CoV-2 must hijack cellular serine proteases TMPRSS2 or CTSL for the maturation of S proteins, a crucial step to trigger membrane fusion so that the viral genome can enter the cytoplasm (Fig. 2c) [31]. TMPRSS2 is predominantly localized to the plasma membrane, therefore allowing the virus to fuse at the cell surface. Alternatively, SARS-CoV-2 can endocytose and subsequently engage endosomal CTSL to fuse in the endosome (Fig. 2c). In other words, SARS-CoV-2 can enter host cells via at least two distinctive routes. In cell lines where both TMPRSS2 and CTSL are functional and sufficiently available, the virus can co-opt either route to enter the cells—rendering these two proteases redundant, hence the genetic screens would fail to detect either of them. However, some cell lines exclusively express one of these serine proteases, forcing the virus to rely solely on a single route to enter host cells. Under such circumstances, the CRISPR screens would be able to identify the only serine protease as a hit (Fig. 2d). Theoretically, it would be illogical to identify both TMPRSS2 and CTSL as top hits in any currently available cell line based genetic KO screening platforms, prompting us to use GeneRaMeN’s ‘Rank Correlation’ to confirm this claim. Consistent with the findings presented in Fig. 2d, analysis of SARS-CoV-2 CRISPR screening datasets revealed a strong negative correlation between the ranks of TMPRSS2 and CTSL, with a correlation coefficient of −0.62. The result posited CTSL among the factors displaying the highest degrees of anti-correlations to TMPRSS2. Notably, most other host factors exhibiting significant anti-correlation with TMPRSS2 were also mainly involved in the endosomal trafficking and acidification processes, including components of vacuolar ATPase (e.g. ATP6AP1 and ATP6V1C1), and CCC complex (e.g. COMMD2; Fig. 2e).

Figure 2 GeneRaMeN empowers integration and meta-analysis of genetic screening hit lists to determine aggregated hits. (a) The top 10 SARS-CoV-2 host factors obtained in each of the genetic screens are shown. The viral receptor—ACE2—is indicated. (b) Rank Aggregation results of SARS-CoV-2 CRISPR screens using the RRA algorithm. The outcomes are displayed for three distinct maximum rank thresholds: N = 1000, N = 2000, and N = 5000. The top 10 host factors are highlighted in each scatter plot. (c) Schematic diagram showcasing the two functionally redundant pathways for SARS-CoV-2 membrane fusion, facilitating the release of viral genome into the host cytoplasm: TMPRSS2 is involved in the maturation of S protein at the plasma membrane, whereas CTSL mediates S protein maturation within the endosomes. (d) Original ranks of TMPRSS2 and CTSL identified in SARS-CoV-2 CRISPR screens, in addition to a representative Rank Aggregation result (RRA, N = 5000). (e) Cluster-gram representation of the original ranks of TMPRSS2 gene and its top 20 anti-correlated hits across SARS-CoV-2 studies.

The recovery of TMPRSS2 and CTSL together as top aggregated hits from CRISPR screens is a unique outcome of the rank aggregation method, underscoring the strength and necessity of conducting meta-analysis by integrating all possible datasets. As such, we showcased the robustness of GeneRaMeN’s ‘Rank Aggregation’ in establishing a single list of meaningful aggregated hits that were otherwise overlooked in individual datasets due to technical/biological challenges such as redundant cellular pathways. Interestingly, many other top aggregated hits were also previously validated host factors for SARS-CoV-2, such as TMEM106B [11, 32], AP1G1 [17], DYRK1A [33, 34], EP300 [35], VPS35, and RAB7A [12] (Supplementary Table 1). Notably, GeneRaMeN also empowers gene ontology and pathway enrichment analyses for the list of aggregated hits. For instance, gene ontology and pathway enrichment analysis revealed that the top 100 aggregated hits for SARS-CoV-2 are mainly associated with endocytosis-related pathways (Supplementary Fig. 1b and c).

Genetic screens identifying flavivirus host factors are heavily influenced by the choice of cell line

The flavivirus genus comprises a group of enveloped, positive-sense, single-stranded RNA viruses that are mainly transmitted through arthropod vectors such as mosquitoes and ticks. Some flaviviruses are important human pathogens, such as DENV, JEV, West Nile virus (WNV), yellow fever virus (YFV), and Zika virus (ZIKV). Each year, these viruses are estimated to cause hundreds of millions of infections in human populations worldwide [36, 37]. We first collated sixteen CRISPR screening datasets and four haploid genetic screening results sourced from seven papers published by different labs, in which DENV, ZIKV, WNV, and YFV host factors were assessed across different conditions [5, 38–43]. To expand the scope of the analysis, we generated new in-house datasets for JEV and DENV (serotype 1) using 293FT cells incorporated with the Brunello CRISPR KO library (Supplementary Table 2). Using GeneRaMeN’s ‘Rank Aggregation’, we integrated all these genetic screening datasets to generate a list of consensus flavivirus host factors (Fig. 3a, Supplementary Table 1). We observed numerous experimentally validated host factors emerging as top aggregated hits, including human genes involved in cellular pathways such as heparan sulfate biosynthesis (e.g. SLC35B2, EXT1, and BSGAT3) [44, 45], endoplasmic reticulum protein translocation (e.g. SPCS1, SEC61B, and SEC61A1) [38, 46], endoplasmic reticulum-associated degradation (e.g. EMC subunits, SEL1L) [47–49], oligosaccharyltransferase (OST) complex (e.g. STT3A, OSTC, and STT3B) [5], and others (e.g. HDLBP and FURIN) [41, 50] (Fig. 3a, Supplementary Fig. 2). In addition, visualizing these top aggregated hits’ original ranks within their respective datasets unveiled that the studies were mainly clustered according to cell lines, demonstrating a clear influence of the cell line of choice on the screening outcomes (Fig. 3b). The observed pattern indicated that the prevalence of these consensus host factors being identified in each of the original studies is predominantly determined by the cell line of choice, rather than other experimental settings such as the virus species. Likewise, all genetic screens performed using HAP1 cells were clustered together, regardless of the other experimental conditions like the different genetic screening platforms (i.e. CRISPR KO versus Haploid Genetic). Similar trends were observed for the SARS-CoV-2 screening datasets, where screens performed using the same cell lines showed a higher tendency to group together, as evidenced by the datasets generated using Calu-3 and A549 cell lines (Supplementary Fig. 1a).

Figure 3 GeneRaMeN’s ‘Rank Aggregation’ analysis of flavivirus screens identifies a substantial impact of the cell line of choice on screening outcome. (a) The Rank Aggregation outcomes from flavivirus genetic screens conducted with the RRA algorithm (N = 5000). Genes sharing membership within identical pathways are depicted in corresponding colors. (b) Cluster-gram representation of the original ranks of the top 50 aggregated hits of flaviviruses, across their respective study ranks. Genes absent from the haploid genetic datasets are indicated in gray.

Identification of virus- and cell line-specific host factors for multiple coronaviruses

Hitherto, there are seven species of human coronaviruses, including SARS-CoV, MERS-CoV, SARS-CoV-2, HCoV-229E, HCoV-NL63, HCoV-HKU1, and HCoV-OC43. Despite the close phylogenetic relationships, the infection biology of these coronaviruses is likely dissimilar. For example, SARS-CoV and SARS-CoV-2 use ACE2 as the entry receptor, but common cold coronaviruses such as HCoV-OC43 and HCoV-229E bind to heparan sulfate proteoglycans and ANPEP, respectively, to enter human cells [51, 52]. To test the function of GeneRaMeN’s ‘Rank Uniqueness’ in pinpointing unique hits by cross-comparing two subsets of datasets, we took advantage of all available and usable HCoV-229E and HCoV-OC43 CRISPR screening datasets [13, 20, 29, 53, 54], including our recent study on HCoV-OC43 [19], as well as an additional previously unpublished in-house HCoV-229E screening dataset (Supplementary Table 2). As a proof of principle concept, we demonstrated that GeneRaMeN’s ‘Rank Uniqueness’ could identify ANPEP as the most unique host factor for HCoV-229E, while several components of the heparan sulfate pathway (e.g. EXT1, EXT2, and UGDH) appeared as top unique host factors for HCoV-OC43 (Supplementary Table 3). Heatmap visualization of the top ten unique hits for each group revealed a distinct clustering precisely aligning with the virus species, confirming the relevance of these identified sets of genes to represent the two groups (Fig. 4a).

Figure 4 GeneRaMeN’s ‘Rank Uniqueness’ identifies unique hits associated with a specific cell type and coronavirus species. (a) Comparison between common cold coronavirus screens performed for HCoV-OC43 versus HCoV-229E, using a one-way two-sample test (N = 5000). The top 10 hits for each group are presented. (b) Contrast of SARS-CoV-2 screens conducted in Huh7 cell lines versus other cell lines using a one-way two-sample test (N = 5000). The top 10 hits for each group are presented. (c) Contrast of SARS-CoV-2 screens conducted in Calu-3 cell lines versus other cell lines using a one-way two-sample test (N = 5000). The top 10 hits for each group are presented.

This approach can be extended to any other attributes (e.g. cell lines, methods, and user-defined parameters) within a particular collection of screening datasets as well. Using SARS-CoV-2 screening datasets as an example, different cell lines (e.g. Huh7, Calu-3, and IGROV1) were used to screen for host factors of the same virus species across multiple studies. By stratifying and comparing screens performed in Huh7 (Fig. 4b) and Calu-3 (Fig. 4c) versus the rest of the cell lines, we identified cell line-specific hits unique to the SARS-CoV-2 screens conducted in those specific cell lines (Supplementary Table 3). For example, EXOC2 and TMEM106B were uniquely found in screens performed in Huh7 cells, while KMT2C and HEATR5B were mainly found in screens using Calu-3 cells. Likewise, hits that were commonly identified across most of the cell lines but consistently absent in studies employing those particular cell lines could also be identified, e.g. CTSL is consistently absent in Calu-3 datasets, as SARS-CoV-2 entry and fusion in this cell line is almost 100% mediated through TMPRSS2 [55, 56] (Fig. 2d).

Identification of specific cancer dependency factors using Cancer Dependency Map datasets

To further examine the performance of GeneRaMeN in analyzing big datasets (>100 ranked gene lists), we capitalized CRISPR screening datasets deposited in the DepMap portal. DepMap harbors numerous CRISPR KO screening datasets that were used to identify essential genes for the growth and proliferation of >1000 human cancer cell lines derived from various human tissues/organs [9]. To identify genes that are uniquely crucial for human lymphoid cancer cell lines but less important in other cancer cell types, we subjected the relevant DepMap datasets to offline GeneRaMeN’s ‘Rank Uniqueness’ analysis. This analysis revealed a list of new candidate genes as well as genes that were previously reported to be associated with lymphoma/leukemia. (Fig. 5a). For instance, IRF4 has been extensively investigated in leukemia, displaying frequent chromosomal translocations in T-cell lymphoma/leukemia tumors [57]. IRF4 is also involved in an oncogene regulatory network reported for these tumors [58]. Similarly, ATIC translocations have been linked with anaplastic large cell lymphoma [59, 60]. POU2AF1 is a B-cell-specific transcriptional coactivator previously associated with chronic lymphocytic leukemia and diffuse large B-cell lymphoma [61, 62]. Notably, many of these lymphoid-specific genes overlapped with the myeloid lineage as well, underscoring the overall similarity between the two lineages. For instance, previous research has identified ATP1BP3 and NMNAT1 as crucial dependency factors in acute myeloid leukemia (AML) [63, 64], while MBNL1 has been implicated in the regulation of essential alternative RNA splicing in mixed lineage leukemia (MLL) [65]. By mapping the original DepMap dependency scores of the top five essential genes across the tissue of origins, we could further illustrate the comparatively heightened reliance of the lymphoma-derived cell lines on these specific genes (Fig. 5b). Since each lineage within DepMap comprises multiple sub-lineages, we opted to determine the predominant sub-lineages contributing to the significance of these identified marker genes. Accordingly, we broke down the lymphoid lineage into sixteen sub-lineages utilizing available information from DepMap, depicting distinct dependency patterns across sub-lineages. For example, POU2AF1 and IRF4 were disproportionately associated with plasma cell myeloma, and NMNAT1 with Hodgkin lymphoma cell lines, while MBNL1 was almost consistently important across all sub-lineages (Fig. 5c). Using the same method, we further generated a collection of unique essential genes associated with cancer types originating from tissues/organs other than lymphoids found in the DepMap database (Supplementary Table 4). Similarly, several established specific markers of various cancer types were identified by GeneRaMeN from the DepMap data including PPARG and PXRA for bladder cancer [66, 67], CTNNB1 (β-catenin) for bowel/colorectal cancers [68, 69], FOXA1 for breast cancer [70], TP63 (p63) for head and neck tumors [71], PAX8 for kidney and ovary cancers [72, 73], NKX2–1 for lung cancer [74, 75], and MYCN for peripheral nervous system cancers [76, 77] (Supplementary Data 2).

Figure 5 Identification of lymphoid lineage-specific gene dependencies from DepMap using GeneRaMeN’s ‘Rank Uniqueness’. (a) Top 10 hits unique to lymphoid lineage identified using a one-way two-sample t-test (N = 5000). (b) Dependency scores of the top 5 hits unique to lymphoid across all available lineages from DepMap. (c) Breakdown of the dependency scores of the top 5 hits unique to lymphoid lineage. Non-lymphoid cell lines are shown for background reference.

To further delineate potential gene dependencies specific to distinct sub-lineages within the lymphoid cell lines, we opted to conduct additional iterations of the GeneRaMeN’s ‘Rank Uniqueness’ analysis, focusing on comparing cancer cell lines within a sub-lineage against all other sub-lineages. This approach aimed to enhance the precision of identifying candidate sub-lineage marker genes since non-lymphoid cell lines were excluded, thereby reducing background noise during comparisons. We performed the analysis for all lymphoid sub-lineages with at least three cell lines, resulting in eight distinct sets of marker genes each corresponding to a sub-lineage (Fig. 6). The original DepMap dependency scores of the top five genes for each sub-lineage, juxtaposed with the scores of the same genes in the rest of the lymphoid sub-lineages as well as the non-lymphoid background are demonstrated in the figure. Remarkably, the GeneRaMeN-identified genes exhibited substantial overlap with known marker genes experimentally associated with each specific sub-lineage. For example, tyrosine phosphatases PTPN1 and PTPN2 have been demonstrated as tumor growth and survival factors in ALK-positive anaplastic large cell lymphoma via CRISPR screens [78]. Concurrently, SBNO2 and STAT3 hyperactivation were also recently associated with malignancies including ALK-positive T-cell anaplastic large cell lymphoma [79]. For Burkitt lymphoma, CYB561A3 encodes a key enzyme mediating cell survival [80], while BCL6 expression has been proposed as a diagnosis marker [81]. In Hodgkin lymphoma, characteristic JAK/STAT signaling factors—recognized hallmarks of this tumor—were prominently featured [82, 83]. Notably, IL13 activity—another classical marker of Hodgkin lymphoma—and its corresponding receptor comprised of IL13RA1 and IL4R dimers, were identified as top hits in the analysis [84, 85]. Regarding plasma cell myeloma, PIM2 is a known cell-specific factor required for plasma cell survival [86]. BCL11B expression is also shown to be pivotal in T-cell lymphoblastic leukemia [87].

Figure 6 Discovery of lymphoid sub-lineage specific marker genes using GeneRaMeN’s ‘Rank Uniqueness’. Each row exhibits the dependency scores of the top 5 marker genes for each individual sub-lineage. The unique ranks are established through one-way two-sample t-tests (N = 5000). In each set, the cell lines within a specific sub-lineage are compared against the rest of lymphoid cell lines. Non-lymphoid cell lines are depicted as background references.

Discussion

Here we presented GeneRaMeN, a free-to-use web tool providing users with a comprehensive analytical framework to assess multiple ranked gene lists, enabling the identification of commonly shared genes across datasets as well as those unique to specific lists and exhibiting lower rankings in others. GeneRaMen provides users with the flexibility to analyze multiple ranked gene lists to find both genes homogenously detected across the lists, and genes that are specific to certain subsets. While our tool, particularly the GeneRaMeN’s ‘Rank Uniqueness’ and ‘Rank Correlation’ functions, stand as the first of their kind, previous endeavors have aimed at aggregating gene lists to discern commonalities in both ranked and unranked datasets, similar to our GeneRaMeN’s ‘Rank Aggregation’ feature. In such an attempt, MagPlotR was developed to aggregate ranked gene lists from CRISPR screens by computing mean enrichment scores across input datasets and re-ranking genes based on these calculated values [88]. However, its application is limited to CRISPR screen outcomes analyzed exclusively with MAGeCK software, restricting its adaptability to diverse dataset inputs. Another tool, meta-analysis by information content (MAIC) developed by Li et al, sought to integrate both ranked and unranked gene lists to identify consistent hits [89]. Despite the availability of these tools, both still require coding skills hindering experimental scientists to adopt them for meta-analysis, thus many recent studies comparing their screening outcomes with previous ranked datasets have largely resorted to conventional Venn diagrams due to their simplicity in implementation.

To identify genes consistently scoring high across multiple lists, GeneRaMeN implements rank aggregation methods including RRA, systematically analyzing all the hit lists and generating comprehensive overall rankings thereafter. The same rank aggregation approach as well as two-sample t-tests can be employed to highlight differences between two subsets, revealing hits unique to specific studies. While the statistical underpinnings of these analyses are not intricate, conducting such assessments without coding expertise remains challenging for most wet lab scientists. Thus, GeneRaMeN is designed to offer a user-friendly interface, easing accessibility and utilization of these analyses. Additionally, GeneRaMeN presents diverse visualization options including scatter plots and heatmaps, downloadable in various publication-ready formats. Furthermore, users can perform gene ontology and pathway enrichment over-representation analyses directly within the web tool using the generated gene lists. Utilizing genetic screen data aimed at uncovering host factors related to flaviviruses and coronaviruses, we have exemplified the efficacy of rank aggregation methodologies in revealing redundant pathways across datasets. Our analysis also underscores the influential role of cell line selection in shaping the outcomes of these screens. Furthermore, we showcased the utility of rank uniqueness analysis in delineating virus and cell line-specific hits within diverse datasets encompassing various screening conditions. Extending the functionality of GeneRaMeN’s rank uniqueness feature to analyze cancer dependency scores has demonstrated the potential of such comparative analytical approaches in studying diverse biological contexts. While in this work we primarily analyzed lymphoid cell lines as an example, these methodologies can be readily applied to any lineage or sub-lineages of cell lines available in DepMap (Supplementary Data 2, Supplementary Table 4), as well as any custom user-defined groups. Such analyses would serve as a robust platform for hypothesis generation, aiding in the identification of specific biological traits or features unique to certain cell lines, and offering valuable insights prior to experimental validation in wet laboratory settings.

An increased volume of screening data is anticipated to enhance the accuracy of consensus hit identification and biomarker discovery within these datasets in the future. While our analyses have revealed the influence of experimental design attributes on current screening data outcomes, prospective computational efforts to mitigate such biases need to be developed in the future. As such, it is currently crucial to fully consider all possible confounders before integrating gene lists through ‘Rank Aggregation’, or contrasting different study groups for ‘Rank Uniqueness’. For instance, in the common cold coronaviruses dataset, a comparison of all screens conducted in the Huh7 cell line versus others would inadvertently lead to an over-representation of HCoV-229E specific factors, since all HCoV-229E screens are exclusively conducted in the Huh7 cell line. Techniques such as supervised and unsupervised batch effect removal across all studies of a dataset, along with weighted analysis wherein certain studies contribute less or more to the final ranking hit list, present promising avenues for mitigating such biases. However, the delineation of these batches and weights is a non-trivial task, given the multiplicity of attributes impacting the results. Assigning weights or transforming data based on a single attribute through batch effect removal methods may inadvertently enlarge biases imposed from another attribute. Therefore, understanding the characteristics and design of the studies is crucial. To address this complexity, future efforts need to aim to comprehensively identify all parameters possibly impacting the screening results and use such parameters in more sophisticated data modeling approaches. As the number of screening datasets continues to grow, machine learning and artificial intelligence methods could contribute significantly to the meta-analysis of these datasets. These advanced techniques could facilitate holistic analyses by integrating screening data with other omics datasets addressing similar biological questions [90]. Ultimately incorporating additional information such as the experimental validation results of the hits would significantly enhance the accuracy of the proposed consensus and unique hits. All candidate hits identified through our methods are thus the best suggestive options solely based on the available input data. Therefore, while they can serve as suitable candidates at the hypothesis generation step, validations through laboratory methods remain necessary.

Key Points

GeneRaMeN is an R Shiny application designed for the integration, comparison, and meta-analysis of ranked gene lists. GeneRaMeN features three major modules, ‘Rank Aggregation’, ‘Rank Uniqueness’, and ‘Rank Correlation’, empowering users to analyze their datasets of interest on the web or locally.

GeneRaMeN’s ‘Rank Aggregation’ allows users to identify high confident hits consistently present in multiple genetic screening hit lists. By analyzing published and new CRISPR screening datasets, we determined consensus lists of flavivirus and SARS-CoV-2 host factors. Additionally, the analyses helped to demonstrate the redundant cellular pathways exploited by SARS-CoV-2 and the choice of cell line influencing screening outcomes for flaviviruses.

GeneRaMeN’s ‘Rank Correlation’ enables rank correlation analyses to determine groups of genes which show similar or opposite trends to a hit of interest, as exemplified by a significant anti-correlation for the components of the two redundant SARS-CoV-2 entry pathways.

GeneRaMeN’s ‘Rank Uniqueness’ permits users to uncover top hits that uniquely associated with a specific subset of the datasets. We identified specific genes associated with the vulnerability of each cancer type via analyses of more than 1000 CRISPR screening datasets from DepMap initiative, covering nineteen major human tissues and organs.

Supplementary Material

Sup_Fig1_bbae452

Sup_Fig2_bbae452

Supplementary_Table_1_bbae452

Supplementary_Table_2_bbae452

Supplementary_table_3_bbae452

Supplementary_Table_4_bbae452

Supplementary_Data_1_bbae452

Supplementary_Data_2_bbae452

Acknowledgements

We would like to thank Kartik Chandran, Nicholas Heaton, Nanhai Jiang, Dennis Ko, Joseph Trimarco, Ben Waldman, Jacob Berrigan, and Kuo-Feng Weng for test running the prototype versions of GeneRaMeN. We are grateful for the fruitful discussions on data dashboarding with Kuan Rong Chan. We would like to express our gratitude to Marcus Mah, Martin Linster, and Camille Arcinas for their technical help with virus propagation. We are grateful for the excellent services provided by Angie Tan and Su Ting Tay of the Duke-NUS Genome Biology Facility. In preparation of this manuscript, we have occasionally used LLM Chatbots in order to paraphrase some of our original sentences and improve the readability. After using this tool, the authors reviewed and edited the content as needed and take full responsibility for the content of the publication.

Funding

YSO is supported by Duke-NUS Medical School (pilot grant Duke/Duke-NUS/RECA(Pilot)/2019/0047), National Research Foundation of Singapore (NRF-MOST Joint Grant MOH-000928), and the Ministry of Education of Singapore (MOE-000095-01 and MOE-T2EP30123-0008). SG is supported by the Louisiana Clinical and Translational Science Center (NIGMS 2U54GM104940) and the Khoo Bridge Fund, Singapore (KBrFA/2022/0060). The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.

Conflict of interest: The authors declare no conflict of interests.

Data availability

GeneRaMeN is powered by Shiny in R, and its deployed web application can be found at (https://ysolab.shinyapps.io/GeneRaMeN/). GeneRaMeN source codes are deposited on GitHub (https://github.com/MeisamYSF/GeneRaMeN/) and can be freely accessed under a GPLv2 license. All NGS data generated in this study are available on Array Express under accession number E-MTAB-13751.

Author contributions

Conceptualization: MY, YSO

Methodology: MY, SG, YSO

Formal analysis: MY

Investigation: MY, WRS, KLAY, WSL, CLY, FF, YSO

Resources: GJDS, EEO, SL, SG, YSO

Writing–original draft: MY, YSO

Writing–review & editing: MY, KLAY, WSL, CLY, FF, GJDS, SG, YSO

Visualization: MY

Supervision: SG, YSO

Project administration: MY, YSO

Funding acquisition: SG, YSO
==== Refs
References

1. Bock C , DatlingerP, ChardonF. et al. High-content CRISPR screening. Nat Rev Methods Primers 2022;2 :1–23. 10.1038/s43586-021-00093-4.
2. Shi J , WangE, MilazzoJP. et al. Discovery of cancer drug targets by CRISPR-Cas9 screening of protein domains. Nat Biotechnol 2015;33 :661–7. 10.1038/nbt.3235.25961408
3. Deans RM , MorgensDW, ÖkesliA. et al. Parallel shRNA and CRISPR-Cas9 screens enable antiviral drug target identification. Nat Chem Biol 2016;12 :361–6. 10.1038/nchembio.2050.27018887
4. Han K , JengEE, HessGT. et al. Synergistic drug combinations for cancer identified in a CRISPR screen for pairwise genetic interactions. Nat Biotechnol 2017;35 :463–74. 10.1038/nbt.3834.28319085
5. Marceau CD , PuschnikAS, MajzoubK. et al. Genetic dissection of Flaviviridae host factors through genome-scale CRISPR screens. Nature 2016;535 :159–63. 10.1038/nature18631.27383987
6. Puschnik AS , MajzoubK, OoiYS. et al. A CRISPR toolbox to study virus–host interactions. Nat Rev Microbiol 2017;15 :351–64. 10.1038/nrmicro.2017.29.28420884
7. Rousset F , CuiL, SiouveE. et al. Genome-wide CRISPR-dCas9 screens in E. coli identify essential genes and phage host factors. PLoS Genet 2018;14 :e1007749. 10.1371/journal.pgen.1007749.30403660
8. Jeng EE , BhadkamkarV, IbeNU. et al. Systematic identification of host cell regulators of legionella pneumophila pathogenesis using a genome-wide CRISPR screen. Cell Host Microbe 2019;26 :551–563.e6. 10.1016/j.chom.2019.08.017.31540829
9. Tsherniak A , VazquezF, MontgomeryPG. et al. Defining a cancer dependency map. Cell 2017;170 :564–576.e16. 10.1016/j.cell.2017.06.010.28753430
10. Behan FM , IorioF, PiccoG. et al. Prioritization of cancer therapeutic targets using CRISPR–Cas9 screens. Nature 2019;568 :511–6. 10.1038/s41586-019-1103-9.30971826
11. Baggen J , PersoonsL, VanstreelsE. et al. Genome-wide CRISPR screening identifies TMEM106B as a proviral host factor for SARS-CoV-2. Nat Genet 2021;53 :435–44. 10.1038/s41588-021-00805-2.33686287
12. Daniloski Z , JordanTX, WesselsH-H. et al. Identification of required host factors for SARS-CoV-2 infection in human cells. Cell 2021;184 :92–105.e16. 10.1016/j.cell.2020.10.030.33147445
13. Wang R , SimoneauCR, KulsuptrakulJ. et al. Genetic screens identify host factors for SARS-CoV-2 and common cold coronaviruses. Cell 2021;184 :106–119.e14. 10.1016/j.cell.2020.12.004.33333024
14. Zhu Y , FengF, HuG. et al. A genome-wide CRISPR screen identifies host factors that regulate SARS-CoV-2 entry. Nat Commun 2021;12 :961. 10.1038/s41467-021-21213-4.33574281
15. Wei J , AlfajaroMM, DeWeirdtPC. et al. Genome-wide CRISPR screens reveal host factors critical for SARS-CoV-2 infection. Cell 2021;184 :76–91.e13. 10.1016/j.cell.2020.10.028.33147444
16. Biering SB , SarnikSA, WangE. et al. Genome-wide bidirectional CRISPR screens identify mucins as host factors modulating SARS-CoV-2 infection. Nat Genet 2022;54 :1078–89. 10.1038/s41588-022-01131-x.35879412
17. Rebendenne A , RoyP, BonaventureB. et al. Bidirectional genome-wide CRISPR screens reveal host factors regulating SARS-CoV-2, MERS-CoV and seasonal HCoVs. Nat Genet 2022;54 :1090–102. 10.1038/s41588-022-01110-2.35879413
18. Israeli M , FinkelY, Yahalom-RonenY. et al. Genome-wide CRISPR screens identify GATA6 as a proviral host factor for SARS-CoV-2 via modulation of ACE2. Nat Commun 2022;13 :2237. 10.1038/s41467-022-29896-z.35469023
19. Yousefi M , LeeWS, ChanWOY. et al. Betacoronaviruses SARS-CoV-2 and HCoV-OC43 infections in IGROV-1 cell line require aryl hydrocarbon receptor. Emerg Microbes Infect 2023;12 :2256416. 10.1080/22221751.2023.2256416.37672505
20. Grodzki M , BluhmAP, SchaeferM. et al. Genome-scale CRISPR screens identify host factors that promote human coronavirus infection. Genome Med 2022;14 :10. 10.1186/s13073-022-01013-1.35086559
21. Pagès H , CarlsonM, FalconS. et al. AnnotationDbi: Manipulation of SQLite-Based Annotations in Bioconductor. R package version 1.66. 0. 2023. https://bioconductor.org/packages/AnnotationDbi.
22. Kolde R , LaurS, AdlerP. et al. Robust rank aggregation for gene list integration and meta-analysis. Bioinformatics 2012;28 :573–80. 10.1093/bioinformatics/btr709.22247279
23. Li W , XuH, XiaoT. et al. MAGeCK enables robust identification of essential genes from genome-scale CRISPR/Cas9 knockout screens. Genome Biol 2014;15 :554. 10.1186/s13059-014-0554-4.25476604
24. Breitling R , ArmengaudP, AmtmannA. et al. Rank products: a simple, yet powerful, new method to detect differentially regulated genes in replicated microarray experiments. FEBS Lett 2004;573 :83–92. 10.1016/j.febslet.2004.07.055.15327980
25. Kolberg L , RaudvereU, KuzminI. et al. gprofiler2—an R package for gene list functional enrichment analysis and namespace conversion toolset g:profiler. F1000Res 2020;9 :709. 10.12688/f1000research.24956.2.
26. Tsagris M , PapadakisM. Forward regression in R: from the extreme slow to the extreme fast. J Data Sci 2021;16 :771–80. 10.6339/JDS.201810_16(4).00006.
27. Yousefi M , LeeWS, YanB. et al. TMEM41B and VMP1 modulate cellular lipid and energy metabolism for facilitating dengue virus infection. PLoS Pathog 2022;18 :e1010763. 10.1371/journal.ppat.1010763.35939522
28. Steiner S , KratzelA, BarutGT. et al. SARS-CoV-2 biology and host interactions. Nat Rev Microbiol 2024;22 :206–25. 10.1038/s41579-023-01003-z.38225365
29. Schneider WM , LunaJM, HoffmannH-H. et al. Genome-scale identification of SARS-CoV-2 and pan-coronavirus host factor networks. Cell 2021;184 :120–132.e14. 10.1016/j.cell.2020.12.006.33382968
30. Ugalde AP , BretonesG, RodríguezD. et al. Autophagy-linked plasma and lysosomal membrane protein PLAC8 is a key host factor for SARS-CoV-2 entry into human cells. EMBO J 2022;41 :e110727. 10.15252/embj.2022110727.36124427
31. Lee WS , YousefiM, YanB. et al. Know your enemy and know yourself - the case of SARS-CoV-2 host factors. Curr Opin Virol 2021;50 :159–70. 10.1016/j.coviro.2021.08.007.34488003
32. Baggen J , JacquemynM, PersoonsL. et al. TMEM106B is a receptor mediating ACE2-independent SARS-CoV-2 cell entry. Cell 2023;186 :3427–3442.e22. 10.1016/j.cell.2023.06.005.37421949
33. Strine MS , CaiWL, WeiJ. et al. DYRK1A promotes viral entry of highly pathogenic human coronaviruses in a kinase-independent manner. PLoS Biol 2023;21 :e3002097. 10.1371/journal.pbio.3002097.37310920
34. Fu Z , XiangY, FuY. et al. DYRK1A is a multifunctional host factor that regulates coronavirus replication in a kinase-independent manner. J Virol 2023;98 :e01239–23. 10.1128/jvi.01239-23.38099687
35. Wei J , AlfajaroMM, CaiWL. et al. The KDM6A-KMT2D-p300 axis regulates susceptibility to diverse coronaviruses by mediating viral receptor expression. PLoS Pathog 2023;19 :e1011351. 10.1371/journal.ppat.1011351.37410700
36. Pierson TC , DiamondMS. The continued threat of emerging flaviviruses. Nat Microbiol 2020;5 :796–812. 10.1038/s41564-020-0714-0.32367055
37. Kuhn RJ , BarrettADT, DesilvaAM. et al. A prototype-pathogen approach for the development of flavivirus countermeasures. J Infect Dis 2023;228 :S398–413. 10.1093/infdis/jiad193.37849402
38. Zhang R , MinerJJ, GormanMJ. et al. A CRISPR screen defines a signal peptide processing pathway required by flaviviruses. Nature 2016;535 :164–8. 10.1038/nature18625.27383988
39. Li Y , MuffatJ, Omer JavedA. et al. Genome-wide CRISPR screen for Zika virus resistance in human neural cells. Proc Natl Acad Sci U S A 2019;116 :9527–32. 10.1073/pnas.1900867116.31019072
40. Labeau A , Simon-LoriereE, HafirassouM-L. et al. A genome-wide CRISPR-Cas9 screen identifies the dolichol-phosphate mannose synthase complex as a host dependency factor for dengue virus infection. J Virol 2020;94 :e01751–19. 10.1128/JVI.01751-19.31915280
41. Ooi YS , MajzoubK, FlynnRA. et al. An RNA-centric dissection of host complexes controlling flavivirus infection. Nat Microbiol 2019;4 :2369–82. 10.1038/s41564-019-0518-2.31384002
42. Hoffmann H-H , SchneiderWM, Rozen-GagnonK. et al. TMEM41B is a pan-flavivirus host factor. Cell 2021;184 :133–148.e20. 10.1016/j.cell.2020.12.005.33338421
43. Ng WC , KwekSS, SunB. et al. A fast-growing dengue virus mutant reveals a dual role of STING in response to infection. Open Biol 2022;12 :220227. 10.1098/rsob.220227.36514984
44. Savidis G , McDougallWM, MeranerP. et al. Identification of Zika virus and dengue virus dependency factors using functional genomics. Cell Rep 2016;16 :232–46. 10.1016/j.celrep.2016.06.028.27342126
45. Gao H , LinY, HeJ. et al. Role of heparan sulfate in the Zika virus entry, replication, and cell death. Virology 2019;529 :91–100. 10.1016/j.virol.2019.01.019.30684694
46. Shah PS , LinkN, JangGM. et al. Comparative flavivirus-host protein interaction mapping reveals mechanisms of dengue and Zika virus pathogenesis. Cell 2018;175 :1931–1945.e18. 10.1016/j.cell.2018.11.028.30550790
47. Lin DL , InoueT, ChenY-J. et al. The ER membrane protein complex promotes biogenesis of dengue and Zika virus non-structural multi-pass transmembrane proteins to support infection. Cell Rep 2019;27 :1666–1674.e4. 10.1016/j.celrep.2019.04.051.31067454
48. Tabata K , ArakawaM, IshidaK. et al. Endoplasmic reticulum-associated degradation controls virus protein homeostasis, which is required for flavivirus propagation. J Virol 2021;95 :e0223420. 10.1128/JVI.02234-20.33980593
49. Ngo AM , ShurtleffMJ, PopovaKD. et al. The ER membrane protein complex is required to ensure correct topology and stable expression of flavivirus polyproteins. Elife 2019;8 :e48469. 10.7554/eLife.48469.31516121
50. Stadler K , AllisonSL, SchalichJ. et al. Proteolytic activation of tick-borne encephalitis virus by furin. J Virol 1997;71 :8475–81. 10.1128/jvi.71.11.8475-8481.1997.9343204
51. Yeager CL , AshmunRA, WilliamsRK. et al. Human aminopeptidase N is a receptor for human coronavirus 229E. Nature 1992;357 :420–2. 10.1038/357420a0.1350662
52. Hulswit RJG , LangY, BakkersMJG. et al. Human coronaviruses OC43 and HKU1 bind to 9- O -acetylated sialic acids via a conserved receptor-binding site in spike protein domain A. Proc Natl Acad Sci U S A 2019;116 :2681–90. 10.1073/pnas.1809667116.30679277
53. Trimarco JD , HeatonBE, ChaparianRR. et al. TMEM41B is a host factor required for the replication of diverse coronaviruses including SARS-CoV-2. PLoS Pathog 2021;17 :e1009599. 10.1371/journal.ppat.1009599.34043740
54. Kratzel A , KellyJN, V’kovski P. et al. A genome-wide CRISPR screen identifies interactors of the autophagy pathway as conserved coronavirus targets. PLoS Biol 2021;19 :e3001490. 10.1371/journal.pbio.3001490.34962926
55. Padmanabhan P , DesikanR, DixitNM. Targeting TMPRSS2 and cathepsin B/L together may be synergistic against SARS-CoV-2 infection. PLoS Comput Biol 2020;16 :e1008461. 10.1371/journal.pcbi.1008461.33290397
56. Pires De Souza GA , Le BideauM, BoschiC. et al. Choosing a cellular model to study SARS-CoV-2. Front Cell Infect Microbiol 2022;12 :1003608. 10.3389/fcimb.2022.1003608.36339347
57. Feldman AL , LawM, RemsteinED. et al. Recurrent translocations involving the IRF4 oncogene locus in peripheral T-cell lymphomas. Leukemia 2009;23 :574–80. 10.1038/leu.2008.320.18987657
58. Wong RWJ , TanTK, AmandaS. et al. Feed-forward regulatory loop driven by IRF4 and NF-κB in adult T-cell leukemia/lymphoma. Blood 2020;135 :934–47. 10.1182/blood.2019002639.31972002
59. Trinei M , LanfranconeL, CampoE. et al. A new variant anaplastic lymphoma kinase (ALK)-fusion protein (ATIC-ALK) in a case of ALK-positive anaplastic large cell lymphoma. Cancer Res 2000;60 :793–8.10706082
60. Colleoni GWB , BridgeJA, GaricocheaB. et al. ATIC-ALK: a novel variant ALK gene fusion in anaplastic large cell lymphoma resulting from the recurrent cryptic chromosomal inversion, inv(2)(p23q35). Am J Pathol 2000;156 :781–9. 10.1016/S0002-9440(10)64945-0.10702393
61. Auer RL , StarczynskiJ, McElwaineS. et al. Identification of a potential role for POU2AF1 andBTG4 in the deletion of 11q23 in chronic lymphocytic leukemia. Genes Chromosomes Cancer 2005;43 :1–10. 10.1002/gcc.20159.15672409
62. González-Rincón J , MéndezM, GómezS. et al. Unraveling transformation of follicular lymphoma to diffuse large B-cell lymphoma. PloS One 2019;14 :e0212813. 10.1371/journal.pone.0212813.30802265
63. Shi X , JiangY, KitanoA. et al. Nuclear NAD+ homeostasis governed by NMNAT1 prevents apoptosis of acute myeloid leukemia stem cells. Sci Adv 2021;7 :eabf3895. 10.1126/sciadv.abf3895.34290089
64. Schneider C , SpainkH, AlexeG. et al. Breaking the pump: targeting the sodium-potassium pump as a therapeutic strategy in acute myeloid leukemia. Blood 2022;140 :4936–7. 10.1182/blood-2022-168226.
65. Itskovich SS , GurunathanA, ClarkJ. et al. MBNL1 regulates essential alternative RNA splicing patterns in MLL-rearranged leukemia. Nat Commun 2020;11 :2369. 10.1038/s41467-020-15733-8.32398749
66. Goldstein JT , BergerAC, ShihJ. et al. Genomic activation of PPARG reveals a candidate therapeutic axis in bladder cancer. Cancer Res 2017;77 :6987–98. 10.1158/0008-5472.CAN-17-1701.28923856
67. Tate T , XiangT, WobkerSE. et al. Pparg signaling controls bladder cancer subtype and immune exclusion. Nat Commun 2021;12 :6160. 10.1038/s41467-021-26421-6.34697317
68. Morin PJ , SparksAB, KorinekV. et al. Activation of β-catenin-Tcf signaling in colon cancer by mutations in β-catenin or APC. Science 1997;275 :1787–90. 10.1126/science.275.5307.1787.9065402
69. Ilyas M , TomlinsonIPM, RowanA. et al. β-Catenin mutations in cell lines established from human colorectal cancers. Proc Natl Acad Sci U S A 1997;94 :10330–4. 10.1073/pnas.94.19.10330.9294210
70. Mehta RJ , JainRK, LeungS. et al. FOXA1 is an independent prognostic marker for ER-positive breast cancer. Breast Cancer Res Treat 2012;131 :881–90. 10.1007/s10549-011-1482-6.21503684
71. Lo Muzio L , SantarelliA, CaltabianoR. et al. p63 overexpression associates with poor prognosis in head and neck squamous cell carcinoma. Hum Pathol 2005;36 :187–94. 10.1016/j.humpath.2004.12.003.15754296
72. Tacha D , ZhouD, ChengL. Expression of PAX8 in normal and neoplastic tissues: a comprehensive immunohistochemical study. Appl Immunohistochem Mol Morphol 2011;19 :293–9. 10.1097/PAI.0b013e3182025f66.21285870
73. Hu Y , HartmannA, StoehrC. et al. PAX8 is expressed in the majority of renal epithelial neoplasms: an immunohistochemical study of 223 cases using a mouse monoclonal antibody. J Clin Pathol 2012;65 :254–6. 10.1136/jclinpath-2011-200508.22135028
74. Moisés J , NavarroA, SantasusagnaS. et al. NKX2–1 expression as a prognostic marker in early-stage non-small-cell lung cancer. BMC Pulm Med 2017;17 :197. 10.1186/s12890-017-0542-z.29237428
75. Mollaoglu G , JonesA, WaitSJ. et al. The lineage-defining transcription factors SOX2 and NKX2-1 determine lung cancer cell fate and shape the tumor immune microenvironment. Immunity 2018;49 :764–779.e9. 10.1016/j.immuni.2018.09.020.30332632
76. Huang M , WeissWA. Neuroblastoma and MYCN. Cold Spring Harb Perspect Med 2013;3 :a014415–5. 10.1101/cshperspect.a014415.24086065
77. Guglielmi L , CinnellaC, NardellaM. et al. MYCN gene expression is required for the onset of the differentiation programme in neuroblastoma cells. Cell Death Dis 2014;5 :e1081–1. 10.1038/cddis.2014.42.24556696
78. Karaca Atabay E , MeccaC, WangQ. et al. Tyrosine phosphatases regulate resistance to ALK inhibitors in ALK+ anaplastic large cell lymphoma. Blood 2022;139 :717–31. 10.1182/blood.2020008136.34657149
79. Brandstoetter T , SchmoellerlJ, GrausenburgerR. et al. SBNO2 is a critical mediator of STAT3-driven hematological malignancies. Blood 2023;141 :1831–45. 10.1182/blood.2022018494.36630607
80. Wang Z , GuoR, TrudeauSJ. et al. CYB561A3 is the key lysosomal iron reductase required for Burkitt B-cell growth and survival. Blood 2021;138 :2216–30. 10.1182/blood.2021011079.34232987
81. Seegmiller AC , GarciaR, HuangR. et al. Simple karyotype and bcl-6 expression predict a diagnosis of Burkitt lymphoma and better survival in IG-MYC rearranged high-grade B-cell lymphomas. Mod Pathol 2010;23 :909–20. 10.1038/modpathol.2010.76.20348878
82. Tiacci E , LadewigE, SchiavoniG. et al. Pervasive mutations of JAK-STAT pathway genes in classical Hodgkin lymphoma. Blood 2018;131 :2454–65. 10.1182/blood-2017-11-814913.29650799
83. Baus D , NonnenmacherF, JankowskiS. et al. STAT6 and STAT1 are essential antagonistic regulators of cell survival in classical Hodgkin lymphoma cell line. Leukemia 2009;23 :1885–93. 10.1038/leu.2009.103.19440213
84. Skinnider BF , KappU, MakTW. The role of interleukin 13 in classical Hodgkin lymphoma. Leuk Lymphoma 2002;43 :1203–10. 10.1080/10428190290026259.12152987
85. Staege MS , KewitzS, BernigT. et al. Prognostic biomarkers for Hodgkin lymphoma. Pediatr Hematol Oncol 2015;32 :433–54. 10.3109/08880018.2015.1071903.26380871
86. Haas M , CaronG, ChatonnetF. et al. PIM2 kinase has a pivotal role in plasmablast generation and plasma cell survival, opening up novel treatment options in myeloma. Blood 2022;139 :2316–37. 10.1182/blood.2021014011.35108359
87. Li K , ChenC, GaoR. et al. Inhibition of BCL11B induces downregulation of PTK7 and results in growth retardation and apoptosis in T-cell acute lymphoblastic leukemia. Biomark Res 2021;9 :17. 10.1186/s40364-021-00270-3.33663588
88. Matía A , LorenzoMM, PengD. MaGplotR: A Software for the Analysis and Visualization of Multiple MaGeCK Screen Datasets through Aggregation. bioRxiv 2023.01.12.523725. 10.1101/2023.01.12.523725.
89. Li B , ClohiseySM, ChiaBS. et al. Genome-wide CRISPR screen identifies host dependency factors for influenza A virus infection. Nat Commun 2020;11 :164. 10.1038/s41467-019-13965-x.31919360
90. Li R , LiL, XuY. et al. Machine learning meets omics: applications and perspectives. Brief Bioinform 2022;23 :bbab460. 10.1093/bib/bbab460.34791021
