
==== Front
Comput Struct Biotechnol J
Comput Struct Biotechnol J
Computational and Structural Biotechnology Journal
2001-0370
Research Network of Computational and Structural Biotechnology

S2001-0370(24)00271-X
10.1016/j.csbj.2024.08.012
Database Article
RCDdb: A manually curated database and analysis platform for regulated cell death
Wang Xiaopeng abc1
Wang Qing d1
Zhao Jun i1
Chen Jiaxin e1
Wu Ruo ab1
Pan Juanjuan f
Li Jiaxin g
Wang Zechang h
Chen Yongchang abc
Guo Wenting guowt@lpbr.cn
ab⁎
Li Yuanyuan liyy@lpbr.cn
ab⁎
a State Key Laboratory of Primate Biomedical Research, Institute of Primate Translational Medicine, Kunming University of Science and Technology, Kunming, Yunnan, 650500, China
b Yunnan Key Laboratory of Primate Biomedical Research, Kunming, Yunnan, 650500, China
c Southwest United Graduate School, Kunming 650092, China
d College of Bioengineering, Graduate School, Chongqing University, Chongqing 400044, China
e School of Life Sciences and Biotechnology, Shanghai Jiao Tong University, Shanghai 200240, China
f Department of Neurology, The Third Affiliated Hospital of Sun Yat-Sen University, Guangzhou 510630, China
g Graduate School of Zunyi Medical University, Zunyi 563000, China
h Economics and Management School of Wuhan University, Wuhan 430060, China
i School of Medical Informatics, Daqing Campus, Harbin Medical University, Daqing 163319, China
⁎ Corresponding authors at: State Key Laboratory of Primate Biomedical Research, Institute of Primate Translational Medicine, Kunming University of Science and Technology, Kunming, Yunnan, 650500, China. guowt@lpbr.cnliyy@lpbr.cn
1 These authors contributed equally to this work and share first authorship.

22 8 2024
12 2024
22 8 2024
23 32113221
30 6 2024
11 8 2024
11 8 2024
© 2024 The Authors
2024
https://creativecommons.org/licenses/by-nc-nd/4.0/ This is an open access article under the CC BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/).
Regulated cell death is a pivotal regulatory mechanism governing the development and homeostasis of multicellular organisms. A comprehensive understanding of RCD's regulatory mechanisms is crucial for developing novel therapeutic strategies against diseases associated with cell death, such as cancer and neurodegenerative diseases. However, existing data repositories support limited types of cell death data and lack comprehensive annotation and analytical functionalities. Thus, establishing an extensive cell death database is an urgent imperative. To address this gap, we developed the Regulated Cell Death Database (RCDdb, chenyclab.com/RCDdb), the first comprehensively manually annotated database designed to support annotations and analytical capabilities across all RCD types. We compiled 3090 marker gene annotations associated with 15 RCD types from 2180 relevant articles. The RCDdb includes annotation data on these marker genes concerning diseases, drugs, pathways, proteins, and gene expressions. Furthermore, it provides 49 diverse visualization methods to present this information. More importantly, the RCDdb features three online analysis tools for identifying and analyzing RCD-related features within user-submitted data. Furthermore, the RCDdb offers a user-friendly interface for querying, browsing, analysis, and visualization of detailed information associated with each RCD category. This resource promises to significantly aid researchers in better understanding the mechanisms of cell death, thereby accelerating progress in research and therapeutic strategies aimed at combating RCD-related diseases.

Keywords

Regulated cell death
Database
Analysis platform
Bioinformatics
==== Body
pmc1 Introduction

Cells actively initiate and execute their death through regulated cell death (RCD) processes in a controlled manner. RCD, which encompasses programmed cell death (PCD), is a biological process wherein cells undergo orchestrated and intentional self-destruction. This mechanism is crucial for maintaining intracellular homeostasis, eliminating unnecessary or damaged cells, and regulating the development and function of tissues and organs. Unlike accidental cell death (ACD), which results from severe physical, chemical, or mechanical stress, RCD involves a series of molecular events mediated by signal transduction pathways. According to the updated guidelines established by the Nomenclature Committee on Cell Death, RCD is classified into various forms, including apoptosis, pyroptosis, necroptosis, autophagy-dependent cell death (ADCD), entotic cell death, NETotic cell death, parthanatos, mitochondrial permeability transition (MPT)-driven necrosis, immunogenic cell death (ICD), lysosome-dependent cell death (LCD), and ferroptosis [1]. In recent years, research has unveiled additional forms of cell death, including alkaliptosis, oxeiptosis, cuproptosis, and disulfidptosis.

Apoptosis, the most well-known form of RCD, plays a crucial role in maintaining normal development and tissue homeostasis [2], [3]. Pyroptosis, an inflammatory form of cell death triggered by certain inflammasomes, leads to the cleavage of gasdermin D (GSDMD) and the activation of inactive cytokines such as IL-18 and IL-1β [4]. Necroptosis is a regulated, caspase-independent cell death mechanism characterized by morphological features similar to necrosis [5]. ADCD is defined as RCD that mechanistically depends on the autophagy machinery, with components of this machinery being selectively utilized or repurposed for ADCD [1], [6], [7]. Entotic cell death is triggered by the detachment of epithelial cells from the extracellular matrix and the loss of integrin signaling [1]. NETotic cell death can promote the release of danger signals, contributing to the pathogenesis of autoimmune disorders and other conditions [1], [8]. Parthanatos is a type of RCD that relies on poly(ADP-ribose) polymerase 1 (PARP-1) and is activated by oxidative stress-induced DNA damage and chromatinolysis [1], [8], [9]. MPT-driven necrosis refers to a form of RCD initiated by the opening of the mitochondrial permeability transition pore (MPTP), resulting in mitochondrial dysfunction and cell death [1]. ICD can drive inflammatory responses by activating cytotoxic T lymphocyte (CTL)-driven adaptive immune responses and establishing long-term immune memory [10]. LCD refers to the form of RCD that is triggered by lysosomal membrane permeabilization (LMP) and the subsequent release of lysosomal contents into the cytosol, driven by lysosomal cathepsin proteases [11]. Ferroptosis is a unique form of RCD driven by iron-dependent phospholipid peroxidation [12]. Alkaliptosis involves intracellular alkalinization and shares many morphological features with primary necrosis, with the activation of the NF-kB pathway and the downregulation of CA9 being key factors promoting alkaliptosis [13]. Oxeiptosis is a novel form of caspase-independent RCD induced by oxygen radicals through activation of the KEAP1–PGAM5–AIFM1 pathway [14]. Cuproptosis is triggered by the direct binding of copper ions to lipoylated components of the tricarboxylic acid (TCA) cycle, resulting in protein aggregation and the loss of iron–sulfur cluster proteins [15]. Finally, disulfidptosis is a novel form of RCD induced by the accumulation of intracellular disulfide molecules under conditions of glucose starvation and abnormal expression of the cystine transporter SLC7A11 [16].

In recent years, RCD has garnered increasing attention from researchers worldwide, leading to the establishment of databases that house gene data associated with various cell death mechanisms. For example, databases such as FerrDb and HAMdb encompass gene information relevant to ferroptosis and autophagy [17], [18]. XDeathDB offers a comprehensive visualization platform for analyzing signaling networks of cell death modes associated with cancer and COVID-19 [19]. The ApoptoProteomics database integrates proteomic data from apoptosis processes, facilitating access to and analysis of dynamic changes and functional impacts of related proteins [20]. DeathBase is dedicated to the structural, evolutionary, and functional analysis of proteins involved in cell death and apoptosis, offering precise information on protein roles and supporting user queries and database browsing [21]. However, current data resources encounter several limitations, including inadequate support for various RCD mechanisms, incomplete data coverage for specific genes associated with different types of RCD, and limited capabilities for disease association analysis and data visualization. To address these gaps, we developed a new database called the Regulated Cell Death database (RCDdb) (chenyclab.com/RCDdb). This database was meticulously curated manually to encompass a comprehensive array of RCD-related data, including all RCD types identified to date. The current version of RCDdb includes 3090 literature-verified annotations and associated information on RCD, encompassing 1850 genes, 52,046 proteins, 69,554 diseases, 20,701 drugs, and 15,236 functional pathways. Moreover, the RCDdb features a user-friendly interface for searching, browsing, and visualizing various RCD-related data resources (Fig. 1). We anticipate that this database, supported by robust literature and meticulous design, will significantly catalyze future research endeavors.Fig. 1 Database content and construction. The RCDdb is the first comprehensive and manually curated database of RCD and its analysis platform, integrating the most extensive and detailed data resources currently available in the field of RCD research. RCDdb provides a user-friendly interface for efficient querying, browsing, analysis, and visualization of detailed information across various RCD datasets. (B) The chronological timeline of the discovery of RCD. (C) The types of data accessible to users and their associated analytical functionalities.

Fig. 1

2 Materials and methods

2.1 Article collection

After integrating data resources from FerrDb, HAMdb, and ncFO [22], we identified 2425 RCD-related genes through a keyword search on PubMed (https://www.ncbi.nlm.nih.gov/pubmed) (Table 1). These genes were manually annotated to include information on gene biotype, Ensembl ID, PMID, RCD type associated with each gene, and the relationship between the gene and RCD as reported in the literature. In total, our annotations encompassed 3090 entries sourced from 2180 articles. The annotated genes included 15 alkaliptosis-related genes, 609 apoptosis-related genes, 876 ADCD-related genes, 27 cuproptosis-related genes, 16 disulfidptosis-related genes, 17 entotic cell death-related genes, 602 ferroptosis-related genes, 34 ICD-related genes, 32 LCD-related genes, 30 MPT-driven necrosis-related genes, 86 necroptosis-related genes, 8 NETotic cell death-related genes, 10 oxeiptosis-related genes, 10 parthanatos-related genes, and 53 pyroptosis genes. These genes comprised 100 protein-coding genes, 152 miRNAs, 32 circRNAs, 55 lncRNAs, 3 transcribed pseudogenes, 1 antisense RNA, 1 pre-miRNA, 1 RNase P RNA, and 1 pseudogene.Table 1 Statistical information on RCD types.

Table 1RCD type	Number of RCD genes	Number of PMIDs	Number of Comments	
Alkaliptosis	15	6	15	
Apoptosis	609	587	608	
Autophagy_dependent_cell_death	876	814	1151	
Cuproptosis	27	1	8	
Disulfidptosis	16	1	16	
Entotic_cell_death	17	11	11	
Ferroptosis	602	585	766	
Immunogenic_cell_death	34	11	19	
Lysosome_dependent_cell_death	32	15	36	
MPT_driven_necrosis	30	18	27	
Necroptosis	86	77	85	
NETotic_cell_death	8	7	9	
Oxeiptosis	10	6	10	
Parthanatos	10	8	11	
Pyroptosis	53	48	56	

2.2 Data collection

The clusterProfiler package [23] in R was used to perform enrichment analysis on genes associated with each RCD type, with P-values < 0.05 indicating significantly enriched pathways. Genes associated with the 15 RCD types were found to be enriched in a total of 14,146 Gene Ontology (GO) terms and 1090 Kyoto Encyclopedia of Genes and Genomes (KEGG) pathways. Individual analysis of 2425 RCD-related genes was conducted using the disgenet2r package [24] in R, selecting data with experimental evidence from the CURATED database, which resulted in 69,554 gene–disease associations based on published literature. Data pertaining to 20,701 gene–drug interactions were downloaded from DGIdb and DrugBank [25], [26] as well as data on 48,943 protein annotations from STRING (v11.5) [27]. In addition, we extracted data on 3103 annotations for all RCD-related genes from UniProt. Cancer analysis data were extracted from the UCSC XENA database [28], including genomic mutation data, transcriptomic sequencing data, and clinical survival data, covering 33 cancer tissues and normal tissues from the Genotype-Tissue Expression (GTEx) cohort. The dataset encompassed 10,327 tumor tissue samples and 11,483 normal tissue samples.

2.3 Enrichment analysis of RCD-related genes

In the RCDdb, enrichment analysis can be performed using 15 RCD reference gene sets. Users can submit gene lists and select parameters according to their requirements. The RCDdb annotates user-submitted genes against these reference sets and calculates the statistical significance of enrichment analysis using hypergeometric tests. The significance (P-value) of gene enrichment within the reference set is calculated as follows:(i) P=1−∑i=0x−1kin−ks−ins

The reference set is defined as comprising n genes, with k being a component of the reference set under study. The query list of genes of interest consists of a total of s genes, with i belonging to the same reference set. Users can modify the number of enriched genes and accordingly set P-values, FDR, and Bonferroni correction thresholds to ensure the accuracy of the analysis.

To evaluate the similarity between query gene set A and reference gene set B, we employed two classical measures: the Jaccard score and the Simpson score. The Jaccard score represents the proportion of elements in the intersection of A and B relative to their union. The Simpson score represents the proportion of elements in the intersection of A and B relative to the smaller of the two sets, A or B.

The Jaccard score is calculated as follows:(ii) JaccardA,B=A∩BA∪B=A∩BA+B−A∩B

The Simpson score is calculated as follows:(iii) SimpsonA,B=A∩Bmin(A|,|B)

These two scores offer additional information for enrichment analysis, allowing users to select additional parameters that enhance their comprehension of the results.

2.4 Analysis of signature RCD-related genes

RCDdb utilizes three classic machine learning methods to extract feature genes from input data: (i) The glmnet package [29] in R is used for least absolute shrinkage and selection operator (LASSO) regression. This method selects non-zero RCD-related genes through 10-fold cross-validation. LASSO regression, a variant of linear regression, incorporates an L1 regularization term to the coefficients, which facilitates feature selection by reducing the coefficients of less important features to zero. (ii) The randomForest package in R implements the random forest algorithm, leveraging 500 decision trees to identify the top 20 key feature genes. Random forest, an ensemble learning method, improves the model's accuracy and robustness by constructing multiple decision trees and aggregating their predictions. (iii) The xgboost package is employed to conduct 500 rounds of iterative analysis to identify the top 20 significant feature genes. This method is supported by SHapley Additive exPlanation (SHAP) values, which elucidate the importance of each gene feature. Extreme Gradient Boosting (XGBoost), a gradient boosting algorithm, incrementally improves the model's predictive ability by iteratively training multiple decision trees. SHAP values explain model predictions, aiding in the comprehension of the contribution of each feature to the model's predictions.

2.5 RCD scoring analysis

The RCDdb integrates 15 reference gene sets, encompassing a total of 2425 genes associated with RCD. The GSVA package [30] in R supports both gene set variation analysis (GSVA) and single-sample gene set enrichment analysis (ssGSEA). GSVA is employed to assess the activity levels of gene sets across samples, while ssGSEA calculates enrichment scores for gene sets in individual samples. These analytical methods enable researchers to accurately identify the types of RCD influencing specific samples, thereby unveiling the complex regulatory mechanisms of RCD across various disease states and biological conditions.

2.6 Other analyses

Statistical analysis: R software was used for statistical analysis, with P-values < 0.05 indicating statistical significance. Differential expression analysis: Differentially expressed genes (DEGs) were identified based on read counts using the DESeq2 package [31] in R. The screening criteria were set to a fold change of ≥ 2 and adjusted P-values of < 0.05, ensuring the inclusion of the RCD gene among the DEGs. Enrichment analysis: The clusterProfiler package in R was employed for GO [32] and KEGG [33] analyses of differentially expressed RCD-related genes. In addition, all RCD-related genes were subjected to GSEA, retaining pathways with P-values of < 0.05. Weighted gene co-expression network analysis (WGCNA) [34]: WGCNA, a weighted analytical method based on gene co-expression networks, was utilized to examine the correlation between genes, classify, and screen them. This method is widely used for analyzing gene expression data. In this study, we extracted the expression data for all RCD-related genes and performed WGCNA for clinical trait-based module analysis. The top 50 RCD-related genes showing the strongest correlation were identified for protein–protein interaction network analysis. Analysis of tumor-related signaling molecules: The PROGENy package [35] in R was used to evaluate the activity of 14 tumor-related signaling pathways based on the expression data of RCD-related genes and numerous publicly available perturbation experiments. Immune cell infiltration analysis: The expression data of RCD-related genes were analyzed using the CIBERSORT algorithm [36] to determine the proportion of immune cells in tumor tissues. CIBERSORT is a computational method that utilizes gene expression data to infer the relative abundance of various cell types in mixed cell populations. It operates by employing established single-cell gene expression data to train a linear model, which can then map gene expression data from mixed-cell populations to estimate the relative abundance of different cell types. Signature gene analysis: A composite sampling method was employed to address the data imbalance between tumor and normal samples. Subsequently, using LASSO, random forest, XGBoost, and multiple support vector machine recursive feature elimination (mSVM-RFE) algorithms [37], we identified RCD-related genes with significant characteristics between tumor and normal samples. Survival analysis as performed using the gepia package in Python (log-rank test) and the survival package in R (univariate Cox proportional hazards regression analysis). Mutation data for somatic mutation analysis were downloaded from the UCSC XENA database. The maftools package [38] in R was used to analyze somatic mutations, retain samples containing signature genes identified using the four machine-learning algorithms, and visualize 15 RCD-related genes with the highest mutation rate.

3 Database use and access

3.1 Search interface for retrieving RCD data

The RCDdb is a robust platform featuring an intuitive search interface, enabling users to easily access RCD-related data. Users can refine their RCD dataset queries through three pathways: "Search by RCD Type (input the RCD of interest)," "Search by Gene Name (input the gene name of interest)," and "Search by Cancer Type (input the cancer type of interest)" (Fig. 2A). For queries based on RCD type, users can navigate details via a drop-down list or by clicking on a network diagram. The search results include the following information: (i) RCD type, literature sources, and gene descriptions (Fig. 2B); (ii) GO and KEGG functional annotations for RCD-related genes (Fig. 2C); (iii) Disease annotations, disease names, related literature, and original text descriptions; (iv) Drug information associated with RCD-related genes, supported by corresponding published literature evidence; (v) Protein annotations and intracellular localization details for RCD-related genes; (vi) Transcriptional expression data from seven databases, including TCGA and ENCODE, for RCD genes. In the "Gene Name" search module, users can input a gene name to access comprehensive details. The RCDdb provides exhaustive information for each gene, encompassing RCD type, related literature with supporting evidence, and links to eight mainstream annotation databases (Fig. 2D). Additionally, users can access information pertaining to the gene's diseases, drugs, and protein annotations, three-dimensional (3D) structural information of the corresponding protein, survival prognosis data of the gene in various cancers, and transcriptional expression data across different tissues (Fig. 2E). This meticulous organization in the RCDdb facilitates an in-depth understanding of the characteristics and functions of RCD-related genes, supplying essential data for research and analysis. Given that cancer is a heterogeneous disease characterized by dysregulated cell death mechanisms, the RCDdb's cancer search module specifically targets RCD genes across all cancer types. It employs more than ten analytical methods for multi-omics data analysis to identify key RCD biomarkers in cancer, advancing the understanding of cancer development mechanisms and exploring potential therapeutic avenues.Fig. 2 Search functionality introduction and user guide. (A) The RCDdb provides three distinct search modes to meet various querying requirements. (B) The search results interface for different RCD types comprises seven major modules: RCD overview, GO pathways, KEGG pathways, diseases, drugs, proteins, and gene expression. (C) Visualization of annotation analysis results for GO and KEGG pathways. (D) The RCD gene query results page includes comprehensive details on genes, disease associations, drug effects, protein characteristics, survival rate analysis, and gene expression, organized into six panels. (E) Graphical representation of results for the modules on diseases, proteins, survival analysis, and gene expression.

Fig. 2

3.2 Effective online analysis tools

In addition to serving as a comprehensive data resource, the RCDdb contains three robust analysis tools that offer extensive coverage from the gene to the sample level (Fig. 3A). The RCD Gene Set Enrichment Analysis tool allows users to input a list of genes of interest. The RCDdb maps these genes to the RCD gene sets for hypergeometric testing and calculates Jaccard and Simpson similarities between the submitted gene list and each type of RCD gene set, thereby identifying RCD types potentially associated with the input genes. The results display RCD types, RCD gene counts, proportions, Jaccard scores, Simpson scores, hypergeometric test P-values, and adjusted P-values. Additionally, the RCDdb provides Venn diagrams, bubble charts, and bar graphs to visually represent the results. Upon selecting an RCD type of interest and clicking the hyperlink in the count column, users can view a network diagram of intersecting genes, facilitating further study of their association with cancer through survival analysis (Fig. 3B).Fig. 3 Analysis functionality explanation and user guide. (A) The RCDdb provides three unique analysis tools designed to meet the diverse requirements of researchers. (B) The RCD gene enrichment analysis allows users to input a specific gene set, with the RCDdb returning result charts based on calculated statistical metrics. (C) The RCD signature gene analysis requires users to input an expression matrix and grouping information, utilizing algorithms such as LASSO, random forest, and XGBoost to screen for and return charts of signature genes. (D) The RCD gene scoring analysis requires users to input an expression matrix, with the RCDdb employing methods such as GSVA and ssGSEA to calculate scores for various RCD indicators within the matrix.

Fig. 3

The RCD Signature Gene Analysis tool employs three classic machine learning algorithms—LASSO, random forest, and XGBoost—to identify RCD-related signature genes. Users submit a standardized expression matrix and sample grouping information, after which the RCDdb identifies RCD signature genes from the matrix using the chosen algorithm. The results present the signature genes and their corresponding RCD types in a sunburst chart, illustrating their hierarchical relationships. Additionally, graphical representations of model training information and expression correlations between signature genes aid users in comprehensively understanding and analyzing the characteristics and associations of RCD signature genes (Fig. 3C).

Using the RCD Gene Set Scoring Analysis tool, users can calculate an RCD score for any sample. The RCDdb supports scoring with two algorithms, GSVA and ssGSEA, requiring only a standardized expression matrix as input. The RCDdb computes scores for the input expression matrix based on RCD gene sets. The analysis results provide sample details, RCD types, and corresponding RCD scores. Additionally, heatmap and violin plot visualizations enable users to intuitively observe the distribution of scores for each RCD type within the samples (Fig. 3D).

3.3 User-friendly interface for browsing RCD data

The "Browse" page features a timeline showcasing 15 RCD findings, accompanied by interactive tables that allow users to quickly search for RCD-related genes and apply customizable filters based on RCD type. When a specific RCD type is selected, the primary table on the right displays information, including the RCD type, gene name, gene description, gene type, PMID, and associated evidence.

3.4 Data download

All gene sets of RCD types have been organized and categorized into individual files available for download in our database. The RCDdb offers these files in two formats: ".csv" and ".txt." Users can download the reference collection as crucial supplementary data for conducting in-depth experimental research.

4 Case studies

4.1 Exploring RCD mechanisms associated with gastric cancer prognosis using RCDdb tools

The RCDdb offers multiple analysis tools for investigating the potential functions of RCD-related genes. To demonstrate the utility of these tools within the RCDdb database, we utilized the PANoptosis genes associated with gastric cancer prognosis, identified by Pan et al., as input data for the "RCD Gene Set Enrichment Analysis" tool (Fig. 4A) [39]. PANoptosis represents a novel and complex mode of cell death that encompasses the combined pathways of necroptosis, pyroptosis, and apoptosis [40]. The analysis results showed significant associations of these genes with three RCD types: necroptosis (P = 6.72E-09), pyroptosis (P = 3.36E-07), and apoptosis (P = 3.40E-03) (Fig. 4B). Further survival analysis indicated that the intersecting gene sets of pyroptosis and apoptosis significantly impact the survival rates of patients with gastric cancer (Fig. 4C). Interleukin 1 alpha (IL1A), identified as a high-risk gene for gastric cancer, exhibits high expression in gastric cancer, associated with a significant reduction in patient survival time (Fig. 4D). In contrast, a functional inactivation point mutation of interferon regulatory factor 1 (IRF1) is related to the development of gastric cancer (Fig. 4E) [41]. These findings are consistent with the conclusions of the study authors. Additionally, we quantified the scores of RCD gene sets for the dataset used by Pan et al. using the "RCD Gene Set Scoring Analysis" tool. The results indicated high scores for RCD types such as necroptosis, pyroptosis, and apoptosis, with other RCD types also demonstrating good performance, especially the recently discovered disulfidptosis, which scored most prominently (Fig. 4F). Research indicates that various types of RCD significantly impact gastric cancer and its microenvironment [42]. These results provide novel insights into exploring gastric cancer from an RCD perspective. Importantly, the validation of these findings further substantiates the practicality and value of the RCDdb in RCD research.Fig. 4 Comprehensive analysis of PANoptosis prognostic genes in gastric cancer. (A) Schematic diagram illustrating the PANoptosis pathway and the associated gene sets input. (B) Results table of RCD type enrichment analysis. (C) Survival analysis results for the intersection genes between apoptosis and necroptosis in gastric cancer. (D) Survival analysis result for IL1A in gastric cancer. (E) Disease-related annotations of IRF1, including its expression levels across pan-cancer and normal tissues. (F) Scoring results of gastric cancer samples across different types of RCD.

Fig. 4

4.2 Multidimensional analysis study of head and neck squamous cell carcinoma (HNSC) from the perspective of RCD

HNSC, a malignant tumor affecting the oral cavity, pharynx, nasal cavity, or larynx, ranks as the sixth most common cancer worldwide [43]. Research indicates that various types of RCD are implicated in the pathogenesis of HNSC [44]. We utilized the "Search by Cancer Type" module to explore the impact of multiple RCD types on HNSC. The results identified 432 DEGs associated with RCD across 14 RCD types (Fig. 5A). These DEGs are predominantly enriched in signaling pathways associated with cell death and immune responses (Fig. 5B). Through WGCNA, we identified gene modules relevant to HNSC and constructed a protein–protein interaction network for the module most strongly associated with HNSC (Fig. 5C). KIF4A emerged with high connectivity within this module and exhibited elevated expression levels in tumor samples. KIF4A, a marker gene for ferroptosis, has been identified as a potential therapeutic target for tumors [45]. To further elucidate the role of RCD in HNSC, we employed four machine learning algorithms—LASSO, random forest, XGBoost, and mSVM-RFE—to identify 80 characteristic genes involved in RCD processes such as apoptosis, ferroptosis, and autophagy (Fig. 5D). Genomic variation analysis in patients with HNSC revealed mutations in approximately 51.38 % (93/181) of these characteristic genes (Fig. 5E). Evaluation of the 16 overlapping genes identified by these algorithms in relation to immune cell abundance demonstrated that MMP9 expression significantly correlates with the presence of various immune cells (Fig. 5F). Further data from the RCDdb revealed that MMP9, a marker gene for apoptosis, is associated with 173 diseases and represents a potential target for 17 drugs. Moreover, the MMP family has been validated as therapeutic targets and prognostic biomarkers for HNSC treatment [46]. Additionally, Cox regression analysis highlighted significant prognostic implications of characteristic genes such as AATF, CHRND, SLC3A2, and CFL1 in HNSC treatment (Fig. 5H) [47], [48], [49], [50]. This in-depth study from various RCD perspectives has advanced our understanding of the pathogenesis of HNSC, offering novel insights and encouraging the exploration of RCD-related biological questions in other cancer types (Fig. 5I).Fig. 5 Research on the identification of key RCD biomarkers in HNSC. (A) Distribution of DEGs in HNSC across various types of RCD (adjusted P-values < 0.05 and |log2FC||= > 1). (B) Results of GO and KEGG enrichment analyses. (C) WGCNA and the protein–protein interaction network within the green module. (D) Distribution of 80 signature genes identified by four machine learning algorithms: LASSO, random forest, XGBoost, and mSVM-RFE, across different RCD types. (E) Oncoplot illustrating signature genes in the TCGA-HNSC cohort. (F) Analysis of immune infiltration using CIBERSORT, including correlation and significance assessments between overlapping signature genes and immune cell populations. (G) Results of univariate Cox regression analysis of overlapping signature genes. (H) Number of tumor samples and normal control samples included in the cancer search module, along with a summary of the proportion of signature genes identified in each cancer type.

Fig. 5

5 System design and implementation

The current iteration of the RCDdb was developed using MySQL 5.7.17 (http://www.mysql.com) and operates on a Linux-based Apache Web server (http://www.apache.org). PHP 7.0 (http://www.php.net) was used for server-side scripting. The interactive interface was designed and built using Bootstrap v3.3.7 (https://v3.bootcss.com) and JQuery v2.1.1 (http://jquery.com). We employed ECharts (https://www.echartsjs.com/) and Highcharts (https://www.highcharts.com.cn/) to construct a graphical visualization framework. The web page features an interactive 3D protein structure viewer based on Mol* viewer [51], which automatically fetches structures from the AlphaFold Protein Structure Database. We recommend the use of modern web browsers that support HTML5 standards, such as Firefox, Google Chrome, Safari, Opera, or IE 9.0 + , for an optimal user interface. Some website icons were sourced from https://www.iconfont.cn/ and https://www.biorender.com/.

6 Conclusions and future extensions

RCD is not an isolated event within organisms but rather exists within a complex network comprising multiple cell death mechanisms that can synergistically act under specific physiological and pathological states of the organism [8]. Various forms of RCD, such as apoptosis, necroptosis, autophagic death, and other recently identified types, closely interact and regulate each other. These interactions play crucial roles in organism development, maintaining tissue homeostasis, and responding to external stimuli. Recent research has emphasized the combined utilization of different RCD types. For example, Zou et al., through a comprehensive analysis of 12 RCD types, effectively predicted clinical prognosis and drug sensitivity in patients with triple-negative breast cancer [52]. Similarly, Wei et al. employed machine learning to predict molecular subtypes, prognosis, and treatment responses in patients with lung adenocarcinoma across 13 RCD patterns [53]. Therefore, investigating the interconnectedness and synergistic mechanisms among different RCD types holds significant scientific and clinical relevance for understanding cell fate decisions, disease progression, and optimizing treatment strategies.

In recent years, RCD has emerged as a research hotspot, leading to the rapid accumulation of RCD data. The RCDdb addresses the limitations of incomplete data in existing resources by providing a reliable platform for investigating interactions and associations across various RCD types, thereby advancing RCD research. The RCDdb offers several advantages, underscoring its importance as a comprehensive database: (i) It consolidates data from 2180 articles, encompassing over 3000 detailed reference annotations and covering 1850 genes associated with 15 RCD types. (ii) It facilitates the classification of individual genes into specific RCD types, supported by extensive literature evidence, thereby aiding the exploration of synergistic mechanisms between different RCD types. (iii) The meticulously curated RCD gene association dataset includes information on diseases, drugs, proteins, and functions, presented using sophisticated visualization techniques. (iv) It integrates various online analysis tools for enriching and annotating RCD features in user data. (v) It conducts in-depth RCD-level analyses across all cancer types, providing valuable resources and novel perspectives for cancer research. (vi) It features a user-friendly interface that supports the download of reference RCD collections with interactive tables. (vii) It enables easy navigation of referenced RCD datasets, complemented by detailed help documentation. However, the RCDdb does have some limitations. For instance, some analysis results are primarily based on transcriptomic data, which may not fully capture RCDs that rely on post-transcriptional activation mechanisms. Therefore, we encourage researchers to incorporate multi-omics data when using the RCDdb to obtain more accurate and comprehensive results.

The RCDdb is the first comprehensive and manually annotated database dedicated to RCD. It was developed in response to the significant demand from researchers in the fields of cellular biology, molecular biology, genetics, and data science for a database that consolidates RCD data comprehensively. The current iteration of the RCDdb houses the most extensive and detailed collection of data within the RCD domain. Moving forward, we are committed to ongoing manual curation and regular updates of the database with newly validated RCD data. Additionally, we plan to incorporate advanced algorithms and develop new annotation tools. We firmly believe that the RCDdb will emerge as an indispensable and invaluable resource in the field of RCD research, fostering scientific discoveries in the biomedical domain.

Ethical approval

This declaration is not applicable.

Funding

This work was supported by the 10.13039/501100001809 National Natural Science Foundation of China (82125008 , 81930121 ), the National Science and Technology Innovation 2030 Major Program (2021ZD0200900 ), the 10.13039/501100012166 National Key Research and Development Program of China (2018YFA0107902 and 2018YFA0801403 ), and the Major Basic Research Project of Science and Technology of Yunnan (202102AA100053 ).

Declaration of Competing Interest

The manuscript is not submitted to print and electronic manuscripts elsewhere, and there is no economic benefit (except for the author's basic academic career) that may lead to the appearance of a conflict of interest. We are glad to take this opportunity to submit our work to show our platform. We are very grateful for your editorial attention and suggestions for this manuscript.

Acknowledgements

We would like to thank KetengEdit (www.ketengedit.com) for its linguistic assistance during the preparation of this manuscript.

Author contributions

Y.L. and T.G. conceived and designed the research. X.W. and Q.W. wrote the manuscript. X.W. contributed to data analysis and website development. Q.W. contributed to data collection and data analysis. J.P., J.L., and Z.W. contributed to data collection. J.C. and R.W. contributed to data visualization and article revision. Y.C. provided financial support.
==== Refs
References

1 Galluzzi L. Vitale I. Aaronson S.A. Abrams J.M. Adam D. Agostinis P. Molecular mechanisms of cell death: recommendations of the Nomenclature Committee on Cell Death 2018 Cell Death Differ 25 2018 486 541 29362479
2 Ketelut-Carneiro N. Fitzgerald K.A. Apoptosis, pyroptosis, and necroptosis-oh my! The many ways a cell can die J Mol Biol 434 2022 167378
3 Tian Y. Xiao H. Yang Y. Zhang P. Yuan J. Zhang W. Crosstalk between 5-methylcytosine and N(6)-methyladenosine machinery defines disease progression, therapeutic response and pharmacogenomic landscape in hepatocellular carcinoma. Mol Cancer 22 2023 5 36627693
4 Fang Y. Tian S. Pan Y. Li W. Wang Q. Tang Y. Pyroptosis: a new frontier in cancer Biomed Pharm 121 2020 109595
5 Jagtap P.G. Degterev A. Choi S. Keys H. Yuan J. Cuny G.D. Structure-activity relationship study of tricyclic necroptosis inhibitors J Med Chem 50 2007 1886 1895 17361994
6 Lu Q. Kou D. Lou S. Ashrafizadeh M. Aref A.R. Canadas I. Nanoparticles in tumor microenvironment remodeling and cancer immunotherapy J Hematol Oncol 17 2024 16 38566199
7 Yang Y. Liu L.X. Tian Y. Gu M.M. Wang Y.N. Ashrafizadeh M. Autophagy-driven regulation of cisplatin response in human cancers: exploring molecular and cell death dynamics Cancer Lett 587 2024
8 Tang D. Kang R. Berghe T.V. Vandenabeele P. Kroemer G. The molecular machinery of regulated cell death Cell Res 29 2019 347 364 30948788
9 Zheng D. Liu J. Piao H. Zhu Z. Wei R. Liu K. ROS-triggered endothelial cell death mechanisms: Focus on pyroptosis, parthanatos, and ferroptosis Front Immunol 13 2022 1039241
10 Galluzzi L. Vitale I. Warren S. Adjemian S. Agostinis P. Martinez A.B. Consensus guidelines for the definition, detection and interpretation of immunogenic cell death J Immunother Cancer 2020 8
11 Aits S. Jaattela M. Lysosomal cell death at a glance J Cell Sci 126 2013 1905 1912 23720375
12 Dixon S.J. Lemberg K.M. Lamprecht M.R. Skouta R. Zaitsev E.M. Gleason C.E. Ferroptosis: an iron-dependent form of nonapoptotic cell death Cell 149 2012 1060 1072 22632970
13 Song X. Zhu S. Xie Y. Liu J. Sun L. Zeng D. JTC801 induces pH-dependent death specifically in cancer cells and slows growth of tumors in mice Gastroenterology 154 2018 1480 1493 29248440
14 Holze C. Michaudel C. Mackowiak C. Haas D.A. Benda C. Hubel P. Oxeiptosis, a ROS-induced caspase-independent apoptosis-like cell-death pathway Nat Immunol 19 2018 130 140 29255269
15 Tsvetkov P. Coy S. Petrova B. Dreishpoon M. Verma A. Abdusamad M. Copper induces cell death by targeting lipoylated TCA cycle proteins Science 375 2022 1254 1261 35298263
16 Liu X. Zhuang L. Gan B. Disulfidptosis: disulfide stress-induced cell death Trends Cell Biol 2023
17 Zhou N. Yuan X. Du Q. Zhang Z. Shi X. Bao J. FerrDb V2: update of the manually curated database of ferroptosis regulators and ferroptosis-disease associations Nucleic Acids Res 51 2023 D571 D582 36305834
18 Wang N.N. Dong J. Zhang L. Ouyang D. Cheng Y. Chen A.F. HAMdb: a database of human autophagy modulators with specific pathway and disease information J Chemin 10 2018 34
19 Gadepalli V.S. Kim H. Liu Y. Han T. Cheng L. XDeathDB: a visualization platform for cell death molecular interactions Cell Death Dis 12 2021 1156 34907160
20 Arntzen M.O. Thiede B. ApoptoProteomics, an integrated database for analysis of proteomics data obtained from apoptotic cells Mol Cell Proteom 11 2012
21 Diez J. Walter D. Munoz-Pinedo C. Gabaldon T. DeathBase: a database on structure, evolution and function of proteins involved in apoptosis and other forms of cell death Cell Death Differ 17 2010 735 736 20383157
22 Zhou S. Huang Y.E. Xing J. Zhou X. Chen S. Chen J. ncFO: a comprehensive resource of curated and predicted ncRNAs associated with ferroptosis Genom Proteom Bioinforma 2022
23 Wu T. Hu E. Xu S. Chen M. Guo P. Dai Z. clusterProfiler 4.0: a universal enrichment tool for interpreting omics data Innov (Camb) 2 2021 100141
24 Pinero J. Ramirez-Anguita J.M. Sauch-Pitarch J. Ronzano F. Centeno E. Sanz F. The DisGeNET knowledge platform for disease genomics: 2019 update Nucleic Acids Res 48 2020 D845 D855 31680165
25 Cannon M. Stevenson J. Stahl K. Basu R. Coffman A. Kiwala S. DGIdb 5.0: rebuilding the drug-gene interaction database for precision medicine and drug discovery platforms Nucleic Acids Res 52 2024 D1227 D1235 37953380
26 Knox C. Wilson M. Klinger C.M. Franklin M. Oler E. Wilson A. DrugBank 6.0: the DrugBank Knowledgebase for 2024 Nucleic Acids Res 52 2024 D1265 D1275 37953279
27 Szklarczyk D. Gable A.L. Lyon D. Junge A. Wyder S. Huerta-Cepas J. STRING v11: protein-protein association networks with increased coverage, supporting functional discovery in genome-wide experimental datasets Nucleic Acids Res 47 2019 D607 D613 30476243
28 Goldman M.J. Craft B. Hastie M. Repecka K. McDade F. Kamath A. Visualizing and interpreting cancer genomics data via the Xena platform Nat Biotechnol 38 2020 675 678 32444850
29 Friedman J. Hastie T. Tibshirani R. Regularization paths for generalized linear models via coordinate descent J Stat Softw 33 2010 1 22 20808728
30 Hanzelmann S. Castelo R. Guinney J. GSVA: gene set variation analysis for microarray and RNA-seq data BMC Bioinforma 14 2013 7
31 Love M.I. Huber W. Anders S. Moderated estimation of fold change and dispersion for RNA-seq data with DESeq2 Genome Biol 15 2014 550 25516281
32 Harris M.A. Clark J. Ireland A. Lomax J. Ashburner M. Foulger R. The Gene Ontology (GO) database and informatics resource Nucleic Acids Res 32 2004 D258 D261 14681407
33 Kanehisa M. Furumichi M. Tanabe M. Sato Y. Morishima K. KEGG: new perspectives on genomes, pathways, diseases and drugs Nucleic Acids Res 45 2017 D353 D361 27899662
34 Langfelder P. Horvath S. WGCNA: an R package for weighted correlation network analysis BMC Bioinforma 9 2008 559
35 Holland C.H. Tanevski J. Perales-Paton J. Gleixner J. Kumar M.P. Mereu E. Robustness and applicability of transcription factor and pathway analysis tools on single-cell RNA-seq data Genome Biol 21 2020 36 32051003
36 Newman A.M. Liu C.L. Green M.R. Gentles A.J. Feng W. Xu Y. Robust enumeration of cell subsets from tissue expression profiles Nat Methods 12 2015 453 457 25822800
37 Duan K.B. Rajapakse J.C. Wang H. Azuaje F. Multiple SVM-RFE for gene selection in cancer classification with expression data IEEE Trans Nanobioscience 4 2005 228 234 16220686
38 Mayakonda A. Lin D.C. Assenov Y. Plass C. Koeffler H.P. Maftools: efficient and comprehensive analysis of somatic variants in cancer Genome Res 28 2018 1747 1756 30341162
39 Pan H. Pan J. Li P. Gao J. Characterization of PANoptosis patterns predicts survival and immunotherapy response in gastric cancer Clin Immunol 238 2022 109019
40 Wang Y. Kanneganti T.D. From pyroptosis, apoptosis and necroptosis to PANoptosis: a mechanistic compendium of programmed cell death pathways Comput Struct Biotechnol J 19 2021 4641 4657 34504660
41 Nozawa H. Oda E. Ueda S. Tamura G. Maesawa C. Muto T. Functionally inactivating point mutation in the tumor-suppressor IRF-1 gene identified in human gastric cancer Int J Cancer 77 1998 522 527 9679752
42 Wang H. Liu M. Zeng X. Zheng Y. Wang Y. Zhou Y. Cell death affecting the progression of gastric cancer Cell Death Discov 8 2022 377 36038533
43 Johnson D.E. Burtness B. Leemans C.R. Lui V.W.Y. Bauman J.E. Grandis J.R. Head and neck squamous cell carcinoma Nat Rev Dis Prim 6 2020 92 33243986
44 Raudenska M. Balvan J. Masarik M. Cell death in head and neck cancer pathogenesis and treatment Cell Death Dis 12 2021 192 33602906
45 Sun X. Chen P. Chen X. Yang W. Chen X. Zhou W. KIF4A enhanced cell proliferation and migration via Hippo signaling and predicted a poor prognosis in esophageal squamous cell carcinoma Thorac Cancer 12 2021 512 524 33350074
46 Liu M. Huang L. Liu Y. Yang S. Rao Y. Chen X. Identification of the MMP family as therapeutic targets and prognostic biomarkers in the microenvironment of head and neck squamous cell carcinoma J Transl Med 21 2023 208 36941602
47 Fu L. Jin Q. Dong Q. Li Q. AATF is overexpressed in human head and neck squamous cell carcinoma and regulates STAT3/survivin signaling Onco Targets Ther 14 2021 5237 5248 34785906
48 Li J. Xu Y. Peng G. Zhu K. Wu Z. Shi L. Identification of the Nerve-Cancer Cross-Talk-Related Prognostic Gene Model in Head and Neck Squamous Cell Carcinoma Front Oncol 11 2021 788671
49 Li C. Chen S. Jia W. Li W. Wei D. Cao S. Identify metabolism-related genes IDO1, ALDH2, NCOA2, SLC7A5, SLC3A2, LDHB, and HPRT1 as potential prognostic markers and correlate with immune infiltrates in head and neck squamous cell carcinoma Front Immunol 13 2022 955614
50 Liu R. Wang C. Sun Z. Shi X. Zhang Z. Luo J. Neuronal CFL1 upregulation in head and neck squamous cell carcinoma enhances tumor-nerve crosstalk and promotes tumor growth Mol Carcinog 2024
51 Sehnal D. Bittrich S. Deshpande M. Svobodova R. Berka K. Bazgier V. Mol* Viewer: modern web app for 3D visualization and analysis of large biomolecular structures Nucleic Acids Res 49 2021 W431 W437 33956157
52 Zou Y. Xie J. Zheng S. Liu W. Tang Y. Tian W. Leveraging diverse cell-death patterns to predict the prognosis and drug sensitivity of triple-negative breast cancer patients after surgery Int J Surg 107 2022 106936
53 Wei Q. Jiang X. Miao X. Zhang Y. Chen F. Zhang P. Molecular subtypes of lung adenocarcinoma patients for prognosis and therapeutic response prediction with machine learning on 13 programmed cell death patterns J Cancer Res Clin Oncol 149 2023 11351 11368 37378675
