
==== Front
Cell Rep Med
Cell Rep Med
Cell Reports Medicine
2666-3791
Elsevier

S2666-3791(24)00375-6
10.1016/j.xcrm.2024.101661
101661
Article
Discovery and validation of a 10-gene predictive signature for response to adjuvant chemotherapy in stage II and III colon cancer
Xu Chaohan 129
Xia Peng 39
Li Jie 49
Lewis Keeli.B. 5
Ciombor Kristen K. 6
Wang Lily 28
Smith J. Joshua smithj5@mskcc.org
710∗
Beauchamp R. Daniel 510
Chen X. Steven steven.chen@miami.edu
281011∗∗
1 College of Bioinformatics Science and Technology, Harbin Medical University, Harbin 150081, China
2 Department of Public Health Sciences, University of Miami Miller School of Medicine, Miami, FL 33136, USA
3 School of Biological Science & Medical Engineering, Southeast University, Nanjing 210096, China
4 Academy of Biomedical Engineering, Kunming Medical University, Kunming 650500, China
5 Section of Surgical Sciences, Department of Surgery, Vanderbilt University Medical Center, Nashville, TN 37232, USA
6 Division of Hematology and Oncology, Department of Medicine, Vanderbilt University Medical Center, Nashville, TN 37232, USA
7 Colorectal Service, Department of Surgery, Memorial Sloan Kettering Cancer Center, New York, NY 10065, USA
8 Sylvester Comprehensive Cancer Center, University of Miami Miller School of Medicine, Miami, FL 33136, USA
∗ Corresponding author smithj5@mskcc.org
∗∗ Corresponding author steven.chen@miami.edu
9 These authors contributed equally

10 These authors contributed equally

11 Lead contact

25 7 2024
20 8 2024
25 7 2024
5 8 1016612 8 2023
30 12 2023
2 7 2024
© 2024 The Author(s)
2024
https://creativecommons.org/licenses/by-nc-nd/4.0/ This is an open access article under the CC BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/).
Summary

Identifying patients with stage II and III colon cancer who will benefit from 5-fluorouracil (5-FU)-based adjuvant chemotherapy is crucial for the advancement of personalized cancer therapy. We employ a semi-supervised machine learning approach to analyze a large dataset with 933 stage II and III colon cancer samples. Our analysis leverages gene regulatory networks to discover an 18-gene prognostic signature and to explore a 10-gene signature that potentially predicts chemotherapy benefits. The 10-gene signature demonstrates strong prognostic power and shows promising potential to predict chemotherapy benefits. We establish a robust clinical assay on the NanoString nCounter platform, validated in a retrospective formalin-fixed paraffin-embedded (FFPE) cohort, which represents an important step toward clinical application. Our study lays the groundwork for improving adjuvant chemotherapy and potentially expanding into immunotherapy decision-making in colon cancer. Future prospective studies are needed to validate and establish the clinical utility of the 10-gene signature in clinical settings.

Graphical abstract

Highlights

• Semi-supervised learning to derive clinically relevant biomarkers

• A new 10-gene signature refines chemotherapy decisions in stage II/III colon cancer

• The signature shows the strong prognostic value and potential for treatment benefits

• Validation using FFPE samples to ensure clinical applicability

Xu et al. develop and validate a 10-gene signature that refines chemotherapy decisions for stage II and III colon cancer, demonstrating its prognostic value and potential to predict treatment benefits using extensive dataset analysis and validation with FFPE samples.

Keywords

colon cancer
prediction
machine learning
chemotherapy
immune checkpoint blockade
prognosis
gene expression
adjuvant therapy
biomarkers
Published: July 25, 2024
==== Body
pmcIntroduction

Colon cancer, a disease characterized by its heterogeneity, is the fourth deadliest cancer worldwide, causing nearly 900,000 deaths each year.1 Colon cancer has been categorized as either localized to bowel (stage I to II), regionally spread to lymph nodes (stage III), or distantly spread to the liver, lung, peritoneum, brain, or distant lymph nodes (stage IV). The five-year survival rates for patients with stage I colon cancer are approximately 90% without chemotherapy but range from 10% to 50% for stage IV disease despite treatment with chemotherapy.2,3 Five-year disease-free survival rates in patients with high-risk stage II colon cancer were reported to be 80.7%–83.9%, and the three-year disease-free survival rate among patients with stage III colon cancer was from 73.4% to 76.3%.4,5 Generally, patients with stage II and III colon cancer can be considered as clinically relevant groups for risk stratification of disease recurrence.6,7,8 For some patients with high-risk stage II colon cancer, and the majority of patients with stage III colon cancer, adjuvant chemotherapy (ACT) consisting of fluoropyrimidine with or without oxaliplatin has been routinely used as an important treatment following surgery with the intent to cure.9

However, the decision about ACT for stage II colon cancer has been challenging and controversial; the results from clinical trials are not supportive of ACT for all patients with stage II colon cancer.10,11 Although ACT has been recommended for all patients with stage III colon cancer, only ∼20% of patients receiving such therapy will benefit from the treatment in terms of a survival advantage.12 Progress in treating this disease has been hampered by the inability to accurately determine which patients are most likely to benefit and which are least likely to benefit from the current systemic treatment options. It is therefore critical to discover better predictive biomarkers to identify the stage II and III patients who will most likely benefit from ACT. Also, it is important to avoid unnecessary toxicity for the rest of the patients who are not as likely to benefit from that therapy and/or to stimulate further research into more effective adjuvant therapy options.

Several biological biomarkers, including microsatellite instability (MSI) and loss of heterozygosity at chromosome 18q (18q LOH), have historically been proposed for prognostic and predictive purposes in colon cancer.13,14 However, only a minority of patients with early-stage colon cancer (less than 15%) are MSI-high (MSI-H), and the prognostic significance of 18q LOH has not been conclusively validated in clinical trials.15,16 Moreover, the currently available commercial gene signature panels for colon or colorectal cancer (CRC), such as Oncotype DX colon cancer,17 ColoPrint,18 Veridex,19 and GeneFx colon,20 have primarily been developed with a focus on prognosis. While these panels have undergone external validation, their results have been somewhat inconsistent, and there is a lack of concrete evidence supporting their predictive value specifically for ACT benefit.21,22

In the realm of colon cancer ACT benefit prediction, a few gene signatures have been proposed.23,24,25 However, compared to general prognostic signatures, these specific signatures targeting chemotherapy benefit prediction have not undergone extensive independent validation. This gap highlights a critical need for reliable ACT benefit biomarkers, not only for stage II and III colon cancer but also for practical implementation in clinical settings using formalin-fixed paraffin-embedded (FFPE) samples. Our study aims to address this gap by focusing on the development and validation of gene signatures tailored for predicting ACT benefits, thereby contributing to the refinement of personalized treatment strategies in colon cancer.

To further improve the identification of prognosis/prediction signatures, we assembled a comprehensive colon cancer stage II and III gene expression dataset, which included 933 samples from six datasets. We then applied an innovative semi-supervised machine learning approach to identify an 18-gene prognosis signature and a 10-gene chemotherapy benefit prediction signature set (Figure 1). The prognosis and chemotherapy-benefit prediction signatures were validated in five and two testing datasets, respectively, including targeted NanoString nCounter data we generated from FFPE samples. We showed that patients predicted to benefit from ACT based on the 10-gene signature had significantly better survival outcomes than those predicted not to benefit using both fresh-frozen and FFPE samples. Additionally, the 10-gene signature displayed the potential to predict the utility of cancer immunotherapy in an adjuvant setting in selected patients.Figure 1 The workflow chart

Schematic of the study design for (A) the development and validation of the 18-hub-gene prognosis signature; and (B) the development and validation of the 10-gene chemotherapy-benefit signature.

Results

Semi-supervised gene network analysis identified an 18-hub-gene set

We first selected GSE39582, which was our only discovery dataset with survival outcomes and is one of the largest gene expression datasets for stage II and III colon cancer, to identify genes associated with survival outcome. A total of 3,091 genes with significant nominal p values (p ≤ 0.05) were selected as candidates. Next, to further pare down a more focused gene set from these survival-associated genes, these survival-associated genes in GSE39582 were combined with five gene expression datasets without survival outcome to construct the large gene network in stage II and III colon cancer. Specifically, we combined GSE39582 with five additional stage II and III colon cancer expression datasets (GSE27854, GSE26906, GSE21510, GSE18088, and GSE2109) to form a large gene expression dataset with 933 samples and 3,091 genes. The SPACE algorithm was then used to construct a gene network, which included 504 nodes and 737 edges (Figure 2A). Since hub genes with more connections in the gene network are more biologically meaningful compared with the p value-ranked gene list of meta-analysis,26,27 we selected the top 18 hub genes with more than 11 connections (degrees) in the hub gene list (Table S2), which included THY1, BUB1, SDPR, NAP1L3, DEPDC1, STON1, DLGAP5, FBN1, ATAD2, HSD17B2, KIF23, CHAC2, FRMD6, MCM10, TPX2, AURKB, CYR61, and FAM84A genes (Figure 2B). The description of gene functions and pathway analysis results are listed in Table S3 and Figure S1, respectively.Figure 2 Identification of the 18-hub-gene signature and validation of the Random Survival Forests (RSF) prognosis model

(A) The topology of the constructed survival-related gene network in colon cancer. A total of 504 genes (nodes) with at least 11 connections with other genes were identified as hub genes and labeled in enlarging.

(B) Composition of the 18 survival-associated hub genes. Degree refers to the number of genes that the hub gene has a direct connection with based on the constructed gene network in Figure A. GISTIC (http://portals.broadinstitute.org/tcga/home) records the information of somatic copy-number variation: “+” indicates the genes with significant amplification and “−” indicates significant deletion in CRC. MethHC (http://methhc.mbc.nctu.edu.tw/php/index.php) records the information on DNA methylation: “+” indicates the genes with significant hypermethylation and “−” indicates significant hypomethylation in CRC. DisGeNET (https://www.disgenet.org/home/) records the information on human gene-disease associations (GDAs): “+” indicates the genes related to CRC. The gene symbols of the ten genes with more than two pieces of evidence that occurred in CRC were highlighted in red.

(C) Maximally selected log rank statistics were applied to the in-bag predicted value of the RSF prognosis model. The optimal cutoff point was determined as 14.38.

(D–G) Validation of the 18-hub-gene signature in 4 independent datasets, including GSE17538, GSE33113, GSE37892, and GSE38832. The high- and low-risk groups were defined based on the 18-hub-gene signature, which was derived from the GSE39582 cohort. The optimal cutoff point, which was determined by the maximally selected rank statistics (maxstat), was used to divide the patients into high- and low-risk groups. Red and dark blue lines indicate predicted high- and low-relapse risk groups. HR compares the RFS of the low- and high-relapse risk groups. p values were obtained by the log rank test.

Prognosis model construction and performance evaluation

To determine if the discovered 18-hub-gene set could predict colon cancer prognosis, we constructed a prognosis prediction model using GSE39582 as our training dataset. The optimal cutoff point based on maximally selected rank statistics was determined as 14.38 for the high- and low-risk group (Figure 2C).28 Next, we evaluated this prognostic model using four additional independent expression datasets with available survival information (GSE17538, GSE33113, GSE37892, and GSE38832). In each testing dataset, the samples were also divided into high- and low-relapse risk groups using the optimal cutoff point in Figure 2C. The hazard ratio (HR) and log rank p value of predicted risk groups in the five datasets were 3.324 (p = 0.005), 3.933 (p = 0.003), 11.543 (p < 0.0001), and 4.745 (p = 0.032), respectively (Figures 2D–2G). Furthermore, the univariable Cox proportional hazard regression analysis was performed and the results suggested that the 18-hub-gene set was significantly associated with relapse-free survival (RFS) outcome in GSE17538 (HR = 3.336, p = 0.008), GSE33113 (HR = 3.967, p = 0.006), GSE37892 (HR = 12.05, p = 3.79 × 10−9), and GSE38832 (HR = 4.795, p = 0.051) (Table S4). Also, the multivariable Cox proportional hazard regression analysis with additional clinical variables including age, gender, and stage was completed, and the results showed that the 18-hub-gene set was an independent factor for RFS outcome in GSE17538 (HR = 4.136, p = 0.003), GSE33113 (HR = 4.117, p = 0.005), GSE37892 (HR = 8.238, p = 2.73 × 10−6), and GSE38832 (HR = 4.577, p = 0.058) (Figure S2A; Table S4). These results demonstrated the 18-hub-gene set provides an effective gene signature for colon cancer prognosis.

18-Hub-gene prognostic model outperformed existing signatures

In the construction of clinical prognostic models, the efficient selection of a gene signature set from a high-dimensional gene list is a crucial step. Typically, in a p value-based gene selection approach, the set of genes with the most significant p values is selected to build the prognostic prediction model. To compare the performances of the p value-based gene selection approaches with our network-based gene selection approach, we chose 18 genes with the most significant p values from the Cox regression analysis and re-constructed a prognosis model of RFS using the same five independent GEO datasets described earlier. We found that the prediction performance of the 18-hub genes obtained with our semi-supervised approach had better performance than those based on the 18 most significant genes. The median AUC of the 18-hub-gene set and the 18-most-significant gene set for predicting 5-year RFS was 0.732 and 0.591, respectively (Figure 3A). Additionally, analysis using pairwise mutual information distance measuring association showed that the 18-hub-gene signature set generated by our semi-supervised approach had much lower redundant information than the 18-top-ranked gene set (two-sample Kolmogorov-Smirnov test, p value < 0.00001, Figure 3A). Moreover, principal component analysis also demonstrated that the 18-hub-gene set captured more variation in patients with colon cancer than the 18-top-ranked gene set (Figure 3B).Figure 3 Comparison of the 18-hub-gene set and 18-most-significant gene set

(A) Boxplot of median AUC for predicting 5-year RFS in four GEO testing datasets (GSE17538, GSE33113, GSE37892, and GSE38832), and pairwise mutual information distance based on expression values in the GSE39582 dataset.

(B) Expression variation across the samples in the GSE39582 dataset based on the principal component analysis.

(C) Gene expression patterns of the two gene sets in the GSE39582 dataset.

Using the expression dataset of 469 stage II and III colon cancer samples in GSE39582, we next compared the expression patterns of two gene sets and found the average expression level of the 18-hub-gene group was significantly higher than the expression level of the 18-top-ranked group (t test, p value < 2.2 × 10−16), demonstrating the effectiveness of our approach for optimally identifying a more relevant colon cancer-related gene signature set for prediction analysis (Figure 3C).

Furthermore, we compared the prediction performance of the 18-hub-gene set with other established gene sets for colon cancer prognosis prediction, including the Oncotype DX,29 ColoGuidePro,30 ColoFinder,31 the 8-gene signature of Peng et al.,32 the 13-gene signature of Tian et al.,33 the 3-gene signature of Cheng et al.,34 the 15-gene signature of Rokavec et al.,35 and the 12-gene signature of Fang et al.36 For each prognostic signature generated by these alternative methods, we also calculated HR, log rank test p values, and 5-year AUC for the four GEO validation cohorts (GSE17538, GSE33113, GSE37892, and GSE38832). The results showed the 18-hub-gene set generated by our semi-supervised approach had the best performance in prognostic prediction (Table S5).

Identification and validation of 5-FU-based ACT chemotherapy-benefit gene signature

5-Fluorouracil (5-FU)-based ACT refers to patients who have received 5-FU treatment alone (either intravenous 5-FU or oral capecitabine) or in combination with other medications, including folinic acid and oxaliplatin. Among all available colon cancer datasets, two datasets (GSE39582 and GSE17538) included 142 and 80 primary tumor samples from patients who subsequently received adjuvant 5-FU-based treatments. We next tried to determine a gene signature that can more accurately predict patients with colon cancer who are likely to benefit from ACT therapy and thus demonstrate improved survival relative to 5-FU-based ACT from the primary 18-hub-gene set. To this end, we identified potential cancer driver genes with either significant genomic aberration or DNA methylation alteration between CRC and normal samples using the GISTIC tool37 for The Cancer Genome Atlas copy number and MethHC database38 (see Figure 1B) and manually interrogated the 18-hub-gene set with these two databases. To not miss some crucial genes that play an important role in colon cancer, we also verified the 18-hub genes in the DisGeNET database,39 which is the largest disease-associated gene database. Finally, the genes occurring more than two times in the GISTIC, MethHC, or DisGeNET databases (Figure 2B) were selected and used to construct a 10-gene chemotherapy-benefit signature set. An RSF chemotherapy-benefit model was re-constructed using these 10 genes, along with 142 samples with 5-FU-based ACT information in the GSE39582 dataset.

The prognostic performance of the 10-gene chemotherapy-benefit signature set was evaluated in the four GEO colon cancer cohorts (GSE17538, GSE33113, GSE37892, and GSE38832). The HRs and log rank test p values calculated to evaluate the prediction performance in four datasets were 2.97 (p = 0.002), 4.467 (p = 0.004), 5.963 (p < 0.0001), and 2.642 (p = 0.122), respectively (Figures S3A–S3D). The results suggested that the 10-gene signature set demonstrated a prognostic performance similar to the 18-hub-gene prognosis signature set.

We then assessed which patients demonstrated an association with survival outcomes relative to 5-FU-based ACT by the 10-gene chemotherapy-benefit prediction model. According to the maximally selected rank statistics (maxstat) provided in the “survminer” package, the samples were divided into ACT-benefit and non-benefit groups using the optimal cutoff point (Figure 4A). The HR and p value validated by another colon cancer cohort, GSE17538 with ACT information, were HR = 3.884 and p = 0.002 (Figure 4B), which demonstrated significant prediction of RFS outcome for the 10-gene chemotherapy-benefit signature set in stage II and III colon cancer. Furthermore, the results of univariable Cox proportional hazard regression analysis suggested that the 10-gene chemotherapy-benefit signature set was significantly associated with RFS of patients with 5-FU-based ACT in GSE17538 (HR = 3.933, p = 0.004) (Table S6); and the results of multivariable Cox proportional hazard regression analysis also showed that the 10-gene chemotherapy-benefit signature set was a significant independent factor for RFS of patients with 5-FU-based ACT in GSE17538 (HR = 4.476, p = 0.003) (Figure S2B; Table S6). Furthermore, we conducted a control experiment in which we randomly selected 5–15 genes 1,000 times and compared their prediction performance with the 10-gene set in GSE17538. As shown in the results (Figure 4C), the AUC of the 10-gene set for predicting 5-year RFS was 0.756, significantly outperforming the random gene set (p = 0.009). These data demonstrate that the 10-gene chemotherapy-benefit signature set selected by the genomic event could more accurately predict patients with colon cancer who would more likely benefit from ACT.Figure 4 Validation of the RSF chemotherapy-benefit model

(A) Maximally selected log rank statistics were applied to the in-bag predicted value of the RSF chemotherapy-benefit model. The optimal cutoff point was determined as 8.88.

(B) Validation of the 10-gene chemotherapy-benefit signature in another independent dataset, GSE17538, which contains only samples with 5-FU-based ACT. The ACT-benefit and non-benefit groups were defined based on the 10-gene chemotherapy-benefit signature, which was derived from samples with 5-FU-based ACT in the GSE39582 cohort. The optimal cutpoint, which was determined by the maximally selected rank statistics (maxstat), was used to divide the patients into the ACT-benefit and non-benefit groups. Dark blue and red lines indicate predicted ACT-benefit and non-benefit groups. HR compares the RFS of the ACT-benefit and non-benefit group. p values were obtained by the log rank test.

(C) The density plot illustrated the AUC values of 1,000 randomly selected gene sets for predicting 5-year RFS. The red dotted line represents an AUC of 0.756 for the 10-gene set in predicting 5-year RFS.

(D and E) The predictive performance evaluation of the prognosis model and chemotherapy-benefit model in the VUMC data. The 18-gene prognosis model divided high- and low-relapse risk groups in stage II to III patients in (D), and the 10-gene chemotherapy-benefit model divided ACT-responsive and non-responsive groups in stage II to III patients in (E).

(F) The mutation pattern of KRAS, BRAF, PIK3CA, PTEN, NRAS, and AKT1 between VUMC ACT-benefit and non-benefit patients.

(G) The boxplot of the 18-gene score and 10-gene score between the different locations of lesions in the VUMC cohort.

(H) The bar plot of the 18-gene and 10-gene groupings according to the patients with MSI-H in the VUMC cohort.

Validation of 18- and 10-gene signatures using FFPE samples

To verify the utility of our 18-gene prognosis model and 10-gene chemotherapy-benefit model, we collected 109 stage II/III colon cancer specimens and detected the expression of 18-hub-gene by the NanoString nCounter platform, and 56 of them had received 5-FU-based ACT (Table 1; Table S8). The 18-gene prognosis model was used to divide these 109 specimens into high- and low-relapse risk groups. The HR and log rank p values were 3.542 (p = 0.009) (Figure 4D). The results of univariable (p = 0.014, HR = 3.559) (Table S4) and multivariable (p = 0.0262, HR = 4.46) (Figure S2A; Table S4) Cox regression analyses suggested that the 18-gene prognosis model is an independent prognosis factor. Then the 10-gene chemotherapy-benefit prediction model was used to divide 56 stage II to III specimens with 5-FU-based ACT into benefit and non-benefit groups. Those patients predicted to be in the non-benefit group showed worse RFS after 5-FU-based ACT. The HR was 2.735 and log rank test p value was p = 0.058 (Figure 4E). The HR and p value of univariable and multivariable Cox regression analyses were 2.765 (p = 0.069) and 4.469 (p = 0.074) (Figure S2B; Table S6). We also investigated the mutation pattern of KRAS, BRAF, PIK3CA, PTEN, NRAS, and AKT1 in the Vanderbilt University Medical Center (VUMC) analysis group of patients within the ACT-benefit and non-benefit groups. The results suggested that the mutation frequencies of KRAS and PIK3CA in the ACT-non-benefit group (40% and 20%) were higher than in the ACT-benefit group (9.8% and 4.9%) (Figure 4F). In addition, we analyzed the distribution of the location of lesions and MSI-H status according to the 18- and 10-gene signature groupings in the VUMC cohort. The results suggested that the 18-gene score and 10-gene score show no significant difference between left- and right-sided lesions in the VUMC cohort; the p value (Wilcoxon rank-sum test) was 0.548 and 0.82, respectively (Figure 4G). For MSI status, the VUMC cohort has 10 patients with MSI-H. Intriguingly, MSI-H status is equally distributed among the high- (n = 5) and low (n = 5)-risk groups according to the 18-gene signature (Figure 4H). However, MSI-H status is distributed only according to the ACT-benefit grouping (n = 6) of the 10-gene signature (Figure 4H). This observation suggests that the 10-gene signature may have implications for identifying patients who are more likely to benefit from ACT, specifically within the MSI-H subgroup.Table 1 Patient and tumor characteristics of the different sets

Characteristics	Type	Mean age (sda, range), years	Gender a(M/F)	TNM stage	Treatment with 5-FUa	Median follow-upa (sd, range), months	Relapsea	
0	I	II	III	IV	NA	Yes	No	N/A	
GSE39582 (n = 566)∗	mRNA, independent	66.9 (13.3, 22–97)	310/256	4	33	264∗	205∗	60	0	142∗	48 (40.4, 0–201)	139	322	8	
GSE27854 (n = 115)∗	mRNA, independent	–	–	0	16	41∗	35∗	23	0	–	–	–	–	–	
GSE26906 (n = 90)∗	mRNA, independent	66.5 (12.6, 31–94)	43/47	0	0	90∗	0∗	0	0	–	–	–	–	–	
GSE21510 (n = 123)∗	mRNA, independent	–	–	0	15	46∗	39∗	23	0	–	–	–	–	–	
GSE18088 (n = 53)∗	mRNA, independent	65.4 (12.2, 34–85)	26/27	0	0	53∗	0∗	0	0	–	–	–	–	–	
GSE2109 (n = 276)∗	mRNA, independent	–	138/137	2	38	87∗	73∗	36	40	–	–	–	–	–	
GSE17538 (n = 232)∗	mRNA, independent	64.7 (13.4, 23–94)	122/110	0	27	72∗	76∗	56	0	80∗	40.7 (26.6, 0.4–118.6)	34	111	3	
GSE33113 (n = 90)∗	mRNA, independent	70.4 (13, 35–95)	42/48	0	0	90∗	0∗	0	0	0	38.8 (26.5, 1.6–118.3)	18	71	1	
GSE37892 (n = 130)∗	mRNA, independent	68.3 (12.7, 22–97)	69/61	0	0	73∗	57∗	0	0	0	42.9 (23.4, 0.6–103.1)	37	93	0	
GSE38832 (n = 122)∗	mRNA, independent	–	–	0	18	35∗	39∗	30	0	0	41.2 (34.1, 0.3–111.4)	9	65	0	
VUMC (n = 136/144)∗	mRNA, independent	65.5 (11.5, 42–86)	–	0	1	65∗	44∗	26	0	56∗	49.9 (51.8, 0.1–241.6)	20	89	0	
N/A, not available; sd, standard deviation; ∗ the datasets used to derive and validate the signature.

a Among patients with stage II to III colon cancer.

Drug sensitivity analysis for 5-FU-based ACT chemotherapy-benefit gene signature

The 10-gene chemotherapy-benefit signature was applied to predict the 5-FU benefit of NCI-60 cell lines, which were collected and downloaded from CellMiner.40 We then assigned 60 cell lines into 14 5-FU-sensitive and 46 5-FU-resistant groups by the 10-gene prediction model, which included 2 and 5 colon cell lines, respectively (Figure 5A). Meanwhile, 1,886 5-FU-related NCI-60 studies were collected from CellMiner, The Wilcoxon rank-sum test for −log10(GI 50) values between 2 5-FU-resistant colon cell lines and 5 5-FU-sensitive colon cell lines was applied. Interestingly, we found the median of GI 50 values in the 5 5-FU-sensitive colon cell lines was significantly higher than that in the 2 5-FU-resistant colon cell lines (p value < 2.2 × 10−16) (Figure 5B), suggesting that 5 ACT-beneficial colon cell lines were more sensitive to 5-FU. In summary, the 10-gene chemotherapy-benefit signature set has a good performance for predicting 5-FU sensitivity vs. resistance for colon cancer cell lines from the NCI-60 data.Figure 5 The performance of the 10-gene chemotherapy-benefit signature set for predicting 5-FU resistance for 60 cancer cell lines from the NCI-60 data

(A) The heatmap plot of NCI-60 cell lines, which were clustered into 22 resistant and 38 sensitive groups based on the drug activity Z scores of the 10-gene chemotherapy-benefit signature.

(B) The violin plot of −log10(GI 50) values between resistant and sensitive groups. Abbreviations are as follows: BR, breast; CNS, central nervous system; CO, colon; LC, non-small cell lung; LE, leukemia; ME, melanoma; OV, ovarian; PR, prostate; and RE, renal.

Pathway enrichment analysis for 10-gene chemotherapy-benefit signature-associated gene networks

To better understand the biology underlying the 10-gene chemotherapy-benefit signature, we next performed pathway enrichment of the ten genes (THY1, BUB1, DEPDC1, STON1, ATAD2, HSD17B2, TPX2, AURKB, CYR61, and FAM84A), along with the 135 genes connected to them on the SPACE gene network (Figure 2A). We tested 743 canonical pathways collected in MSigDB (Molecular Signature Database from Broad Institute) that contained genes overlapping with the 145 genes described earlier (Table S7). The results showed a total of 147 pathways with nominal p values less than 0.05. Next, we computed a pathway enrichment score for each sample using ssGSEA (single-sample gene set enrichment analysis) and compared these scores for the ACT-benefit and non-benefit groups in the testing cohort GSE17538 (Table S7).41 Figure S4 shows a heatmap including pathways with significantly different ssGSEA scores between the two groups (adjusted p value less than 0.1). Among these significant pathways, 49 pathways showed high expression activities in the ACT-non-benefit group, and notably, many of them are tumor microenvironment-related pathways, such as REACTOME_INTEGRIN_CELL_SURFACE_INTERACTIONS, REACTOME_EXTRACELLULAR_MATRIX_ORGANIZATION, and others.

The potential predicted response to immune checkpoint blockade of 10-gene chemotherapy-benefit signature

A study explored the roles of fatty acid metabolism in CRC and developed a prognostic risk score model. The study utilized the Tumor Immune Dysfunction and Exclusion (TIDE) score alongside a SubMap algorithm to demonstrate that high-risk patients, as identified by the model, are more suitable for immunotherapy. This suggests that the TIDE score can help identify patients with CRC who might benefit from specific immunotherapies.42 Given our findings relative to the patients with MSI-H and the 10-gene signature, we investigated whether the 10-gene chemotherapy-benefit signature could also predict response to immune checkpoint blockade (ICB). In our study, we assessed the relationship between the TIDE score43 and the 10-gene chemotherapy-benefit score for each sample in our datasets. We then examined the Pearson correlation between these two scores. Our analysis, involving 442 tumor samples across four datasets (GSE17538, GSE33113, GSE37892, and GSE38832), revealed a significant yet moderate positive correlation between the TIDE score and the 10-gene chemotherapy-benefit score (r = 0.33, p = 4.97 × 10−13, Figure 6A). However, this correlation varied across datasets. Specifically, in the GSE37892 dataset, the correlation was relatively weak and not statistically significant (r = 0.13, p = 0.13, Figure 6D), while in the GSE17538 (r = 0.42, p = 1.05 × 10−6, Figure 6B), GSE33113 (r = 0.42, p = 3.77 × 10−5, Figure 6C), and GSE38832 (r = 0.53, p = 9.2 × 10−7, Figure 6E) datasets, a moderate and significant positive correlation was observed. Our results, therefore, underscore the need for further investigation, particularly in clinical settings where immunotherapy is administered, to better understand and validate the potential predictive role of our 10-gene signature in the context of ICB.Figure 6 Association of the 10-gene chemotherapy-benefit signature with TIDE score

The Pearson correlation between the TIDE score and the 10-gene chemotherapy-benefit score in the merged four datasets (A), GSE17538 (B), GSE33113 (C), GSE37892 (D), and GSE38832 (E).

Discussion

In our pursuit of personalized cancer therapy, identifying patients with stage II and III colon cancer most likely to benefit from 5-FU-based ACT remains a crucial goal. However, developing gene signatures related to prognosis or ACT benefit poses challenges due to limited training datasets, variability in gene expression data across different platforms, and heterogeneity in patient stages, treatments, and tumor molecular characteristics. To overcome these challenges, we employed a semi-supervised machine learning approach, identifying genes significantly associated with survival and highly interconnected within gene networks. Utilizing a discovery dataset of 933 stage II and III colon cancer samples, the largest of its kind, we aimed to maximize the robustness of our model.

Gene regulatory network modeling, which studies how genes interact with each other and work together as modules to perform biological functions, has been extensively applied to genomic data analysis and shown to be an effective approach for understanding complex biological systems.44,45,46,47,48 It also has been demonstrated that compared with the genes selected by meta-analysis p values, highly connected hub genes in the network have more biological functional importance and are more predictive of outcome.27 The hub genes have been used as effective biomarkers for the prognosis and prediction of survival in different cancer types.49,50,51,52,53,54 To select relevant genes for the construction of a gene network, we first applied the Cox regression model to select 3,091 genes significantly associated with patient survival outcomes in a subset of the discovery cohort with available survival information (n = 566) and then constructed a gene network using the survival-associated genes with all 933 samples in the discovery cohort.

The 18-hub-gene selected from our semi-supervised gene network, in which each gene has more than 11 connections with its neighbors in the gene co-expression network, led to coordinated changes at the system level. We found that compared to the 18-top-ranked gene set, our 18-hub-gene set has higher information content as well as more robust prognostic performances across different datasets. Further genomic and DNA methylation analysis also developed a 10-gene chemotherapy-benefit signature, which was successfully validated in another independent fresh-frozen tissues dataset we generated previously (GSE17538). Moreover, we demonstrated both 18- and 10-gene signatures to have independent prognostic and predictive value beyond the clinical variables including tumor stage. These results suggested that the gene signatures we developed were not just prognostic for patients with colon cancer but, more importantly, were also predictive of patients’ benefit from 5-FU-based chemotherapy in terms of RFS as in previous work.55,56

The pathway analyses of the 10-ACT benefit gene signature and its network revealed that for the ACT-non-benefit group, a series of the tumor microenvironment and cell surface interaction-related pathways such as integrin 5 and integrin 2 pathways are activated, which might provide a survival advantage for tumor cells against chemotherapy treatment.57,58,59 It is also well known that increased TGFβ signaling in the microenvironment is associated with adverse outcomes in patients with CRC.60,61 On the other hand, the ACT-benefit group is enriched with cell cycle-related pathways. Thus, the 10-ACT benefit gene signature captures the distinct biological functions and activities associated with chemotherapy responses.

ICB is an effective strategy for enhancing anti-tumor T cell activity and therapeutic effects. ICB therapy has resulted in unprecedented and durable responses in a sizable proportion of patients with diverse cancers such as melanoma, lung, kidney, bladder, and others.62,63,64 ICB drugs such as pembrolizumab and nivolumab, which are inhibitors of programmed cell death protein 1 (PD-1), have been authorized by the Food and Drug Administration to treat patients with metastatic CRC whose tumors are MSI-H or deficient mismatch repair (dMMR). Recently PD-1 blockade was applied to patients with locally advanced rectal cancer with dMMR in a neoadjuvant setting, and all 12 patients achieved a striking clinical complete response.65 However, MSI-H/dMMR CRC only comprises around 15% of CRC, and patients with microsatellite stable (MSS)/mismatch repair-proficient (pMMR) CRC show poor responses to ICB treatments due to “immune-cold” tumor immune microenvironment.66,67,68,69,70

Although overall patients with MSS/pMMR CRC lack strong responses to ICB therapy, the heterogeneity nature of CRC indicates that the different subgroups may have divergent treatment responses.71,72,73 A recent large-scale transcriptome analysis identified a subgroup of “immune-hot” MSS/pMMR rectal cancer with favorable outcomes to neoadjuvant therapy, which highlights the potential for effective response to ICB therapy.74 On the other hand, there is a small yet notable fraction of patients with MSI-H/dMMR metastatic CRC who unfortunately show early disease progression after undergoing ICB treatment.75,76 Therefore, there is an urgent need to identify potential biomarkers for ICB response.

TIDE score has been shown to capture tumor immune evasion by combining gene signatures estimating T cell dysfunction and exclusion to predict ICB response.43 However, TIDE score calculation may need a large number of genes and is focused on the state of T cells. We observed a significant, although moderate, correlation (r = 0.33, p = 4.97E−13, Figure 6A) between the 10-gene signature and TIDE score using 442 colon tumors. While this finding is intriguing, we acknowledge that it does not strongly indicate a predictive capability for checkpoint inhibition response. This association, therefore, should be interpreted with caution, and we do not assert it as a definitive predictor of immunotherapy benefit. Further investigation in populations receiving immunotherapy is required to explore the potential relevance of this correlation.

One of the key bottlenecks for translating biomarkers discovered from the microarray/RNA sequencing transcriptomic data based on fresh-frozen tissue to routine clinical practice is the development of accurate and robust clinical tests, especially on FFPE samples. Previously, we systematically evaluated the gene expression pattern of colon cancer prognostic gene signature between Affymetrix microarray data on fresh-frozen tissue and NanoString nCounter data on matched FFPE samples and showed the moderate correlation between these two measurements, which indicated the feasibility of translating the discovery from fresh-frozen tissues to clinical tests using FFPE samples for colon cancer gene signatures.77,78 In this study, we developed the clinical assay containing 18 genes using NanoString nCounter platform and tested the 18- and 10-gene signatures in VUMC stage II to III colon cancer samples that were sectioned from FFPE tissue blocks. The 18-gene prognosis signature performed exceptionally well in the FFPE validation cohort (log rank p value = 0.009). Despite the relatively small sample size of the chemotherapy-benefit validation cohort (56 patients with 13 recurrent events), the 10-gene chemotherapy-benefit signature also achieved marginal significance (log rank p value = 0.0583), indicating a promising trend that distinguishes 5-FU-based ACT-beneficial and non-beneficial groups in terms of survival outcomes. These results demonstrated that our 10-gene signature set could provide useful information for identifying patients with colon cancer who are most likely to benefit from 5-FU-based ACT.

The development of our custom NanoString panel represents a significant advancement in translating our gene signature from the research setting into clinical practice. This innovative approach enables the practical application of our 10-gene signature, potentially changing how patients with stage II and III colon cancer could be managed in terms of 5-FU-based ACT. By providing a robust and clinically applicable tool, we anticipate that our signature could aid in stratifying patients more accurately, thereby enabling personalized treatment strategies. The panel’s utility in distinguishing patients likely to benefit from ACT could lead to more precise and effective treatments, reducing unnecessary exposure to chemotherapy and its associated side effects for those unlikely to benefit.

Furthermore, the integration of our panel into clinical workflows necessitates a series of next steps to fully realize its potential. Prospective clinical trials are imperative to validate the predictive power of the 10-gene signature in a real-world setting and assess its impact on patient outcomes. These trials should also explore the utility of the signature in conjunction with other emerging diagnostic tools, like circulating tumor DNA, to further refine treatment strategies. Additionally, the broader applicability of the signature in various stages of colon cancer and in different treatment scenarios, such as immune checkpoint blockade therapy, should be investigated. This expanded research could uncover more comprehensive applications of the signature, thereby contributing to the overarching goal of precision oncology. The economic aspects of implementing this technology in routine clinical settings also warrant thorough evaluation to ensure its cost-effectiveness and accessibility. Ultimately, these steps will be crucial in transitioning our research findings from the bench to the bedside, offering new hope for optimized cancer treatment strategies.

In our study, we also conducted a simulation analysis to assess the adequacy of our sample size for validating the 10-gene signature for prospective studies. Adhering to the methodological framework suggested by Riley et al.,79 we targeted a calibration slope with narrow confidence. Our simulations across varying sample sizes (N = 250 to 3,000) demonstrated a consistent trend of decreasing SE with increasing sample size, underscoring the enhancement in the precision of our model’s calibration (Table S9). Notably, the SE decreased from 0.259 at N = 250 to 0.075 at N = 3,000, highlighting the improved precision in larger cohorts. This analysis not only substantiates the reliability of our gene signature in smaller cohorts but also suggests increased robustness in larger ones. These findings, pivotal in contextualizing the clinical applicability of our model, support the notion that while larger sample sizes offer enhanced precision, even smaller sample sizes in our study provide valuable insights, contributing significantly to the nuanced understanding of gene signature validations in oncological research. This alignment of our results with statistical principles and clinical relevance underlines the robustness of our study’s findings and reinforces its potential impact in guiding ACT decisions in colon cancer.

Our 10-gene signature offers promise, especially with its strong prognostic value and hints at predicting the benefits of ACT. Pathway analysis of this signature has uncovered specific biological functions related to chemotherapy responses, highlighting its potential to deepen our understanding of treatment impacts. In summary, our research introduces a 10-gene signature that could play a significant role in predicting ACT benefits. Looking ahead, its value in clinical decision-making could greatly enhance how we tailor ACT and immunotherapy for patients with colon cancer. The next steps involve more extensive validation through prospective clinical trials to firmly establish its predictive power.

Limitations of the study

Our study presents a 10-gene signature for predicting ACT benefit in stage II and III colon cancer, yet it is not without limitations. The primary challenge was our inability to benchmark this signature against others due to the unique focus on chemotherapy response, a contrast to the prevalent prognostic nature of existing signatures. This limitation, combined with the absence of comparable signatures and suitable datasets, restricts direct comparative analysis. In addition, future validation of the 10-gene signature is essential to confirm its clinical utility. This should involve prospective clinical trials and diverse datasets to ensure its efficacy across various patient groups. Exploring and benchmarking against alternative computational models could enrich the understanding of its relative performance.

Moreover, it is important to note that our study was based solely on transcriptome data. While our approach provides valuable insights into gene expression profiles related to chemotherapy response, it does not capture the full spectrum of biological mechanisms that might influence treatment outcomes. The generation of multi-omics data from future clinical trials could offer a more comprehensive understanding of the biological landscape of colon cancer. Incorporating data from genomics, proteomics, metabolomics, and other omics layers will not only enhance our biological insight but also facilitate the discovery of biomarkers from different omics data, potentially leading to more robust and predictive biomarkers for chemotherapy benefit.

STAR★Methods

Key resources table

REAGENT or RESOURCE	SOURCE	IDENTIFIER	
Software and algorithms	
	
R (v3.6.3)	R CRAN	https://cran.r-project.org/	
frma R package (v1.38.0)	McCall et al. (2010)80	https://git.bioconductor.org/packages/frma	
sva R package (v3.34.0)	Leek et al. (2012)81	https://git.bioconductor.org/packages/sva	
space R package (v 0.1-1.1)	Peng et al. (2009)82	https://cran.r-project.org/src/contrib/Archive/space/	
randomForestSRC R package (v2.7.0)	Ishwaran et al. (2008)83	https://cran.r-project.org/web/packages/randomForestSRC/index.html	
clusterProfiler R package (v3.14.3)	Yu et al. (2012)84	https://git.bioconductor.org/packages/clusterProfiler	
GSVA R package (v1.34.0)	Hanzelmann et al. (2013)85	https://git.bioconductor.org/packages/GSVA	
hgu133plus2.db R package (v3.2.3)	Carlson (2016)86	https://bioconductor.org/packages/release/data/annotation/html/hgu133plus2.db.html	
survminer R package (v0.4.9)	Alboukadel et al. (2021)87	https://cran.r-project.org/web/packages/survminer/index.html	
NanoStringNorm R package (v1.2.1)	Waggott et al. (2012)88	https://cran.r-project.org/src/contrib/Archive/NanoStringNorm/	
GISTIC	Mermel et al. (2011)37	https://portals.broadinstitute.org/tcga/home	
MethHC	Huang et al. (2021)38	https://methhc.mbc.nctu.edu.tw/php/index.php	
DisGeNET	Pinero et al. (2020)39	https://www.disgenet.org/home/	
CellMiner	Reinhold et al. (2012)40	https://discover.nci.nih.gov/cellminer/	
TIDE	Jiang et al. (2018)43	https://tide.dfci.harvard.edu/	
	
Deposited data	
	
Microarray of colon cancer expression profiles	Marisa et al. (2013)89	GEO: GSE39582	
Microarray of colon cancer expression profiles	Smith et al. (2010)90	GSE17538	
Microarray of colon cancer expression profiles	Tripathi et al. (2014)91	GSE38832	
Microarray of colon cancer expression profiles	De Sousa et al. (2011)92	GSE33113	
Microarray of colon cancer expression profiles	Laibe et al. (2012)93	GSE37892	
Microarray of colon cancer expression profiles	Kikuchi et al. (2012)94	GSE27854	
Microarray of colon cancer expression profiles	Birnbaum et al. (2011)95	GSE26906	
Microarray of colon cancer expression profiles	Tsukamoto et al. (2011)96	GSE21510	
Microarray of colon cancer expression profiles	Gröne et al. (2011)97	GSE18088	
Microarray of colon cancer expression profiles	Expression Project for Oncology (http://intgen.org)	GSE2109	
The VUMC NanoString nCounter® data	This study	https://github.com/PoShine/CRC-5-FU	

Resource availability

Lead contact

Further information and requests for resources and reagents should be directed to and will be fulfilled by the lead contact, X. Steven Chen (steven.chen@miami.edu).

Materials availability

This study did not generate new unique reagents.

Data and code availability

• All public datasets used in this study are available in the Gene Expression Omnibus (https://www.ncbi.nlm.nih.gov/geo/). NanoString nCounter® data is available in the Supplementary Data S1.

• The R code for the computational refinement of gene signatures was deployed https://github.com/PoShine/CRC-5-FU. The DOI at Zenodo was https://doi.org/10.5281/zenodo.10748699.

• Any additional information required to reanalyze the data reported in this paper is available from the lead contact upon request.

Experimental model and study participant details

Patients and ethics statements

The FFPE samples of surgically resected tumor tissues were obtained from colon cancer patients undergoing surgery at Vanderbilt University Medical Center (Nashville, USA) from 1990 to 2013 after informed consent was provided. All cases were confirmed by histopathology. This research was approved by the Institutional Review Board (IRB) and Ethics Committee of Vanderbilt University Medical Center (Nashville, USA). Each patient provided written informed consent. The patient demographic information is available in the Supplementary Data S1.

Method details

Patients and tumor samples

To ensure the most comprehensive coverage of probesets in the human genome, as well as concordance in the microarray platform, currently available colon cancer expression profiles detected by "Affymetrix Human Genome U133 Plus 2.0 Array" platform (GPL570), were collected and used for analysis. All gene expression data of stage I and IV patients were removed before signature derivation/verification. As a result, a total of 1375 stage II to III colon cancer frozen tissue specimens from 10 independent gene expression datasets were collected from the Gene Expression Omnibus (GEO) microarray data repository, including GSE39582, GSE17538, GSE38832, GSE33113, GSE37892, GSE27854, GSE26906, GSE21510, GSE18088, and GSE2109 (Table 1). Among them, the first five datasets provide clinical information about relapse-free survival (RFS) time (Table 1); and GSE39582 and GSE17538 contain the samples of patients who had received 5-fluorouracil-based (5-FU) adjuvant chemotherapy (Table S8).

Data pre-processing

First, we applied the Frozen Robust Multi-array Analysis (fRMA) algorithm to each microarray dataset for data normalization.80 The probeset identifiers (IDs) were mapped to gene symbols using the R package "hgu133plus2.db". When multiple probesets correspond to one gene, the probeset with the most significant p-value associated with survival outcome was selected to represent the gene. For the discovery cohort, batch effects within GSE39582 and among six different GEO datasets (GSE39582, GSE27854, GSE26906, GSE21510, GSE18088, and GSE2109) were removed using the ComBat method as implemented in the R package 'sva'.81

Construction of survival-related gene network and identification of hub genes

Currently, for colon cancer, there are limited publicly available gene expression datasets with complete clinical annotation, such as survival outcomes. To maximally leverage information in gene expression datasets without survival information, we developed and validated our prognosis prediction models using a new semi-supervised gene-network based strategy. Within the largest stage II to III colon cancer collection, GSE39582 was the only dataset with RFS information. We first used GSE39582 to select genes significantly associated with RFS outcome based on the Cox proportional hazards model (p < 0.05). Next, we further filtered these genes using a network-based strategy. We hypothesized that hub genes, which have many interactions with other genes in the survival-related network, are more likely to be predictive of prognosis. To obtain a robust survival-related network, we combined all stage II and III colon cancer samples in GSE39582, together with the other five expression datasets: GSE27854, GSE26906, GSE21510, GSE18088, and GSE2109, all of which did not have RFS information, into an aggregated gene expression matrix. In the combined gene expression dataset, survival-associated genes were used to construct the survival-related gene network using the Sparse Partial Correlation Estimation (SPACE) algorithm,82 which was shown to be one of the best-performing gene network approaches to discover hub genes.98 Finally, we selected genes with more than 11 connections to their neighbor nodes in the network as hub genes.

Prognosis model construction and evaluation

Given the selected set of hub genes, we next used the "randomForestSRC" R package to construct the random survival forest prognosis model, using GSE39582 which has RFS time as our training samples, and validated the prediction model in four additional independent datasets with available RFS time, including GSE17538, GSE33113, GSE37892, and GSE38832. The random survival forests approach is particularly advantageous for its ability to handle complex interactions between genes and its robustness in managing high-dimensional genomic data, making it an excellent tool for survival analysis without the need for assumptions on the data’s distribution.83,99 The maximally selected rank statistic (maxstat) provided in the "survminer" package was used to determine the optimal cut-off point for discriminating high-risk vs. low-risk groups. The prediction performance of the RSF prognosis model was assessed using the hazard ratio (HR) and log rank test p-value. The detailed workflow for our proposed semi-supervised gene network-based prognosis model-building strategy is illustrated in Figure 1A. The other existing colon cancer prognosis signatures, including OncotypeDX,29 ColoGuidePro,30 and ColoFinder,31 were used to compare the 18-gene prognosis model.

Determining a robust chemotherapy-benefit signature set from adjuvant chemotherapy (ACT) samples

To additionally identify a gene signature set for stage II to III colon cancer patients who are most likely to benefit from the 5-FU-based ACT, we further refined the 18-hub-gene set and investigated its ability to predict colon cancer outcomes following 5-FU-based ACT. More specifically, we identified potential cancer driver genes with either significant genomic aberration or DNA methylation alteration between colon cancer and normal samples using the GISTIC tool37 for TCGA copy number and MethHC database38 (see Figure 1B). We also verified the 18-hub-genes in the DisGeNET database.39 An RSF chemotherapy-benefit model based on these 10 genes was then re-constructed using 142 samples with 5-FU-based ACT information in the GSE39592 dataset. Finally, we accessed the prediction performance of this chemotherapy-benefit signature using two datasets generated from our group-microarray data based on fresh frozen tissues in GSE17538 and NanoString nCounter data using FFPE samples from Vanderbilt University Medical Center (VUMC), which included expression data of colon cancer patients who received 5-FU-based ACT samples (Figure 1B).

NanoString nCounter assay development

To facilitate the clinical application of the gene signatures, we designed a custom nCounter assay (NanoString Technologies, Seattle, WA) for quantitative assessment of the expression of 25 gene elements using a multiple probe approach. The final codeset contained the 18-hub-gene set identified from the discovery cohort based on the microarray platform using fresh frozen tissue. Seven housekeeping genes derived from the least variant expression elements in pilot nCounter studies were also included. To provide a more robust measurement, three probes were designed for each of the 18 hub genes and 6 of the housekeeping genes. Only two probes were designed for RPL38 due to proximity issues between the probes. A full list of gene symbols and the probe design for these 25 genes can be found in Table S1. This multiplexed assay is reported to detect expression of up to 800 transcripts at very low mRNA concentrations (0.1fM/1 copy per cell) even in RNA samples that are significantly degraded, as long as at least 20% of the sample has RNA fragments of 300 base pairs or longer. Therefore, the 144 CRC samples with 20%–76% (median value is 43%) of the RNA sample containing fragments of 300 base pairs or longer were hybridized to the custom nCounter assay.

Evaluation of the predictive performance in the VUMC NanoString nCounter data

To analyze stage II to III colon cancer specimens using the NanoString nCounter platform, total RNAs were isolated from 144 FFPE preserved colon cancer specimens (with associated RFS prognosis information), of which 109 specimens were stage II to III colon cancer patients, 56 of which had received 5-fluorouracil-based (5-FU) adjuvant chemotherapy (Table 1). We used the R package 'NanoStringNorm' to normalize the raw expression data, summarized the CodeCount (positive) and SampleContent (housekeeping) controls by the geometric mean, and used the 'mean' option for background correction. Normalized data were log2-transformed for further analyses. Batch effects were removed using the ComBat method as implemented in the R package 'sva'.81 Finally, when multiple probesets were targeted to the same gene, the probeset with the maximum expression value was used to represent the gene. A total of 109 stage II to III samples were utilized to assess the RFS prognosis model constructed by the 18-hub-gene set, and 56 samples that received 5-fluorouracil-based (5-FU) adjuvant chemotherapy were utilized to validate the chemotherapy-benefit 10-gene prediction model. Because of the difference in microarray platforms, when the NanoString data were utilized to evaluate the efficiency of our prediction model, for each gene, we computed Z score transformations for samples in the training set and samples in the testing set separately. The NanoString data is available in the Supplementary Data S1.

Drug sensitivity analyses

Gene transcript z scores data for the chemotherapy-benefit gene signature set and their corresponding -log10(GI 50) values treated by 5-FU in NCI-60 cell lines were obtained from CellMiner database (https://discover.nci.nih.gov/cellminer/).40 GI 50 represents the concentration of a drug required to reduce cancer cell proliferation by 50%. Using Z score data from the chemotherapy-benefit gene signature set, complete-linkage hierarchical clustering was performed to categorize all cell lines into resistant and sensitive groups. The Wilcoxon rank-sum test was then carried out to evaluate the difference between resistant and sensitive cell lines by their median -log10(GI 50) values.

Pathway analysis

The five pathway collections including BioCarta, KEGG, PID, Reactome, and WikiPathways from MSigDB C2 Canonical pathway database were used for pathway analysis. We used the enricher function in the Bioconductor package clusterProfiler for pathway enrichment analysis.84 The ssGSEA method was implemented in the GSVA R package and was applied to compare ACT benefit and non-benefit groups and heatmap visualization for samples from the GSE17538 dataset.85

Tumor immune dysfunction and exclusion (TIDE) analysis

The TIDE score computed for each tumor sample can serve as a surrogate biomarker to predict response to immune checkpoint blockade, including anti-PD1 and anti-CTLA.43 We utilized the TIDE online tool (http://tide.dfci.harvard.edu/) to calculate the TIDE score for each tumor sample in four datasets: GSE17538, GSE33113, GSE37892, and GSE38832. Then, we calculated a 10-gene chemotherapy-benefit score for each sample in the four datasets, irrespective of whether the sample had received 5FU chemotherapy. The TIDE scores and the 10-gene chemotherapy-benefit scores were further compared.

Quantification and statistical analysis

The two-sample Kolmogorov-Smirnov test or Wilcoxon rank-sum test was appropriately used to compare groups. The Kaplan–Meier method was used to analyze the RFS of patients with different risk stratification. All analyses were performed using R software (version 3.6.3).

Supplemental information

Document S1. Figures S1‒S4 and Tables S1–S7 and S9

Table S8. The detailed 5-FU-based ACT chemotherapy regimen of the cohort GSE39582, GSE17538, and VUMC, related to Figure 4

Data S1. The NanoString nCounter data and clinical information of the VUMC validation cohort, related to Figure 4

Document S2. Article plus supplemental information

Acknowledgments

This work was supported by 10.13039/100018896 Sylvester Comprehensive Cancer Center (10.13039/100000054 NCI /10.13039/100000002 NIH BOOST-2023-02 ) funding. A select number of illustrations were drawn using BioRender. This paper is dedicated to Dr. R. Daniel Beauchamp in memory of his contributions to colorectal cancer research.

Author contributions

X.S.C., C.X., and R.D.B. conceived and designed the study. C.X. led the data analysis. P.X. and J.L. contributed to the bioinformatics and statistical analysis. K.B.L. conducted experiments for NanoString validation. K.K.C., L.W., and J.J.S. contributed to the interpretation of the results. C.X. and X.S.C. drafted the paper. X.S.C., J.J.S., and R.D.B. supervised the study. All of the authors have read and approved the final version before publication.

Declaration of interests

The authors declare no competing interests.

Supplemental information can be found online at https://doi.org/10.1016/j.xcrm.2024.101661.
==== Refs
References

1 Dekker E. Tanis P.J. Vleugels J.L.A. Kasi P.M. Wallace M.B. Colorectal cancer Lancet 394 2019 1467 1480 10.1016/S0140-6736(19)32319-0 31631858
2 O'Connell J.B. Maggard M.A. Ko C.Y. Colon cancer survival rates with the new American Joint Committee on Cancer sixth edition staging J. Natl. Cancer Inst. 96 2004 1420 1425 10.1093/jnci/djh275 15467030
3 White A. Joseph D. Rim S.H. Johnson C.J. Coleman M.P. Allemani C. Colon cancer survival in the United States by race and stage (2001-2009): Findings from the CONCORD-2 study Cancer 123 2017 5014 5036 10.1002/cncr.31076 29205304
4 Meyerhardt J.A. Shi Q. Fuchs C.S. Meyer J. Niedzwiecki D. Zemla T. Kumthekar P. Guthrie K.A. Couture F. Kuebler P. Effect of Celecoxib vs Placebo Added to Standard Adjuvant Therapy on Disease-Free Survival Among Patients With Stage III Colon Cancer: The CALGB/SWOG 80702 (Alliance) Randomized Clinical Trial JAMA 325 2021 1277 1286 10.1001/jama.2021.2454 33821899
5 Iveson T.J. Sobrero A.F. Yoshino T. Souglakos I. Ou F.S. Meyers J.P. Shi Q. Grothey A. Saunders M.P. Labianca R. Duration of Adjuvant Doublet Chemotherapy (3 or 6 months) in Patients With High-Risk Stage II Colorectal Cancer J. Clin. Oncol. 39 2021 631 641 10.1200/JCO.20.01330 33439695
6 Gramont A. Adjuvant therapy of stage II and III colon cancer Semin. Oncol. 32 2005 11 14 10.1053/j.seminoncol.2005.06.004
7 Tsikitis V.L. Larson D.W. Huebner M. Lohse C.M. Thompson P.A. Predictors of recurrence free survival for patients with stage II and III colon cancer BMC Cancer 14 2014 336 10.1186/1471-2407-14-336 24886281
8 de Gramont A. Tournigand C. Andre T. Larsen A.K. Louvet C. Adjuvant therapy for stage II and III colorectal cancer Semin. Oncol. 34 2007 S37 S40 10.1053/j.seminoncol.2007.01.004 17449351
9 Benson A.B. Venook A.P. Al-Hawary M.M. Arain M.A. Chen Y.J. Ciombor K.K. Cohen S. Cooper H.S. Deming D. Farkas L. Colon Cancer, Version 2.2021, NCCN Clinical Practice Guidelines in Oncology J. Natl. Compr. Canc. Netw. 19 2021 329 359 10.6004/jnccn.2021.0012 33724754
10 Fang S.H. Efron J.E. Berho M.E. Wexner S.D. Dilemma of stage II colon cancer and decision making for adjuvant chemotherapy J. Am. Coll. Surg. 219 2014 1056 1069 10.1016/j.jamcollsurg.2014.09.010 25440029
11 Kannarkatt J. Joseph J. Kurniali P.C. Al-Janadi A. Hrinczenko B. Adjuvant Chemotherapy for Stage II Colon Cancer: A Clinical Dilemma J. Oncol. Pract. 13 2017 233 241 10.1200/JOP.2016.017210 28399381
12 Auclin E. Zaanan A. Vernerey D. Douard R. Gallois C. Laurent-Puig P. Bonnetain F. Taieb J. Subgroups and prognostication in stage III colon cancer: future perspectives for adjuvant therapy Ann. Oncol. 28 2017 958 968 10.1093/annonc/mdx030 28453690
13 Sargent D.J. Marsoni S. Monges G. Thibodeau S.N. Labianca R. Hamilton S.R. French A.J. Kabat B. Foster N.R. Torri V. Defective mismatch repair as a predictive marker for lack of efficacy of fluorouracil-based adjuvant therapy in colon cancer J. Clin. Oncol. 28 2010 3219 3226 10.1200/JCO.2009.27.1825 20498393
14 Dienstmann R. Salazar R. Tabernero J. Personalizing colon cancer adjuvant therapy: selecting optimal treatments for individual patients J. Clin. Oncol. 33 2015 1787 1796 10.1200/JCO.2014.60.0213 25918287
15 Goldstein J. Tran B. Ensor J. Gibbs P. Wong H.L. Wong S.F. Vilar E. Tie J. Broaddus R. Kopetz S. Multicenter retrospective analysis of metastatic colorectal cancer (CRC) with high-level microsatellite instability (MSI-H) Ann. Oncol. 25 2014 1032 1038 10.1093/annonc/mdu100 24585723
16 Bertagnolli M.M. Redston M. Compton C.C. Niedzwiecki D. Mayer R.J. Goldberg R.M. Colacchio T.A. Saltz L.B. Warren R.S. Microsatellite instability and loss of heterozygosity at chromosomal location 18q: prospective evaluation of biomarkers for stages II and III colon cancer--a study of CALGB 9581 and 89803 J. Clin. Oncol. 29 2011 3153 3162 10.1200/JCO.2010.33.0092 21747089
17 O'Connell M.J. Lavery I. Yothers G. Paik S. Clark-Langone K.M. Lopatin M. Watson D. Baehner F.L. Shak S. Baker J. Relationship between tumor gene expression and recurrence in four independent studies of patients with stage II/III colon cancer treated with surgery alone or surgery plus adjuvant fluorouracil plus leucovorin J. Clin. Oncol. 28 2010 3937 3944 10.1200/JCO.2010.28.9538 20679606
18 Salazar R. Roepman P. Capella G. Moreno V. Simon I. Dreezen C. Lopez-Doriga A. Santos C. Marijnen C. Westerga J. Gene expression signature to improve prognosis prediction of stage II and III colorectal cancer J. Clin. Oncol. 29 2011 17 24 10.1200/JCO.2010.30.1077 21098318
19 Jiang Y. Casey G. Lavery I.C. Zhang Y. Talantov D. Martin-McGreevy M. Skacel M. Manilich E. Mazumder A. Atkins D. Development of a clinically feasible molecular assay to predict recurrence of stage II colon cancer J. Mol. Diagn. 10 2008 346 354 10.2353/jmoldx.2008.080011 18556775
20 Kennedy R.D. Bylesjo M. Kerr P. Davison T. Black J.M. Kay E.W. Holt R.J. Proutski V. Ahdesmaki M. Farztdinov V. Development and independent validation of a prognostic assay for stage II colon cancer using formalin-fixed paraffin-embedded tissue J. Clin. Oncol. 29 2011 4620 4626 10.1200/JCO.2011.35.4498 22067406
21 Park Y.Y. Lee S.S. Lim J.Y. Kim S.C. Kim S.B. Sohn B.H. Chu I.S. Oh S.C. Park E.S. Jeong W. Comparison of prognostic genomic predictors in colorectal cancer PLoS One 8 2013 e60778 10.1371/journal.pone.0060778 23626670
22 Di Narzo A.F. Tejpar S. Rossi S. Yan P. Popovici V. Wirapati P. Budinska E. Xie T. Estrella H. Pavlicek A. Test of four colon cancer risk-scores in formalin fixed paraffin embedded microarray gene expression data J. Natl. Cancer Inst. 106 2014 dju247 10.1093/jnci/dju247 25246611
23 Gao S. Tibiche C. Zou J. Zaman N. Trifiro M. O'Connor-McCourt M. Wang E. Identification and Construction of Combinatory Cancer Hallmark-Based Gene Signature Sets to Predict Recurrence and Chemotherapy Benefit in Stage II Colorectal Cancer JAMA Oncol. 2 2016 37 45 10.1001/jamaoncol.2015.3413 26502222
24 Zhao W. Chen B. Guo X. Wang R. Chang Z. Dong Y. Song K. Wang W. Qi L. Gu Y. A rank-based transcriptional signature for predicting relapse risk of stage II colorectal cancer identified with proper data sources Oncotarget 7 2016 19060 19071 10.18632/oncotarget.7956 26967049
25 Tong M. Zheng W. Li H. Li X. Ao L. Shen Y. Liang Q. Li J. Hong G. Yan H. Multi-omics landscapes of colorectal cancer subtypes discriminated by an individualized prognostic signature for 5-fluorouracil-based chemotherapy Oncogenesis 5 2016 e242 10.1038/oncsis.2016.51 27429074
26 van Dam S. Vosa U. van der Graaf A. Franke L. de Magalhaes J.P. Gene co-expression analysis for functional classification and gene-disease predictions Brief. Bioinform. 19 2018 575 592 10.1093/bib/bbw139 28077403
27 Langfelder P. Mischel P.S. Horvath S. When is hub gene selection better than standard meta-analysis? PLoS One 8 2013 e61505 10.1371/journal.pone.0061505 23613865
28 Delgado J. Pereira A. Villamor N. Lopez-Guillermo A. Rozman C. Survival analysis in hematologic malignancies: recommendations for clinicians Haematologica 99 2014 1410 1420 10.3324/haematol.2013.100784 25176982
29 Gray R.G. Quirke P. Handley K. Lopatin M. Magill L. Baehner F.L. Beaumont C. Clark-Langone K.M. Yoshizawa C.N. Lee M. Validation study of a quantitative multigene reverse transcriptase-polymerase chain reaction assay for assessment of recurrence risk in patients with stage II colon cancer J. Clin. Oncol. 29 2011 4611 4619 10.1200/JCO.2010.32.8732 22067390
30 Sveen A. Agesen T.H. Nesbakken A. Meling G.I. Rognum T.O. Liestol K. Skotheim R.I. Lothe R.A. ColoGuidePro: a prognostic 7-gene expression signature for stage III colorectal cancer patients Clin. Cancer Res. 18 2012 6001 6010 10.1158/1078-0432.CCR-11-3302 22991413
31 Shi M. He J. ColoFinder: a prognostic 9-gene signature improves prognosis for 871 stage II and III colorectal cancer patients PeerJ 4 2016 e1804 10.7717/peerj.1804 26989635
32 Peng J. Wang Z. Chen W. Ding Y. Wang H. Huang H. Huang W. Cai S. Integration of genetic signature and TNM staging system for predicting the relapse of locally advanced colorectal cancer Int. J. Colorectal Dis. 25 2010 1277 1285 10.1007/s00384-010-1043-1 20706727
33 Tian X. Zhu X. Yan T. Yu C. Shen C. Hu Y. Hong J. Chen H. Fang J.Y. Recurrence-associated gene signature optimizes recurrence-free survival prediction of colorectal cancer Mol. Oncol. 11 2017 1544 1560 10.1002/1878-0261.12117 28796930
34 Cheng X. Hu M. Chen C. Hou D. Computational analysis of mRNA expression profiles identifies a novel triple-biomarker model as prognostic predictor of stage II and III colorectal adenocarcinoma patients Cancer Manag. Res. 10 2018 2945 2952 10.2147/CMAR.S170502 30214289
35 Rokavec M. Ozcan E. Neumann J. Hermeking H. Development and Validation of a 15-gene Expression Signature with Superior Prognostic Ability in Stage II Colorectal Cancer Cancer Res. Commun. 3 2023 1689 1700 10.1158/2767-9764.CRC-22-0489 37654625
36 Fang Z. Xu S. Xie Y. Yan W. Identification of a prognostic gene signature of colon cancer using integrated bioinformatics analysis World J. Surg. Oncol. 19 2021 13 10.1186/s12957-020-02116-y 33441161
37 Mermel C.H. Schumacher S.E. Hill B. Meyerson M.L. Beroukhim R. Getz G. GISTIC2.0 facilitates sensitive and confident localization of the targets of focal somatic copy-number alteration in human cancers Genome Biol. 12 2011 R41 10.1186/gb-2011-12-4-r41 21527027
38 Huang H.Y. Li J. Tang Y. Huang Y.X. Chen Y.G. Xie Y.Y. Zhou Z.Y. Chen X.Y. Ding S.Y. Luo M.F. MethHC 2.0: information repository of DNA methylation and gene expression in human cancer Nucleic Acids Res. 49 2021 D1268 D1275 10.1093/nar/gkaa1104 33270889
39 Pinero J. Ramirez-Anguita J.M. Sauch-Pitarch J. Ronzano F. Centeno E. Sanz F. Furlong L.I. The DisGeNET knowledge platform for disease genomics: 2019 update Nucleic Acids Res. 48 2020 D845 D855 10.1093/nar/gkz1021 31680165
40 Reinhold W.C. Sunshine M. Liu H. Varma S. Kohn K.W. Morris J. Doroshow J. Pommier Y. CellMiner: a web-based suite of genomic and pharmacologic tools to explore transcript and drug patterns in the NCI-60 cell line set Cancer Res. 72 2012 3499 3511 10.1158/0008-5472.CAN-12-1370 22802077
41 Barbie D.A. Tamayo P. Boehm J.S. Kim S.Y. Moody S.E. Dunn I.F. Schinzel A.C. Sandy P. Meylan E. Scholl C. Systematic RNA interference reveals that oncogenic KRAS-driven cancers require TBK1 Nature 462 2009 108 112 10.1038/nature08460 19847166
42 Ding C. Shan Z. Li M. Chen H. Li X. Jin Z. Characterization of the fatty acid metabolism in colorectal cancer to guide clinical therapy Mol. Ther. Oncolytics 20 2021 532 544 10.1016/j.omto.2021.02.010 33738339
43 Jiang P. Gu S. Pan D. Fu J. Sahu A. Hu X. Li Z. Traugh N. Bu X. Li B. Signatures of T cell dysfunction and exclusion predict cancer immunotherapy response Nat. Med. 24 2018 1550 1558 10.1038/s41591-018-0136-1 30127393
44 Segal E. Shapira M. Regev A. Pe'er D. Botstein D. Koller D. Friedman N. Module networks: identifying regulatory modules and their condition-specific regulators from gene expression data Nat. Genet. 34 2003 166 176 10.1038/ng1165 12740579
45 Margolin A.A. Nemenman I. Basso K. Wiggins C. Stolovitzky G. Dalla Favera R. Califano A. ARACNE: an algorithm for the reconstruction of gene regulatory networks in a mammalian cellular context BMC Bioinf. 7 2006 S7 10.1186/1471-2105-7-S1-S7
46 Zhang B. Horvath S. A general framework for weighted gene co-expression network analysis Stat. Appl. Genet. Mol. Biol. 4 2005 Article17 10.2202/1544-6115.1128 16646834
47 Schafer J. Strimmer K. An empirical Bayes approach to inferring large-scale gene association networks Bioinformatics 21 2005 754 764 10.1093/bioinformatics/bti062 15479708
48 Huynh-Thu V.A. Irrthum A. Wehenkel L. Geurts P. Inferring regulatory networks from expression data using tree-based methods PLoS One 5 2010 e12776 10.1371/journal.pone.0012776 20927193
49 Tang H. Xiao G. Behrens C. Schiller J. Allen J. Chow C.W. Suraokar M. Corvalan A. Mao J. White M.A. A 12-gene set predicts survival benefits from adjuvant chemotherapy in non-small cell lung cancer patients Clin. Cancer Res. 19 2013 1577 1586 10.1158/1078-0432.CCR-12-2321 23357979
50 Cai Y. Mei J. Xiao Z. Xu B. Jiang X. Zhang Y. Zhu Y. Identification of five hub genes as monitoring biomarkers for breast cancer metastasis in silico Hereditas 156 2019 20 10.1186/s41065-019-0096-6 31285741
51 Li S. Liu X. Liu T. Meng X. Yin X. Fang C. Huang D. Cao Y. Weng H. Zeng X. Wang X. Identification of Biomarkers Correlated with the TNM Staging and Overall Survival of Patients with Bladder Cancer Front. Physiol. 8 2017 947 10.3389/fphys.2017.00947 29234286
52 Li T. Gao X. Han L. Yu J. Li H. Identification of hub genes with prognostic values in gastric cancer by bioinformatics analysis World J. Surg. Oncol. 16 2018 114 10.1186/s12957-018-1409-3 29921304
53 Yang Y. Han L. Yuan Y. Li J. Hei N. Liang H. Gene co-expression network analysis reveals common system-level properties of prognostic genes across cancer types Nat. Commun. 5 2014 3231 10.1038/ncomms4231 24488081
54 Zhou Z. Cheng Y. Jiang Y. Liu S. Zhang M. Liu J. Zhao Q. Ten hub genes associated with progression and prognosis of pancreatic carcinoma identified by co-expression analysis Int. J. Biol. Sci. 14 2018 124 136 10.7150/ijbs.22619 29483831
55 Iacopetta B. Kawakami K. Watanabe T. Predicting clinical outcome of 5-fluorouracil-based chemotherapy for colon cancer patients: is the CpG island methylator phenotype the 5-fluorouracil-responsive subgroup? Int. J. Clin. Oncol. 13 2008 498 503 10.1007/s10147-008-0854-3 19093176
56 Shigeta K. Ishii Y. Hasegawa H. Okabayashi K. Kitagawa Y. Evaluation of 5-fluorouracil metabolic enzymes as predictors of response to adjuvant chemotherapy outcomes in patients with stage II/III colorectal cancer: a decision-curve analysis World J. Surg. 38 2014 3248 3256 10.1007/s00268-014-2738-1 25167895
57 Desgrosellier J.S. Cheresh D.A. Integrins in cancer: biological implications and therapeutic opportunities Nat. Rev. Cancer 10 2010 9 22 10.1038/nrc2748 20029421
58 Aoudjit F. Vuori K. Integrin signaling in cancer cell survival and chemoresistance Chemother. Res. Pract. 2012 2012 283181 10.1155/2012/283181 22567280
59 Bergonzini C. Kroese K. Zweemer A.J.M. Danen E.H.J. Targeting Integrins for Cancer Therapy - Disappointments and Opportunities Front. Cell Dev. Biol. 10 2022 863850 10.3389/fcell.2022.863850 35356286
60 Tauriello D.V.F. Palomo-Ponce S. Stork D. Berenguer-Llergo A. Badia-Ramentol J. Iglesias M. Sevillano M. Ibiza S. Canellas A. Hernando-Momblona X. TGFβ drives immune evasion in genetically reconstituted colon cancer metastasis Nature 554 2018 538 543 10.1038/nature25492 29443964
61 Calon A. Lonardo E. Berenguer-Llergo A. Espinet E. Hernando-Momblona X. Iglesias M. Sevillano M. Palomo-Ponce S. Tauriello D.V. Byrom D. Stromal gene expression defines poor-prognosis subtypes in colorectal cancer Nat. Genet. 47 2015 320 329 10.1038/ng.3225 25706628
62 Hargadon K.M. Johnson C.E. Williams C.J. Immune checkpoint blockade therapy for cancer: An overview of FDA-approved immune checkpoint inhibitors Int. Immunopharmacol. 62 2018 29 39 10.1016/j.intimp.2018.06.001 29990692
63 Wei S.C. Duffy C.R. Allison J.P. Fundamental Mechanisms of Immune Checkpoint Blockade Therapy Cancer Discov. 8 2018 1069 1086 10.1158/2159-8290.CD-18-0367 30115704
64 Postow M.A. Callahan M.K. Wolchok J.D. Immune Checkpoint Blockade in Cancer Therapy J. Clin. Oncol. 33 2015 1974 1982 10.1200/JCO.2014.59.4358 25605845
65 Cercek A. Lumish M. Sinopoli J. Weiss J. Shia J. Lamendola-Essel M. El Dika I.H. Segal N. Shcherba M. Sugarman R. PD-1 Blockade in Mismatch Repair-Deficient, Locally Advanced Rectal Cancer N. Engl. J. Med. 386 2022 2363 2376 10.1056/NEJMoa2201445 35660797
66 Andre T. de Gramont A. Vernerey D. Chibaudel B. Bonnetain F. Tijeras-Raballand A. Scriva A. Hickish T. Tabernero J. Van Laethem J.L. Adjuvant Fluorouracil, Leucovorin, and Oxaliplatin in Stage II to III Colon Cancer: Updated 10-Year Survival and Outcomes According to BRAF Mutation and Mismatch Repair Status of the MOSAIC Study J. Clin. Oncol. 33 2015 4176 4187 10.1200/JCO.2015.63.4238 26527776
67 Marisa L. Svrcek M. Collura A. Becht E. Cervera P. Wanherdrick K. Buhard O. Goloudina A. Jonchere V. Selves J. The Balance Between Cytotoxic T-cell Lymphocytes and Immune Checkpoint Expression in the Prognosis of Colon Tumors J. Natl. Cancer Inst. 110 2018 10.1093/jnci/djx136
68 Miller K.D. Nogueira L. Devasia T. Mariotto A.B. Yabroff K.R. Jemal A. Kramer J. Siegel R.L. Cancer treatment and survivorship statistics, 2022 CA. Cancer J. Clin. 72 2022 409 436 10.3322/caac.21731 35736631
69 Lee V. Murphy A. Le D.T. Diaz L.A. Jr. Mismatch Repair Deficiency and Response to Immune Checkpoint Blockade Oncol. 21 2016 1200 1211 10.1634/theoncologist.2016-0046
70 Grasso C.S. Giannakis M. Wells D.K. Hamada T. Mu X.J. Quist M. Nowak J.A. Nishihara R. Qian Z.R. Inamura K. Genetic Mechanisms of Immune Evasion in Colorectal Cancer Cancer Discov. 8 2018 730 749 10.1158/2159-8290.CD-17-1327 29510987
71 Stintzing S. Wirapati P. Lenz H.J. Neureiter D. Fischer von Weikersthal L. Decker T. Kiani A. Kaiser F. Al-Batran S. Heintges T. Consensus molecular subgroups (CMS) of colorectal cancer (CRC) and first-line efficacy of FOLFIRI plus cetuximab or bevacizumab in the FIRE3 (AIO KRK-0306) trial Ann. Oncol. 30 2019 1796 1803 10.1093/annonc/mdz387 31868905
72 Buikhuisen J.Y. Torang A. Medema J.P. Exploring and modelling colon cancer inter-tumour heterogeneity: opportunities and challenges Oncogenesis 9 2020 66 10.1038/s41389-020-00250-6 32647253
73 Marisa L. Blum Y. Taieb J. Ayadi M. Pilati C. Le Malicot K. Lepage C. Salazar R. Aust D. Duval A. Intratumor CMS Heterogeneity Impacts Patient Prognosis in Localized Colon Cancer Clin. Cancer Res. 27 2021 4768 4780 10.1158/1078-0432.CCR-21-0529 34168047
74 Chatila W.K. Kim J.K. Walch H. Marco M.R. Chen C.T. Wu F. Omer D.M. Khalil D.N. Ganesh K. Qu X. Genomic and transcriptomic determinants of response to neoadjuvant therapy in rectal cancer Nat. Med. 28 2022 1646 1655 10.1038/s41591-022-01930-z 35970919
75 Le D.T. Kim T.W. Van Cutsem E. Geva R. Jager D. Hara H. Burge M. O'Neil B. Kavan P. Yoshino T. Phase II Open-Label Study of Pembrolizumab in Treatment-Refractory, Microsatellite Instability-High/Mismatch Repair-Deficient Metastatic Colorectal Cancer: KEYNOTE-164 J. Clin. Oncol. 38 2020 11 19 10.1200/JCO.19.02107 31725351
76 Overman M.J. McDermott R. Leach J.L. Lonardi S. Lenz H.J. Morse M.A. Desai J. Hill A. Axelson M. Moss R.A. Nivolumab in patients with metastatic DNA mismatch repair-deficient or microsatellite instability-high colorectal cancer (CheckMate 142): an open-label, multicentre, phase 2 study Lancet Oncol. 18 2017 1182 1191 10.1016/S1470-2045(17)30422-9 28734759
77 Chen X. Deane N.G. Lewis K.B. Li J. Zhu J. Washington M.K. Beauchamp R.D. Comparison of Nanostring nCounter® Data on FFPE Colon Cancer Samples and Affymetrix Microarray Data on Matched Frozen Tissues PLoS One 11 2016 e0153784 10.1371/journal.pone.0153784 27176004
78 Zhu J. Deane N.G. Lewis K.B. Padmanabhan C. Washington M.K. Ciombor K.K. Timmers C. Goldberg R.M. Beauchamp R.D. Chen X. Evaluation of frozen tissue-derived prognostic gene expression signatures in FFPE colorectal cancer samples Sci. Rep. 6 2016 33273 10.1038/srep33273 27623752
79 Riley R.D. Collins G.S. Ensor J. Archer L. Booth S. Mozumder S.I. Rutherford M.J. van Smeden M. Lambert P.C. Snell K.I.E. Minimum sample size calculations for external validation of a clinical prediction model with a time-to-event outcome Stat. Med. 41 2022 1280 1295 10.1002/sim.9275 34915593
80 McCall M.N. Bolstad B.M. Irizarry R.A. Frozen robust multiarray analysis (fRMA) Biostatistics 11 2010 242 253 10.1093/biostatistics/kxp059 20097884
81 Leek J.T. Johnson W.E. Parker H.S. Jaffe A.E. Storey J.D. The sva package for removing batch effects and other unwanted variation in high-throughput experiments Bioinformatics 28 2012 882 883 10.1093/bioinformatics/bts034 22257669
82 Peng J. Wang P. Zhou N. Zhu J. Partial Correlation Estimation by Joint Sparse Regression Models J. Am. Stat. Assoc. 104 2009 735 746 10.1198/jasa.2009.0126 19881892
83 Ishwaran H. K. U.B. Blackstone E.H. Lauer M.S. Random survival forests Ann. Appl. Stat. 2 2008 841 860
84 Yu G. Wang L.G. Han Y. He Q.Y. clusterProfiler: an R package for comparing biological themes among gene clusters OMICS 16 2012 284 287 10.1089/omi.2011.0118 22455463
85 Hanzelmann S. Castelo R. Guinney J. GSVA: gene set variation analysis for microarray and RNA-seq data BMC Bioinf. 14 2013 7 10.1186/1471-2105-14-7
86 Carlson M. hgu133plus2.db: Affymetrix Human Genome U133 Plus 2.0 Array annotation data (chip hgu133plus2) R package 2016 version 3.2.3
87 Alboukadel K. Kosinski M. Przemyslaw B. Fabian S. survminer: Drawing Survival Curves using 'ggplot2 R package 2021 version 0.4.9
88 Waggott D. Chu K. Yin S. Wouters B.G. Liu F.F. Boutros P.C. NanoStringNorm: an extensible R package for the pre-processing of NanoString mRNA and miRNA data Bioinformatics 28 2012 1546 1548 10.1093/bioinformatics/bts188 22513995
89 Marisa L. de Reynies A. Duval A. Selves J. Gaub M.P. Vescovo L. Etienne-Grimaldi M.C. Schiappa R. Guenot D. Ayadi M. Gene expression classification of colon cancer into molecular subtypes: characterization, validation, and prognostic value PLoS Med. 10 2013 e1001453 10.1371/journal.pmed.1001453 23700391
90 Smith J.J. Deane N.G. Wu F. Merchant N.B. Zhang B. Jiang A. Lu P. Johnson J.C. Schmidt C. Bailey C.E. Experimentally derived metastasis gene expression profile predicts recurrence and death in patients with colon cancer Gastroenterology 138 2010 958 968 10.1053/j.gastro.2009.11.005 19914252
91 Tripathi M.K. Deane N.G. Zhu J. An H. Mima S. Wang X. Padmanabhan S. Shi Z. Prodduturi N. Ciombor K.K. Nuclear factor of activated T-cell activity is associated with metastatic capacity in colon cancer Cancer Res. 74 2014 6947 6957 10.1158/0008-5472.CAN-14-1592 25320007
92 de Sousa E.M.F. Colak S. Buikhuisen J. Koster J. Cameron K. de Jong J.H. Tuynman J.B. Prasetyanti P.R. Fessler E. van den Bergh S.P. Methylation of cancer-stem-cell-associated Wnt target genes predicts poor prognosis in colorectal cancer patients Cell Stem Cell 9 2011 476 485 10.1016/j.stem.2011.10.008 22056143
93 Laibe S. Lagarde A. Ferrari A. Monges G. Birnbaum D. Olschwang S. Project C.O.L. A seven-gene signature aggregates a subgroup of stage II colon cancers with stage III A seven-gene signature aggregates a subgroup of stage II colon cancers with stage III 16 2012 560 565 10.1089/omi.2012.0039
94 Kikuchi A. Ishikawa T. Mogushi K. Ishiguro M. Iida S. Mizushima H. Uetake H. Tanaka H. Sugihara K. Identification of NUCKS1 as a colorectal cancer prognostic marker through integrated expression and copy number analysis Int. J. Cancer 132 2013 2295 2302 10.1002/ijc.27911 23065711
95 Birnbaum D.J. Laibe S. Ferrari A. Lagarde A. Fabre A.J. Monges G. Birnbaum D. Olschwang S. Project C.O.L. Expression Profiles in Stage II Colon Cancer According to APC Gene Status Transl. Oncol. 5 2012 72 76 10.1593/tlo.11325 22496922
96 Tsukamoto S. Ishikawa T. Iida S. Ishiguro M. Mogushi K. Mizushima H. Uetake H. Tanaka H. Sugihara K. Clinical significance of osteoprotegerin expression in human colorectal cancer Clin. Cancer Res. 17 2011 2444 2450 10.1158/1078-0432.CCR-10-2884 21270110
97 Grone J. Lenze D. Jurinovic V. Hummel M. Seidel H. Leder G. Beckmann G. Sommer A. Grutzmann R. Pilarsky C. Molecular profiles and clinical outcome of stage UICC II colon cancer patients Int. J. Colorectal Dis. 26 2011 847 858 10.1007/s00384-011-1176-x 21465190
98 Allen J.D. Xie Y. Chen M. Girard L. Xiao G. Comparing statistical methods for constructing large scale gene networks PLoS One 7 2012 e29348 10.1371/journal.pone.0029348 22272232
99 Chen X. Ishwaran H. Random forests for genomic data analysis Genomics 99 2012 323 329 10.1016/j.ygeno.2012.04.003 22546560
