==== Front Res Sq ResearchSquare Research Square American Journal Experts 37398194 10.21203/rs.3.rs-2925426/v1 10.21203/rs.3.rs-2925426 preprint 1 Article Proposal of a new genomic framework for categorization of pediatric acute myeloid leukemia associated with prognosis Umeda Masayuki 18 Ma Jing 18 http://orcid.org/0000-0002-1000-6698 Westover Tamara 1 Ni Yonghui 2 Song Guangchun 1 Maciaszek Jamie L. 1 http://orcid.org/0000-0002-5363-1848 Rusch Michael 3 Rahbarinia Delaram 3 Foy Scott 3 http://orcid.org/0000-0001-6996-0833 Huang Benjamin J. 4 Walsh Michael P. 1 Kumar Priyadarshini 1 Liu Yanling 3 Fan Yiping 5 http://orcid.org/0000-0002-1678-5864 Wu Gang 15 Baker Sharyn D. 6 http://orcid.org/0000-0002-6233-2145 Ma Xiaotu 3 Wang Lu 1 http://orcid.org/0000-0001-9885-3527 Rubnitz Jeffrey E. 7 http://orcid.org/0000-0002-9167-2114 Pounds Stanley 2 http://orcid.org/0000-0003-2961-6960 Klco Jeffery M. 19 1 Department of Pathology, St. Jude Children’s Research Hospital, Memphis, TN, USA. 2 Department of Biostatistics, St. Jude Children’s Research Hospital, Memphis, TN, USA. 3 Department of Computational Biology, St. Jude Children’s Research Hospital, Memphis, TN, USA. 4 Department of Pediatrics, University of California San Francisco, San Francisco, CA, US 5 Center for Applied Bioinformatics, St. Jude Children’s Research Hospital, Memphis, TN, USA. 6 Division of Pharmaceutics and Pharmacology, College of Pharmacy, Comprehensive Cancer Center, The Ohio State University, Columbus, OH, USA. 7 Department of Oncology, St. Jude Children’s Research Hospital, Memphis, TN, USA. 8 These authors contributed equally 9 Corresponding author: Contact information: Jeffery M. Klco: Mail Stop 342, Room D4047B, St. Jude Children’s Research Hospital, 262 Danny Thomas Place, Memphis, TN 38105-3678, Phone: (901) 595-6807, Fax: (901) 595-5947, jeffery.klco@stjude.org 29 5 2023 rs.3.rs-2925426https://creativecommons.org/licenses/by/4.0/ This work is licensed under a Creative Commons Attribution 4.0 International License, which allows reusers to distribute, remix, adapt, and build upon the material in any medium or format, so long as attribution is given to the creator. The license allows for commercial use. nihpp-rs2925426v1.pdf Recent studies on pediatric acute myeloid leukemia (pAML) have revealed pediatric-specific driver alterations, many of which are underrepresented in the current classification schemas. To comprehensively define the genomic landscape of pAML, we systematically categorized 895 pAML into 23 molecular categories that are mutually distinct from one another, including new entities such as UBTF or BCL11B, covering 91.4% of the cohort. These molecular categories were associated with unique expression profiles and mutational patterns. For instance, molecular categories characterized by specific HOXA or HOXB expression signatures showed distinct mutation patterns of RAS pathway genes, FLT3, or WT1, suggesting shared biological mechanisms. We show that molecular categories were strongly associated with clinical outcomes using two independent cohorts, leading to the establishment of a prognostic framework for pAML based on molecular categories and minimal residual disease. Together, this comprehensive diagnostic and prognostic framework forms the basis for future classification of pAML and treatment strategies. ==== Body pmcIntroduction Pediatric acute myeloid leukemia (pAML) is characterized by aberrant clonal expansion of hematopoietic progenitors with differentiation defects1–4. Although pAML shares many clinical and pathological characteristics with adult AML, genetic differences have also been appreciated5–7. Notably, t(11;x), resulting in KMT2A rearrangements, are more common in pAML, and adult AML frequently harbors mutations in DNMT3A and splicing factor genes, while core binding factor (CBF) AMLs are common across the age spectrum5. Additionally, progress in diagnostic technologies has led to the identification of cryptic fusions of NUP988 and GLIS family9 members and complex UBTF tandem duplications10 that are enriched in pAML. Recent updates in the WHO classification11 (WHO5th) and the International Consensus Classification (ICC)12 define AMLs with KMT2A and NUP98 rearrangements with various partners as distinct disease entities. However, a substantial number of more recently discovered recurrent driver alterations in pAML continue to be categorized as “acute myeloid leukemia with other defined genetic alterations” or “AML, not otherwise specified (NOS)”, confirming the need to understand both the biological features of pAMLs with these other driver alterations and their clinical significance. Accumulation of clinical outcomes associated with gene alterations enabled the risk stratification of adult AML according to detailed mutational profiling. The ELN2022 risk stratification13 incorporated various fusions, NPM1 mutations, FLT3-ITD, and myelodysplastic syndrome-related changes, including chromosomal alterations and somatic mutations. In contrast, risk stratification for pAML is still developing, and various strategies are utilized in clinical trials14,15. This is partly due to differences in the genetic background between adult and pediatric AML5, the rarity of the disease16, and a shortage of clinical outcome studies related to genetic alterations. To clarify the genomic landscape of pAML and its association with clinical outcomes, we characterized 895 cases of pAML by transcriptome and genome profiling. These analyses resulted in 23 molecular categories, defined by mutually exclusive gene alterations and specific expression profiles, that show unique biological characteristics and mutational backgrounds. We further determined that these molecular categories have predictive value regarding clinical outcomes that can be leveraged to establish a framework for diagnosis and outcome prediction by investigating an independent clinical study cohort. Results Comprehensive genetic characterization of pAML Pediatric AML samples were collected from previously published studies5,9,10,17–26 or clinical trials at St. Jude Children’s Research Hospital, resulting in a cohort of 895 unique pAMLs either at diagnosis (n=786, 87.8%) or at relapse (n=109, 12.2%) (Fig.1A and Fig.S1A, Table.S1). This pAML cohort showed a wide age distribution at diagnosis (range: 0–23.5, median 9.3), with peaks in infancy and adolescence (Fig.S1B). We first assessed the genetic landscape of these AMLs using RNA sequencing (RNA-Seq) data to detect fusions, internal or partial tandem duplications (ITD/PTD), copy number variants (CNV), as well as single nucleotide variants (SNV) and insertions and deletions (Indel) (Fig.1A–E, Table S2–9). For 671 cases (75.0%) with either whole genome sequencing (WGS, 59.2%) or whole exome sequencing (WES, 43.7%), we also collected processed data from publications or performed de novo calling for new cases included in this study (Fig.1A and Fig.S1C). Pathogenic fusions or structural variants (SV) were identified in 627 patients (70.1%). Most of these are recurrent and class-defining in pAML (e.g., KMT2Ar: 20.2%, RUNX1::RUNX1T1: 12.3%, Fig.1B, Fig.1E, Table.S6), whereas we also found fusions recurrent in other leukemias, such as SET::NUP21427 (n=1) or SFPQ::ZFP36L228 (n=1). Mutational profiling revealed 1,947 pathogenic or likely pathogenic somatic mutations in 757 (84.6%) patients, including class-defining NPM1 (68 patients: 7.6%) and CEBPA (49 patients: 5.5%) mutations (Fig.1C, Fig.1E, Table.S7–8). The majority of mutations were in genes involved in signaling pathways (n=874), epigenetics (n=313), and transcription factors (n=440). RAS pathway mutations were most frequent, with 37.5% (336/895) having at least one RAS-related mutation and 21.1% of those (71/336) having mutations in multiple RAS pathway genes. Among CNVs, we frequently observed gains of chromosome 8 (7.2%) or chromosome 21 (6.5%) and loss of the long arm of chromosome 5 (5q-: 1.5%) or chromosome 7 (3.9%) (Fig. 1D, Fig.S1E Table.S9). Enrichment of focal deletions involving RB126 (13q14: 2.9%), ETV629 (12p13: 2.1%), NF130 (17q11: 2.0%), and TP53 (17p13: 2.0%) were also observed in this cohort. In addition, underappreciated focal gains involving AKT331 and FH32 (1q43: 3.0%) or ABCA transporters (17q24: 2.3%) were identified, suggesting their possible importance in leukemogenesis. GRIN (genomic random interval) analysis33 identified 143 genes significantly altered in the entire cohort (Fig.1E, Fig.S2A-B, Table.S10). Consistent with previous reports, RAS-related mutations or FLT3-ITD with variable variant allele frequencies (VAFs) were highly co-occurring with class-defining alterations (Fig.1E, Fig.S2). In contrast, mutations in UBTF or CBFB genes were predominantly found in cases without a defining driver alteration, as previously shown10,34, suggesting that these alterations define subgroups with distinct molecular characteristics. Based on these collective data, we classified pAMLs using current WHO and ICC systems (Fig.1E–F, Fig.S1E-F, Table.S1), and the frequencies of major classifications are consistent with cytogenetic profiles of European pediatric AML cohorts35,36. In our pAML cohort, 68.3% of cases had specified genetic alterations in WHO5th, 10.7% of cases were defined as “acute myeloid leukemia, myelodysplasia-related” (AML-MR), and the remaining cases with rare fusions or no defining alteration were classified as “acute myeloid leukemia with other defined genetic alterations” (15.9%) or by differentiation stages (3.4%). In contrast, 95.0% of adult AMLs can be classified either by specific gene alteration (67.1%) or as AML-MR (27.8%)37, emphasizing the need for a more comprehensive classification of pAML based on the unique biology of childhood AML. Molecular categories of pAML defined by mutually exclusive gene alterations We and others have shown that class-defining driver alterations are associated with specific expression patterns10,38,39 or that allele-specific and outlier expression of MECOM40,41, BCL11B42, or MNX143 by SVs can define disease subtypes. We then integrated the mutational landscape with expression profiling to define granular molecular categories for pAML (Table.S11). UMAP analysis44,45 of transcriptional data revealed tight clustering of classes defined in WHO5th, including KMT2A rearrangement (KMT2Ar), NUP98 rearrangement (NUP98r), RUNX1::RUNX1T1, CBFB::MYH11, NPM1 mutation, CEBPA mutation, and DEK::NUP214, suggesting subtype-specific expression patterns (Fig.2A, Fig.S3A-B). We noted that the clustering is also driven in part by differentiation status represented by marker gene expression, FAB (French-American-British) classification, or cellular hierarchy46 (Fig.S3C-E), contributing to heterogeneity within large categories such as KMT2Ar or NUP98r (Fig.2A, Fig.S3A, S4A). Diffusion maps47 confirmed similar patterns of clustering and differentiation status (Fig.S3A-E). Cases with NPM1 fusions or indels outside the N-terminus48 clustered with canonical NPM1 mutations, and thus we assigned them to the NPM1 category (Fig.S4A), and similarly, TBL1XR1::RARB49 to the APL (acute promyelocytic leukemia) category. For the MECOM category, we noted a case showing outlier and allele-specific expression of MECOM without any evidence of rearrangements, leaving other possible dysregulation mechanisms, such as enhancer amplification42 (Table.S12). Among the remaining cases without class-defining alterations, we found that the following alterations were also mutually exclusive with each other and thus were defined as independent molecular categories: UBTF tandem duplications10, GLIS family (GLIS2–3) fusions9, fusions of FET and ETS family genes50,51 (e.g., FUS::ERG), BCL11B structural variants42, PICALM::MLLT1052, KAT6A rearrangements53, MNX1 structural variants54, RUNX1 fusion with CBFA2T2–355,56 (RUNX1::RUNX1T1-like), and newly reported CBFB insertions (CBFB-GDXY)34 (Fig.2A–C). GATA1 fusions (e.g., MYB::GATA1) or mutations57, rearrangements involving HOX cluster genes26, and PTD of KMT2A58 could rarely co-occur with the above-mentioned category-defining alterations (Fig.2B). However, they are still predominantly found in cases without category-defining alterations, and cases were assigned to these categories only with consistent expression patterns and without previously explained driver alterations (Fig.2A–C). In contrast, defining mutations of AML-MR in WHO5th were overall rare (range: 0.1–2.1%), frequently co-occurred with other defining alterations (e.g., EZH2 in PICALM::MLLT10), and could be found in various clusters rather than as a distinct group (Fig.S3A, F), leading to its exclusion as a defining category for pAML. These 23 molecular categories with defining driver alterations and characteristic expression profiles covered 91.4% (818/895) of the pediatric AML cohort (Fig.2D, Table.S1). Clinical, mutational, and transcriptional characterization of the molecular categories Establishing updated molecular categories for pAML allowed for the investigation of clinicopathological associations. Categories with acute megakaryoblastic or erythroid leukemia (AMKL/AEL) phenotypes are clearly enriched in infants, whereas CBF leukemias and mutation-defining leukemias (e.g., UBTF, NPM1, CEBPA) were enriched in adolescents and young adults (Fig.3A, Fig.S4A-B). Notably, among KMT2A fusion partners, MLLT3 and MLLT10 were found in both monocytic AML and AMKL; however, these fusions preferentially show AMKL phenotypes in infants (Fig.S4A), suggesting that AMKL phenotypes are defined both by types of driver alterations and by the developmental stages as discussed in GATA1-driven AMKL in Down syndrome patients59 or GLIS2::CBFA2T3-driven AMKL60. Overall, however, each molecular category showed variable morphological features represented by FAB classification except categories with APL (M3) or AMKL (M7) phenotypes (Fig.3A). Likewise, conventional cytogenetics also had a limited role. Complex karyotypes, which also define AML-MR11, were frequently observed in MNX1, HOXr, and PICALM::MLLT10 categories. Additionally, as many of these category-defining alterations are cytogenetically cryptic (e.g., NUP98r8 or GLIS family9,61) or somatic mutations (e.g., CEBPA, UBTF, or GATA1), sequencing approaches are needed for the appropriate molecular diagnosis of pAML. We next explored the association between initiating driver alterations and cooperating mutations, as some cooperating mutations co-occur and act synergistically with specific driver events,5,62–64. Signaling alterations were broadly found in 66.6% of patients, whereas each mutation showed distinct patterns among molecular categories with variable VAFs (Fig.1E, Fig.3B, Table.S7). Among RAS mutations, NRAS mutations were broadly found and enriched in CBFB::MYH11, whereas KRAS mutations were enriched in KMT2Ar. Similarly, FLT3-ITD showed strong enrichment in NUP98r, NPM1, UBTF, KMT2A-PTD, DEK::NUP214, APL, and BCL11B categories, accounting for 76.3% of FLT3-ITD+ cases, whereas 74.0% of FLT3-TKD (tyrosine kinase domain) were found in KMT2Ar, NPM1, and CBF-AMLs. Similarly, WT1 mutations were specifically enriched in NUP98r, UBTF, and BCL11B and highly co-occurring with FLT3-ITD (Fig.3B, Fig.S2A). We further evaluated gene expression signatures among molecular categories. Top variable genes across the cohort enriched genes in development (e.g., HOX genes), differentiation, or inflammation (Fig.S5A, Table.S13), consistent with previous reports that the heterogeneity of AML can be partly attributed to differentiation status1,46,65. Gene set enrichment analysis (GSEA) first confirmed that expression profiles of major categories were congruent with previous reports38,66,67 (Fig.3C, Table.S14). The new categories we propose in this study show similarities and differences with clearly defined categories. For example, UBTF showed expression signatures similar to NPM1 and DEK::NUP214, while KAT6Ar was similar to KMT2Ar, suggesting shared biological mechanisms. In addition, genes involved in signaling pathways, immunity, or drug resistance showed unique enrichment across categories (Fig.3C). Weighted gene co-expression network analysis (WGCNA)68 confirmed that each category showed characteristic patterns of active gene networks associated with specific biological functions (Fig.S5B, Table.S15). Given recent adult AML-focused studies using expression profiling to uncover the associations of cellular stemness69,70 or hierarchy46,71 with prognosis or drug response, we investigated these features in our pAML dataset. We observed unique patterns of stemness and cellular hierarchy scores in each category. Molecular categories known to have a good prognosis (RUNX1::RUNX1T135,36, CBFB::MYH1135,36, and CEBPA72) tended to have high GMP scores (median >0.20) (Fig.3D, Fig.S5C), with the exception of the low GMP scores (median: 0.047) and mid-high stemness-related scores in NPM1. Also, KMT2Ar, associated with poor prognosis35,73,74, showed low stemness-related scores and variable differentiation-related scores. Various prognostic scores (e.g., LSC1769, iScore65) also correlated with molecular categories (Fig.S5D). These data collectively demonstrate that molecular categories are associated with unique pathophysiological characteristics, whose clinical association needs to be assessed based on detailed molecular categories. Superfamilies defined by HOX gene expression profiles These molecular categories also showed inter-categorical similarities, forming large clusters of AMKL/AEL, immature AML, CBF leukemias, CEBPA, and two clusters demarcated by HOXA and HOXB cluster gene expression (Fig.2A, Fig.4A–B). The cluster with high HOXA gene expression and low HOXB expression consisted mainly of KMT2Ar and KAT6Ar (herein referred to as HOXA categories), and the other cluster characterized by high expression of both HOXA and HOXB genes included NPM1, NUP98r, UBTF, KMT2A-PTD, and DEK::NUP214 (HOXB group) (Fig.4A, Fig.S6A). Overall, HOXA and HOXB groups, not including those with AMKL features, account for 17.8% and 22.4% of the cohort, respectively. Differential gene expression analyses revealed that pAMLs in HOXB group had high expression of stemness-related genes (PRDM16 and NKX2–3) or differentiation genes (CD96 and WT1) (Fig.4C–D, Table.S16). In contrast, HOXA group cases showed high expression of monocyte or signaling-related genes. GRIN analysis also revealed striking differences in mutational patterns between HOXA and HOXB groups (Fig.4E–F, Table.S17). FLT3 was significantly altered in both HOX groups but with different mutation types; FLT3-TKD was dominant in HOXA group, and FLT3-ITD was prevalent in HOXB group, accounting for 67.3% of FLT3-ITD+ patients (Fig.4F, Fig.S6B). WT1 mutations were preferentially found in HOXB group (56.6%). FLT3-ITD75 and WT1 mutations17,76 have been associated with poor prognosis: however, these data suggest that these mutations highly confound with specific driver alterations that converge on a common expression signature. KRAS mutations were strongly associated with HOXA group and rare in HOXB group (22.2% and 4.2%, respectively). In comparison, NRAS mutations were prevalent in both HOXA and HOXB group (22.7% and 19.3%) (Fig.4F); however, among NRAS mutations, NRAS p.G12–13 mutations were comparable in both categories, while NRAS p.Q61 mutations were more frequent in HOXA group (Fig.4E, Table.S7). It is well-established that each RAS mutation has preferential distribution among cancer subtypes77. Expression levels and differences in the downstream signaling are postulated as the possible mechanisms, and similarly, between FLT3-ITD and TKD78, while at the RNA levels, these genes were homogenously expressed in this pAML cohort (Fig.S6B). These molecular category-dependent mutational patterns may reflect different signal dependencies, potentially offering less toxic targeted therapies guided by these biological insights. Along with the global distinction between HOXA and HOXB groups, we also noted heterogeneity within each HOX cluster. The HOXA cluster consisted of subclusters characterized by MECOM or LAMP5 expression (Fig.S7A-C, Table.S18), harboring most KMT2Ar cases (120/181; 66.3%). Notably, the larger subcluster expressed XAGE1 family genes specifically (Fig.S7B-C), which encode members of testis-specific proteins postulated as therapeutic targets in various tumor types79. Also, the remaining KMT2Ar cases were clustered with other categories with HOXB expression or AMKL less frequently (Fig.S4A and S7A). These clustering patterns were associated with fusion partners (e.g., KMT2A::ELL in the HOXB cluster) or age (younger age with AMKL and outliers), but the associations were not exclusive (Fig.S7D-E). Among KMT2Ar, fusion partners and MECOM expression have been reported to be prognostic73,74; however, our data suggest considerable heterogeneity in expression patterns not explained by only fusion partners or MECOM expression. The HOXB cluster showed similar heterogeneity represented by cellular hierarchies (Fig.S7F-G). These heterogeneities were occasionally associated with molecular categories or somatic mutations but were not exclusive (Fig.S7G-H), with possible factors, including cell-extrinsic factors65,80 to be further investigated. Molecular basis of AML without defining gene alterations Seventy-seven Unclassified cases remained after assignments into these 23 molecular categories. Twenty-one cases had recurrent driver alterations previously reported in the literature, whose sample sizes are insufficient to assess the biological characteristics (Fig.5A, Table.S19). These included rare or novel in-frame RUNX1 fusions (n=2: USP4281, n=1: EVX182 and ZEB2) and rare MLLT10 fusions (n=1: DDX3X83, TEC84, and MAP2K226), with the possibility of further categorization in a larger cohort. Also, in addition to hi-allelic burden JAK2 p.V617F mutation (n=1), we found candidate driver somatic mutations of MLLT1 p.C119>SPAR (n=1) and H3F3A p.K28M (n=1) in pAML with HOX gene expression (Fig.5A, Fig.S8A, Table.S7). These mutations resemble recurrent mutations in other pediatric cancer types with HOX gene expression and immature phenotypes (MLLT1 p.C118QPPG in Wilms tumor85 or H3F3A p.H28M in high-grade glioma86), postulating a shared mechanism of tumorigenesis among these pediatric neoplasms. This series of genomic characterization did not find any pathogenic alterations in 9 cases of the remaining 56 Unclassified cases, partly attributed to the lack of WGS data for 8 of these cases. The rest had at least one pathogenic but not subtype-defining alteration enriched in ETV6, RUNX1, TP53, and RAS pathway genes (Fig.5B–C, Table.S19–20), in addition to complex karyotypes or monosomy 7 (Fig. 3A). Of note, complex karyotypes or NRAS mutations were found broadly in various clusters, whereas non-canonical ETV6 and RUNX1 alterations were found preferentially in clusters comprised of FAB M0–1 cases. These clusters are associated with immature or T cell-like signatures (Fig.5D, Fig.S8B, Table.S21), consistent with a new entity of AMTL87. Although various ETV6 or RUNX1 alterations can co-occur with other defining alterations or be class-defining (e.g., RUNX1::RUNX1T1), the mutations in the Unclassified category are commonly loss-of-function (Fig.5E). Given that germline mutations of RUNX1 or ETV6 are associated with leukemia with incomplete penetrance88,89, these data suggest somatic alterations of these genes also require additional mutations for leukemia development, which may cooperatively define the immature leukemic phenotypes in the absence of other defining alterations. Further accumulation of genomic data and experimental models will be necessary to understand the stepwise leukemogenesis and the clinicopathological features of immature pAML with these mutations. Clinical association of molecular categories Although the association between KMT2Ar or NUP98r with poor outcomes is well-appreciated, the clinical associations of new molecular categories have been discussed only in separate studies10,26. To address this deficiency and translate them into a clinical framework, we investigated the outcomes of these molecular categories using the COG AAML1031 clinical study15 (n=1,034, Table.S22). Analyses of the AAML1031 RNA-Seq data using the same pipeline revealed similar clustering of molecular categories (Fig. 6A) and the overall category frequencies (Fig. 6B). The AAML1031 cohort confirmed the association of molecular categories with age and FLT3-ITD status (Fig.6C) and showed variable MRD (minimal residual disease) positivity among molecular categories. We first assessed the clinical association of the categories using recursive partitioning models90 for censored event time data, which revealed three groups with distinctive prognoses (Fig.6D, Fig.S9A-B). The grouping of major categories aligns with previous reports (e.g., RUNX1::RUNX1T1 (n=141), CBFB::MYH11 (n=102), and CEBPA (n=63) in low-risk35,36,72) except DEK::NUP214 (n=17) in low-risk, historically regarded as high-risk36,91. We confirmed the known association of GLISr61 (n=20), MECOM92 (n=11), PICALM::MLLT1093 (n=8), and KAT6Ar93 (n=7) with poor outcomes, while new categories of MNX1 (n=4), RUNX1::RUNX1T1-like (n=4), and CBFB-GDXY (n=4) were assigned to low-risk. Univariate analyses of other risk factors revealed that age and FLT3-ITD were not prognostic, whereas MRD positivity and cellular hierarchy scores (GMP-like, cDC-like, and cycling LSPC) were associated with the overall survival (Fig.6E, Fig.S9C-D, Table.S23). A Cox proportional hazards model using risk groups and prognostic factors showed that hierarchy scores did not significantly contribute to prognosis (Fig.S9E), whereas risk groups and MRD positivity were independently prognostic (Table.S23). These data led us to establish a simple predictive framework solely based on molecular categories and MRD positivity, resulting in six risk strata with granular outcome prediction (Fig.6F and Fig.S9F-G), whose prognostic values were validated using the AML08 trial14 (n=211, Fig.S9H-J, Table.S24). With these transcriptional and outcome data, we also investigated the clinical association of transcriptional heterogeneity within major molecular categories. Among KMT2Ar, fusion partners or MECOM expression73,74 also confound in the AAML1031 cohort (Fig.6G–H). Cox hazard models showed that both fusion partners and expression clusters are prognostic (P =0.00052 and 0.0015, respectively), with fusions with SEPTIN family and ELL or immature expression patterns associated with favorable outcomes (Fig.6I). Bootstrapping showed that the association of fusion partners or expression clusters with prognosis did not significantly differ (difference in C-index of 95% bootstrap interval for fusions and expression clusters: −0.025–0.093). Although HOXB categories of NUP98r, NPM1, and UBTF also showed heterogeneity of expression patterns, their outcomes were not associated with UMAP clusters or FLT3-ITD status (Fig.S10), and further mutational profiling outside the clinical standard (e.g., WT1 mutations) may be required to risk-stratify within these categories further. Discussion In addition to known enrichment of chromosomal events like t(11,x) in pediatric patients with AML, advances in sequencing technology have identified additional pediatric-specific driver alterations9,10,34. This prompted us to comprehensively investigate the increasingly complex genomic landscape of pAML in the context of the latest classification systems for hematological malignancies (WHO5th,11 and ICC12) and to develop a pAML-focused categorization schema based on the unique disease biology of childhood AML. In this study, we systematically categorized our pAML cohort of 895 patients using an RNA-Seq-based approach, resulting in 23 molecular categories defined by mutually-exclusive driver alterations, covering 91.4% of the entire cohort. Of these 23 categories, 12 are not currently defined by WHO5th. This includes common categories like UBTF, GLISr and GATA1, which were otherwise categorized as “acute myeloid leukemia, myelodysplasia-related” or “acute myeloid leukemia, other defined gene alteration” in the current WHO classification. Notably, myelodysplasia-related (MR) mutations or chromosomal alterations often co-occur with many pAML category-defining alterations and override them in WHO5th despite these MR alterations not driving consistent patterns of gene expression. Considering that the current classification systems are mainly based on evidence from adult AML94,95 and pediatric myelodysplasia syndrome (MDS) is rare22, we propose an alternative framework for pAML to better reflect the unique genetic and clinical landscape. These molecular categories show unique expression and mutational profiles, whereas some categories also show critical similarities, which can suggest common molecular mechanisms and potentially therapeutic options. In particular, we noticed two large clusters characterized by unique HOXA-B expression profiles. Molecular categories with HOXB signatures were strongly associated with FLT3-ITD and WT1 mutations, whereas those with HOXA signatures were associated with RAS mutations. Considering that AMLs with KMT2Ar, NUP98r, and NPM1 have been shown to be dependent on KMT2A/Menin96–98, and that a Menin inhibitor (SNDX-5613) targeting KMT2Ar and NPM1 AML is in a clinical trial99 (NCT04065399), our data suggest that other subtypes marked by HOX expression, such as UBTF or DEK::NUP214, may also be candidates for Menin inhibitors. Thus, this approach may cover nearly 50% of pediatric AML. Also, the high frequency of FLT3-ITD in categories with HOXB expression implies that FLT3 signaling is closely related to biology and that a data-driven implementation of FLT3 inhibitors to HOXB subtypes can be effective. Some cases without category-defining alterations could be characterized by rare fusion or mutations, which need further evidence to establish as a disease entity, including MLLT1 and H3F3A mutations that are frequent and class-defining in Wilms tumor85 and glioma86, respectively. Considering that AML and Ewing sarcoma also share ETS family fusions51 (e.g., EWSR1::ERG), it would be intriguing to incorporate knowledge of these solid tumors to understand the biology behind pAML with these rare alterations. Also, enrichment of RUNX1 or ETV6 loss of function alterations in immature AML implies that these can be class-defining in the absence of other defining alterations and likely with specific cooperating mutations. These findings further suggest a continuum with other immature leukemias, such as early T-cell precursor (ETP)-ALL and mixed phenotype acute leukemias (T/My) which have similar mutational features82,100. We further investigated the clinical outcomes of these molecular categories using two independent cohorts --the COG AAML1031 study and the St. Jude AML08 study. Using both cohorts, we show a strong association of new molecular categories with outcomes (e.g., PICALM::MLLT10 and KAT6Ar as high-risk, CBFB-GDXY as low-risk). These analyses also revealed confounding variables like molecular categories and known prognostic factors like FLT3-ITD status or cellular hierarchy scores. With this comprehensive profiling recognizing new pAML subtypes, we established a simple risk stratification using molecular categories and MRD. This strategy, however, heavily relies on the analysis of next-generation sequencing data. While the WHO classification requires targeted sequencing or WGS, we propose a diagnostic pipeline utilizing RNA-Seq which is highly sensitive for canonical and cryptic fusion calling, allows for categorization based on gene expression signatures, including outlier and allele-specific expression (MECOM, BCL11B, and MNX1), and provides limited but sufficiently sensitive mutation calling to enable our comprehensive molecular categorization strategy to newly diagnosed pAML.This approach is favored over current commercial panels commonly used for pAML, which either lack coverage of all the defining genes (e.g., UBTF) or are not suitable to detect complex structural variations that drive aberrant expression of MECOM or BCL11B. Given that clinical sequencing is not readily available globally and these molecular analyses require substantial expertise, robust and easy pipelines are needed for future and broad application of this framework for pAML in the general clinical setting. Online Methods Subject cohorts and sample details. Tumor samples from patients with AML from the St. Jude Children’s Research Hospital tissue resource core facility were obtained with written informed consent using a protocol approved by the St. Jude Children’s Research Hospital institutional review board (IRB). Studies were conducted in accordance with the International Ethical Guidelines for Biomedical Research Involving Human Subjects. Samples for RNA sequencing (RNA-Seq: n=221), whole genome sequencing (WGS: n=54), and whole exome sequencing (WES: n=6) are newly sequenced in this study, and the rest of the data were obtained from previous publications5,9,10,17–26 or public databases (see details in Data availability and Table S1). For samples with multiple available data points, we included one representative time point with a high tumor purity and good RNA-Seq data quality. Cases were assigned to current WHO and ICC by board-certified hematopathologists (PK and JMK). Sample processing, library preparation, and sequencing. For newly sequenced samples with low tumor purity (below 60%), the leukemic cell population was enriched either by flow cytometric sorting or T cell depletion by magnetic beads (EasySep Human CD3 Positive Selection Kit II, 17851, StemCell Technologies). For flow cytometric sorting, CD45dimCD33dim positive population was sorted using anti-CD45 PerCP-Cyanine5.5 (eBioscience cat# 8045–9459-120) and anti-CD33 APC (eBioscience cat# 17–0338-42). CD34 gating using anti-CD34 PE (Beckman cat# IM1459U) was added depending on the positivity of each patient sample. Enrichment of the tumor population was confirmed flow cytometric analysis of the post-sorting samples (generally > 90%). Libraries were constructed using the TruSeq Stranded Total RNA Kit, with Ribozero Gold (20020598, Illumina) for RNA-Seq, the TruSeq DNA PCR-Free Library Prep Kit (20015963, Illumina) for WGS, and the TruSeq Exome Kit v1 (20020614, Illumina) for WES according to the manufacturer’s instructions. After library quality and quantity assessment, samples were sequenced on HiSeq2000 or 2500 (Illumina, RRID:SCR_020132, RRID:SCR_016383) instruments with paired-end (2 × 101 bp, 2 × 126 bp, or 2 × 151 bp) sequencing using TruSeq SBS Kit v3-HS (FC-401–3001, Illumina) or TruSeq Rapid SBS Kit (FC-402–4023, Illumina). RNA-Seq mapping, fusion detection, and large-scale copy number variant calling. RNA reads from newly sequenced samples and from publications were mapped to the GRCh37/hg19 human genome assembly using the StrongARM pipeline101. Chimeric fusion detection was carried out using CICERO102 (v0.3.0). For the cases with only RNA-Seq data, RNAseqCNV103 (v1.2.1) was used to call large-scale copy number variants (CNV). Somatic mutation calling from RNA-Seq. We initially performed RNA-based variant calling methods for detecting Single-nucleotide variants (SNV) and Insertions and deletions (Indel). Specifically, we applied the following approach to simultaneously account for germline polymorphisms (without germline control) and sequencing artifacts specific to RNA-Seq on a panel of 86 predefined genes previously reported to be significantly mutated in pediatric AML5 and myelodysplastic syndrome (MDS; Table.S5). Briefly, candidate SNVs/Indels were called by Bambino104 (v1.07) or RNAindel105,106 (v3.0.4), annotated by VEP107 (v95), and in turn, classified for putative pathogenicity with PeCanPie/MedalCeremony108. Candidate variants with putative pathogenicity were considered germline or artifacts if present in >5% of the cases. Candidate variants were further filtered if the number of supporting reads was ≤5 or if the variant allele fraction (VAF) was ≤5%. UBTF tandem duplications were detected by soft-clipped read counting in addition to CICERO and RNAindel as we previously described10. Whole genome and whole exome sequencing data analysis. The previous genomic lesion calls for the cases (WGS; n=394, WES; n=284) from published studies5,9,10,17,19–21,24,26 were collected from their respective publications. For the unpublished cases with DNA data (WGS; n=136, WES; n=107), DNA reads were mapped using BWA109,110(WGS: v0.7.15-r1140 and v0.5.9-r26-dev; WES: v0.5.9-r26-dev and v0.5.9, RRID:SCR_010910) to the GRCh37/hg19 human genome assembly. Aligned files were merged, sorted, and de-duplicated using Picard tools 1.65 (broadinstitute.github.io/picard/). SNVs and Indels were called using Bambino104. Candidate SNVs and Indels were similarly classified for putative pathogenicity with PeCanPie/MedalCeremony as in somatic mutation calling from RNA-Seq. The counting of somatic mutations included all the pathogenic or likely pathogenic mutations detected by WGS, whereas mutation detection from cases with only RNA-Seq data is limited to the 86 preselected genes. Structural variations (SV) were analyzed using CREST111 (v1.0), and CNVs were analyzed using CONSERTING112 on the WGS data. CNVs were also called on cases with only WES DNA data using the following methods. Briefly, Samtools113 mpileup command was used to generate a mpileup file from matched germline and tumor BAM files with duplicates removed. If a matched germline was not available, a high-quality normal sample was used to pair with the tumor sample. VarScan114 (v2.3.5) was then used to take the mpileup file to call somatic CNVs after adjusting for normal/tumor sample read coverage depth and GC content. Circular Binary Segmentation algorithm115 implemented in the DNAcopy R package (v1.52.0) was used to identify the candidate CNVs for each sample. B-allele frequency info was also used to assess allelic imbalance. GRIN analysis for significantly mutated genes. For the 895 AML cases, the genomic random interval (GRIN; v2.0) model33 was used to evaluate the statistical significance of the number of subjects with each type of lesion: fusions, CNVs (amplifications and deletions), copy neutral loss of heterozygosity (CN-LOH), SNV/indels, and tandem duplications in each gene. For each type of lesion, robust false discovery estimates were computed from P values using Storey’s q value116 with the Pounds-Cheng estimator of the proportion of hypothesis tests with a true null hypothesis117. FDR cutoff of <0.05 for the number of subjects with any one type of lesion overlapping the gene locus was used to obtain significantly mutated genes, where we focused on protein-coding genes and genes that are known or likely to be pathogenic in leukemia. We also excluded genes that are part of a large chromosomal gain, loss, or CN-LOH but not the target of the CNVs based on the GISTIC (Genomic Identification of Significant Targets in Cancer) analysis. Subgroup GRIN analyses for HOXA categories (n=167), HOXB categories (n=211) categories and the Unclassified category (n=77) were also performed using the same methods. GISTIC analysis for significant recurring copy-number alterations. We used GISTIC (v2.0.23, RRID:SCR_000151) 118,119 to identify genomic regions that are significantly amplified or deleted across our 895 samples. Each aberration was assigned a G-score that considered the amplitude of the aberration as well as the frequency of its occurrence across samples. False discovery rate q values were then calculated for the aberrant regions, and regions with q values ≤0.25 were considered significant. A “peak region” was identified for each significant region with the greatest amplitude and frequency of alteration. In addition, a “wide peak” was determined using a leave-one-out algorithm to allow for errors in the boundaries in a single sample. The “wide peak” boundaries were more robust for identifying the most likely gene targets in the region. Each significantly aberrant region was also tested to determine whether it resulted primarily from broad or focal events (A broad event was set as >90% of the chromosome arm, whereas a focal event was ≤90%). Allele-specific expression estimation for MNX1, BCL11B, and MECOM categories. For cases with both WGS and RNA-Seq available, SNP (single-nucleotide polymorphism) markers in the respective gene locus with ≥10x coverage that are heterozygous (defined as 0.2≤VAF≤0.8) in WGS and also present in RNA-Seq were extracted and a two-sided binomial test (with probability of success P=0.5) was performed on each marker for allelic imbalance in RNA expression. The median of binomial P values was used to assess allele-specific expression (ASE). For RNA-Seq only cases, SNP markers in the respective gene locus with ≥10x coverage and allelic imbalance (VAF≤0.2 or VAF≥0.8) support ASE. Germline variant curation methods. We focused on 15 candidate genes relevant to AML that define specific categories in WHO5th (Table.S25) and scanned for germline mutations in the cases with WGS or WES germline BAM files available (WGS n=367; WES n=354). For cases with germline mutation called in previously published studies10,22, we collected calls from the studies. For the remaining cases, the putative germline variants were called using Bambino104, annotated by VEP107, and in turn, classified for putative pathogenicity with PeCanPie/MedalCeremony108. We then used the following criteria to obtain the candidate germline variants: gnomAD (v2.1.1, RRID: SCR_014964)120 population allele frequency ≤0.001; read coverage SNV≥20 and Indel≥15; for SNV, variant allele frequency between 0.2 and 0.8; for Indel, ≥3 reads supporting the alternative allele. All candidate germline variants were comprehensively reviewed and classified as pathogenic, likely pathogenic, of uncertain significance, likely benign, or benign based on recommendations from the American College of Medical Genetics and Genomics and the Association for Molecular Pathology121 and the Clinical Genome Resource122–125 by a variant scientist (JLM). Gene expression data summarization, batch correction, dimension reduction, and clustering. Reads from aligned RNA-Seq BAM files were assigned to genes and counted using HTSeq126 (v0.11.2, RRID: SCR_005514) with the GENCODE (RRID: SCR_014966) human release 19 gene annotation. The gene count matrix was generated, and for a gene to be considered as expressed, we required that at least 5 samples should have ≥10 read counts per million reads sequenced. The count data were transformed to log2-counts per million (log2CPM) using Voom127 available from R package Limma128(v3.50.3, RRID: SCR_010943). We corrected for library strand (stranded total RNA vs. unstranded mRNA) and batch effect between St. Jude and TARGET cases using the ComBat method available from R package SVA129(v3.42.0, RRID:SCR_012836). The R package Seurat130–133(v4.1.0, RRID:SCR_016341) was used for dimension reduction and sample clustering. Briefly, the top variable genes were selected using the “vst” method. The expression data were then scaled, and PCA (Principal Component Analysis) was performed on the scaled data using the top 320 variable genes. Dimension reduction was performed using UMAP44,45 (Uniform Manifold Approximation and Projection, RRID:SCR_018217) with the top 100 principal components, n_neighbors =15 and min_dist =0.2. Samples were clustered using the top 100 principal components by first constructing a K nearest-neighbor graph and then iteratively optimizing the modularity using Louvain algorithm with resolution=3.5. Dimension reduction was also performed by Diffusion maps 47,134 algorithm available in the R package destiny135 (v 3.10.0) using the same 320 genes with the default setting except for number of principal components n_pcs=50. Differential gene expression analysis was performed by Limma128, and we set Log2 CPM = 0 if it is < 0 based on the Log2 CPM data distribution. P values were adjusted by the Benjamini-Hochberg method to calculate the false discovery rate (FDR) using R function p.adjust. Genes with absolute fold change > 2 and FDR < 0.05 were regarded as significantly differentially expressed. Gene Set Enrichment Analysis (GSEA)136 was performed by GSEA (v4.2.3, RRID: SCR_003199) using MSigDB gene sets c2.all (v7.5.1), comparing each category with the rest of the categories. Permutations were done 1000 times among gene sets with sizes between 15 and 1500 genes. Normalized enrichment scores (NES) and FDR for arbitrary gene sets representing hematopoiesis, leukemia phenotype, biological processes, and drug responses were extracted from the entire results and shown in a heatmap. Weighted correlation network analysis (WGCNA) was done by R package WGCNA68 (v1.70–3, RRID:SCR_003302) using top 2000 variable genes and default setting with the exception of block-wide module calculation with reassignThreshold = 0 and mergeCutHeight = 0.25. Functional annotation of top 320 variable genes, differentially expressed genes, and genes in WGCNA modules were performed with DAVID137 (v6.8), and results for GO term, biological process (GOTERM_BP_DIRECT) were exported. Inference of cellular hierarchy by CIBERSORT138 (RRID:SCR_016955) was performed by the web interface of CIBERSORTx in absolute mode with S-mode batch correction without a permutation as previously reported46. TPM values were used as input data, and Malignant Signature Matrix and Malignant Single Cell Reference Samples were used as previously described46, and the malignant cell populations were normalized to 1 to calculate the relative fraction scores, which were shown in UMAP space or violin plots. Prognostic scores of LSC1769, pLSC670, ADE-RS139, and iScore65 were calculated as reportedly. Hierarchical clustering (RRID:SCR_014673) of expression data, mutual-exclusivity matrix, and GSEA scores were performed using the Euclidian distance and Ward method. Statistical test. For discrete values of the molecular category and the mutation frequency in cohorts, statistical significance and mutual exclusivity were assessed by two-sided Fisher’s exact test and Pearson’s correlation. Adjustment of multiple testing was performed by the Benjamini-Hochberg method using p.adjust function on R when appropriate. For survival data, decision trees were established by a recursive partitioning method using R library rpart90 (v4.1.19, RRID:SCR_021777). Kaplan–Meier curves for the probability of overall survival (OS) and event-free survival (EFS) were constructed using R package survival (v3.3–1, RRID:SCR_021137). Events in the probability of EFS calculations were defined as relapse, death in remission by any cause, and non-response, which was included as an event at the date of diagnosis. The Cox proportional hazards model was used to calculate the statistical significance of individual prognostic factors by univariate analyses first, and significant factors were included in a multivariate analysis. Clinical association of the molecular categories were first assessed using the AAML1031 study (n=1034), and the results were validated using the AML08 cohort (n=211, independent from the AAML1031, a part of this study cohort). R statistical environment (R v4.0.2, RRID:SCR_001905) was used for statistical tests. Visualization. Mutational heatmaps and mutations on individual genes were visualized using ProteinPaint (https://proteinpaint.stjude.org/). Heatmaps of expression data, mutual-exclusivity matrix, and GSEA scores were created by pheatmap function of R library pheatmap (v1.0.12, RRID:SCR_016418). Other data visualizations were performed by ggplot function of R library ggplot2 (v3.3.6, RRID:SCR_014601), survminer (v0.4.9), and base plot function in R statistical environment. Figures are incorporated and edited using Adobe Illustrator (2021, RRID:SCR_010279). Annotation of genes in mutational heatmaps depend on common knowledges, and the definition of RAS pathway genes included causative genes of Noonan or Noonan-like syndrome140 (NRAS, KRAS, PTPN11, NF1, CBL, LZTR1, RIT1, BRAF, SOS1, and HRAS). Data availability. The genomic data and expression data newly generated in this study (RNA-Seq: n=221, WGS: n=54, WES: n=6) have been deposited in the European Genome-Phenome Archive (EGA, RRID:SCR_004944), which is hosted by the European Bioinformatics Institute (EBI), under accession EGAS00001005760. For the remaining RNA-Seq data for 588 cases, 401 are St. Jude cases, of which 274 cases with data from the publications are available either on EGA or St. Jude Cloud9,10,18,20–24,26 or on the original publication25. For the other 127 published cases19, we downloaded the BAM files from EGA (EGAS00001004701). Unpublished data (n=86) are also available on St. Jude Cloud under the PCGP study (https://permalinks.stjude.cloud/permalinks/PCGP, n=8) and the RTCG study (https://platform.stjude.cloud/data/cohorts?dataset_accession=SJC-DS-1007, n=78). Of the remaining WGS data for 394 cases, 207 are St. Jude cases, of which 115 cases with data from the original publications9,10,20,21,24,26 are available on either EGA or St. Jude Cloud, and for the other 92 published cases19, we downloaded the BAM files from EGA (EGAS00001004701). Unpublished WGS data (n=82) are also available on St. Jude Cloud under the RTCG study. For the remaining WES data for 314 cases, 275 are St. Jude cases, of which 155 with data from the original publications9,10,18,20–24,26 are available either on St. Jude Cloud or EGA, and for the other 120 published cases19, we downloaded the BAM files from EGA (EGAS00001004701). Unpublished WES data (n=101) are also available on St. Jude Cloud under the PCGP study (n=2) and the RTCG study (n=99). The data generated by the TARGET initiative5,17 (n=187) is also available under accession phs000218 (TARGET-AML) and phs000465 (TARGET sub-study, data is available as a part of phs000218), managed by the NCI. Information about TARGET can be found at http://ocg.cancer.gov/programs/target. Other data generated in this study are available in the Supplemental tables or upon request to the corresponding author. Acknowledgements We thank all the patients and their families at St. Jude Children’s Research Hospital (SJCRH) for their contribution of the biological specimens used in this study. We also thank the Biorepository, the Flow Cytometry and Cell Sorting Core, and the Hartwell Center for Bioinformatics and Biotechnology at SJCRH for their essential services. This work was funded by the American Lebanese and Syrian Associated Charities of St. Jude Children’s Research Hospital and grants from the NIH (P30 CA021765, Cancer Center Support Grant and a Developmental Fund Award, to J.M. Klco and X. Ma). The content, however, does not necessarily represent the official views of the NIH and is solely the responsibility of the authors. This work was also supported in part by the Fund for Innovation in Cancer Informatics (www.the-ici-fund.org, to X. Ma and J.M. Klco). J.M. Klco holds a Career Award for Medical Scientists from the Burroughs Wellcome Fund and is a previous recipient of the V Foundation Scholar Award (Pediatric). Figure 1: Comprehensive genetic characterization of pediatric acute myeloid leukemia (pAML) A. Study cohort of pediatric AML (n=895) and study design B. Recurrent pathogenic or likely pathogenic in-frame fusions (blue) and structural variants (SV: gray) detected in the entire cohort. Fusions included only in-frame fusions, and SVs included out-of-frame fusions resulting loss of the C-terminus of the protein and alterations detected from WGS data using CREST. C. Recurrent pathogenic or likely pathogenic somatic mutations. Colors represent types of mutations. Bars in Fig.1B–C represent the total number of alterations in the cohort. D. Results of GISTIC analysis for focal chromosomal events (shorter than 90% of the chromosome arm). The left panel shows the enrichment of focal gains, and the right panel shows the enrichment of focal loses. Green lines show a significance threshold for q values (0.25). Representative genes in enriched regions are highlighted. E. The genomic landscape and the WHO classification of pAML. Representative genes from GRIN analysis or defining alterations are shown. F. Summary of the WHO classification of the entire cohort. Figure 2: Molecular categories defined by mutually exclusive gene alterations A. UMAP plot of the entire pAML cohort (n=895) and cord blood CD34+ cells (normal controls: n=5) using top 320 variable genes. The colors of each dot denote the molecular categories of the samples. Representative category names are shown, and large clusters are highlighted in circles. B. A heatmap showing frequencies of defining gene alterations represented by the color. Statistical significance was assessed by two-sided Fisher’s exact test to calculate p values of co-occurrence, followed by the Benjamini-Hochberg adjustment for multiple testing to calculate q values (*P<0.05, **q<0.05). C. Definition of molecular categories and diagnostic flow. Molecular categories not defined in WHO5th are highlighted in red. D. A ribbon plot showing the association between WHO classification and molecular categories. Colors represent molecular categories of samples Figure 3: Clinical and molecular profiles of molecular categories A. Clinical background of molecular categories. Upper row. Violin plots showing age distribution within each molecular category. Large dots and bars represent the median and the range of 2.5~97.5 percentiles, respectively. Small dots represent individual patients’ ages. Bottom row. Frequency of FAB and karyotype in individual categories. B. Mutational heatmap showing mutation frequencies in each molecular category. The color of each panel represents the frequency of a mutation in each molecular category, and the statistical significance was assessed by two-sided Fisher’s exact test to calculate p values of co-occurrence followed by the Benjamini-Hochberg adjustment for multiple testing to calculate q values (*P<0.05, **q<0.05 after the adjustment). Bars on the top panel show the frequency of mutations in the entire cohort, and the colors represent mutation types. Molecular categories are clustered according to Ward clustering using the Euclidean distance of the frequency matrix. Genes are grouped according to the functional annotations. C. A heatmap showing normalized enrichment scores (NES) and false discovery rates (FDR) of gene set enrichment analysis (GSEA) of each molecular category. Colors denote NES, and asterisks show FDR (*FDR<0.05, **FDR<0.01, ***FDR<0.001) D. Violin plots showing cellular hierarchy scores in each molecular category inferred by CIBERSORT. Lines of the box represent 25% quantile, median, and 75% quantile. The upper whisker represents the higher value of maxima or 1.5 × interquartile range (IQR), and the lower whisker represents the lower value of minima or 1.5 × interquartile range (IQR). Dots show outliers. LSPC stands for leukemic stem and progenitor cells. Figure 4: Categories demarcated by HOXA and HOXB cluster expression A. UMAP plot showing groups of molecular categories based on UMAP clustering and HOX cluster gene expression profiles. B. HOXA9 and HOXB5 expression on UMAP plot. The dot colors represent the relative expression of the genes. C. A volcano plot showing differentially expressed genes (DEG) between HOXA and HOXB groups. Genes with absolute fold change > 2 and FDR < 0.05 are considered DEGs. Representative gene names are shown. D. GO term analyses of genes with significantly high expression in each HOX group by DAVID. Bars represent logged FDR. E. Plots showing results of GRIN analyses in HOXA group (horizontal axis) and HOXB group (vertical axis). Genes with FDR<0.1 in either HOXA or HOXB groups are shown. Red or blue dots show genes enriched only in either HOXA or HOXB groups, respectively. The dotted lines represent thresholds for statistical significance (FDR<0.05). F. A mutational heatmap comparing patterns between HOXA and HOXB groups. Colors represent mutation types, and molecular categories are annotated on the top. Bar plots on the right show frequencies of mutations in HOXA and HOXB groups. Statistical significance of GRIN analysis in HOXA and HOXB groups (*FDR<0.05) and two-sided Fisher’s exact test between HOXA and HOXB groups (*P<0.05, **q<0.05 after the Benjamini-Hochberg adjustment) are also shown. GRIN results for FLT3 are for the entire gene, while Fisher’s tests were performed separately for ITD, TKD, and non-TKD mutations. Figure 5: Characterization of cases without category-defining alterations A. UMAP plot showing cases without category-defining alterations. Red dots represent cases with rare recurrent gene alterations, blue dots represent cases for which no pathogenic alteration was found, and black dots represent cases with at least one gene alteration not defining the phenotype. B. A plot showing the FDR of GRIN analysis for the Unclassified category (horizontal axis) and relative enrichment of the alteration in the Unclassified category (vertical axis). The dot sizes and colors denote the Unclassified category’s frequency, which included fusions, mutations, copy number loss and gain, and copy-neutral heterozygosity. C. A mutational heatmap of the Unclassified cases, including complex karyotypes and monosomy 7. Patients’ age, FAB, and UMAP clustering are annotated on the top. Colors represent mutation types. D. UMAP plots showing FAB (top-left), CD34 or CD3D expression (bottom-left), and cases with ETV6 alterations (top-right) and RUNX1 alteration (bottom-right). E. Patterns of alteration in ETV6 (left) and RUNX1 (right). Category-defining fusions are shown in the top row, alterations co-occurring with category-defining alterations in the middle row, and alterations in the Unclassified category in the bottom row. Bars represent a relative fraction of alteration in each group; the colors denote the alteration types. Figure 6: Clinical association of molecular categories A. UMAP plot of transcriptome data of the AAML1031 cohort (n=1,034) using top340 variable genes. The dot colors denote molecular categories assigned to the samples according to genomic profiling using the same pipeline as this study cohort. Representative category names are shown, and large clusters are highlighted in circles. B. Frequency of molecular categories in the AAML1031 cohort. Asterisks denote the statistical significance of the frequency of each category assessed by two-sided Fisher’s exact test followed by the Benjamini-Hochberg adjustment (*P<0.05, **q<0.05, blue: fewer and black: more in the AAML1031). C. Clinical features of molecular categories showing age at diagnosis (left), FLT3-ITD status (mid), and MRD (minimal residual disease) positivity at the end of induction (right). Molecular category names associated with megakaryocytic phenotypes are highlighted in red. Lines of the box represent 25% quantile, median, and 75% quantile. The upper whisker represents the higher value of maxima or 1.5 × interquartile range (IQR), and the lower whisker represents the lower value of minima or 1.5 × interquartile range (IQR). D. Grouping of molecular categories into Low, Intermediate, and High-risk groups by recursive partitioning (top) and Kaplan-Meier curves of overall survival of patients in each risk group (bottom). E. Kaplan-Meier curves and statistical significance of overall survival of patients with known prognostic factors (FLT3-ITD status: top-left, age: bottom-left, MRD positivity at the end of the induction I: top-right). F. Kaplan-Meier curves of overall survival of patients in six risk strata using risk groups (Low-Intermediate-High) and MRD positivity. G. Distribution of KMT2Ar cases among transcriptional clusters on UMAP plot, colors representing fusion partners (left) and XAGE1A and MECOM expression, colors representing relative expression (right) on UMAP plot. H. The association of fusion partners of KMT2Ar among different clusters. I. Kaplan-Meier curves of overall survival of patients with each fusion (left) and in each cluster (right). For survival curves in D, E, F, and I, statistical significance was assessed by Cox Proportional-Hazards models, and P values are shown in the plot. For the validity of prediction by KMT2Ar fusion partners and clusters in I, c-index scores assessed by bootstrapping were shown below the plots. For I, statistical significance of the enrichment and exclusivity were assessed by two-sided Fisher’s exact test followed by the Benjamini-Hochberg adjustment (*P<0.05, **q<0.05, blue: exclusive, black: enriched). Additional Declarations: There is NO Competing Interest. ==== Refs References 1 Miles L. A. Single-cell mutation analysis of clonal evolution in myeloid malignancies. Nature 587 , 477–482, doi:10.1038/s41586-020-2864-x (2020).33116311 2 Cicconi L. & Lo-Coco F. Current management of newly diagnosed acute promyelocytic leukemia. Ann Oncol 27 , 1474–1481, doi:10.1093/annonc/mdw171 (2016).27084953 3 Tenen D. G. Disruption of differentiation in human cancer: AML shows the way. Nat Rev Cancer 3 , 89–101, doi:10.1038/nrc989 (2003).12563308 4 Morita K. Clonal evolution of acute myeloid leukemia revealed by high-throughput single-cell genomics. Nat Commun 11 , 5327, doi:10.1038/s41467-020-19119-8 (2020).33087716 5 Bolouri H. The molecular landscape of pediatric acute myeloid leukemia reveals recurrent structural alterations and age-specific mutational interactions. Nat Med 24 , 103–112, doi:10.1038/nm.4439 (2018).29227476 6 Cancer Genome Atlas Research, N. Genomic and epigenomic landscapes of adult de novo acute myeloid leukemia. N Engl J Med 368 , 2059–2074, doi:10.1056/NEJMoa1301689 (2013).23634996 7 Papaemmanuil E. Genomic Classification and Prognosis in Acute Myeloid Leukemia. N Engl J Med 374 , 2209–2221, doi:10.1056/NEJMoa1516192 (2016).27276561 8 Jaju R. J. A novel gene, NSD1, is fused to NUP98 in the t(5;11)(q35;p15.5) in de novo childhood acute myeloid leukemia. Blood 98 , 1264–1267, doi:10.1182/blood.v98.4.1264 (2001).11493482 9 Gruber T. A. An Inv(16)(p13.3q24.3)-encoded CBFA2T3-GLIS2 fusion protein defines an aggressive subtype of pediatric acute megakaryoblastic leukemia. Cancer Cell 22 , 683–697, doi:10.1016/j.ccr.2012.10.007 (2012).23153540 10 Umeda M. Integrated Genomic Analysis Identifies UBTF Tandem Duplications as a Recurrent Lesion in Pediatric Acute Myeloid Leukemia. Blood Cancer Discov 3 , 194–207, doi:10.1158/2643-3230.BCD-21-0160 (2022).35176137 11 Khoury J. D. The 5th edition of the World Health Organization Classification of Haematolymphoid Tumours: Myeloid and Histiocytic/Dendritic Neoplasms. Leukemia 36 , 1703–1719, doi:10.1038/s41375-022-01613-1 (2022).35732831 12 Arber D. A. International Consensus Classification of Myeloid Neoplasms and Acute Leukemias: integrating morphologic, clinical, and genomic data. Blood 140 , 1200–1228, doi:10.1182/blood.2022015850 (2022).35767897 13 Mrozek K. Outcome prediction by the 2022 European LeukemiaNet genetic-risk classification for adults with acute myeloid leukemia: an Alliance study. Leukemia, doi:10.1038/s41375-023-01846-8 (2023). 14 Rubnitz J. E. Clofarabine Can Replace Anthracyclines and Etoposide in Remission Induction Therapy for Childhood Acute Myeloid Leukemia: The AML08 Multicenter, Randomized Phase III Trial. J Clin Oncol 37 , 2072–2081, doi:10.1200/JCO.19.00327 (2019).31246522 15 Pollard J. A. Sorafenib in Combination With Standard Chemotherapy for Children With High Allelic Ratio FLT3/ITD+ Acute Myeloid Leukemia: A Report From the Children’s Oncology Group Protocol AAML1031. J Clin Oncol 40 , 2023–2035, doi:10.1200/JCO.21.01612 (2022).35349331 16 Puumala S. E. , Ross J. A. , Aplenc R. & Spector L. G. Epidemiology of childhood acute myeloid leukemia. Pediatr Blood Cancer 60 , 728–733, doi:10.1002/pbc.24464 (2013).23303597 17 McNeer N. A. Genetic mechanisms of primary chemotherapy resistance in pediatric acute myeloid leukemia. Leukemia 33 , 1934–1943, doi:10.1038/s41375-019-0402-3 (2019).30760869 18 Iacobucci I. Genomic subtyping and therapeutic targeting of acute erythroleukemia. Nat Genet 51 , 694–704, doi:10.1038/s41588-019-0375-1 (2019).30926971 19 Fornerod M. Integrative Genomic Analysis of Pediatric Myeloid-Related Acute Leukemias Identifies Novel Subtypes and Prognostic Indicators. Blood Cancer Discov 2 , 586–599, doi:10.1158/2643-3230.BCD-21-0049 (2021).34778799 20 Newman S. Genomes for Kids: The Scope of Pathogenic Mutations in Pediatric Cancer Revealed by Comprehensive DNA and RNA Sequencing. Cancer Discov 11 , 3008–3027, doi:10.1158/2159-8290.CD-20-1631 (2021).34301788 21 Rusch M. Clinical cancer genomic profiling by three-platform sequencing of whole genome, whole exome and transcriptome. Nat Commun 9 , 3962, doi:10.1038/s41467-018-06485-7 (2018).30262806 22 Schwartz J. R. The genomic landscape of pediatric myelodysplastic syndromes. Nat Commun 8 , 1557, doi:10.1038/s41467-017-01590-5 (2017).29146900 23 Andersson A. K. The landscape of somatic mutations in infant MLL-rearranged acute lymphoblastic leukemias. Nat Genet 47 , 330–337, doi:10.1038/ng.3230 (2015).25730765 24 Faber Z. J. The genomic landscape of core-binding factor acute myeloid leukemias. Nat Genet 48 , 1551–1556, doi:10.1038/ng.3709 (2016).27798625 25 Buelow D. R. Uncovering the Genomic Landscape in Newly Diagnosed and Relapsed Pediatric Cytogenetically Normal FLT3-ITD AML. Clin Transl Sci 12 , 641–647, doi:10.1111/cts.12669 (2019).31350825 26 de Rooij J. D. Pediatric non-Down syndrome acute megakaryoblastic leukemia is characterized by distinct genomic subsets with varying outcomes. Nat Genet 49 , 451–456, doi:10.1038/ng.3772 (2017).28112737 27 Van Vlierberghe P. The recurrent SET-NUP214 fusion as a new HOXA activation mechanism in pediatric T-cell acute lymphoblastic leukemia. Blood 111 , 4668–4680, doi:10.1182/blood-2007-09-111872 (2008).18299449 28 Liu Y. The genomic landscape of pediatric and young adult T-lineage acute lymphoblastic leukemia. Nat Genet 49 , 1211–1218, doi:10.1038/ng.3909 (2017).28671688 29 Andreasson P. Deletions of CDKN1B and ETV6 in acute myeloid leukemia and myelodysplastic syndromes without cytogenetic evidence of 12p abnormalities. Genes Chromosomes Cancer 19 , 77–83, doi:10.1002/(sici)1098-2264(199706)19:2<77::aid-gcc2>3.0.co;2-x (1997).9171997 30 Balgobind B. V. Leukemia-associated NF1 inactivation in patients with pediatric T-ALL and AML lacking evidence for neurofibromatosis. Blood 111 , 4322–4328, doi:10.1182/blood-2007-06-095075 (2008).18172006 31 Turner K. M. Genomically amplified Akt3 activates DNA repair pathway and promotes glioma progression. Proc Natl Acad Sci U S A 112 , 3421–3426, doi:10.1073/pnas.1414573112 (2015).25737557 32 Zhang H. , Ju Q. , Ji J. & Zhao Y. Pan-Cancer Analysis Reveals FH as a Potential Prognostic and Immunological Biomarker in Lung Adenocarcinoma. Dis Markers 2021 , 8554844, doi:10.1155/2021/8554844 (2021).34737838 33 Pounds S. A genomic random interval model for statistical analysis of genomic lesion data. Bioinformatics 29 , 2088–2095, doi:10.1093/bioinformatics/btt372 (2013).23842812 34 Ryland G. L. Description of a novel subtype of acute myeloid leukemia defined by recurrent CBFB insertions. Blood 141 , 800–805, doi:10.1182/blood.2022017874 (2023).36179268 35 von Neuhoff C. Prognostic impact of specific chromosomal aberrations in a large group of pediatric patients with acute myeloid leukemia treated uniformly according to trial AML-BFM 98. J Clin Oncol 28 , 2682–2689, doi:10.1200/JCO.2009.25.6321 (2010).20439630 36 Harrison C. J. Cytogenetics of childhood acute myeloid leukemia: United Kingdom Medical Research Council Treatment trials AML 10 and 12. J Clin Oncol 28 , 2674–2681, doi:10.1200/JCO.2009.24.8997 (2010).20439644 37 Huber S. AML classification in the year 2023: How to avoid a Babylonian confusion of languages. Leukemia, doi:10.1038/s41375-023-01909-w (2023). 38 Ross M. E. Gene expression profiling of pediatric acute myelogenous leukemia. Blood 104 , 3679–3687, doi:10.1182/blood-2004-03-1154 (2004).15226186 39 Cheng W. Y. Transcriptome-based molecular subtypes and differentiation hierarchies improve the classification framework of acute myeloid leukemia. Proc Natl Acad Sci U S A 119 , e2211429119, doi:10.1073/pnas.2211429119 (2022).36442087 40 Groschel S. A single oncogenic enhancer rearrangement causes concomitant EVI1 and GATA2 deregulation in leukemia. Cell 157 , 369–381, doi:10.1016/j.cell.2014.02.019 (2014).24703711 41 Schwartz J. R. The acquisition of molecular drivers in pediatric therapy-related myeloid neoplasms. Nat Commun 12 , 985, doi:10.1038/s41467-021-21255-8 (2021).33579957 42 Montefiori L. E. Enhancer Hijacking Drives Oncogenic BCL11B Expression in Lineage-Ambiguous Stem Cell Leukemia. Cancer Discov 11 , 2846–2867, doi:10.1158/2159-8290.CD-21-0145 (2021).34103329 43 Tosi S. Paediatric acute myeloid leukaemia with the t(7;12)(q36;p13) rearrangement: a review of the biological and clinical management aspects. Biomark Res 3 , 21, doi:10.1186/s40364-015-0041-4 (2015).26605042 44 McInnes L , H. J. , & Melville J (2018). UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction. Retrieved from http://arxiv.org/abs/1802.03426. 45 Becht E. Dimensionality reduction for visualizing single-cell data using UMAP. Nat Biotechnol, doi:10.1038/nbt.4314 (2018). 46 Zeng A. G. X. A cellular hierarchy framework for understanding heterogeneity and predicting drug response in acute myeloid leukemia. Nat Med 28 , 1212–1223, doi:10.1038/s41591-022-01819-x (2022).35618837 47 Haghverdi L. , Buettner F. & Theis F. J. Diffusion maps for high-dimensional single-cell analysis of differentiation data. Bioinformatics 31 , 2989–2998, doi:10.1093/bioinformatics/btv325 (2015).26002886 48 Martelli M. P. Novel NPM1 exon 5 mutations and gene fusions leading to aberrant cytoplasmic nucleophosmin in AML. Blood 138 , 2696–2701, doi:10.1182/blood.2021012732 (2021).34343258 49 Zhang X. , Sun J. , Yu W. & Jin J. Current views on the genetic landscape and management of variant acute promyelocytic leukemia. Biomark Res 9 , 33, doi:10.1186/s40364-021-00284-x (2021).33957999 50 Panagopoulos I. Fusion of the FUS gene with ERG in acute myeloid leukemia with t(16;21)(p11;q22). Genes Chromosomes Cancer 11 , 256–262, doi:10.1002/gcc.2870110408 (1994).7533529 51 Thomsen C. , Grundevik P. , Elias P. , Stahlberg A. & Aman P. A conserved N-terminal motif is required for complex formation between FUS, EWSR1, TAF15 and their oncogenic fusion proteins. FASEB J 27 , 4965–4974, doi:10.1096/fj.13-234435 (2013).23975937 52 Dreyling M. H. The t(10;11)(p13;q14) in the U937 cell line results in the fusion of the AF10 gene and CALM, encoding a new member of the AP-3 clathrin assembly protein family. Proc Natl Acad Sci U S A 93 , 4804–4809, doi:10.1073/pnas.93.10.4804 (1996).8643484 53 Borrow J. The translocation t(8;16)(p11;p13) of acute myeloid leukaemia fuses a putative acetyltransferase to the CREB-binding protein. Nat Genet 14 , 33–41, doi:10.1038/ng0996-33 (1996).8782817 54 von Bergh A. R. High incidence of t(7;12)(q36;p13) in infant AML but not in infant ALL, with a dismal outcome and ectopic expression of HLXB9. Genes Chromosomes Cancer 45 , 731–739, doi:10.1002/gcc.20335 (2006).16646086 55 Gamou T. The partner gene of AML1 in t(16;21) myeloid malignancies is a novel member of the MTG8(ETO) family. Blood 91 , 4028–4037 (1998).9596646 56 Guastadisegni M. C. CBFA2T2 and C20orf112: two novel fusion partners of RUNX1 in acute myeloid leukemia. Leukemia 24 , 1516–1519, doi:10.1038/leu.2010.106 (2010).20520637 57 Hitzler J. K. , Cheung J. , Li Y. , Scherer S. W. & Zipursky A. GATA1 mutations in transient leukemia and acute megakaryoblastic leukemia of Down syndrome. Blood 101 , 4301–4304, doi:10.1182/blood-2003-01-0013 (2003).12586620 58 Caligiuri M. A. Molecular rearrangement of the ALL-1 gene in acute myeloid leukemia without cytogenetic evidence of 11q23 chromosomal translocations. Cancer Res 54 , 370–373 (1994).8275471 59 Li Z. Developmental stage-selective effect of somatically mutated leukemogenic transcription factor GATA1. Nat Genet 37 , 613–619, doi:10.1038/ng1566 (2005).15895080 60 Lopez C. K. Ontogenic Changes in Hematopoietic Hierarchy Determine Pediatric Specificity and Disease Phenotype in Fusion Oncogene-Driven Myeloid Leukemia. Cancer Discov 9 , 1736–1753, doi:10.1158/2159-8290.CD-18-1463 (2019).31662298 61 de Rooij J. D. Recurrent abnormalities can be used for risk group stratification in pediatric AMKL: a retrospective intergroup study. Blood 127 , 3424–3430, doi:10.1182/blood-2016-01-695551 (2016).27114462 62 Patel J. P. Prognostic relevance of integrated genetic profiling in acute myeloid leukemia. N Engl J Med 366 , 1079–1089, doi:10.1056/NEJMoa1112304 (2012).22417203 63 Vassiliou G. S. Mutant nucleophosmin and cooperating pathways drive leukemia initiation and progression in mice. Nat Genet 43 , 470–475, doi:10.1038/ng.796 (2011).21441929 64 Yun H. Mutational synergy during leukemia induction remodels chromatin accessibility, histone modifications and three-dimensional DNA topology to alter gene expression. Nat Genet 53 , 1443–1455, doi:10.1038/s41588-021-00925-9 (2021).34556857 65 Lasry A. An inflammatory state remodels the immune microenvironment and improves risk stratification in acute myeloid leukemia. Nat Cancer 4 , 27–42, doi:10.1038/s43018-022-00480-0 (2023).36581735 66 Takeda A. , Goolsby C. & Yaseen N. R. NUP98-HOXA9 induces long-term proliferation and blocks differentiation of primary human CD34+ hematopoietic cells. Cancer Res 66 , 6628–6637, doi:10.1158/0008-5472.CAN-06-0458 (2006).16818636 67 Mullighan C. G. Pediatric acute myeloid leukemia with NPM1 mutations is characterized by a gene expression profile with dysregulated HOX gene expression distinct from MLL-rearranged leukemias. Leukemia 21 , 2000–2009, doi:10.1038/sj.leu.2404808 (2007).17597811 68 Langfelder P. & Horvath S. WGCNA: an R package for weighted correlation network analysis. BMC Bioinformatics 9 , 559, doi:10.1186/1471-2105-9-559 (2008).19114008 69 Ng S. W. A 17-gene stemness score for rapid determination of risk in acute leukaemia. Nature 540 , 433–437, doi:10.1038/nature20598 (2016).27926740 70 Elsayed A. H. A six-gene leukemic stem cell score identifies high risk pediatric acute myeloid leukemia. Leukemia 34 , 735–745, doi:10.1038/s41375-019-0604-8 (2020).31645648 71 Bottomly D. Integrative analysis of drug response and clinical outcome in acute myeloid leukemia. Cancer Cell 40 , 850–864 e859, doi:10.1016/j.ccell.2022.07.002 (2022).35868306 72 Matsuo H. Prognostic implications of CEBPA mutations in pediatric acute myeloid leukemia: a report from the Japanese Pediatric Leukemia/Lymphoma Study Group. Blood Cancer J 4 , e226, doi:10.1038/bcj.2014.47 (2014).25014773 73 Groschel S. Deregulated expression of EVI1 defines a poor prognostic subset of MLL-rearranged acute myeloid leukemias: a study of the German-Austrian Acute Myeloid Leukemia Study Group and the Dutch-Belgian-Swiss HOVON/SAKK Cooperative Group. J Clin Oncol 31 , 95–103, doi:10.1200/JCO.2011.41.5505 (2013).23008312 74 Bill M. Mutational landscape and clinical outcome of patients with de novo acute myeloid leukemia and rearrangements involving 11q23/KMT2A. Proc Natl Acad Sci U S A 117 , 26340–26346, doi:10.1073/pnas.2014732117 (2020).33020282 75 Meshinchi S. Clinical implications of FLT3 mutations in pediatric AML. Blood 108 , 3654–3661, doi:10.1182/blood-2006-03-009233 (2006).16912228 76 Ho P. A. Prevalence and prognostic implications of WT1 mutations in pediatric acute myeloid leukemia (AML): a report from the Children’s Oncology Group. Blood 116 , 702–710, doi:10.1182/blood-2010-02-268953 (2010).20413658 77 Prior I. A. , Lewis P. D. & Mattos C. A comprehensive survey of Ras mutations in cancer. Cancer Res 72 , 2457–2467, doi:10.1158/0008-5472.CAN-11-2612 (2012).22589270 78 Takahashi S. Downstream molecular pathways of FLT3 in the pathogenesis of acute myeloid leukemia: biology and therapeutic implications. J Hematol Oncol 4 , 13, doi:10.1186/1756-8722-4-13 (2011).21453545 79 Mahmoud A. M. Cancer testis antigens as immunogenic and oncogenic targets in breast cancer. Immunotherapy 10 , 769–778, doi:10.2217/imt-2017-0179 (2018).29926750 80 Chen P. Single-Cell and Spatial Transcriptomics Decodes Wharton’s Jelly-Derived Mesenchymal Stem Cells Heterogeneity and a Subpopulation with Wound Repair Signatures. Adv Sci (Weinh) 10 , e2204786, doi:10.1002/advs.202204786 (2023).36504438 81 Jeandidier E. A cytogenetic study of 397 consecutive acute myeloid leukemia cases identified three with a t(7;21) associated with 5q abnormalities and exhibiting similar clinical and biological features, suggesting a new, rare acute myeloid leukemia entity. Cancer Genet 205 , 365–372, doi:10.1016/j.cancergen.2012.04.007 (2012).22867997 82 Zhang J. The genetic basis of early T-cell precursor acute lymphoblastic leukaemia. Nature 481 , 157–163, doi:10.1038/nature10725 (2012).22237106 83 Nilius-Eliliwi V. Broad genomic workup including optical genome mapping uncovers a DDX3X: MLLT10 gene fusion in acute myeloid leukemia. Front Oncol 12 , 959243, doi:10.3389/fonc.2022.959243 (2022).36158701 84 Chisholm K. M. Pathologic, cytogenetic, and molecular features of acute myeloid leukemia with megakaryocytic differentiation: A report from the Children’s Oncology Group. Pediatr Blood Cancer, e30251 , doi:10.1002/pbc.30251 (2023). 85 Perlman E. J. MLLT1 YEATS domain mutations in clinically distinctive Favourable Histology Wilms tumours. Nat Commun 6 , 10013, doi:10.1038/ncomms10013 (2015).26635203 86 Schwartzentruber J. Driver mutations in histone H3.3 and chromatin remodelling genes in paediatric glioblastoma. Nature 482 , 226–231, doi:10.1038/nature10833 (2012).22286061 87 Gutierrez A. & Kentsis A. Acute myeloid/T-lymphoblastic leukaemia (AMTL): a distinct category of acute leukaemias with common pathogenesis in need of improved therapy. Br J Haematol 180 , 919–924, doi:10.1111/bjh.15129 (2018).29441563 88 Brown A. L. RUNX1-mutated families show phenotype heterogeneity and a somatic mutation profile unique to germline predisposed AML. Blood Adv 4 , 1131–1144, doi:10.1182/bloodadvances.2019000901 (2020).32208489 89 Feurstein S. & Godley L. A. Germline ETV6 mutations and predisposition to hematological malignancies. Int J Hematol 106 , 189–195, doi:10.1007/s12185-017-2259-4 (2017).28555414 90 Breiman L. , Friedman J. H. & Olshen R. A. Classification and regression trees. (New York (N.Y.) : Chapman and Hall, 1984). 91 Tarlock K. Significant Improvements in Survival for Patients with t(6;9)(p23;q34)/DEK-NUP214 in Contemporary Trials with Intensification of Therapy: A Report from the Children’s Oncology Group. Blood 138 , 519, doi:10.1182/blood-2021-147576 (2021). 92 Lugthart S. Clinical, molecular, and prognostic significance of WHO type inv(3)(q21q26.2)/t(3;3)(q21;q26.2) and various other 3q abnormalities in acute myeloid leukemia. J Clin Oncol 28 , 3890–3898, doi:10.1200/JCO.2010.29.2771 (2010).20660833 93 Creutzig U. Diagnosis and management of acute myeloid leukemia in children and adolescents: recommendations from an international expert panel. Blood 120 , 3187–3205, doi:10.1182/blood-2012-03-362608 (2012).22879540 94 Lindsley R. C. Acute myeloid leukemia ontogeny is defined by distinct somatic mutations. Blood 125 , 1367–1376, doi:10.1182/blood-2014-11-610543 (2015).25550361 95 Gao Y. Distinct Mutation Landscapes Between Acute Myeloid Leukemia With Myelodysplasia-Related Changes and De Novo Acute Myeloid Leukemia. Am J Clin Pathol 157 , 691–700, doi:10.1093/ajcp/aqab172 (2022).34664628 96 Krivtsov A. V. A Menin-MLL Inhibitor Induces Specific Chromatin Changes and Eradicates Disease in Models of MLL-Rearranged Leukemia. Cancer Cell 36 , 660–673 e611, doi:10.1016/j.ccell.2019.11.001 (2019).31821784 97 Uckelmann H. J. Therapeutic targeting of preleukemia cells in a mouse model of NPM1 mutant acute myeloid leukemia. Science 367 , 586–590, doi:10.1126/science.aax5863 (2020).32001657 98 Heikamp E. B. The menin-MLL1 interaction is a molecular dependency in NUP98-rearranged AML. Blood 139 , 894–906, doi:10.1182/blood.2021012806 (2022).34582559 99 Issa G. C. The menin inhibitor revumenib in KMT2A-rearranged or NPM1-mutant leukaemia. Nature 615 , 920–924, doi:10.1038/s41586-023-05812-3 (2023).36922593 100 Alexander T. B. The genetic basis and cell of origin of mixed phenotype acute leukaemia. Nature 562 , 373–379, doi:10.1038/s41586-018-0436-0 (2018).30209392 101 Wu G. The genomic landscape of diffuse intrinsic pontine glioma and pediatric non-brainstem high-grade glioma. Nat Genet 46 , 444–450, doi:10.1038/ng.2938 (2014).24705251 102 Tian L. CICERO: a versatile method for detecting complex and diverse driver fusions using cancer RNA sequencing data. Genome Biol 21 , 126, doi:10.1186/s13059-020-02043-x (2020).32466770 103 Jakubek Y. A. Large-scale analysis of acquired chromosomal alterations in non-tumor samples from patients with cancer. Nat Biotechnol 38 , 90–96, doi:10.1038/s41587-019-0297-6 (2020).31685958 104 Edmonson M. N. Bambino: a variant detector and alignment viewer for next-generation sequencing data in the SAM/BAM format. Bioinformatics 27 , 865–866, doi:10.1093/bioinformatics/btr032 (2011).21278191 105 Hagiwara K. RNAIndel: discovering somatic coding indels from tumor RNA-Seq data. Bioinformatics 36 , 1382–1390, doi:10.1093/bioinformatics/btz753 (2020).31593214 106 Hagiwara K. , Edmonson M. N. , Wheeler D. A. & Zhang J. indelPost: harmonizing ambiguities in simple and complex indel alignments. Bioinformatics, doi:10.1093/bioinformatics/btab601 (2021). 107 McLaren W. The Ensembl Variant Effect Predictor. Genome Biol 17 , 122, doi:10.1186/s13059-016-0974-4 (2016).27268795 108 Edmonson M. N. Pediatric Cancer Variant Pathogenicity Information Exchange (PeCanPIE): a cloud-based platform for curating and classifying germline variants. Genome Res 29 , 1555–1565, doi:10.1101/gr.250357.119 (2019).31439692 109 Li H. & Durbin R. Fast and accurate short read alignment with Burrows-Wheeler transform. Bioinformatics 25 , 1754–1760, doi:10.1093/bioinformatics/btp324 (2009).19451168 110 Li H. & Durbin R. Fast and accurate long-read alignment with Burrows-Wheeler transform. Bioinformatics 26 , 589–595, doi:10.1093/bioinformatics/btp698 (2010).20080505 111 Wang J. CREST maps somatic structural variation in cancer genomes with base-pair resolution. Nat Methods 8 , 652–654, doi:10.1038/nmeth.1628 (2011).21666668 112 Chen X. CONSERTING: integrating copy-number analysis with structural-variation detection. Nat Methods 12 , 527–530, doi:10.1038/nmeth.3394 (2015).25938371 113 Li H. The Sequence Alignment/Map format and SAMtools. Bioinformatics 25 , 2078–2079, doi:10.1093/bioinformatics/btp352 (2009).19505943 114 Koboldt D. C. VarScan 2: somatic mutation and copy number alteration discovery in cancer by exome sequencing. Genome Res 22 , 568–576, doi:10.1101/gr.129684.111 (2012).22300766 115 Olshen A. B. , Venkatraman E. S. , Lucito R. & Wigler M. Circular binary segmentation for the analysis of array-based DNA copy number data. Biostatistics 5 , 557–572, doi:10.1093/biostatistics/kxh008 (2004).15475419 116 Storey J. D. A direct approach to false discovery rates. Journal of the Royal Statistical Society: Series B (Statistical Methodology) 64 , 479–498, doi:10.1111/1467-9868.00346 (2002). 117 Pounds S. & Cheng C. Robust estimation of the false discovery rate. Bioinformatics 22 , 1979–1987, doi:10.1093/bioinformatics/btl328 (2006).16777905 118 Beroukhim R. Assessing the significance of chromosomal aberrations in cancer: methodology and application to glioma. Proc Natl Acad Sci U S A 104 , 20007–20012, doi:10.1073/pnas.0710052104 (2007).18077431 119 Mermel C. H. GISTIC2.0 facilitates sensitive and confident localization of the targets of focal somatic copy-number alteration in human cancers. Genome Biol 12 , R41, doi:10.1186/gb-2011-12-4-r41 (2011).21527027 120 Karczewski K. J. The mutational constraint spectrum quantified from variation in 141,456 humans. Nature 581 , 434–443, doi:10.1038/s41586-020-2308-7 (2020).32461654 121 Richards S. Standards and guidelines for the interpretation of sequence variants: a joint consensus recommendation of the American College of Medical Genetics and Genomics and the Association for Molecular Pathology. Genet Med 17 , 405–424, doi:10.1038/gim.2015.30 (2015).25741868 122 Abou Tayoun A. N. Recommendations for interpreting the loss of function PVS1 ACMG/AMP variant criterion. Hum Mutat 39 , 1517–1524, doi:10.1002/humu.23626 (2018).30192042 123 Lee K. Specifications of the ACMG/AMP variant curation guidelines for the analysis of germline CDH1 sequence variants. Hum Mutat 39 , 1553–1568, doi:10.1002/humu.23650 (2018).30311375 124 Luo X. ClinGen Myeloid Malignancy Variant Curation Expert Panel recommendations for germline RUNX1 variants. Blood Adv 3 , 2962–2979, doi:10.1182/bloodadvances.2019000644 (2019).31648317 125 Gelb B. D. ClinGen’s RASopathy Expert Panel consensus methods for variant interpretation. Genet Med 20 , 1334–1345, doi:10.1038/gim.2018.3 (2018).29493581 126 Anders S. , Pyl P. T. & Huber W. HTSeq--a Python framework to work with high-throughput sequencing data. Bioinformatics 31 , 166–169, doi:10.1093/bioinformatics/btu638 (2015).25260700 127 Law C. W. , Chen Y. , Shi W. & Smyth G. K. voom: Precision weights unlock linear model analysis tools for RNA-seq read counts. Genome Biol 15 , R29, doi:10.1186/gb-2014-15-2-r29 (2014).24485249 128 Ritchie M. E. limma powers differential expression analyses for RNA-sequencing and microarray studies. Nucleic Acids Res 43 , e47, doi:10.1093/nar/gkv007 (2015).25605792 129 Leek J. T. , Johnson W. E. , Parker H. S. , Jaffe A. E. & Storey J. D. The sva package for removing batch effects and other unwanted variation in high-throughput experiments. Bioinformatics 28 , 882–883, doi:10.1093/bioinformatics/bts034 (2012).22257669 130 Satija R. , Farrell J. A. , Gennert D. , Schier A. F. & Regev A. Spatial reconstruction of single-cell gene expression data. Nat Biotechnol 33 , 495–502, doi:10.1038/nbt.3192 (2015).25867923 131 Butler A. , Hoffman P. , Smibert P. , Papalexi E. & Satija R. Integrating single-cell transcriptomic data across different conditions, technologies, and species. Nat Biotechnol 36 , 411–420, doi:10.1038/nbt.4096 (2018).29608179 132 Stuart T. Comprehensive Integration of Single-Cell Data. Cell 177 , 1888–1902 e1821, doi:10.1016/j.cell.2019.05.031 (2019).31178118 133 Hao Y. Integrated analysis of multimodal single-cell data. Cell 184 , 3573–3587 e3529, doi:10.1016/j.cell.2021.04.048 (2021).34062119 134 Haghverdi L. , Buttner M. , Wolf F. A. , Buettner F. & Theis F. J. Diffusion pseudotime robustly reconstructs lineage branching. Nat Methods 13 , 845–848, doi:10.1038/nmeth.3971 (2016).27571553 135 Angerer P. destiny: diffusion maps for large-scale single-cell data in R. Bioinformatics 32 , 1241–1243, doi:10.1093/bioinformatics/btv715 (2016).26668002 136 Subramanian A. Gene set enrichment analysis: a knowledge-based approach for interpreting genome-wide expression profiles. Proc Natl Acad Sci U S A 102 , 15545–15550, doi:10.1073/pnas.0506580102 (2005).16199517 137 Huang da W. , Sherman B. T. & Lempicki R. A. Systematic and integrative analysis of large gene lists using DAVID bioinformatics resources. Nat Protoc 4 , 44–57, doi:10.1038/nprot.2008.211 (2009).19131956 138 Steen C. B. , Liu C. L. , Alizadeh A. A. & Newman A. M. Profiling Cell Type Abundance and Expression in Bulk Tissues with CIBERSORTx. Methods Mol Biol 2117 , 135–157, doi:10.1007/978-1-0716-0301-7_7 (2020).31960376 139 Elsayed A. H. A 5-Gene Ara-C, Daunorubicin and Etoposide (ADE) Drug Response Score As a Prognostic Tool to Predict AML Treatment Outcome. Blood 134 , 1429, doi:10.1182/blood-2019-128787 (2019). 140 Tartaglia M. , Gelb B. D. & Zenker M. Noonan syndrome and clinically related disorders. Best Pract Res Clin Endocrinol Metab 25 , 161–179, doi:10.1016/j.beem.2010.09.002 (2011).21396583