
==== Front
Plant Commun
Plant Commun
Plant Communications
2590-3462
Elsevier

S2590-3462(24)00214-1
10.1016/j.xplc.2024.100944
100944
Research Article
The genomes of seven economic Caesalpinioideae trees provide insights into polyploidization history and secondary metabolite biosynthesis
Chen Rong 124
Meng Sihan 24
Wang Anqi 2
Jiang Fan 2
Yuan Lihua 23
Lei Lihong 23
Wang Hengchao 2
Fan Wei fanwei@caas.cn
2∗
1 College of Agronomy, Qingdao Agricultural University, Qingdao 266109, China
2 Guangdong Laboratory for Lingnan Modern Agriculture (Shenzhen Branch), Genome Analysis Laboratory of the Ministry of Agriculture and Rural Affairs, Agricultural Genomics Institute at Shenzhen, Chinese Academy of Agricultural Sciences, Shenzhen, Guangdong 518120, China
3 State Key Laboratory of Crop Stress Adaptation and Improvement, School of Life Sciences, Henan University, Kaifeng 475004, China
∗ Corresponding author fanwei@caas.cn
4 These authors contributed equally to this article.

10 5 2024
09 9 2024
10 5 2024
5 9 10094424 1 2024
29 3 2024
8 5 2024
© 2024 The Authors
2024
https://creativecommons.org/licenses/by-nc-nd/4.0/ This is an open access article under the CC BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/).
The Caesalpinioideae subfamily contains many well-known trees that are important for economic sustainability and human health, but a lack of genomic resources has hindered their breeding and utilization. Here, we present chromosome-level reference genomes for the two food and industrial trees Gleditsia sinensis (921 Mb) and Biancaea sappan (872 Mb), the three shade and ornamental trees Albizia julibrissin (705 Mb), Delonix regia (580 Mb), and Acacia confusa (566 Mb), and the two pioneer and hedgerow trees Leucaena leucocephala (1338 Mb) and Mimosa bimucronata (641 Mb). Phylogenetic inference shows that the mimosoid clade has a much higher evolutionary rate than the other clades of Caesalpinioideae. Macrosynteny comparison suggests that the fusion and breakage of an unstable chromosome are responsible for the difference in basic chromosome number (13 or 14) for Caesalpinioideae. After an ancient whole-genome duplication (WGD) shared by all Caesalpinioideae species (CWGD, ∼72.0 million years ago [MYA]), there were two recent successive WGD events, LWGD-1 (16.2–19.5 MYA) and LWGD-2 (7.1–9.5 MYA), in L. leucocephala. Thereafter, ∼40% gene loss and genome-size contraction have occurred during the diploidization process in L. leucocephala. To investigate secondary metabolites, we identified all gene copies involved in mimosine metabolism in these species and found that the abundance of mimosine biosynthesis genes in L. leucocephala largely explains its high mimosine production. We also identified the set of all potential genes involved in triterpenoid saponin biosynthesis in G. sinensis, which is more complete than that based on previous transcriptome-derived unigenes. Our results and genomic resources will facilitate biological studies of Caesalpinioideae and promote the utilization of valuable secondary metabolites.

This study reports chromosome-level genome assemblies of seven Caesalpinioideae species: Gleditsia sinensis, Biancaea sappan, Albizia julibrissin, Delonix regia, Acacia confusa, Leucaena leucocephala, and Mimosa bimucronata. It presents phylogeny, palaeo-polyploidization, and macrosynteny analyses, that provides insights into polyploidization history and chromosome evolution in the Caesalpinioideae subfamily, and the identification of a comprehensive set of genes potentially involved in mimosine metabolism and triterpenoid saponin biosynthesis.

Key words

Caesalpinioideae
hybridization origin
chromosome rearrangement
mimosine biosynthesis genes
triterpenoid saponins
Published: May 10, 2024
==== Body
pmcIntroduction

As the second largest subfamily in Fabaceae (Leguminosae) (Zhao et al., 2021), Caesalpinioideae does not contain any well-known food crops but is instead best known for containing many economically important, ornamental, and rapidly growing shrubs and trees. Important food and industrial trees in this subfamily include Gleditsia sinensis (Zaojiao) and Biancaea sappan (sappan wood). Seeds of G. sinensis are a popular health food for maintaining beauty and youth, similar to peach gum and tragacanth gum, and fruits of G. sinensis can also be used for extraction of natural cleaning materials that are widely used in the shampoo and cosmetics industries (Zhu et al., 2014). The dry heartwood of B. sappan is a natural dye for coloring food, beverages, and clothes (Li et al., 2020). Famous shade and ornamental trees include Albizia julibrissin (silk tree) (Zhang et al., 2021), Delonix regia (flame tree) (Zeng et al., 2020), and Acacia confusa (Taiwan Acacia) (Lin et al., 2018). These trees are widely grown in gardens and on roadsides because of their beautiful leaves and flowers. Familiar pioneer and hedgerow trees include Leucaena leucocephala (white popinac) and Mimosa bimucronata. Because of its fast growth rate, L. leucocephala is widely used as a forage plant and a material for the paper production industry (Rajarajan et al., 2022), and M. bimucronata has traditionally been used as firewood. However, L. leucocephala has shown high invasiveness (Global Invasive Species Database, www.iucngisd.org/gisd), and M. bimucronata has become a notorious invasive plant known for destroying original ecosystems in many areas of the world (Xie et al., 2023). Both utilization and control of Caesalpinioideae plants are therefore necessary. Fabaceae has traditionally been divided into three subfamilies (Papilionoideae, Caesalpinioideae, and Mimosoideae), mainly on the basis of flower structures (Steyermark et al., 1995; Kenicer, 2005; Smýkal et al., 2014). Recent molecular analyses using chloroplast sequences (Azani et al., 2017), plastomes (Zhang et al., 2020), or nuclear genes from sequenced transcriptomes and genomes (Koenen et al., 2020) support a new classification system, LPWG-2017 (The Legume Phylogeny Working Group, year 2017), that contains six subfamilies (Papilionoideae, Caesalpinioideae, Cercidoideae, Detarioideae, Dialioideae, and Duparquetioideae) (Azani et al., 2017; Koenen et al., 2020; Zhang et al., 2020). Caesalpinioideae (ca. 148 genera and 4400 species) includes a large clade, the mimosoids (ca. 41 genera and 3300 species), that was previously the subfamily Mimosoideae. Caesalpinioideae is a monophyletic subfamily of Leguminosae (Azani et al., 2017). The majority of Caesalpinioideae are found in tropical areas, but a small fraction extend to the temperate zone (Legume Data Portal, www.legumedata.org). Leaf types vary across the subfamily, including bipinnate leaves, pinnate leaves, and phyllodes, and inflorescences and flowers are also highly variable (www.legumedata.org). Interestingly, when species in the genera Neptunia and Mimosa are touched, their leaves move and close quickly through a calcium-mediated mechanism (Hagihara et al., 2022), and this feature has helped Mimosa pudica become a model plant for the study of plant movement. More importantly, most Caesalpinioideae plants can fix nitrogen from the atmosphere through symbiotic root nodules, thus playing an essential role in terrestrial ecosystems (Desbrosses and Stougaard, 2011; Huisman and Geurts, 2020).

Compared to that of the largest subfamily, Papilionoideae, the number of sequenced genomes for the subfamily Caesalpinioideae lags far behind. To date, the genomes of only eight species from the mimosoid clade in Caesalpinioideae have been published: Prosopis cineraria (Sudalaimuthuasari et al., 2022), Prosopis alba (Kong et al., 2023), Acacia pycnantha (McLay et al., 2022), Acacia crassicarpa (Massaro et al., 2024), Acacia pachyceras (Habibi et al., 2023), Mimosa pudica (Griesmann et al., 2018; Libourel et al., 2023), Faidherbia albida (Chang et al., 2019), and Entada phaseoloides (Lin et al., 2022). Besides these, the genomes of three species from the Cassia clade in Caesalpinioideae are also available: Chamaecrista fasciculata (Griesmann et al., 2018), Senna siamea (Yang et al., 2023), and Senna tora (Kang et al., 2020). This lack of genomic data greatly hinders comparative genomics and in-depth biological studies of Caesalpinioideae plants. Here, we present chromosome-scale reference genomes of seven economically important Caesalpinioideae trees, G. sinensis, B. sappan, A. julibrissin, D. regia, A. confusa, L. leucocephala, and M. bimucronata, to expand the comparative genomics of Caesalpinioideae, reveal polyploidization events in their evolutionary history, and identify genes involved in their secondary metabolite production.

Results

Haplotype-resolved chromosome-level assembly of seven Caesalpinioideae trees

We generated 93 G (102×), 82 G (94×), 103 G (147×), 94 G (161×), 75 G (132×), 309 G (231×), and 86 G (135×) of HiFi data for G. sinensis, B. sappan, A. julibrissin, D. regia, A. confusa, L. leucocephala, and M. bimucronata, respectively (Supplemental Tables 1 and 2). On the basis of k-mer frequency distribution (Liu et al., 2013), we found that the heterozygosity of G. sinensis, A. julibrissin, D. regia, A. confusa, and M. bimucronata was much higher than that of B. sappan and L. leucocephala (Supplemental Figure 1). Similar to the estimated genome sizes obtained by k-mer analysis (Liu et al., 2013), the total assembled contig sizes of G. sinensis, B. sappan, A. julibrissin, D. regia, A. confusa, L. leucocephala and M. bimucronata were 921 Mb, 872 Mb, 705 Mb, 580 Mb, 566 Mb, 1338 Mb, and 641 Mb, with contig N50 sizes of 42 Mb, 52 Mb, 43 Mb, 36 Mb, 43 Mb, 19 Mb, and 40 Mb, respectively. We used Hi-C data to anchor 96.4%, 97.6%, 96.3%, 93.8%, 97.1%, 97.7%, and 98.8% of the contig sequences onto 14, 12, 13, 14, 13, 56, and 13 chromosome-scale scaffolds for the seven species (Supplemental Figure 2 and Supplemental Table 3). Overall, the mapping rate of HiFi reads to the genome assemblies was greater than 98%, the percentage of complete benchmarking universal single-copy orthologs (BUSCOs) was greater than 99%, and the quality value (QV) calculated by Merqury was greater than 60 for all seven species, indicating the very high accuracy of the assemblies (Table 1; Supplemental Tables 4 and 5).Table 1 Statistics of genome assembly and annotation.

Genomic feature	G. sinensis	B. sappan	A. julibrissin	D. regia	A. confusa	L. leucocephala	M. bimucronata	
Genome assembly	
	
Estimated genome size (Mb)	916	889	729	571	580	1338	643	
Total assembly size (bp)	921 423 504	871 958 265	705 260 536	580 467 956	565 633 385	1 338 184 995	640 542 845	
Contig N50 (bp)	42 062 168	52 252 475	43 102 934	35 809 352	42 518 953	19 491 048	39 843 024	
Scaffold N50 (bp)	68 211 918	69 792 917	55 539 299	37 307 771	42 518 953	23 912 662	49 009 160	
Anchored to chromosomes (%)	96.43	97.61	96.34	93.75	97.06	97.66	98.77	
Complete BUSCOs in genome (%)	98.90	99.30	99.40	99.40	99.20	99.70	99.40	
LAI (LTR assembly index)	12.90	25.22	12.67	9.25	12.89	15.18	13.57	
QV calculated by Merqury	67.02	64.11	71.72	67.72	67.96	69.65	70.46	
	
Genome annotation	
	
Tandem repeats (%)	9.82	6.96	6.64	14.33	5.44	6.98	7.10	
TE sequences (%)	74.41	69.97	66.52	56.21	60.66	58.56	64.36	
No. of tRNA genes	631	873	847	580	780	1785	701	
No. of rRNA genes	4873	2886	2552	9696	1809	6108	1530	
No. of protein-coding genes	39 698	40 555	40 866	39 976	38 699	98 316	33 062	
Total CDS percentage in genome (%)	4.96	5.08	6.08	7.32	7.39	8.79	5.95	
Complete BUSCOs in gene set (%)	98.20	97.40	98.00	98.40	98.90	99.60	98.00	

In total, 39 698, 40 555, 40 866, 39 976, 38 699, 98 316, and 33 062 protein-coding gene models were predicted, with coding sequences (CDSs) covering 46 Mb (4.96%), 44 Mb (5.08%), 43 Mb (6.08%), 43 Mb (7.32%), 42 Mb (7.39%), 118 Mb (8.79%), and 38 Mb (5.95%) of the G. sinensis, B. sappan, A. julibrissin, D. regia, A. confusa, L. leucocephala, and M. bimucronata genomes, respectively (Supplemental Tables 6–8). The percentages of complete BUSCOs in the gene sets of the seven species (97.4%–99.6%) were comparable to those in the genome assemblies (98.9%–99.7%) (Simao et al., 2015), suggesting the high completeness of the gene annotations. To facilitate functional analysis, 34 775, 35 087, 34 578, 35 243, 33 670, 91 513, and 30 564 genes in the seven species were annotated by at least one of the National Center for Biotechnology Information Non-Redundant (NCBI-NR), Kyoto Encyclopedia of Genes and Genomes (KEGG), InterPro, and Gene Ontology (GO) databases (Supplemental Table 9). We also identified 631, 873, 847, 580, 780, 1785, and 701 tRNA genes and 4873, 2886, 2552, 9696, 1809, 6108, and 1530 rRNA genes in the seven species, respectively (Table 1; Supplemental Tables 10 and 11).

Mimosoids have evolved faster than other clades in Caesalpinioideae

To study the evolution of Caesalpinioideae (Manzanilla and Bruneau, 2012), we clustered the reference genes of G. sinensis, B. sappan, A. julibrissin, D. regia, A. confusa, L. leucocephala, M. bimucronata, and seven published Caesalpinioideae species (P. cineraria, A. pycnantha, M. pudica, F. albida, E. phaseoloides, C. fasciculata, and S. tora) into 37 030 orthologous groups (OGs), each of which contained at least two genes (Supplemental Tables 12 and 13), and 639 single-copy OGs were used for phylogenetic inference. Although the topology of the resulting phylogenetic tree was consistent with previous studies (Manzanilla and Bruneau, 2012), the branch lengths of mimosoids were much longer than those of other clades in Caesalpinioideae, suggesting that mimosoid species may have evolved at a much faster rate than the other Caesalpinioideae species (Figure 1A). This may have been caused by differences in mutation rates, generation times, population sizes, and other factors.Figure 1 Phylogenetic and time tree of sequenced species in Caesalpinioideae.

(A) Species tree built from 639 single-copy orthologous groups.

(B) Time tree constructed by the RelTime-ML method in MEGA. The outgroup Vitis vinifera is not shown, and the seven species sequenced in this study are marked with stars. The numbers next to internal nodes indicate divergence times. The blue circles represent inferred whole-genome duplication (WGD) events.

Using two calibration constraints, we estimated the divergence times among the seven sequenced species in this study (Figure 1B and Supplemental Figure 3). Within the mimosoid clade, the two Acacia species A. confusa and A. pycnantha diverged from each other at 5.0–12.5 million years ago (MYA). The genus Albizia is a sister clade to Acacia, diverging at 6.7–15.2 MYA. The two Mimosa species M. bimucronata and M. pudica diverged from each other at 9.1–17.3 MYA. L. leucocephala diverged from the latest common ancestor of Acacia, Albizia, and Mimosa species at 17.5–33.2 MYA. Within the subfamily Caesalpinioideae, D. regia, B. sappan, and G. sinensis diverged from the mimosoid clade at 36.9–66.6 MYA, 44.1–69.1 MYA, and 54.0–84.5 MYA, respectively.

All seven plants share an ancient WGD at the origin of Caesalpinioideae

Whole-genome duplication (WGD), a key evolutionary event that drives phenotypic complexity, functional novelty, and ecological adaptation, is widely found in flowering plants (Van de Peer et al., 2009). For Papilionoideae, the largest subfamily in Fabaceae, a lineage-specific WGD event termed PWGD around 55.0 MYA has been reported in several studies (Cannon et al., 2015). A species-specific WGD event termed the G-LD has also been found for soybean, the most important plant in Papilionoideae (Schmutz et al., 2010). However, because of the small number of sequenced genomes and poor assembly quality for the second largest subfamily, Caesalpinioideae, the existence of a shared WGD event in all Caesalpinioideae species remains controversial.

In 2022, Lin et al. reported an ancient WGD event in E. phaseoloides at about 72.8 MYA on the basis of a chromosome-scale genome assembly and an analysis of 4DTV distribution (Lin et al., 2022). Subsequently, McLay et al. observed two Ks peaks in three Caesalpinioideae plants: A. pycnantha and P. alba both had a Ks peak at ∼0.80, and S. tora had a Ks peak at ∼0.60. However, the authors could not conclude whether these three species shared the same WGD event (McLay et al., 2022). In the present study, we observed Ks peaks at ∼0.80 for all four mimosoids (A. julibrissin, A. confusa, L. leucocephala, and M. bimucronata), consistent with the Ks peaks previously reported for two other mimosoids (A. pycnantha and P. alba) (McLay et al., 2022). We also observed Ks peaks at ∼0.50 for both B. sappan and D. regia and a Ks peak at ∼0.40 for G. sinensis (Figure 2A). Taking these data together, the evolutionary rates of S. tora, B. sappan, D. regia, and G. sinensis are much lower than those of the mimosoid plants, resulting in differences in Ks peak values among these species that are proportional to the tree branch lengths estimated from protein substitutions (Figure 1A). Therefore, all these Ks peaks are likely to originate from the same WGD event. This Caesalpinioideae lineage-specific WGD (CWGD) at the origin of Caesalpinioideae occurred at ∼72.0 MYA, after the divergence of Papilionoideae and Caesalpinioideae at about 81.9–93.6 MYA (Kang et al., 2020). Furthermore, the large-scale intra-species macrosynteny observed for all seven Caesalpinioideae plants using MCScanX strongly confirms the existence of the CWGD (Figure 2B–2H and Supplemental Table 14). Ks distribution and macrosynteny analyses also revealed the presence of this CWGD in other published genomes (Supplemental Figure 4), indicating that the CWGD is shared in Caesalpinioideae.Figure 2 Whole-genome duplication event in Caesalpinioideae.

(A) Distribution of synonymous substitution rates (Ks) for orthologous genes in seven species sequenced in this study and Glycine max (black dashed line). One subgenome (A1) of L. leucocephala was selected for analysis. CWGD represents a shared WGD event in the Caesalpinioideae subfamily, G-LD is a self-duplication event in Glycine max, and PWGD denotes a shared WGD event in the Papilionoideae subfamily.

(B–H) Circular collinearity plots for the seven species, using collinear segments with ten or more homologous gene pairs; each segment is depicted in a random color.

In addition to WGD, dispersed duplication also increases the number of gene copies related to specific functions (Nezamivand-Chegini et al., 2021). Here, we identified dispersed duplicated genes from the MCScanX results (Wang et al., 2012) and performed GO enrichment analyses using TBtools (Chen et al., 2023). For each species, the GO terms were sorted by their adjusted p values, and the top 50 most enriched biological process GO terms in the dispersed duplicated genes were selected (Supplemental Figure 5). Interestingly, genes involved in the organonitrogen compound biosynthetic process or the cellular nitrogen compound metabolic process were significantly expanded in the dispersed duplicated genes for all seven species, indicating that these genes may influence the nitrogen-fixing ability of Caesalpinioideae plants.

Two recent successive WGD events produced octoploid L. leucocephala

The Leucaena genus, a member of the mimosoid clade in Caesalpinioideae, includes 19 tetraploid species (2n = 52 or 56) and 5 octoploid species (2n = 104 or 112). However, most mimosoid species are diploid (2n = 26 or 28), suggesting that the haploid base number (x) for mimosoids is 13 or 14 (Abair et al., 2019). Our assembly of L. leucocephala included 112 haplotype-resolved chromosomes, which was eight times the mimosoid base number of 14, suggesting that L. leucocephala is an octoploid. After removing the heterozygous redundancy, we obtained 56 chromosomes to represent the haploid reference genome. By intra-species alignment, it was easy to cluster the 56 haploid chromosomes into 14 groups, each containing four homologous chromosomes. High levels of macrosynteny were observed among the four homologous chromosomes in each group, with only several local inversions in some chromosomes (Figure 3A). We then analyzed the distribution of Ks values for intra-species paralogous gene pairs. In addition to the Ks peak at ∼0.80 (ancient CWGD) that was shared by all Caesalpinioideae plants, two relatively recent Ks peaks at ∼0.15 and ∼0.05 were observed, indicating that two independent WGD events occurred in the Leucaena lineage (LWGD-1 and LWGD-2), thus explaining the octoploidy of L. leucocephala (Figure 3B).Figure 3 Hybridization origin of L. leucocephala.

(A) Collinearity analysis of the four subgenomes (A1B1A2B2) of L. leucocephala, with intra-chromosomal inversions highlighted in gray.

(B) Ks distribution of homologous genes for L. leucocephala (red line) and G. max (black dashed line). The two peaks on the left side of L. leucocephala represent two WGD events specific to L. leucocephala.

(C) Divergence time tree of four syntenic chromosomes in L. leucocephala. Prosopis cineraria and Senna tora were used as outgroups. The range of divergence time for 95% HPD is shown for each inner node.

To determine the relationships among the four homologous chromosomes, we reconstructed the phylogeny and estimated the divergence times based on a concatenated multiple sequence alignment of single-copy genes for each of the 14 chromosome groups, using P. cineraria and S. tora as outgroups. The results showed that the four homologous chromosomes were assigned to a phylogenetic topology of ((A1,B1),(A2,B2)), which was formed by LWGD-1 and LWGD-2 successively (Figure 3C and Supplemental Figure 5). LWGD-1 occurred at ∼16.2–19.5 MYA, possibly around the origin of Leucaena, making the Leucaena ancestor a tetraploid derived from a diploid. We noticed that Donovan et al. sequenced another species in this genus, Leucaena trichandra (2n = 52), and they also reported a WGD event at the Botany 2021 conference (http://2021.botanyconference.org). Thus, the WGD event found in L. trichandra is probably equivalent to LWGD-1 in L. leucocephala, suggesting that this shared WGD event possibly predated the diversification of extant Leucaena species. LWGD-2 occurred at ∼7.1–9.5 MYA, making L. leucocephala an octoploid species derived from a tetraploid ancestor. However, we are unable to determine whether LWGD-1 and LWGD-2 were auto-WGD or allo-WGD events using the limited evidence available.

Genome-size contraction and gene loss during diploidization of L. leucocephala

Although the haploid genome size of L. leucocephala is as large as 1.3 G, this represents the sum of four chromosome sets corresponding to 56 chromosomes. Considering only one set of 14 chromosomes, the basic genome size is only ∼325 Mb, which, on average, is half of the haploid genome size of the other six studied plants. We next investigated the content of transposable elements (TEs), which make the largest contribution to the genome sizes of most plants (Galindo-González et al., 2017), and found D. regia and G. sinensis had the lowest and highest amount of TEs, respectively. The distributions of TE families and subfamilies were somewhat similar among the studied species, with long terminal repeat (LTR) TEs, especially Gypsy elements, showing the greatest differences. Interestingly, we found that LTR TEs mostly determined the genome sizes of all six diploid species except for L. leucocephala (Figure 4A–4C; Supplemental Tables 15 and 16). The TE content of L. leucocephala was not the lowest: it was comparable to that of A. confusa and even higher than that of D. regia. Analysis of LTR insertion times also showed that L. leucocephala has still experienced high LTR activity in the recent historical period after the two successive WGD events (Figure 4D). Therefore, TEs are not the major reason for the genome-size contraction of L. leucocephala.Figure 4 Factors influencing genome-size variation.

(A) Relationship between TE length and genome size for the sequenced species.

(B) Content distribution of TEs across various classes (LTR, DNA, TIR, LINE, MITE, SINE).

(C) Content distribution of different subclasses of long terminal repeat (LTR) elements (Gypsy, Copia, unknown).

(D) Insertion times of LTR elements in the sequenced species.

(E) Relationship between the number of genes and genome size of the sequenced species.

(F) Numbers of genes in sequenced species.

For L. leucocephala, only one subgenome was used for all analyses.

Previous studies have shown that all chromosomes of L. leucocephala pair into bivalents at meiosis (Boff and Schifino-Wittmann, 2003), indicating that this species has successfully completed diploidization from the polyploid stage. During the diploidization process, random gene losses and biased retention of genes on each ploid offer a unique evolutionary opportunity for the evolution of key innovations (Van De Peer et al., 2017; Yuan and Song, 2023). We found that the basic genome (325 Mb) of L. leucocephala encodes only ∼25 000 protein-coding genes, many fewer than those of the other six diploid species (∼33 000–41 000) (Figure 4E and 4F), indicating that ∼40% of the genes in L. leucocephala have been lost during the diploidization process after the two successive WGD events. We next identified pseudogenes in each species by aligning their protein sequences to the genome using Exonerate and choosing gene models with frameshifts (Slater and Birney, 2005). We found that L. leucocephala had many more pseudogenes than the other six species. The ratio of pseudogenes to genes in L. leucocephala (50%) was much higher than those of the other species (16%–23%) (Supplementary Table 17), indicating that a certain degree of gene loss was caused by pseudogenization in addition to gene deletion. Considering that gene regions occupy only a small part of the genome and that DNA fragment deletion often exceeds gene regions, we inferred that gene deletion together with deletion of neighboring DNA fragments was likely to be the major reason for the drastic contraction in genome size of L. leucocephala.

Fusion and breakage of an unstable chromosome explain differences in basic chromosome number

All species in the mimosoid clade and most species in Caesalpinioideae have a base chromosome number of 13 or 14 (Abair et al., 2019), but the reason for this discrepancy has been unclear. With chromosome-level assemblies for a broad range of representative Caesalpinioideae species, we were able to compare macrosynteny within and between the three species with base number 13 (A. confusa, A. julibrissin, and M. bimucronata) and the three species with base number 14 (L. leucocephala, D. regia, and G. sinensis). In general, the extent of chromosome structural conservation was inversely proportional to divergence time. Because species in the mimosoid clade diverged from each other much later than from D. regia, G. sinensis, and B. sappan, chromosome structure is much more conserved within the mimosoids than between the mimosoid clade and other clades in Caesalpinioideae. Most chromosomes of all these Caesalpinioideae species have one-to-one relationships with only intra-chromosomal local inversions, indicating that the structure of these chromosomes is more conserved in comparison with the other chromosomes that do not have one-to-one relationships. Interestingly, we found that two chromosomes in each of the three species with base number 14 completely corresponded to one chromosome in each of the three species with base number 13, suggesting that a single chromosome fusion or breakage event occurred when the two species groups diverged (Figure 5). Considering that the three species with base number 14 diverged much earlier in Caesalpinioideae than did the three species with base number 13, we inferred that the Caesalpinioideae ancestor possibly had 14 chromosomes and that two chromosomes fused into one around the origin of the mimosoid clade.Figure 5 Chromosome conservation and rearrangements in Caesalpinioideae.

Macrosyntenic blocks among 13 chromosomes of A. confusa, A. julibrissin, and M. bimucronata, 14 chromosomes of L. leucocephala, D. regia, and G. sinensis, and 12 chromosomes of B. sappan. The process of changes in chromosome numbers of 13 and 14 is highlighted in gray, and the changes in chromosome numbers of 12 and 14 are highlighted in yellow.

Within the mimosoids (A. confusa, A. julibrissin, M. bimucronata, and L. leucocephala), most species have the basic chromosome number of 13, and their chromosome structures are highly similar and well conserved. Although the basic chromosome number of our assembled L. leucocephala genome is 14, basic chromosome numbers of 13 and 14 are both present throughout the Leucaena genus (Abair et al., 2019), suggesting that the fused chromosome at the origin of the mimosoid clade is not stable and may have again broken into two chromosomes in some later lineages, thus explaining the different basic chromosome numbers of mimosoid species. Many more intra-chromosomal inversions and occasional inter-chromosomal translocations have occurred in the three non-mimosoid species (D. regia, G. sinensis, and B. sappan). Moreover, two pairs of chromosomes from the 14 Caesalpinioideae ancestral chromosomes have merged, leading to the 12 chromosomes of B. sappan (Figure 5; Supplemental Figures 7 and 8).

Abundant copies of mimosine biosynthesis genes in L. leucocephala contribute to high mimosine production

Widely distributed in mimosoid species, L-mimosine is a non-protein amino acid with defense functions against invading viruses and allelopathic effects that inhibit other plants (Kato-Noguchi and Kurniadie, 2022). Over the past years, mimosine has also shown various biological activities that are valuable for human health, such as anti-cancer, anti-inflammatory, anti-fibrosis, and anti-influenza activities (Nguyen and Tawata, 2016). L-mimosine metabolism involves three key enzymes: (1) serine acetyltransferase (SAT) (EC 2.3.1.30) catalyzes the formation of O-acetyl-L-serine from serine and acetyl-coenzyme A (Howarth et al., 2003); (2) mimosine synthase (EC 2.5.1.52) catalyzes the formation of L-mimosine from O-acetyl-L-serine and 3,4-dihydroxypyridin (Harun-Ur-Rashid et al., 2018); and (3) mimosinase (EC 4.3.3.8) degrades L-mimosine into 3H4P and pyruvate (Negi et al., 2014) (Figure 6A). A few genes involved in the biosynthesis and degradation of mimosine have been reported in M. pudica and L. leucocephala (Negi et al., 2014; Harun-Ur-Rashid et al., 2018); however, most genes have not been identified in most mimosoid species.Figure 6 Genes associated with mimosine biosynthesis and decomposition.

(A) Roles of serine acetyltransferase, mimosine synthase, and mimosinase in L-mimosine metabolism.

(B) Numbers of serine acetyltransferase, mimosine synthase, and mimosinase genes in the sequenced species.

(C and D) Phylogenetic trees of serine acetyltransferase and mimosine synthase for the sequenced species, using homologous genes from Vitis vinifera (grape) as outgroups. The star represents L. leucocephala.

In this study, we identified 16, 4, 4, 4, 3, 3, and 2 SAT genes and 18, 12, 10, 9, 8, 7, and 6 mimosine synthase genes in L. leucocephala, A. confusa, A. julibrissin, M. bimucronata, G. sinensis, D. regia, and B. sappan, respectively (Figure 6B and Supplemental Table 16). Overall, the four mimosoid species had more copies of SAT and mimosine synthase genes than the three non-mimosoid species. In particular, L. leucocephala had the highest gene number contributed by its two recent successive WGD events (Figure 6C and 6D). Although mimosine is widely distributed in mimosoid species, mimosine content differs widely. In L. leucocephala, mimosine constitutes 3%–5% of the seeds and 3%–6% of the leaves, much more than in other mimosoid plants. Therefore, L. leucocephala has been chosen for the industrial extraction of mimosine (Soedarjo and Borthakur, 1998). The abundance of gene copies involved in mimosine biosynthesis may be the major explanation for the high mimosine production of L. leucocephala. Furthermore, 3, 7, 4, and 3 mimosinase genes were identified in L. leucocephala, A. confusa, A. julibrissin, and M. bimucronata, respectively (Supplemental Figures 9 and 10; Supplemental Table 18). No mimosinase genes were identified in the three non-mimosoid plants, indicating that the mimosinase gene seems to play a decisive role in whether mimosoid substances are produced, consistent with the fact that mimosine metabolism occurs only in the mimosoid clade of Caesalpinioideae. L. leucocephala had the lowest copy number (3) of mimosinase genes compared with the other three mimosoid species (3–7). The lower mimosinase gene number in L. leucocephala may decrease mimosine degradation and help to retain more mimosine in L. leucocephala. Taken together, the higher number of SAT genes and mimosine synthase genes and the lower number of mimosine genes may help to explain why L. leucocephala has the highest mimosine content in the mimosoid clade.

A more complete set of potential genes for triterpenoid saponin biosynthesis

Saponins, well known for their surfactant properties, are one of the most important classes of secondary metabolites produced by plants (Yao et al., 2020). Triterpenoid saponins, found mainly in leguminous plants, are composed of a triterpene aglycone linked to one, two, or three saccharide chains of varying size and complexity. The fruit of G. sinensis contains a large amount of triterpenoid saponins (Zhang et al., 1999), making it a useful source of natural cleaning materials. The biosynthesis of triterpenoid saponins involves three sequential steps: (1) synthesis of the triterpenoid backbone β-amyrin, based on isopentenyl diphosphate, by the sequential activities of isopentenyl-diphosphate δ-isomerase (IDI, EC 5.3.3.2), geranyl-diphosphate synthase (GPPS, EC 2.5.1.1), farnesyl-diphosphate synthase (FDPS, EC 2.5.1.10), squalene synthase (SS, EC 2.5.1.21), squalene epoxidase (SE, EC 1.14.14.17), and β-amyrin synthase (AS, EC 5.4.99.39); (2) modification of triterpenoids via hydroxylation by cytochrome P450 monooxygenases (P450s) to increase structural diversity; and (3) glycosidation by UDP-glycosyltransferases (UGTs). Currently, a fraction of the genes involved in biosynthesis of triterpenoid saponins have been identified from transcriptome data in G. sinensis; however, the sequence data have not been released (Kuwahara et al., 2019).

In this study, we identified 125 potential genes associated with triterpenoid saponin biosynthesis in G. sinensis. Among them, 36 (29%) belong to multiple gene families responsible for β-amyrin biosynthesis, 54 (43%) belong to the P450 family, and 35 (28%) belong to the UGT family (Figure 7A; Supplemental Tables 19 and 20). Although the number of β-amyrin biosynthesis genes is comparable to that in the previous transcriptome study, the numbers of P450 genes and UGT genes were 2.3-times and 3.5-times greater, respectively (Kuwahara et al., 2019). Notably, a new P450 subfamily, CYP93E, containing eight genes, was identified only in our study. In summary, owing to the huge advantage of complete genome data versus transcriptome data, we identified a more complete set of potential genes involved in triterpenoid saponin biosynthesis (Figure 7B and 7C; Supplemental Figures 11 and 12).Figure 7 Genes associated with triterpenoid saponin biosynthesis.

(A) Comparison of the number of potential P450 and UGT genes identified here with the number previously identified from transcriptome data.

(B) Expression heatmap of potential P450 genes.

(C) Expression heatmap of potential UGT genes.

Different genes exhibit various expression levels in different plant parts (fruit, leaf, root, stem, and thorn), with high expression observed in fruit for most genes.

Discussion

In this study, we generated chromosome-level genome assemblies and high-quality gene annotations for seven economically important Caesalpinioideae trees, including six diploids and one octoploid. The assembly qualities are nearly telomere-to-telomere level, with only a few gaps in each constructed chromosome. For the six diploid species, the genome sizes ranging from 560 to 920 Mb were mostly determined by LTR TEs. Phylogenetic inference for the studied plants was performed using single-copy gene families, showing that the mimosoid clade has a much higher evolutionary rate than other clades of Caesalpinioideae. By macrosynteny comparisons within and between two groups of representative species with base chromosome numbers of 13 and 14, we found that fusion and breakage of a single chromosome was mostly responsible for differences in base chromosome numbers among Caesalpinioideae, suggesting that chromosome rearrangements are not random. These genomic resources will promote evolutionary and comparative genomics studies in Caesalpinioideae.

The Ks distribution and intra-species macrosynteny of the seven studied plants both showed that all Caesalpinioideae species share an ancient WGD (CWGD) that occurred at ∼72.0 MYA at the origin of the Caesalpinioideae lineage. Besides the CWGD event, L. leucocephala has undergone two recent successive WGDs, LWGD-1 (16.2–19.5 MYA) and LWGD-2 (7.1–9.5 MYA), resulting in an octoploid genome. With the current genomic evidence, it is difficult to know whether LWGD-1 and LWGD-2 were auto-WGD or allo-WGD events. A previous study with a few genetic markers suggested that L. leucocephala originated from hybridization of two parental tetraploids, Leucaena pulverulenta (2n = 56, maternal) and Leucaena cruziana (2n = 52, paternal) (Govindarajulu et al., 2011). If that is true, LWGD-2 should be an allo-WGD event. However, much more solid evidence is needed to validate this hypothesis. During the diploidization process, the gene number decreased from ∼40 000 to ∼25 000 in each basic ploid (14 chromosomes). Gene loss was caused by pseudogenization and deletion, and gene deletion together with deletion of other genomic regions largely explains the genome-size contraction of L. leucocephala. Random gene loss and biased retention of genes on each ploid offer a unique evolutionary opportunity for the evolution of key innovations, improving the adaptability of L. leucocephala to changing environments.

Caesalpinioideae plants produce many valuable secondary metabolites such as mimosine and triterpenoid saponins. Here, we identified all the gene copies encoding SAT, mimosine synthase, and mimosinase in the seven studied plants and found that the high abundance of mimosine biosynthesis genes in L. leucocephala may largely explain its high mimosine production. In addition, we identified 125 genes potentially involved in triterpenoid saponin biosynthesis in G. sinensis, a more complete gene set than that identified in a previous transcriptome study. G. sinensis, B. sappan, and A. julibrissin are traditional medicinal plants used in East Asia for thousands of years. Thorns of G. sinensis, which are found on the trunk and branches, have been widely used as an antimicrobial and antipruritic drug; the dry heartwood of B. sappan can improve blood circulation and has therefore been widely used in wound recovery, hemoptysis, postpartum hemorrhage, and as an emmenagogue; and the flowers and stem bark of A. julibrissin can treat anxiety, depression, and sleep problems. Thus, the genomic resources in this study will also facilitate studies of the biosynthesis of plant medicinal components, which will benefit human health.

Methods

Plant materials and sequencing

We selected a wild M. bimucronata tree growing on the riverside near the chicken-raising houses on the farm of the Agriculture Genomics Institute of Shenzhen (AGIS) in Dapeng District (Shenzhen, Guangdong, China), a cultivated D. regia tree growing in the green belt of Pengfei road in Dapeng District, and a cultivated L. leucocephala tree growing on the roadside of Ruipeng Dadao in Dapeng District. We obtained an A. confusa (Guangdong cultivar) seedling from Guangdong Academy of Forestry (Guangzhou, Guangdong, China), a G. sinensis (huge-thorn cultivar) seedling from Chuanyu Garden company (Xuzhou, Jiangsu, China), and a B. sappan (Guangxi cultivar) seedling from Leaf Horticulture company (Baise, Guangxi, China). We obtained the seeds of A. julibrissin (Chinese cultivar) from Qingfeng seed industry company (Bangbu, Anhui, China) and grew the seeds in a plant growth chamber. We ensured the accuracy of species identification by morphological observations of leaves, stems, flowers, and seeds, communication with the specimen providers, and consistency in genome size and chromosome number between genome assembly statistics and experimental karyotyping results. All the certificated specimens for these sequenced plants were kept in the AGIS, and the specimens are freely available from AGIS for verification of research results or use in downstream studies.

Genomic DNA was extracted from young leaves of a single plant from each species using the Hi-DNAsecure Plant Kit (TIANGEN DP350, China). High-quality genomic DNA was used to prepare 20-kb-insert sequencing libraries with the SMRTbell Express Template Prep Kit 2.0 (PacBio, USA), which were then sequenced on the Sequel II platform in CCS mode (PacBio). Young leaves from the same plants were also used for Hi-C sequencing on the Illumina NovaSeq platform in PE150 mode. Total RNA was extracted from roots, stems, and leaves of each species using the RNeasy Plant Mini Kit (QIAGEN, Germany). The extracted RNA was pooled, and full-length cDNAs were sequenced on the Sequel II platform in Iso-Seq mode (PacBio).

Genome assembly and annotation

K-mer frequency counting and genome-size estimation were performed by genome character estimation (Binghang et al., 2013), with the input of all generated HiFi reads for each species. HiFi reads were filtered to retain those with an overall accuracy above 0.99 and a length greater than 6 kb, and Hi-C reads were filtered to retain those with an overall accuracy above 0.99. Hifiasm v.0.19.5 (Feng et al., 2022) with the parameters “-h1 -h2” was used to assemble the contigs. Minimap2 (Li, 2018) with the parameters “--secondary = no -cx asm20” was used to align the downloaded chloroplast and mitochondrial genomes from NCBI to the assembled contigs, and unaligned contigs were considered to be nuclear contigs. Hi-C reads were mapped to the nuclear contigs using HiC-Pro v.3.1.0 (Servant et al., 2015) to generate a Hi-C contact matrix among contig bins. On the basis of Hi-C linkage information, nuclear contigs were assembled into chromosome-level scaffolds using Endhic v.1.0 (Wang et al., 2022). Assembly quality was evaluated using BUSCO v.5.0 (Manni et al., 2021) with the 1614 conserved genes of the embryophyta lineage; QV was calculated using Merqury v.1.3 (Rhie et al., 2020); and the LTR assembly index (LAI) was calculated using LTR_retriever v2.9.0 (Ou and Jiang, 2018).

Tandem repeats were annotated using TRF v.4.07 (Benson, 1999) with the parameters “2 5 7 80 10 50 2000 -h -d ”. TE annotation involved three steps. First, EDTA v.2.0.0 (Ou et al., 2019) was used to predict structurally intact transposon components (TEs, including LTR-RTs, DNA transposons, Helitrons, etc.) and simultaneously generate a complete TE library. Second, with the aforementioned complete TE library and the Repbase database v.26.05, RepeatMasker v.4.1.2 (http://repeatmasker.org/RepeatMasker/) was used to identify homologous TEs in the genome assembly. Finally, a de novo TE library was constructed using RepeatModeler v.2.0.2 (Flynn et al., 2020) and classified with TERL v.1.0 (Da Cruz et al., 2021). RepeatMasker was used with the de novo TE library to identify additional TEs in the genome assembly.

Protein-coding gene models were predicted with Augustus version 3.4.0 (Stanke et al., 2006) from the TE-masked genome, in which TEs with lengths greater than 80 bp were soft-masked. The training set for Augustus was generated by BUSCO evaluation of the genome. Merged transcripts from root, stem, and leaf tissues were aligned to their corresponding genomes using GMAP v.2020-10-27 (Wu and Watanabe, 2005) to generate support hint files for full-length transcripts. Publicly available protein sequences from species in the Caesalpinioideae (A. pycnantha, C. fasciculata, E. phaseoloides, F. albida, M. pudica, P. cineraria, and S. tora) were aligned independently to the seven genomes in this study using Exonerate v.2.2.0 (Slater and Birney, 2005) with the parameters “--percent 50 --showcigar TRUE --showtargetgff TRUE”, and only the best two alignment results were selected as support hint files for homologous proteins. The quality of the predicted gene set was assessed using BUSCO v.5.0 (Manni et al., 2021) with the embryophyta_odb10 database. Functional annotations of protein-coding genes were obtained by searching the NCBI-NR and KEGG databases using DIAMOND v.0.9.24 (Buchfink et al., 2015) and the InterPro database using InterProScan v.5.52-86 (Jones et al., 2014). rRNA genes were predicted with RNAmmer v.1.2 (Lagesen et al., 2007), and tRNA genes were predicted with tRNAScan-se version 2.0 (Lowe and Chan, 2016). Pseudogenes were identified by aligning the known protein sequences of seven species to the genomic sequences using Exonerate v.2.2.0. Predicted genes with internal stop codons and frameshifts were classified as pseudogenes.

Phylogeny and paleo-polyploidization analysis

OrthoFinder v.2.5.2 (Emms and Kelly, 2019) with the parameters “-M msa -A mafft -1 -y” was used to construct OGs using the reference protein sets of the seven species sequenced in the current study (A. confusa, A. julibrissin, B. sappan, D. regia, G. sinensis, M. bimucronata, and L. leucocephala), seven other species in Caesalpinioideae (A. pycnantha (NCBI: ASM2556357v1), C. fasciculata (NCBI: ASM325492v1), E. phaseoloides (CNSA: CNA0029377), F. albida (ORCAE), M. pudica (Mimpud_MpudA1P6v1), P. cineraria (NCBI: ASM2901754v1), and S. tora (NCBI: ASM1485142v1), and the outgroup Vitis vinifera. Single-copy OGs were selected from all OGs and were required to contain four genes from L. leucocephala (octoploid), two from M. pudica (tetraploid), and one each from the other species (diploid). One representative gene from the single-copy OGs was randomly chosen for L. leucocephala and M.pudica . The protein sequences in each single-copy OG were then aligned using Muscle v.3.8.31 (Edgar, 2004), and all alignment results were concatenated. The concatenated alignment was used to construct a species tree using Raxml-NG v.1.0.3 with the GTR+G model (Kozlov et al., 2019). Divergence times were estimated using the RelTime-ML method in MEGA11 with all other options in default settings (Tamura et al., 2021). Two calibration constraints were used: 35.0–45.0 MYA between S. tora and C. fasciculata (Kumar et al., 2022) and 26.4–46.5 MYA between M. pudica and P. cineraria (www.timetree.org).

WGD events in the Caesalpinioideae subfamily were investigated by analyzing macrosynteny and the Ks distribution of homologous gene pairs. DIAMOND v.0.9.24 (Buchfink et al., 2015) was used for intra-species alignment of the proteomes of the seven studied species, and the results were used as input for MCScanX (Wang et al., 2012). Syntenic gene blocks were filtered by requiring 10 or more genes. The Ks values for homologous gene pairs in the syntenic blocks were calculated using KaKs_Calculator v.2.0 (Wang et al., 2010) with the GMYN model. A circular syntenic plot of homologous genes for each of the seven species was generated using the ggplot2 package in R. To identify the monosomic chromosomes of L. leucocephala, Minimap2 (Li, 2018) with the parameters “--secondary = no -x asm20” was used to align all 56 L. leucocephala chromosomes, and the results clearly divided the chromosomes into 14 highly syntenic groups. For each syntenic group, each homologous chromosome was treated as a virtual species, and OrthoFinder (Emms and Kelly, 2019) was used to identify more conserved single-copy genes using P. cineraria and S. tora as outgroups. The concatenated alignment for all single-copy genes was then used to construct a phylogenetic tree for each syntenic group. Divergence times were estimated by the RelTime-ML method in MEGA11 with other options at default settings (Tamura et al., 2021). Two calibration points were used, 46.6–60.0 MYA between P. cineraria and S. tora and 27.3–38.7 MYA between P. cineraria and L. leucocephala, which were calculated in this study (Figure 1B).

Identification of genes encoding mimosine synthases and decomposition enzymes

Publicly available gene sequences of key enzymes involved in L-mimosine metabolism were downloaded from the NCBI and UniProt databases. They included five SAT (EC 2.3.1.30) isogenes from Arabidopsis thaliana (sp|Q39218|SAT3_ARATH, sp|Q42588|SAT1_ARATH, sp|Q42538|SAT5_ARATH, sp|Q8S895|SAT2_ARATH, sp|Q8W2B8|SAT4_ARATH), two mimosine synthase (EC 2.5.1.52) genes from L. leucocephala (lcl|KF754356.1_cds_AHG97874.1_1, lcl|LC306827.1_cds_BBE37478.1_1), and one mimosinase (EC 4.3.3.8) gene from M. pudica (tr|U6BYK3|U6BYK3_MIMPU). Using DIAMOND (Buchfink et al., 2015) with the parameter “--more-sensitive”, the protein sequences were aligned to the reference proteomes of the seven species sequenced in this study, and matching genes were considered to be the primary candidate genes. All primary candidate genes were assigned to one OG in the OrthoFinder results. Therefore, all genes in this OG were taken as candidate genes. Hmmer (Potter et al., 2018) was then used to check whether the candidate genes contained specific domains: SAT genes (PF06426), mimosine synthase genes (PF00291), and mimosinase genes (PF01053). Only genes with the relevant specific domain were retained. The LG model of FastTree v.2.0 was used to construct phylogenetic trees for SAT, mimosine synthase, and mimosinase. Domain prediction involved aligning the protein sequences to the SMART (Simple Modular Architecture Research Tool) database (Letunic et al., 2021) and then visualizing the results using TBtools.

Identification of all candidate triterpenoid saponin genes in Gleditsia sinensis

The P450 proteins from three Fabaceae plants (Lotus japonicus, Cicer arietinum, and Cajanus cajanifolius) were downloaded from the Cytochrome P450 web page (https://drnelson.uthsc.edu/CytochromeP450.html). The P450 proteins for the CYP71A (Castillo et al., 2013), CYP71D (Ma et al., 2021), CYP72A (Fukushima et al., 2013), CYP88D (Wang et al., 2019), CYP93E (Moses et al., 2014), and CYP716A (Fukushima et al., 2011) families and the UGT protein sequences for the UGT71 (Achnine et al., 2005), UGT73 (Shibuya et al., 2010), UGT74 (Meesapyodsuk et al., 2007), and UGT91 (Shibuya et al., 2010) families, which are reported to participate in saponin biosynthesis, were downloaded from UniProt (https://www.uniprot.org/). The protein sequences of IDI, GPPS, FDPS, SS, SE, and AS genes from A. thaliana and Fabaceae species were downloaded from UniProt (https://www.uniprot.org/). The proteins of G. sinensis were then aligned to all these downloaded sequences using DIAMOND-v0.8.28 (Buchfink et al., 2015) with the parameters “--more-sensitive --evalue 1E−5.” Matching G. sinensis genes were considered to be candidate genes responsible for triterpenoid saponin biosynthesis. P450 candidate genes were required to match to both the Cytochrome P450 web page dataset and the UniProt dataset. The cytochrome P450 (PF00067) domain and UDP-glucuronosyl and UDP-glucosyltransferase (PF08244) domain predicted by Hmmer v.3.1b2 (Potter et al., 2018) were also required for P450 and UGT candidate genes, respectively. Protein sequences were aligned to the Batch CD-Search database for domain prediction (Wang et al., 2023), and the results were visualized with TBtools.

To confirm the triterpenoid saponin candidate genes, we collected 12 G. sinensis RNA-seq datasets from five tissues (root, stem, leaf, thorn, and fruit), with fruit from two different growth stages (early and late). The RNA-seq data were mapped to the G. sinensis genome with HISAT v.2.2.1 (Kim et al., 2019), and SAMtools v.1.3 (Li et al., 2009) with the parameters “-f 2-F 256-q 30” was used to filter the mapped BAM files. Using the gene annotation GFF file of G. sinensis and the filtered BAM files as input, StringTie v.2.0 (Pertea et al., 2015) was used to calculate gene expression levels as fragments per kilobase of transcript per million mapped fragments (FPKM). Candidate genes with low expression (FPKM < 1) in all tissues were removed.

Data and code availability

The genomic and transcriptomic sequencing reads generated in this study have been deposited at the NCBI SRA under accession numbers PRJNA1005071, PRJNA1005083, PRJNA1005079, PRJNA1005069, PRJNA1005080, PRJNA1005077, and PRJNA1005073 for G. sinensis, B. sappan, A. julibrissin, D. regia, A. confusa, L. leucocephala, and M. bimucronata, respectively. The data have also been uploaded to the China National Genomics Data Center with data number PRJCA025694. The corresponding genome assembly and annotations have been deposited at NCBI GenBank under accession numbers JAYWLF000000000, JAYWLH000000000, JAYWLI000000000, JAYWLE000000000, JAYMNN000000000, JAYWLJ000000000, and JAYWLG000000000, respectively. They are also available at Figshare (https://doi.org/10.6084/m9.figshare.25037543, https://doi.org/10.6084/m9.figshare.25037618, https://doi.org/10.6084/m9.figshare.25037561, https://doi.org/10.6084/m9.figshare.25000856, https://doi.org/10.6084/m9.figshare.25000859, https://doi.org/10.6084/m9.figshare.25037573, https://doi.org/10.6084/m9.figshare.25037582).

Funding

This work was supported by the Shenzhen Science and Technology Program (JCYJ20190814163805604 , KQTD20180411143628272 ), the Fund of Key Laboratory of Shenzhen (ZDSYS20141118170111640 ), and The Agricultural Science and Technology Innovation Program.

Author contributions

R.C., F.J., and A.W. prepared the genomic and transcriptomic sequencing samples. R.C. performed assembly, annotation, and evolutionary, polyploidization, and mimosine analyses, and S.M. performed triterpenoid saponin analysis. The other authors provided helpful suggestions. W.F., R.C., and S.M. wrote the manuscript, and all authors approved the final version of the manuscript. W.F. supervised the project.

Supplemental information

Document S1. Supplemental Figures 1–12 and Supplemental Tables 1–20

Document S2. Article plus supplemental information

Acknowledgments

We thank Prof. Shifeng Cheng for giving helpful suggestions. No conflict of interest is declared.

Published by the Plant Communications Shanghai Editorial Office in association with Cell Press, an imprint of Elsevier Inc., on behalf of CSPB and CEMPS, CAS.

Supplemental information is available at Plant Communications Online.
==== Refs
References

Abair A. Hughes C.E. Bailey C.D. The evolutionary history of Leucaena: Recent research, new genomic resources and future directions Trop Grassl-Forrajes 7 2019 65 73 10.17138/Tgft(7)65-73
Achnine L. Huhman D.V. Farag M.A. Sumner L.W. Blount J.W. Dixon R.A. Genomics-based selection and functional characterization of triterpene glycosyltransferases from the model legume Medicago truncatula Plant J. 41 2005 875 887 10.1111/j.1365-313X.2005.02344.x 15743451
Azani N. Babineau M. Bailey C.D. Banks H. Barbosa A.R. Pinto R.B. Boatwright J.S. Borges L.M. Brown G.K. Bruneau A. A new subfamily classification of the Leguminosae based on a taxonomically comprehensive phylogeny Taxon 66 2017 44 77 10.12705/661.3
Benson G. Tandem repeats finder: a program to analyze DNA sequences Nucleic Acids Res. 27 1999 573 580 9862982
Binghang L. Shi Y. Yuan J. Galaxy Y. Zhang H. Li N. Li Z. Chen Y. Mu D. Fan W. Estimation of genomic characteristics by analyzing k-mer frequency in de novo genome projects 2013
Boff T. Schifino-Wittmann M.T. Segmental allopolyploidy and paleopolyploidy in species of Leucaena Benth:: evidence from meiotic behaviour analysis Hereditas 138 2003 27 35 10.1034/j.1601-5223.2003.01646.x 12830982
Buchfink B. Xie C. Huson D.H. Fast and sensitive protein alignment using DIAMOND Nat. Methods 12 2015 59 60 25402007
Cannon S.B. McKain M.R. Harkess A. Nelson M.N. Dash S. Deyholos M.K. Peng Y. Joyce B. Stewart C.N. Rolf M. Multiple Polyploidy Events in the Early Radiation of Nodulating and Nonnodulating Legumes Mol. Biol. Evol. 32 2015 193 210 10.1093/molbev/msu296 25349287
Castillo D.A. Kolesnikova M.D. Matsuda S.P.T. An effective strategy for exploring unknown metabolic pathways by genome mining J. Am. Chem. Soc. 135 2013 5885 5894 10.1021/ja401535g 23570231
Chang Y. Liu H. Liu M. Liao X. Sahu S.K. Fu Y. Song B. Cheng S. Kariba R. Muthemba S. The draft genomes of five agriculturally important African orphan crops GigaScience 8 2019 10.1093/gigascience/giy152
Chen C. Wu Y. Li J. Wang X. Zeng Z. Xu J. Liu Y. Feng J. Chen H. He Y. Xia R. TBtools-II: A "one for all, all for one" bioinformatics platform for biological big-data mining Mol. Plant 16 2023 1733 1742 10.1016/j.molp.2023.09.010 37740491
Da Cruz M.H.P. Domingues D.S. Saito P.T.M. Paschoal A.R. Bugatti P.H. TERL: classification of transposable elements by convolutional neural networks Briefings Bioinf. 22 2021 bbaa185
Desbrosses G.J. Stougaard J. Root Nodulation: A Paradigm for How Plant-Microbe Symbiosis Influences Host Developmental Pathways Cell Host Microbe 10 2011 348 358 10.1016/j.chom.2011.09.005 22018235
Edgar R.C. MUSCLE: multiple sequence alignment with high accuracy and high throughput Nucleic Acids Res. 32 2004 1792 1797 15034147
Emms D.M. Kelly S. OrthoFinder: phylogenetic orthology inference for comparative genomics Genome Biol. 20 2019 238 31727128
Feng X. Cheng H. Portik D. Li H. Metagenome assembly of high-fidelity long reads with hifiasm-meta Nat. Methods 19 2022 671 674 35534630
Flynn J.M. Hubley R. Goubert C. Rosen J. Clark A.G. Feschotte C. Smit A.F. RepeatModeler2 for automated genomic discovery of transposable element families Proc. Natl. Acad. Sci. USA 117 2020 9451 9457 Proceedings of the National Academy of Sciences 32300014
Fukushima E.O. Seki H. Sawai S. Suzuki M. Ohyama K. Saito K. Muranaka T. Combinatorial biosynthesis of legume natural and rare triterpenoids in engineered yeast Plant Cell Physiol. 54 2013 740 749 10.1093/pcp/pct015 23378447
Fukushima E.O. Seki H. Ohyama K. Ono E. Umemoto N. Mizutani M. Saito K. Muranaka T. CYP716A subfamily members are multifunctional oxidases in triterpenoid biosynthesis Plant Cell Physiol. 52 2011 2050 2061 10.1093/pcp/pcr146 22039103
Galindo-González L. Mhiri C. Deyholos M.K. Grandbastien M.A. LTR-retrotransposons in plants: Engines of evolution Gene 626 2017 14 25 10.1016/j.gene.2017.04.051 28476688
Govindarajulu R. Hughes C.E. Alexander P.J. Bailey C.D. The Complex Evolutionary Dynamics of Ancient and Recent Polyploidy in POLYPLOIDY IN LEUCAENA (Leguminosae; Mimosoideae) Am. J. Bot. 98 2011 2064 2076 10.3732/ajb.1100260 22130273
Griesmann M. Chang Y. Liu X. Song Y. Haberer G. Crook M.B. Billault-Penneteau B. Lauressergues D. Keller J. Imanishi L. Phylogenomics reveals multiple losses of nitrogen-fixing root nodule symbiosis Science 361 2018 10.1126/science.aat1743
Habibi N. Al Salameen F. Vyas N. Rahman M. Kumar V. Shajan A. Zakir F. Razzack N.A. Al Doaij B. Genome survey and genetic characterization of Acacia pachyceras O. Schwartz Front. Plant Sci. 14 2023 1062401 10.3389/fpls.2023.1062401
Hagihara T. Mano H. Miura T. Hasebe M. Toyota M. Calcium-mediated rapid movements defend against herbivorous insects in Mimosa pudica Nat. Commun. 13 2022 6412 10.1038/s41467-022-34106-x 36376294
Harun-Ur-Rashid M. Iwasaki H. Parveen S. Oogai S. Fukuta M. Hossain M.A. Anai T. Oku H. Cytosolic Cysteine Synthase Switch Cysteine and Mimosine Production in Leucaena leucocephala Appl. Biochem. Biotechnol. 186 2018 613 632 10.1007/s12010-018-2745-z 29691793
Howarth J.R. Domínguez-Solís J.R. Gutiérrez-Alcalá G. Wray J.L. Romero L.C. Gotor C. The serine acetyltransferase gene family in Arabidopsis thaliana and the regulation of its expression by cadmium Plant Mol. Biol. 51 2003 589 598 10.1023/A:1022349623951 12650624
Huisman R. Geurts R. A Roadmap toward Engineered Nitrogen-Fixing Nodule Symbiosis Plant Commun. 1 2020 100019 10.1016/j.xplc.2019.100019
Jones P. Binns D. Chang H.-Y. Fraser M. Li W. McAnulla C. McWilliam H. Maslen J. Mitchell A. Nuka G. InterProScan 5: genome-scale protein function classification Bioinformatics 30 2014 1236 1240 24451626
Kang S.H. Pandey R.P. Lee C.M. Sim J.S. Jeong J.T. Choi B.S. Jung M. Ginzburg D. Zhao K. Won S.Y. Genome-enabled discovery of anthraquinone biosynthesis in Senna tora Nat. Commun. 11 2020 5875 10.1038/s41467-020-19681-1 33208749
Kato-Noguchi H. Kurniadie D. Allelopathy and Allelochemicals of Leucaena leucocephala as an Invasive Plant Species Plants-Basel 2022 10.3390/plants11131672 11ARTN
Kenicer G. Legumes of the World. Edited by G. Lewis, B. Schrire, B. MacKinder & M. Lock. Royal Botanic Gardens, Kew. 2005 Edinburgh J. Bot. 62 2005 195 196 10.1017/S0960428606190198
Kim D. Paggi J.M. Park C. Bennett C. Salzberg S.L. Graph-based genome alignment and genotyping with HISAT2 and HISAT-genotype Nat. Biotechnol. 37 2019 907 915 10.1038/s41587-019-0201-4 31375807
Koenen E.J.M. Kidner C. de Souza É.R. Simon M.F. Iganci J.R. Nicholls J.A. Brown G.K. de Queiroz L.P. Luckow M. Lewis G.P. Hybrid capture of 964 nuclear genes resolves evolutionary relationships in the mimosoid legumes and reveals the polytomous origins of a large pantropical radiation Am. J. Bot. 107 2020 1710 1735 33253423
Kong W. Liu M. Felker P. Ewens M. Bessega C. Pometti C. Wang J. Xu P. Teng J. Wang J. Genome and evolution of Prosopis alba Griseb., a drought and salinity tolerant tree legume crop for arid climates Plants, People, Planet 5 2023 933 947 10.1002/ppp3.10404
Kozlov A.M. Darriba D. Flouri T. Morel B. Stamatakis A. RAxML-NG: a fast, scalable and user-friendly tool for maximum likelihood phylogenetic inference Bioinformatics 35 2019 4453 4455 31070718
Kumar S. Suleski M. Craig J.M. Kasprowicz A.E. Sanderford M. Li M. Stecher G. Hedges S.B. TimeTree 5: An Expanded Resource for Species Divergence Times Mol. Biol. Evol. 39 2022 10.1093/molbev/msac174
Kuwahara Y. Nakajima D. Shinpo S. Nakamura M. Kawano N. Kawahara N. Yamazaki M. Saito K. Suzuki H. Hirakawa H. Identification of potential genes involved in triterpenoid saponins biosynthesis in Gleditsia sinensis by transcriptome and metabolome analyses J. Nat. Med. 73 2019 369 380 10.1007/s11418-018-1270-2 30547286
Lagesen K. Hallin P. Rødland E.A. Stærfeldt H.-H. Rognes T. Ussery D.W. RNAmmer: consistent and rapid annotation of ribosomal RNA genes Nucleic Acids Res. 35 2007 3100 3108 17452365
Letunic I. Khedkar S. Bork P. SMART: recent updates, new developments and status in 2020 Nucleic Acids Res. 49 2021 D458 D460 10.1093/nar/gkaa937 33104802
Li H. Minimap2: pairwise alignment for nucleotide sequences Bioinformatics 34 2018 3094 3100 29750242
Li H. Handsaker B. Wysoker A. Fennell T. Ruan J. Homer N. Marth G. Abecasis G. Durbin R. 1000 Genome Project Data Processing Subgroup The Sequence Alignment/Map format and SAMtools Bioinformatics 25 2009 2078 2079 10.1093/bioinformatics/btp352 19505943
Li L.M. Fu J.X. Song X.Q. Complete plastome sequence of Caesalpinia sappan Linnaeus, a dyestuff and medicinal species Mitochondrial DNA. B Resour. 5 2020 2535 2536 10.1080/23802359.2020.1778579 33457853
Libourel C. Keller J. Brichet L. Cazalé A.C. Carrère S. Vernié T. Couzigou J.M. Callot C. Dufau I. Cauet S. Comparative phylotranscriptomics reveals ancestral and derived root nodule symbiosis programmes Nat. Plants 9 2023 1067 1080 10.1038/s41477-023-01441-w 37322127
Lin H.Y. Chang T.C. Chang S.T. A review of antioxidant and pharmacological properties of phenolic compounds in Acacia confusa J. Tradit. Complement. Med. 8 2018 443 450 10.1016/j.jtcme.2018.05.002 30302324
Lin M. Jian J.B. Zhou Z.Q. Chen C.H. Wang W. Xiong H. Mei Z.N. Chromosome-level genome of Entada phaseoloides provides insights into genome evolution and biosynthesis of triterpenoid saponins Mol. Ecol. Resour. 22 2022 3049 3067 10.1111/1755-0998.13662 35661414
Liu B. Shi Y. Yuan J. Hu X. Zhang H. Li N. Li Z. Chen Y. Mu D. Fan W. Estimation of genomic characteristics by analyzing k-mer frequency in de novo genome projects Quantitative Biology 35 2013 62 67
Lowe T.M. Chan P.P. tRNAscan-SE On-line: integrating search and context for analysis of transfer RNA genes Nucleic Acids Res. 44 2016 W54 W57 27174935
Ma Y. Cui G. Chen T. Ma X. Wang R. Jin B. Yang J. Kang L. Tang J. Lai C. Expansion within the CYP71D subfamily drives the heterocyclization of tanshinones synthesis in Salvia miltiorrhiza Nat. Commun. 12 2021 685 10.1038/s41467-021-20959-1 33514704
Manni M. Berkeley M.R. Seppey M. Simão F.A. Zdobnov E.M. BUSCO Update: Novel and Streamlined Workflows along with Broader and Deeper Phylogenetic Coverage for Scoring of Eukaryotic, Prokaryotic, and Viral Genomes Mol. Biol. Evol. 38 2021 4647 4654 34320186
Manzanilla V. Bruneau A. Phylogeny reconstruction in the Caesalpinieae grade (Leguminosae) based on duplicated copies of the sucrose synthase gene and plastid markers Mol. Phylogenet. Evol. 65 2012 149 162 10.1016/j.ympev.2012.05.035 22699157
Massaro I. Poethig R.S. Sinha N.R. Leichty A.R. Chromosome-level genome of the transformable northern wattle Acacia crassicarpa. G3 (Bethesda) 14 2024 10.1093/g3journal/jkad284
McLay T.G.B. Murphy D.J. Holmes G.D. Mathews S. Brown G.K. Cantrill D.J. Udovicic F. Allnutt T.R. Jackson C.J. A genome resource for Acacia, Australia's largest plant genus PLoS One 17 2022 e0274267 10.1371/journal.pone.0274267
Meesapyodsuk D. Balsevich J. Reed D.W. Covello P.S. Saponin biosynthesis in Saponaria vaccaria. cDNAs encoding beta-amyrin synthase and a triterpene carboxylic acid glucosyltransferase Plant Physiol. 143 2007 959 969 10.1104/pp.106.088484 17172290
Moses T. Thevelein J.M. Goossens A. Pollier J. Comparative analysis of CYP93E proteins for improved microbial synthesis of plant triterpenoids Phytochemistry 108 2014 47 56 10.1016/j.phytochem.2014.10.002 25453910
Negi V.S. Bingham J.P. Li Q.X. Borthakur D. A Carbon-Nitrogen Lyase from Leucaena leucocephala Catalyzes the First Step of Mimosine Degradation Plant Physiol. 164 2014 922 934 10.1104/pp.113.230870 24351687
Nezamivand-Chegini M. Ebrahimie E. Tahmasebi A. Moghadam A. Eshghi S. Mohammadi-Dehchesmeh M. Kopriva S. Niazi A. New insights into the evolution of SPX gene family from algae to legumes; a focus on soybean BMC Genom. 22 2021 915 10.1186/s12864-021-08242-5
Nguyen B.C.Q. Tawata S. The Chemistry and Biological Activities of Mimosine: A Review Phytother Res. 30 2016 1230 1242 10.1002/ptr.5636 27213712
Ou S. Jiang N. LTR_retriever: A Highly Accurate and Sensitive Program for Identification of Long Terminal Repeat Retrotransposons Plant Physiol. 176 2018 1410 1422 29233850
Ou S. Su W. Liao Y. Chougule K. Agda J.R.A. Hellinga A.J. Lugo C.S.B. Elliott T.A. Ware D. Peterson T. Benchmarking transposable element annotation methods for creation of a streamlined, comprehensive pipeline Genome Biol. 20 2019 275 10.1186/s13059-019-1905-y 31843001
Pertea M. Pertea G.M. Antonescu C.M. Chang T.C. Mendell J.T. Salzberg S.L. StringTie enables improved reconstruction of a transcriptome from RNA-seq reads Nat. Biotechnol. 33 2015 290 295 10.1038/nbt.3122 25690850
Potter S.C. Luciani A. Eddy S.R. Park Y. Lopez R. Finn R.D. HMMER web server: 2018 update Nucleic Acids Res. 46 2018 W200 W204 29905871
Rajarajan K. Uthappa A.R. Handa A.K. Chavan S.B. Vishnu R. Shrivastava A. Handa A. Rana M. Sahu S. Kumar N. Genetic diversity and population structure of Leucaena leucocephala (Lam.) de Wit genotypes using molecular and morphological attributes Genet. Resour. Crop Evol. 69 2022 71 83 10.1007/s10722-021-01203-7
Rhie A. Walenz B.P. Koren S. Phillippy A.M. Merqury: reference-free quality, completeness, and phasing assessment for genome assemblies Genome Biol. 21 2020 245 10.1186/s13059-020-02134-9 32928274
Schmutz J. Cannon S.B. Schlueter J. Ma J. Mitros T. Nelson W. Hyten D.L. Song Q. Thelen J.J. Cheng J. Genome sequence of the palaeopolyploid soybean Nature 463 2010 178 183 10.1038/nature08670 20075913
Servant N. Varoquaux N. Lajoie B.R. Viara E. Chen C.-J. Vert J.-P. Heard E. Dekker J. Barillot E. HiC-Pro: an optimized and flexible pipeline for Hi-C data processing Genome Biol. 16 2015 259
Shibuya M. Nishimura K. Yasuyama N. Ebizuka Y. Identification and characterization of glycosyltransferases involved in the biosynthesis of soyasaponin I in Glycine max FEBS Lett. 584 2010 2258 2264 10.1016/j.febslet.2010.03.037 20350545
Simao F.A. Waterhouse R.M. Ioannidis P. Kriventseva E.V. Zdobnov E.M. BUSCO: assessing genome assembly and annotation completeness with single-copy orthologs Bioinformatics 31 2015 3210 3212 10.1093/bioinformatics/btv351 26059717
Slater G.S.C. Birney E. Automated generation of heuristics for biological sequence comparison BMC Bioinf. 6 2005 31 10.1186/1471-2105-6-31
Smýkal P. Coyne C.J. Ambrose M.J. Maxted N. Schaefer H. Blair M.W. Berger J. Greene S.L. Nelson M.N. Besharat N. Legume Crops Phylogeny and Genetic Diversity for Science and Breeding Crit. Rev. Plant Sci. 34 2014 43 104 10.1080/07352689.2014.897904
Soedarjo M. Borthakur D. Mimosine, a toxin produced by the tree-legume Leucaena provides a nodulation competition advantage to mimosine-degrading Rhizobium strains Soil Biol. Biochem. 30 1998 1605 1613 10.1016/S0038-0717(97)00180-6
Stanke M. Schöffmann O. Morgenstern B. Waack S. Gene prediction in eukaryotes with a generalized hidden Markov model that uses hints from external sources BMC Bioinf. 2006 10.1186/1471-2105-7-62 7
Steyermark J.A. Berry P.E. Holst B.K. Yatskievych K. Flora of the Venezuelan Guayana 1995 Missouri Botanical Garden St. Louis
Sudalaimuthuasari N. Ali R. Kottackal M. Rafi M. Al Nuaimi M. Kundu B. Al-Maskari R.S. Wang X. Mishra A.K. Balan J. The Genome of the Mimosoid Legume Prosopis cineraria, a Desert Tree Int. J. Mol. Sci. 23 2022 8503
Tamura K. Stecher G. Kumar S. MEGA11: Molecular Evolutionary Genetics Analysis Version 11 Mol. Biol. Evol. 38 2021 3022 3027 33892491
Van de Peer Y. Maere S. Meyer A. OPINION The evolutionary significance of ancient genome duplications Nat. Rev. Genet. 10 2009 725 732 10.1038/nrg2600 19652647
Van De Peer Y. Mizrachi E. Marchal K. The evolutionary significance of polyploidy Nat. Rev. Genet. 18 2017 411 424 10.1038/nrg.2017.26 28502977
Wang C. Su X. Sun M. Zhang M. Wu J. Xing J. Wang Y. Xue J. Liu X. Sun W. Chen S. Efficient production of glycyrrhetinic acid in metabolically engineered Saccharomyces cerevisiae via an integrated strategy Microb. Cell Factories 18 2019 95 10.1186/s12934-019-1138-5
Wang D. Zhang Y. Zhang Z. Zhu J. Yu J. KaKs_Calculator 2.0: A Toolkit Incorporating Gamma-Series Methods and Sliding Window Strategies Dev. Reprod. Biol. 8 2010 77 80
Wang J. Chitsaz F. Derbyshire M.K. Gonzales N.R. Gwadz M. Lu S. Marchler G.H. Song J.S. Thanki N. Yamashita R.A. The conserved domain database in 2023 Nucleic Acids Res. 51 2023 D384 D388 10.1093/nar/gkac1096 36477806
Wang S. Wang H. Jiang F. Wang A. Liu H. Zhao H. Yang B. Xu D. Zhang Y. Fan W. EndHiC: assemble large contigs into chromosome-level scaffolds using the Hi-C links from contig ends BMC Bioinf. 23 2022 528
Wang Y. Tang H. DeBarry J.D. Tan X. Li J. Wang X. Lee T.h. Jin H. Marler B. Guo H. MCScanX: a toolkit for detection and evolutionary analysis of gene synteny and collinearity Nucleic Acids Res. 40 2012 e49 22217600
Wu T.D. Watanabe C.K. GMAP: a genomic mapping and alignment program for mRNA and EST sequences Bioinformatics 21 2005 1859 1875 15728110
Xie C. Li M. Jim C.Y. Liu D. Spatio-temporal patterns of an invasive species Mimosa bimucronata (DC.) Kuntze under different climate scenarios in China Front for Glob Chang 6 2023 10.3389/ffgc.2023.1144829 6ARTN 1144829
Yang J. Liu M. Sahu S.K. Li R. Wang G. Guo X. Liu J. Cheng L. Jiang H. Zhao F. Chromosome-scale genomes of five Hongmu species in Leguminosae Sci. Data 10 2023 710 10.1038/s41597-023-02593-2 37848504
Yao L. Lu J. Wang J. Gao W.Y. Advances in biosynthesis of triterpenoid saponins in medicinal plants Chin. J. Nat. Med. 18 2020 417 424 10.1016/S1875-5364(20)30049-2 32503733
Yuan J.Y. Song Q.X. Polyploidy and diploidization in soybean Mol. Breed. 43 2023 51 10.1007/s11032-023-01396-y
Zeng Q. Lai Q. Gu S. Ma Z. The complete plastid genome of Delonix regia (Hook.) Mitochondrial DNA. B Resour. 5 2020 784 785 10.1080/23802359.2020.1715886 33366749
Zhang C. Zhang T. Luebert F. Xiang Y. Huang c.-h. Hu Y. Rees M. Frohlich M.W. Qi J. Weigend M. Ma H. Asterid Phylogenomics/Phylotranscriptomics Uncover Morphological Evolutionary Histories and Support Phylogenetic Placement for Numerous Whole-Genome Duplications Mol. Biol. Evol. 37 2020 3188 3210 32652014
Zhang J. Huang H. Qu C. Meng X. Meng F. Yao X. Wu J. Guo X. Han B. Xing S. Comprehensive analysis of chloroplast genome of Albizia julibrissin Durazz. (Leguminosae sp.) Planta 255 2021 26 10.1007/s00425-021-03812-z 34940902
Zhang Z. Koike K. Jia Z. Nikaido T. Guo D. Zheng J. Triterpenoidal saponins from Gleditsia sinensis Phytochemistry 52 1999 715 722 10.1016/S0031-9422(99)00238-1 10570830
Zhao Y. Zhang R. Jiang K.W. Qi J. Hu Y. Guo J. Zhu R. Zhang T. Egan A.N. Yi T.S. Nuclear phylotranscriptomics and phylogenomics support numerous polyploidization events and hypotheses for the evolution of rhizobial nitrogen-fixing symbiosis in Fabaceae Mol. Plant 14 2021 748 773 10.1016/j.molp.2021.02.006 33631421
Zhu L. Zhang Y. Guo W. Wang Q. Gleditsia sinensis: transcriptome sequencing, construction, and application of its protein-protein interaction network BioMed Res. Int. 2014 2014 404578 10.1155/2014/404578
