
==== Front
iScience
iScience
iScience
2589-0042
Elsevier

S2589-0042(24)01923-0
10.1016/j.isci.2024.110698
110698
Article
Genome of the most noxious weed water hyacinth (Eichhornia crassipes) provides insights into plant invasiveness and its translational potential
Bisht Manohar S. 1
Singh Mitali 1
Chakraborty Abhisek 1
Sharma Vineet K. vineetks@iiserb.ac.in
12∗
1 MetaBioSys Group, Department of Biological Sciences, Indian Institute of Science Education and Research Bhopal, Bhopal 462066, Madhya Pradesh, India
∗ Corresponding author vineetks@iiserb.ac.in
2 Lead contact

10 8 2024
20 9 2024
10 8 2024
27 9 11069818 10 2023
7 5 2024
6 8 2024
© 2024 Published by Elsevier Inc.
2024

https://creativecommons.org/licenses/by-nc-nd/4.0/ This is an open access article under the CC BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/).
Summary

The invasive character of Eichhornia crassipes (water hyacinth) is a major threat to global biodiversity and ecosystems. To investigate the genomic basis of invasiveness, we performed the genome and transcriptome sequencing of E. crassipes and reported the genome of 1.11 Gbp size with 63,299 coding genes and N50 of 1.98 Mb. We confirmed a recent whole genome duplication event in E. crassipes that resulted in high intraspecific collinearity and significant expansion in gene families. Further, the orthologs gene clustering analysis and comparative evolutionary analysis with 14 other aquatic invasive and non-invasive angiosperm species revealed adaptive evolution in genes associated with plant-pathogen interaction, hormone signaling, abiotic stress tolerance, heavy metals sequestration, photosynthesis, and cell wall biosynthesis with highly expanded gene families, which contributes toward invasive characteristics of the water hyacinth. However, these characteristics also make water hyacinth an excellent candidate for biofuel production, phytoremediation, and other translational applications.

Graphical abstract

Highlights

• Whole genome and transcriptome assembly of the most noxious weed Eichhornia crassipes

• Recent whole genome duplication in E. crassipes leading to tetraploidy

• Genomic insights into pathways contributing to its invasiveness and translational potential

Genomics; Plant biology; Plant Genetics

Subject areas

Genomics
Plant biology
Plant Genetics
Published: August 10, 2024
==== Body
pmcIntroduction

Invasive species pose a potential threat to non-native ecosystems and are known to be among the major contributors to the loss of biodiversity and global food security.1,2 Intense competition between invasive species and native flora for resources not only leads to ecological perturbance but also paves the way for other non-native species to invade (invasion meltdown).2,3,4 Several hypotheses including enemy release (ERH), novel weapon (NWH), and empty niche (EN), have been proposed to understand how invasive species outcompete native biodiversity.2,5 The various evolutionary events including speciation, adaptation, hybridizations, and parallel evolution that play a key role in invasion biology are of intriguing interest, and a comprehensive understanding of invasion biology is crucial to manage invasive species. However, only a few studies are yet available that explores these aspects.6,7,8,9

Eichhornia crassipes, commonly known as “water hyacinth” (family Pontederiaceae, 2n = 4x = 32) is known for its rapid invasion across the globe and is also listed in the Global Invasive Species Database.10,11 It is a perennial freshwater flowering plant native to South America10,12 (Figure 1A). In India, it was introduced in the early 19th century due to its high aesthetic ornamental value; later it spread out so invasively across the country, especially in the state of Bengal, that anecdotally, it is now known as the “terror of Bengal”.13Figure 1 E. crassipes evolutionary history

(A) Morphology of E. crassipes (water hyacinth).

(B) Phylogenetic tree with gene family evolution, teal bars denote the confidence intervals for estimated divergence times at each internal node. Calibrated nodes are highlighted with red dots. MRCA—most recent common ancestor.

(C) Synonymous substitution rate (Ks) distributions of paralogs and orthologs of E. crassipes, E. paniculata, and M. acuminata.

(D) Timing of whole-genome duplications, boxes indicate WGD events. Blue boxes indicate WGD events analyzed in this paper.

(E) Stacked columns chart showing the number of gene duplications in various duplicated types and gene duplication-induced highly expanded gene numbers, PD; proximal duplication, TD; tandem duplication, DD; dispersed duplication, SD; segmental duplication.

(F) KEGG enrichment of genes overlapping between highly expanded gene families and types of gene duplications.

Water hyacinth grows vegetatively through stolon, rapidly covering the surface of the aquatic bodies as dense mat, which leads to sunlight cut-off and depletion in the level of oxygen that causes the death of the resident organisms in the aquatic bodies. The high quantities of decomposing biomass further deplete the dissolved oxygen, decrease the pH, and elevate the carbon dioxide and hydrogen sulfide that disrupts the ecosystem, resulting in further loss of biodiversity.14 Morphologically, the mature plant body consists of a tuft of roots, rhizomes, stolons, leaves with swollen petioles, spiked inflorescence, and fruits. Anatomically, the plant shows various peculiar features governing its aquatic adaptations such as reduced vasculature, amphistomatic leaves with thin cuticle, presence of aerenchyma tissue and airspace in the roots, rhizome, and petiole that provide buoyancy.10,15

Various biological, chemical, and physical methods have been employed to control the spread of this invasive species, but effective control strategies are still lacking. However, over the last decade, water hyacinth has valorized from being notorious to a wonder plant with wide applications. A wealth of experimental studies has proposed the utilization of water hyacinth as a biochar and in biofuel production due to their low lignin content and high amount of cellulose and hemicellulose, which are the main substrates for producing fermentable sugars.16,17,18,19 Also, water hyacinth has been considered as a great reservoir for carbon sequestration in the era of global warming due to its high biomass accumulation20,21 and as a phytoremediator due to its potential heavy metal sequestration ability.22

To gain insights into the invasive character of species like water hyacinth and to identify the factors pertaining to their successful niche expansion, genome sequences are much needed. However, only a few genomes of such invasive aquatic plant species are yet available.23 More importantly, in the case of water hyacinth that poses potential threats to the ecosystem due to its invasive nature and also has enormous translational applications, the genome sequence along with comparative evolutionary and functional insights are much needed. Therefore, we performed the genome sequencing of E. crassipes using Oxford Nanopore long reads and Illumina short reads along with the transcriptome from the root, leaf, and stolon tissues. We inferred the genome of E. crassipes to be complex with a recent whole genome duplication (WGD) event. In addition, the phylogenetic position of E. crassipes was resolved along with gene family evolution analysis, orthologous gene clustering, and comparative evolutionary analysis to gain comprehensive insights into the genomic clues responsible for the invasive properties of water hyacinth and for its translational utilization in biofuel production, phytoremediation and other applications.

Results

Genomic and transcriptomic assembly

The estimated genome size of E. crassipes was found to be 1.23 Gbp using SGA-preqc24 and Genomescope2.0,25 which is concordant with the reported Plant DNA C-values Database (https://cvalues.science.kew.org/). Further, E. crassipes genome was confirmed to be an allotetraploid due to the high proportion of aabb k-mers (3.32%) compared to aaab k-mers (0.001%). Allotetraploid would expect a high proportion of aabb and a low proportion of aaab since preferential pairing would ensure that two homologs from the first subgenome and two homologs from the second subgenome are present after recombination25 (Figure S1). A total of 212.4 Gb (172.6X coverage) of Illumina short-read data, and 36.8 Gb (30X coverage) of long-read data from Oxford Nanopore technologies were generated. In addition, a total of 37 Gb transcriptomic data were also generated from the root, leaf, and stolon tissue (Table S1).

The final genome assembly had a size of 1.11 Gbp, which is similar to the estimated genome size. The assembly had a scaffold N50 value of 1.98 Mb, with the largest scaffold of 11.6 Mb and GC of 36.39% (Tables 1 and S2). A total of 98.63% and 94.14% of the Illumina short reads and Nanopore long reads, respectively, could be mapped to the final genome assembly, and an average of 82.19% of the RNA-seq data from three tissues (leaf; 84.45%, root; 83.75%, and stolon; 78.38%) could be mapped to the final genome assembly. The genome assembly had an LAI (LTR Assembly Index) of 10.50, which indicates a “reference” level quality of the genome,26 which was further supported by the complete BUSCO score of 95.3% (Table S3).Table 1 Genome assembly and annotation statistics of E. crassipes

Feature	E. crassipes	
Final assembly	
	
Total length (bp)	1,113,236,536	
Number of scaffolds	1,634	
Longest scaffold (Mb)	11.6	
Scaffold N50 (Mb)	1.98	
Scaffold L50	167	
Scaffold GC content	36.39%	
Scaffold N content	0.09%	
BUSCO (complete)	95.3%	
LTR Assembly Index (LAI)	10.50	
	
Gene models	
	
Number of coding genes obtained from MAKER	73,622	
Number of coding genes with AED value <0.3	63,321	
Final number of coding genes after length-based filtering (High confidence gene set)	63,299	
Repeat content of the genome (%)	55.91	
BUSCO (complete)	88.7%	
	
Non-protein-coding genes	
	
Number of rRNAs	515	
Number of tRNAs (decoding standard amino acids)	1,226	
Number of miRNAs (homology-based)	165	

Genome annotation

Repeat masking using de novo repeat library identified a total of 55.91% repetitive sequences, of which interspersed repeats constituted 46.82%, which is similar to most plant genomes.27 The major type of transposable elements (TEs) was long terminal repeats (LTR) elements that accounted for 25.16% of the genome, including Gypsy/DIRS1 (15.22%) and Ty1/Copia (8.84%) retroelements (Table S4). The insertion time of most LTR-RTs were found to be 0.9 million years ago (mya) (Figure S2).

Before gene set construction, transcriptomic quality-filtered reads were de novo assembled, which resulted in 1,534,604 assembled transcripts that were later used as an empirical evidence in the MAKER pipeline (Table S5). A total of 63,299 high-confidence genes were retained after AED value-based (AED values <0.3) and length-based (≥150 bp) filtering of the MAKER-derived gene models. The coding genes distribution in KEGG pathways and COG categories are mentioned in Tables S6 and S7.

Further, the high-confidence gene set showed the presence of 88.7% complete BUSCOs (Table S3). Overall, 97.12% of genes were annotated using three public databases—NCBI-nr, Pfam-A, and Swiss-Prot (Table S8). Non-coding RNAs were also identified in E. crassipes genome assembly (Table 1).

Phylogeny construction and gene family evolution analysis

Phylogenetic analysis based on 168 one-to-one fuzzy orthogroups of 15 angiosperm species revealed the phylogenetic position of E. crassipes that formed a monophyletic group with Musa acuminata belonging to the order Zingiberales, with an estimated divergence time of ∼88.7 (96.10–83.3) mya (Figure 1B).

The inferred gene clusters from OrthoFinder and tree topology were used for CAFÉ analysis to identify the gene families with significant expansion and contraction. Among the total 8,410 gene families detected in E. crassipes genome, 4,709 showed expansion and 958 showed contraction. Among these expanded gene families, 261 were highly expanded (≥10 expanded genes) (Data S1). The KEGG pathway analysis of the highly expanded gene families revealed that genes from these families were mainly involved in plant secondary-metabolism pathway, MAPK signaling pathway—plant, plant hormone signal transduction, photosynthesis, etc. (Table S9).

Genome synteny and whole-genome duplication analysis

Whole-genome duplication (WGD) or polyploidy is among the significant evolutionary mechanisms for plant adaptation and biodiversification.28,29 The intra-specific genome synteny analysis revealed 60.94% collinearity (38,572 out of 63,299 genes) in E. crassipes genome. Most of the duplicated genes were contributed by whole genome or segmental duplications (60.94%), followed by three other types: dispersed (31.73%), proximal (2.36%), and tandem (2.23%) duplications. Further, the intragenomic paralogs Ks (number of substitutions per synonymous site) distribution of both Eichhornia species revealed a common peak of Ks at 0.44, which was older than their divergence peak (E. crassipes—E. paniculata) Ks at 0.35, suggesting that both Eichhornia species have experienced a common WGD event around ∼34 mya prior to their divergence around ∼27 mya (Figures 1C, 1D, and S3). Additionally, E. crassipes paranome distribution showed a younger peak at Ks 0.05, suggesting that E. crassipes underwent a recent WGD (∼4 mya) independently after divergence that led to their tetraploid state, which was also supported by the high anchor pairs at Ks 0.05 and numerous densely aggregated anchor points with Ks range below <1 in whole genome Ks dot plot of E. crassipes (Figures S4 and S5).

Furthermore, the ortholog Ks distribution between E. crassipes—M. acuminata and E. paniculata—M. acuminata revealed that Commelinales diverged from the Zingiberales around ∼87 mya (Ks ∼1.12) (Figure 1C), which is in agreement with the phylogeny (Figure 1B). Additionally, the anchor pairs in the paralogs Ks distribution of E. crassipes, E. paniculata and M. acuminata at Ks 1.3 (∼100 mya) suggest a shared WGD event between Commelinales and Zingiberales before their divergence (Figure S6), which was further supported by the high collinear genes within the range of Ks 1–2, in the whole genome Ks dot plot between E. crassipes—M. acuminata and E. paniculata—M. acuminata (Figure S7).

Due to the large number of duplicated genes in the genome of E. crassipes, which was also supported by complete duplicated BUSCOs (Table S3), we analyzed the duplication-induced expansion in the highly expanded gene families by identifying genes that were common between highly expanded gene families and each type of duplicated genes. Compared with other duplicate types, segmental duplication (SD) events contributed to the highest proportion of expansion of the gene families (Figure 1E; Table S10). Further, KEGG pathways analysis of these highly expanded genes induced by various types of duplication revealed that these genes were mainly enriched in pathways like plant hormone signaling transduction, photosynthesis, plant-pathogen interaction, MAPK signaling, and other metabolic pathways. (Figure 1F; Tables S11–S14).

Orthologous gene clustering and comparative evolutionary analysis reveal genes with signatures of adaptive evolution

Orthologous gene cluster analysis including invasive and non-invasive aquatic angiosperm species (three of each) revealed 5,896 core gene clusters (Figure S8A). Among the invasive species, 77 core gene clusters were found, whereas 2,596 gene clusters were unique to E. crassipes (Figure S8B). The core gene clusters among the invasive species, and the gene clusters specific to E. crassipes were found to be involved in plant growth development, signal transduction, reproduction, biotic and abiotic stress tolerance, and other metabolic processes, which might cf. their invasive characteristics (Data S2 and S3).

In total, 4,734 orthogroups were identified across the 14 aquatic invasive and non-invasive angiosperm species, including E. crassipes. 24 genes had higher rates of evolution, i.e., higher branch length distance in E. crassipes in comparison to the other species. Further, 1,272 genes had unique amino acid substitution with functional impact, and 255 genes were positively selected (p value <0.05) in E. crassipes (Tables S15–S17 and Data S4). Among these genes, 114 showed more than one signature of adaptive evolution (MSA genes). The COG categories of MSA genes are mentioned in Table S18.

Evolutionary signatures of invasiveness in E. crassipes

The collective outcome of genes that are highly expanded and duplicated, and genes that exhibits signs of adaptive evolution, suggest an association with four key pathways in E. crassipes, namely plant-pathogen interaction, plant hormone signal transduction, photosynthesis, and abiotic stress response. These pathways appear to be contributing significantly to the invasive traits of water hyacinth, since these are also known to be crucial for the establishment in a non-native or challenging environments as per the previous studies.30,31,32,33,34,35,36

Plant-pathogen interaction

Response to pathogens is a crucial trait in plants to survive and adapt to a new habitat and thus plays a critical role in invasiveness. Genes involved in pattern-triggered immunity (PTI) and effector triggered immunity (ETI) in E. crassipes showed various signs of adaptive evolution and high expansion of gene families corresponding with different duplication types (Figure 2). Genes responsible for hypersensitive response and stomatal closure; CNGCs, CDPK, CaM/CML, and RBOH were found to be highly expanded with duplications. CDPK and RBOH also showed unique amino acid substitution among these genes, and CaM/CML showed positive selection. Further, mitogen-activated protein kinase (MAPK) cascade genes, such as MEKK1/, were expanded with segmental and dispersed duplication, and MPK3/6 had unique substitution with functional impact. Further, the downstream-acting defense-related genes of the MAPK pathway, NHO1, showed positive selection, and PR1 was highly expanded. Also, among the genes involved in the molecular chaperone complex, a core modulator in plant immunity, HSP90 was highly expanded with segmental, dispersed duplications, and SGT1 showed unique amino acid substitution.37 Details of functions and roles of the genes involved in plant-pathogen interaction are mentioned in Data S5.Figure 2 Evolutionary signatures in plant-pathogen interaction

PTI; pattern-triggered immunity, ETI; effector-triggered immunity, PRRs; pattern recognition receptors, PAMPs; pathogen-associated molecular patterns, elf18; elongation factor Tu, flg22; flagellin, elf18; elongation factor 18, RBOH; respiratory burst oxidase homologs, CNGC; cyclic nucleotide-gated channel, CDPK; calcium-dependent protein kinases, CaM/CML; calmodulin (CaM) and calmodulin-like (CML), NOS; nitric oxide synthase, NO; nitric oxide, MEKK1; mitogen-activated protein kinase kinase kinase, MKK1/2/4/5; mitogen-activated protein kinase 1/2/4/5, MPK3/6; mitogen-activated protein kinase 3/6, MPK4; mitogen-activated protein kinase 4, FRK1; flg22-induced receptor-like kinase 1, NHO1; nonhost 1, PR1; pathogen-related 1, RIN; RPM1-interacting protein, RPM1; resistance to pseudomonas syringae pv maculicola 1, RPS; resistance to pseudomonas syringae protein 2, PBS1; AVRPPHB SUSCEPTIBLE1, HSP90; heat shock protein 90, SGT1; suppressor of G2 allele of skp1, RAR1; required forMla12 resistance, KCS1/10; 3-ketoacyl-CoA synthase 1/10, HCD1; 3-hydroxyacyl-coA dehydratase 1, PIK1; phosphatidylinositol 4-kinase 1, AvR; avirulence, P; phosphorylation; Ub; Ubiquitination, ?; Receptors are not characterized (created with BioRender.com).

Plant hormone signal transduction

Plant growth and response to environmental cues are largely regulated by plant hormones known as phytohormones, which provide a multitude of roles in an invasive species in terms of their vigorous growth, development, and defense.30,31 The genes involved in the signaling of these hormones in E. crassipes were found to be highly expanded, along with multiple signs of adaptive evolution. For example, Auxin is one of the most important plant growth regulators governing the responses to light and gravity, cell division, and expansion.38,39 In E. crassipes both TIR1, which is an F box protein responsible for the reception of auxin and its target AUX/IAA proteins were found to be highly expanded with segmental and dispersed duplication types. In addition, ARF which are known to be key transcription factor for auxin response were found to have unique amino acid substitution and highly expanded with segmental, dispersed, and proximal duplication types. Similar, to the auxin signaling pathway, genes involved in other phytohormone signaling were also expanded and showed evolutionary adaptation (Figure 3, Data S6).Figure 3 Evolutionary signatures in plant-hormone signaling pathways

PYR/PYL; pyrabactin resistance/PYR-like, PP2C; plant protein phosphatase 2C, SnRK2; serine/threonine protein kinase, ABF; ABRE binding factors, NPR1; nonexpresser of pathogen-related genes 1, TGA; TGACG sequence specific binding transcription factors, PR-1; pathogen related 1, CRE1; cytokinin response 1, AHP; arabidopsis histidine phosphotransfer proteins, B-ARR; type-B arabidopsis response regulator, BRI1; brassinosteroid insensitive 1, BAK; brassinosteroid insensitive 1-associated receptor kinase 1, BSK; brassinosteroid signaling kinase, BKI1; BRI1 kinase inhibitor 1, BSU1; BRI1 suppressor 1, BIN2; brassinosteroid insensitive 2, BZR1/2; brassinosteroid resistant 1/2, CYCD3; cyclin D3, GID1; gibberellin insensitive dwarf 1, GID2; gibberellin insensitive dwarf 2, TF; transcription factors, AUX1; auxin influx transporter, TIR1; transport inhibitor response 1, AUX/IAA; auxin/indole-3-acetic acid, ARF; auxin response factor, GH3; auxin-responsive Gretchen Hagen3, SAUR; small auxin upregulated RNA, JAR1; jasmonate resistant 1, JA-Ile; jasmonyl-L-isoleucine, COI1; coronatine-insensitive 1, JAZ; jasmonate ZIM-domain, MYC2; myelocytomatosis 2, ETR; ethylene receptor, CTR1; constitutive triple response, MPK6; mitogen-activated protein kinase 6, EIN2; ethylene insensitive 2, EIN3; ethylene insensitive 3, EBF1/2; EIN3-binding factor 1/2, ERF1/2; ethylene response factor 1/2, P; phosphorylation; Ub; Ubiquitination (created with BioRender.com).

Photosynthetic efficiency

Photosynthesis is an essential process in plants for growth, development, and biomass accumulation. In E. crassipes both the gene families Lhca and Lhcb that are associated with the light harvesting complex I and II, respectively, were found to be highly expanded (except Lhc6 gene) induced with various duplication types (Table S19). In addition, 11 out of 15 genes involved in carbon fixation through the C3 (Calvin-Benson) pathway showed adaptive evolution with high expansion, unique amino acid substitution with functional impact and/or positive selection (Figure 4). For example, the key enzyme Rubisco that catalyzes the first step of carbon assimilation in photosynthesis through the C3 pathway, was found to be highly expanded with all types of duplications and unique amino acid substitution with a functional impact. Further, all the genes of the C3 cycle with evolutionary signatures were found to have high TPM values (>20) in leaf tissue. Details of the functions and roles of the genes involved in the C3 cycle are mentioned in Data S7.Figure 4 Schematic representation of C3 (Calvin-Benson) pathway showing evolutionary signatures

Rubisco; ribulose bisphosphate carboxylase oxygenase, PGK; phosphoglycerate kinase, GAPDH; glyceraldehyde 3-phosphate dehydrogenase, FBPA; fructose-1,6-bisphosphate aldolase, FBPase; fructose-1,6-bisphosphatases, TK; transketolase, TPI; triosephosphate isomerase, SBPase; sedoheptulose-1,7-bisphosphatase, RPE; ribulose-phosphate 3-epimerase, RPI; ribose 5-phosphate isomerase A, PRK; phosphoribulokinase (created with BioRender.com).

Abiotic stress tolerance

Gene families involved in heavy metal sequestration, such as heavy metal ATPase (HMA), ATP-binding cassette (ABC) transporters, showed high expansion with segmental, dispersed, and tandem duplications with TPM >1 in root tissues. In which ABC also had proximal duplication with MSA. Further, among these heavy metal transporter gene families, natural resistance-associated macrophage proteins (NRAMPs) showed MSA.40 Additionally, genes that play a central role in scavenging the reactive oxygen species (ROS), like catalase (CAT), and superoxide dismutase (SOD), showed unique amino acid substitutions with functional impact, and peroxidase (POX) were among the MSA.41

Cell wall biosynthesis gene families

Cellulose synthase (CesA) and cellulose synthase-like (Csl) are the major gene families involved in the biosynthesis of the cell wall in plants. A total of 54 and 39 genes encoding CesA and Csl, respectively, were identified in E. crassipies, which was larger than that of maize (CesA; 20, Csl; 33). Additionally, these gene families were found to be highly expanded, majorly with segmental and dispersed duplications. However, the gene family CesA was filtered out in the gene family evolution analysis as it contained ≥100 gene copies, and all such families with more than 100 gene copies in one or more species were excluded from the gene family evolution analysis (see methodology; phylogenetic, and gene family evolution analysis).

Discussion

In this study, we present the whole genome sequence of E. crassipes. Though the genome of E. crassipes was complex due to its tetraploid nature and high repetitive content, the quality of the reported genome assembly was of “reference” standard as indicated by the complete BUSCO score and LAI value (Table 1). Further, the study established the phylogenetic position of E. crassipes among the 14 representative seed plant genomes, where it formed a sister clade with M. acuminata (family Musaceae; order Zingiberales), similar with APG IV classification42 (Figure 1B).

Whole genome duplication (WGD) in land plants has been recognized to be the major driving force for evolution and diversification. The resulting gene duplications endow genes with potential sub-functionalization and neo-functionalization, contributing to better adaptivity.28,29,43 Further, there exist various types of duplicates that show different evolutionary patterns in the protein-coding sequences; for example, tandem and proximal duplications have more rapid functional divergence as compared to more conserved duplication types (dispersed and whole genome/segmental duplications).44,45 In this study, from our WGD analysis, we inferred that E. crassipes had experienced a recent WGD event after diverging from its diploid species E. paniculata and both Eichhornia species shared one WGD event before diverging. Further, we concluded that both Commenlinales and Zingiberals have undergone the common WGD event. The recent WGD event (∼4 mya) in E. crassipes appears to be a prominent reason for the presence of a large number of coding genes with high D score (complete and duplicated BUSCOs) and a high number of whole genome/segmental duplications. Further, WGD along with other types of duplication events such as (dispersed, proximal, and tandem duplicates), were noted to be the major contributors in significant gene family expansion, particularly in the pathways that are associated with invasiveness, and which apparently helped in conferring the invasive character to the species43 (Figures 1E and 1F).

The invasive plant species have the ability to move aggressively into a habitat and monopolizing resources such as light, nutrients, water, and space and it mostly does all these with physical and physiological changes that are governed by selection pressure or sometimes predisposed in the genome.8,43,46 From our orthologous gene cluster and comparative evolutionary analysis with other, we found that the genes involved in plant-pathogen interaction, hormone signaling, abiotic stress tolerance, photosynthesis, and cell wall biosynthesis were highly evolved in E. crassipes genome that are important for its survival and proliferation in different habitats and contributes to its invasive character.

Invasive plants encounter novel biotic and abiotic stresses across the introduction-naturalization-invasion continuum that can act as a barrier to invasion.33 Interestingly in E. crassipes, the genes involved in plant-pathogen interaction were found to be evolved for both pattern-triggered immunity and effector-triggered immunity (Figure 2). Further, the gene families involved in heavy metal tolerance were highly expanded with MSA, and it was also noted that the adaptive evolution in these gene families was accompanied by the adaptive evolution in the genes responsible for ROS scavenging, which is an important observation because this correlation has been recognized as a critical mechanism by which plants can negate the detrimental effects of ROS produced from heavy metal accumulation (abiotic stress) without compromising the growth.34,35 Collectively, these evolutionary adaptations support the notion that E. crassipes have evolved to withstand biotic and abiotic stresses that are crucial in the process of colonizing new habitats (Data S8).

The evolution of increased growth rate, high productivity through evolved photosynthetic machinery, and the lack of trade-offs between growth and defenses are some advantageous traits of invasive species.36,47 Water hyacinth stands out as one of the fastest-growing plant with an estimated growth rate of 0.26 T dry substance d−1 (hm2)−148 and has a high photosynthetic efficiency of approximately 1.12 g m−2 day−1.35,48 We found evolutionary signatures in the genes involved in the plant hormone signaling, light harvest complexes (LHC), and C3 cycle in water hyacinth. The phytohormones are associated with growth and development along with tolerance to various biotic and abiotic stresses, and the genes involved in the light-harvesting complexes (LHCs) and C3 cycle are known to cf. high photosynthetic efficiency.49,50,51 Conclusively, these evolutionary adaptations could be among the major driving factors in the observed rapid growth and high photosynthetic efficiency of water hyacinth30,31 (Figures 3 and 4).

Cellulose is the major building block of cell wall in green plants that aids in providing mechanical strength, growth, and maintaining osmotic pressure.52,53 In E. crassipes, both cellulose synthase (CesA) and cellulose synthase-like (Csl) families were highly expanded due to duplications by WGD, which could be the reason for the high cellulose and hemicellulose content in E. crassipes. The high cellulose provides more flexibility in aquatic plants in contrast to lignin, which is found in high amount in terrestrial plants that have a rigid body structure.54,55 The presence of high cellulose and hemicellulose along with the rapid growth of water hyacinth due to its invasive character provides enormous capability to this plant for biofuel production as the conversion of these carbohydrates to fermentable sugars and subsequently into ethanol is easier as compared to lignin-rich plants.16,18,56

Taken together, it is tempting to speculate that the recent WGD in water hyacinth induced the expansion of gene families, and along with adaptive evolution in key pathways like plant-pathogen interaction, hormone signaling, abiotic stress tolerance, and photosynthesis, helped water hyacinth in becoming a globally invasive species. Nevertheless, the traits associated with its invasive character like heavy metal sequestration, good photosynthetic efficiency, high biomass accumulation, and high cellulose and hemicellulose content, make water hyacinth an excellent candidate for multiple commercial purposes, including its utilization as a renewable source of energy such as biofuel, biochar and as a phytoremediator for the treatment of wastewater. Further, the high photosynthetic efficiency of water hyacinth that act as great reservoir for carbon sequestration could be an alternative in tackling the impacts of climate change due to global warming. In summary, the genomic analysis of E. crassipes and insights into plant invasiveness, along with the translational potential of water hyacinth, will act as a valuable resource for understanding the invasive biology of plants with their utilization and management.

Limitations of the study

In this study, we performed the genome sequencing of E. crassipes with the aims of exploring genomic dynamics, functional analysis, and its invasive properties. Due to the allotetraploid nature of the species, the subgenome-specific analyses are desired to decipher functional redundancy and divergence at subgenome level. Further, Commelinales order of flowering plants comprises more than 800 species, and thus, the availability of genomic data and insights from more Commelinales will help in better understanding of shared WGD with Zingiberales.

STAR★Methods

Key resources table

REAGENT or RESOURCE	SOURCE	IDENTIFIER	
Biological samples	
	
Eichhornia crassipes	Bhopal, India (23.25°N 77.34° E)	N/A	
	
Deposited data	
	
Eichhornia crassipes genome sequencing	This study	SRA BioProject: PRJNA1028095	
Eichhornia crassipes transcriptome sequencing	This study	SRA BioProject: PRJNA1028095	
	
Oligonucleotides	
	
Primer for rbcl gene for species identification (Forward: 5’- ATGTCACCACAAACAGAGACTAAAGC-3’; Reverse: 5’- GAAACGGTCTCTCCAACGCAT-3’)	This study	N/A	
Primer for MatK gene for species identification (Forward: 5’- CGATCTATTCATTCAATATTTC -3’; Reverse: 5’- TCTAGCACACGAAAGTCGAAGT -3’)	This study	N/A	
	
Software and algorithms	
	
Trimmomatic	Bolger et al., 201457	http://www.usadellab.org/cms/?page=trimmomatic	
SGA-preqc	Simpson et al., 201424	https://github.com/jts/sga/tree/master/src/SGA	
Jellyfish	Marçais et al., 201158	https://github.com/gmarcais/Jellyfish	
GenomeScope2.0	Ranallo-Benavidez et al., 202025	http://genomescope.org/genomescope2.0/	
Guppy	N/A	https://community.nanoporetech.com/downloads/guppy/release_notes	
Porechop	N/A	https://github.com/rrwick/Porechop	
Wtdbg2	Ruan and Li, 2020	https://github.com/ruanjue/wtdbg2	
Canu	Koren et al., 201759	https://github.com/marbl/canu	
Flye	Kolmogorov et al., 201960	https://github.com/fenderglass/Flye	
Quast	Gurevich et al., 201361	http://bioinf.spbau.ru/quast	
Pilon	Walker et al., 201462	https://github.com/broadinstitute/pilon	
Platanus-allee	Kajitani et al., 201963	http://platanus.bio.titech.ac.jp/platanus2	
LINKS	Warren et al., 201564	https://github.com/bcgsc/LINKS	
AGOUTI	Zhang et al., 201665	https://github.com/svm-zhang/AGOUTI	
SPAdes	Prjibelski et al., 202066	https://github.com/ablab/spades	
Quickmerge	Chakraborty et al., 201667	https://github.com/mahulchak/quickmerge	
LR_Gapcloser	Xu et al., 201868	https://github.com/CAFS-bioinformatics/LR_Gapcloser	
TGS-GapCloser	Xu et al., 202069	https://github.com/BGI-Qingdao/TGS-GapCloser	
Purge haplotig	Roach et al., 201870	https://bitbucket.org/mroachawri/purge_haplotigs/src/master/	
BUSCO	Simão et al., 201571	https://busco.ezlab.org/	
MiniMap2	Li et al., 201872	https://github.com/lh3/minimap2	
BWA	Li, 201373	http://bio-bwa.sourceforge.net/	
SAMtools	Li et al., 200974	http://www.htslib.org/	
LTR_retriever	Ou et al., 201826	https://github.com/oushujun/LTR_retriever	
RepeatModeler	Flynn et al., 202075	http://www.repeatmasker.org/RepeatModeler/	
CD-HIT	Li and Godzik,2006	http://weizhong-lab.ucsd.edu/cd-hit/	
TEclass2	Bickmann et al., 202376	https://www.bioinformatics.uni-muenster.de/tools/teclass2/index.pl?	
RepeatMasker	N/A	http://www.repeatmasker.org	
MAKER	Cantarel et al., 200877	https://www.yandell-lab.org/software/maker.html	
Trinity	Haas et al., 201378	https://github.com/trinityrnaseq/trinityrnaseq/	
BLAST	Altschul et al., 199079	https://blast.ncbi.nlm.nih.gov/Blast.cgi	
Exonerate	N/A	https://github.com/nathanweeks/exonerate	
AUGUSTUS	Stanke et al., 200680	http://bioinf.uni-greifswald.de/augustus/	
SNAP	Korf, 200481	https://github.com/KorfLab/SNAP	
Kallisto	Bray et al., 201682	https://pachterlab.github.io/kallisto/	
tRNAscan	Chan et al., 202183	http://lowelab.ucsc.edu/tRNAscan-SE/	
Barrnap	N/A	https://github.com/tseemann/barrnap	
miRBase	Kozomara et al., 201984	https://www.mirbase.org/	
OrthoFinder	Emms and Kelly, 2019	https://github.com/davidemms/OrthoFinder	
KinFin	Laetsch and Blaxter, 2017	https://github.com/DRL/kinfin	
MAFFT	Katoh and Standley, 2013	https://www.ebi.ac.uk/Tools/msa/mafft/	
BeforePhylo	N/A	https://github.com/qiyunzhu/BeforePhylo	
RAxML	Stamatakis, 201485	https://cme.h-its.org/exelixis/web/software/raxml/	
PAML	Yang, 200786	http://abacus.gene.ucl.ac.uk/software/paml.html	
MCMCtree	Yang, 200786	http://abacus.gene.ucl.ac.uk/software/paml.html	
FigTree	N/A	http://tree.bio.ed.ac.uk/software/figtree/	
CAFÉ	De Bie et al., 2006	https://hahnlab.github.io/CAFE/index.html	
MCL	Van Dongen and Abreu-Goodger, 2012	https://micans.org/mcl/	
MCScanx	Wang et al., 201287	https://github.com/wyp1125/MCScanX	
wgd2	Chen et al., 202488	https://github.com/heche-psb/wgd	
OrthoVenn3	Sun et al., 202389	https://orthovenn3.bioinfotoolkits.net/start/db	
adephylo	Jombart and Dray, 2010	https://cran.r-project.org/web/packages/adephylo/index.html	
SIFT	Ng and Henikoff, 2003	https://sift.bii.a-star.edu.sg/	
HMMER	Finn et al., 201190	http://hmmer.org/	
KAAS	Moriya et al., 200791	https://github.com/PEHGP/kaas	
eggNOG-mapper	Huerta-Cepas et al., 201792	https://github.com/eggnogdb/eggnog-mapper	
shinyGO	Ge et al., 2020	http://bioinformatics.sdstate.edu/go/	
MMseqs2	Steinegger et al., 201793	https://github.com/soedinglab/MMseqs2	

Resource availability

Lead contact

Further information and requests for resources should be directed to and will be fulfilled by the Lead Contact, Vineet K. Sharma (vineetks@iiserb.ac.in).

Materials availability

This study did not generate new unique reagents.

Data and code availability

• Data associated with this study has been deposited at NCBI SRA database under the BioProject accession – PRJNA1028095, BioSample accession – SAMN37810889.

• This study does not report original code.

• We have provided all detailed information to reanalyze this study in STAR Methods to the best of our knowledge. Any additional information is available from the lead contact upon request.

Experimental model and study participant details

This study does not include experimental model or subjects.

Method details

Plant material, library construction, and sequencing

The floating rosette of the plant was collected from Bhopal Lake, India (23.25°N 77.34°E), transported to the laboratory, snap frozen, and stored at -20°C till further processing. DNA was extracted from leaves for species identification using Qiagen DNeasy Plant Mini Kit. The rbcL (ribulose bisphosphate carboxylase) and matK (maturase K) genes were amplified using PCR for the molecular identification of the plant. The amplified products were sequenced on the Sanger platform, and the sequences were BLAST searched on the NCBI database for species identification. After the species confirmation, high molecular weight DNA was extracted using CTAB-based lysis buffer (100 mM Tris HCl, 2% CTAB, 1.4 M NaCl, 20 mM EDTA, 1% PEG8000, pH 9) supplemented with β-mercaptoethanol. The leaves were crushed in liquid nitrogen and lysed using the buffer with subsequent treatment of Proteinase K (Qiagen) and RNase A (PureLink, Invitrogen). The lysate was extracted in the chloroform: isoamyl alcohol (24:1) solution twice, and the aqueous layer was transferred to a fresh tube. 0.7x volume isopropanol was added to the aqueous extract, and the DNA was precipitated by gently inverting the tube several times. The DNA was finally dissolved in nuclease-free water after washing with 70% ethanol. The DNA was purified, and longer fragment size DNA was selected using the Ampure XP beads for Nanopore long-read sequencing. The library was prepared using the ligation sequencing kit, SQK-LSK110 (Oxford Nanopore, UK), loaded onto FLO-MIN106 flow cell, and sequenced on MinION (Oxford Nanopore, UK) using MinKNOW software (version 22.08.6). For the short-read sequencing, the Illumina library was prepared with two insert sizes (350bp and 550bp) using the TruSeq DNA Nano LP kit. Final libraries were quantified using Qubit fluorometer using DNA HS assay kit and analyzed on TapeStation 4150 (Agilent) utilizing High sensitive D1000 ScreenTapes to determine the insert size. The library was sequenced on NovaSeq 6000 platform (Illumina, Inc., United States) for 150 bp paired-end reads. For the transcriptome, RNA was extracted from the root, leaf and stolon tissue using the TRIzol reagent and purified on the Qiagen RNeasy mini kit spin column after the on-column DNA digestion. The transcriptome library preparation was carried out using TruSeq Stranded Total RNA kit, and the final library was quantified on Qubit 4.0 fluorometer using DNA HS assay kit. The insert size of the library was determined by running on TapeStation 4150 (Agilent) utilizing High sensitive D1000 ScreenTapes following the manufacturer’s protocol. The library was sequenced on NovaSeq 6000 platform (Illumina, Inc., United States) for 150 bp paired-end reads.

Genomic data preprocessing, de novo genome assembly and assessment

Illumina raw reads obtained from two different libraries were quality filtered using Trimmomatic v0.3957 with parameters; LEADING:20 TRAILING:20 SLIDINGWINDOW:4:20 MINLEN:50. The genome size of E. crassipes was estimated using these filtered short read data using a k-mer distribution method implemented in SGA-preqc.24 Further, 21-kmers with jellyfish v2.3.158 software were calculated using the filtered Illumina reads and estimated the genome characteristics using GenomScope2.025 (Parameters -k 21 -p 4).

Long read assembly

Raw data obtained from the Oxford Nanopore Technology were base called using Guppy v3.2.1 (Oxford Nanopore Technologies), and Porechop v0.2.4 (Oxford Nanopore Technologies) was used to remove the adapter. De novo assembly was performed with different genome assemblers: wtdbg2,94 Canu v2.2,59 and Flye v2.9.160 using default parameters. Based on the assembly statistics obtained from Quast v5.2.0,61 the genome assembly from the Flye assembler was selected and polished three times with high-quality short read data using Pilon v1.23.62 The resulting contig level assembly was first scaffolded by Platanus-allee v2.2.263 “consensus” with default parameters, using filtered Illumina reads and adapter-filtered Nanopore reads. Further, scaffolding was done with the adapter-filtered Nanopore reads using LINKS v2.0.0.64 The second round of scaffolding was performed by AGOUTI v0.3.365 using trimmed and filtered paired-end transcriptome data generated using Trimmomatic v0.3957 with parameters; LEADING:15 TRAILING:15 SLIDINGWINDOW:4:15 MINLEN:50. The resulting assembly (V1) quality was evaluated by Quast v5.2.0.61

Short read assembly

The quality filtered short read from both the libraries were assembled together using SPAdes v3.15.366 with, K-mer values of 21, 33, 55, 77, 99, 121 and “--only-assembler” option. The obtained assembly was then scaffolded using Platanus-allee v2.2.263 “consensus” with default parameters and with the same data as used for long-read assembly scaffolding using Platanus-allee. Further scaffolding was performed using adapter-filtered Nanopore long-read data by LINKS v2.0.0.64 To further improve the contiguity of the assembly, quality-filtered RNA-Seq paired-end reads were used by AGOUTI v0.3.3.65 The assembly (V2) quality was then assessed using Quast v5.2.0.61

Assembly post-processing, final genome assembly, and assessment

The scaffold-level assemblies V1 and V2 were merged using Quickmerge v0.367 to obtain a more contiguous genome assembly. Gap-closing on this genome assembly was done by LR_Gapcloser68 (five iterations), and TGS-GapCloser v1.2.169 using pre-processed Nanopore reads. The gap-closed assembly was polished using Pilon v1.2362 to fix any small indels, base errors, and misassemblies. Purge haplotig v1.1.270 with default parameters (-a 75) was used to remove the redundant heterozygous region that contributed to the increase in the assembly size. Further, the obtained draft assembly was filtered to retain the scaffolds with length ≥ 5kbp (assembly V3).

To assess the final assembly (V3) quality, Quast v5.2.061 was used to obtain the assembly statistics and BUSCO v5.4.371 with the embryophyta_odb10 dataset to check the completeness of the genome. Further, Nanopore raw reads and Illumina reads (from both the libraries), were mapped to the final draft assembly using MiniMap2 v2.17,72 and BWA-MEM v0.7.1773 respectively, and mapping percentages were calculated by SAMtools v1.1374 “flagstat” utility. To evaluate the contiguity of intergenic and repetitive regions of genome assemblies by long terminal repeat (LTR) Assembly Index (LAI), LTR_retriever v2.9.026 was used in default parameters.

Gene set construction and annotation

The obtained genome assembly of E. crassipes from the previous step was used to construct the de novo Transposable Elements (TEs) library using RepeatModeler v2.0.3.75 These repeats were then clustered to eliminate the redundant sequences with a seed size of 8 bp and 90 % of sequence identity using CD-HIT-EST v4.8.1,95 further TEclass276 software was used to classify the unknown TEs with a probability threshold of >0.65 and the resulting library was utilized by RepeatMasker v4.1.2 (http://www.repeatmasker.org) to soft mask the repeats present in the genome of E. crassipes. The insertion time of intact LTR retrotransposon was calculated using LTR_retriever with -u 6.5 X 10-9.96,97

The soft-masked genome was used for gene set construction using MAKER v2.31.1077 in three iterations by employing both ab initio and evidence-based gene prediction models. In the first round, Quality-filtered transcriptome reads from three tissues (Leaf, Root and Stolon) were assembled using Trinity v2.14.078 with default parameters. The resulting transcriptome assembly and protein sequence of seven monocot species (Musa acuminata, Oryza sativa, Zea mays, Triticum aestivum, Sorghum bicolor, Avena sativa and Saccharum spontaneum; obtained from Ensembl Plant release 56) were used as EST and protein evidence, respectively, for the construction of gene set using BLAST79 and Exonerate v2.2.0 (https://github.com/nathanweeks/exonerate). The output of the first round of MAKER was then used for ab initio gene prediction using AUGUSTUS v3.2.380 and SNAP v1.0.81 The gene set obtained after the third round of MAKER pipeline was further filtered based on Annotation Edit Distance (AED) values and length-based criteria to retain only the genes with AED values < 0.3, and length (≥150 bp) in the final high confidence gene set of E. crassipes. The final gene set was then evaluated for its completeness using BUSCO v5.4.371 with the embryophyta_odb10 dataset. Gene expression values in terms of transcript per million (TPM) were also estimated using Kallisto v0.48.082 by mapping the quality-filtered RNA-Seq data of E. crassipes onto the coding gene set (nucleotides).

Additionally, tRNAscan v2.0.9,83 Barrnap v0.9 (https://github.com/tseemann/barrnap) and miRBase database84 (sequence identity 80%, e-value 10-3) for tRNA, rRNA and miRNA predictions, respectively.

Phylogenetic and gene family evolution analysis

Protein sequences of Asparagus officinalis, Zea mays, Echinochloa crus-galli, Oryza sativa, Ananas comosus, Musa acuminata, Vitis vinifera, Manihot esculenta, Arabidopsis thaliana, Glycine max, Nicotiana attenuata, Cynara cardunculus, Nymphaea colorata, and Amborella trichopoda were retrieved from Ensembl plant release 56 to construct the phylogeny with respect to E. crassipes. The protein sequences of the longest isoforms were extracted, and orthogroups were constructed using OrthoFinder v2.5.4.98 From these orthogroups, fuzzy one-to-one orthogroups containing the protein sequences were extracted using KinFin v1.2,99 which were further aligned by using MAFFT v7.467.100 The obtained multiple sequence alignments were concatenated and processed to remove empty sites using BeforPhylo v0.9.0 (https://github.com/qiyunzhu/BeforePhylo). RAxML v8.2.12 103 was used to construct the maximum likelihood-based species phylogenetic tree using the ‘PROTGAMMAAUTO’ amino acid substitution model with 100 bootstrap values. For the divergence time estimation, maximum likelihood estimates of branch lengths, the gradient vector, and the Hessian matrix were initially calculated using the codeml program implemented in PAML,86 using WAG+Gamma model of amino acid substitution matrix with five calibration points (Table S20). These calculations were further used to run MCMCtree program86 with -burnin 100000 -nsample 1000000 parameters, and the time-tree was visualized using FigTree (http://tree.bio.ed.ac.uk/software/figtree/) .

The evolution of gene families was estimated using CAFÉ v5,101 to model gene gain and loss across the obtained phylogenetic tree from the above step. At first, the species phylogenetic tree was converted to an ultrametric tree using the calibration point of 196 million years between E. crassipes and A. trichopoda obtained from the TimeTree database v5.0 (http://www.timetree.org/). BLASTP in all-by-all mode was performed to search the similar protein sequences of the longest isoform, which was followed by clustering using MCL v14-137102 and removal of those gene families that contained genes in less than two species of the specified clades and ≥ 100 gene copies for at least one species. The filtered gene families output and the ultrametric species tree were used to analyse the expansion and contraction of gene families by employing a two-lambda (λ) model. The expanded gene families that consist of ≥10 genes for E. crassipes were extracted and annotated.

Genome synteny and whole-genome duplication analysis

To investigate the large number of predicted high confidence protein-coding genes in E. crassipes, and high duplicated BUSCO score (D score), Whole-genome duplication (WGD) events were analyzed by identification of intra-species syntenic blocks in E. crassipes genome using MCScanx87 with at least five gene pairs required per syntenic block. Further, to classify the duplicated genes in various categories (whole genome/segmental, tandem, proximal and dispersed), we used duplicate_gene_classifier implemented in MCScanx87 using the same parameter. Before performing WGD analysis, genome assembly of diploid Eichhornia species (Eichhornia paniculata; CoGe ID: 64302) was annotated, with a similar approach employed for E. crassipes (Table S21). Whole genome dot plots and whole paranome and orthologs Ks (number of synonymous substitutions per synonymous site) distributions of E. crassipes, E. paniculata, and M. acuminata genome were constructed using wgd2.88 Further, the rough dating of the genome duplication events was done using the Ks values. The peaks of Ks profiles were determined and translated into divergence time (T) in millions of years using the following formula: T = Ks/2r, where T is the time of divergence and r is the synonymous-site mutation rate (6.5 X 10-9 sites/year).96,97

Orthologous gene clustering and comparative evolutionary analysis of invasive and non-invasive species

To identify the genes responsible for the invasiveness of E. crassipes, gene clustering analysis was performed between invasive (Eichhornia crassipes, Phragmites australis, Lemna minuta) and non-invasive (Brasenia schreberi, Lemna minor, Ceratophyllum demersum) aquatic angiosperm using the OrthoVennn3.89 Further, to gain evolutionary insights into the invasive characteristics of E. crassipes in comparison to other invasive and non-invasive aquatic plants, a comparative evolutionary analysis was performed with selected nine invasive and five non-invasive aquatic angiosperms (Table S22).

Proteome files were downloaded for these 14 species from various genomic databases (Table S22). The ortholog gene sets were constructed using OrthoFinder v2.5.4 100. Orthogroups containing protein sequences of the longest isoforms from all 14 species were identified and extracted.

Genes with high divergence rates

The protein sequences of orthogroups across 15 species were individually aligned using MAFFT v7.467102, which was further used to construct maximum likelihood-based phylogenetic trees using RAxML v8.2.12103 for each orthogroup with “PROTGAMMAAUTO” model with 100 bootstrap value. To calculate the root-to-tip branch length distance values, R package ‘adephylo’103 was used, and the genes of E. crassipes that showed higher values in comparison to other species were extracted and categorized as genes with higher nucleotide divergence.

Genes with unique amino acid substitution with functional impact

The aligned protein sequences obtained from the previous step were used to identify genes in E. crassipes that had different amino acids in position compared to other species (ten amino acids around any gap in the alignments were excluded). These genes with different amino acid compositions were then analyzed using SIFT with UniProt as a database104 for the identification of any functional impact governed by the change.

Genes with positive selection

The nucleotide sequences corresponding to all orthologous gene set across the 15 species were aligned using MAFFT v7.467102. Further, a species phylogenetic tree was constructed for these 10 species using RAxML v8.2.12103. The resulted PHYLIP format containing nucleotide alignments and species phylogenetic tree were used to identify E. crassipes genes with positive selection using PAML v4.10.6105 “codeml” program with branch-site model. Likelihood-ratio tests were performed on the calculated log-likelihood values and genes termed positively selected when they passed the null model threshold with FDR-corrected p-values of less than 0.05 (obtained from chi-square analysis). In addition, Bayes Empirical Bayes (BEB) analysis was performed to identify the codon sites of these positively selected genes that showed >95% probability for the foreground lineage.

The genes with more than one sign of adaptive evolution; higher nucleotide divergence rate, unique substitution with functional impact, and positive selection, were considered to be the genes of multiple signs of adaptive evolution (MSA genes).105,106

Identification of gene families involved in cell-wall-biosynthesis

Cell wall genomics (https://cellwall.genomics.purdue.edu/families/) database was used to download the amino acid sequences of known cellulose synthase (CesA) and cellulose synthase-like (Csl) genes of maize. The obtained sequences were used for BLAST against the protein data set of E. crassipes with e-value 1e-9. Further, HMMER was used to search the HMM profiles of CesA (Pfam: PF03552) and Csl (Pfam: PF00535) obtained from Pfam database (http://pfam.xfam.org) to search against protein data set of E. crassipes with E-value 1e-5. The sequences that were common in both the BLAST and HMMER results were retained, such that every selected candidate sequence will have the BLAST query and own either of the two Pfam domains.

Functional annotation

The high confidence gene set was functionally annotated against the public databases like NCBI Non-Redundant (nr) database using BLASTP (e-value of 10−5), Pfam-A v32.0107 database using HMMER v3.3 116 (e-value of 10−5), and Swiss-Prot108 database using BLASTP (e-value of 10−5). Further, E. crassipes complete coding gene set, genes with signs of adaptive evolution and gene families with high expansion (≥10 genes) were assigned to KEGG Orthology (KO) categories and KEGG pathways using KAAS v2.191 genome annotation server. Coding genes and MSA genes were also annotated to different Clusters of Orthologous Groups (COG) categories using eggNOG-mapper v2.1.12.92 In addition, genes that were highly expanded by different types of duplication in E. crassipes were annotated for major KEGG pathways using KASS v2.191 and further KEGG enrichment analysis was performed using shinyGO v0.77109 by comparing these genes of E. crassipes to closely related species M. acuminata using MMseqs2 v13.4593 with the maximum sensitivity (-s 7.0).

Quantification and Statistical analysis

Statistical tests in positive selection were performed using PAML v4.10.6 (Yang, 2007). TPM values of the coding genes for E. crassipes were quantified using Kallisto v0.48.0 (Bray et al., 2016).

Additional resources

This study does not report additional resources.

Supplemental information

Document S1. Figures S1–S8 and Tables S1–S22

Data S1. Annotation of highly expanded gene families (≥10 genes), related to STAR Methods

Data S2. Annotation of invasive species core gene clusters, related to Figure S8B

Data S3. Annotation of E. crassipes specific gene clusters, related to Figure S8B

Data S4. List of positively selected genes in E. crassipes, related to STAR Methods

Data S5. Detail information of genes involved in plant-pathogen interaction, related to Figure 2

Data S6. Detail information of genes involved in plant hormone signaling, related to Figure 3

Data S7. Detail information of genes involved C3 pathway, related to Figure 4

Data S8. List of MSA genes involved in biotic and abiotic stress tolerance, related to Table S18

Acknowledgments

MSB thanks Ministry of Education, Govt. of India for Prime Minister Research Fellowship (PMRF). M.S. and A.C. thank Council of Scientific and Industrial Research (CSIR) for the fellowship. The authors also thank the Sanger sequencing and NGS facility at IISER Bhopal and the intramural research funds provided by 10.13039/501100010713 IISER Bhopal .

Author contributions

V.K.S. conceived and coordinated the project. M.S.B. collected the plant samples. M.S. performed DNA-RNA extraction, prepared the samples for sequencing, performed Nanopore sequencing, and species identification. M.S.B. and V.K.S. designed the computational framework of the study. M.S.B. performed all the computational analyses presented in the study with inputs from A.C. M.S.B. constructed all the figures. M.S.B., A.C., and V.K.S. interpreted the results. M.S.B. and M.S. annotated the genes in highly expanded gene families and MSA. M.S.B., M.S., and V.K.S. wrote the manuscript. All the authors have read and approved the final version of the manuscript.

Declaration of interests

The authors declare no competing interests.

Supplemental information can be found online at https://doi.org/10.1016/j.isci.2024.110698.
==== Refs
References

1 Courchamp F. Fournier A. Bellard C. Bertelsmeier C. Bonnaud E. Jeschke J.M. Russell J.C. Invasion Biology: Specific Problems and Possible Solutions Specific Difficulties Related to Invasion Biology Trends Ecol. Evol. 32 2017 13 22 10.1016/j.tree.2016.11.001 27889080
2 Kumar Rai P. Singh J.S. Invasive alien plant species: Their impact on environment, ecosystem services and human health Ecol. Indicat. 111 2020 106020 10.1016/J.ECOLIND.2019.106020
3 Pyšek P. Richardson D.M. Invasive Species, Environmental Change and Management, and Health Annu. Rev. Environ. Resour. 35 2010 25 55 10.1146/annurev-environ-033009-095548
4 Simberloff D. Von Holle B. Positive interactions of nonindigenous species: Invasional meltdown? Biol. Invasions 1 1999 21 32 10.1023/A:1010086329619/METRICS
5 Keane R. Crawley M.J. Exotic plant invasions and the enemy release hypothesis Trends Ecol. Evol. 17 2002 164 170 10.1016/S0169-5347(02)02499-0
6 Lee C.E. Gelembiuk G.W. Evolutionary origins of invasive populations Evol. Appl. 1 2008 427 448 10.1111/J.1752-4571.2008.00039.X 25567726
7 Ottenburghs J. The genic view of hybridization in the Anthropocene Evol. Appl. 14 2021 2342 2360 10.1111/EVA.13223 34745330
8 North H.L. McGaughran A. Jiggins C.D. Insights into invasive species from whole-genome resequencing Mol. Ecol. 30 2021 6289 6308 10.1111/MEC.15999 34041794
9 Oh D.H. Kowalski K.P. Quach Q.N. Wijesinghege C. Tanford P. Dassanayake M. Clay K. Novel genome characteristics contribute to the invasiveness of Phragmites australis (common reed) Mol. Ecol. 31 2022 1142 1159 10.1111/MEC.16293 34839548
10 Rojas-Sandoval J. Acevedo-Rodríguez P. Eichhornia crassipes (Water Hyacinth). Invasive Species Compendium, Centre for Agriculture and Biosciences International, Wallingford http://www.cabi.org/isc/datasheet/20544
11 GISD http://www.iucngisd.org/gisd/speciesname/Eichhornia+crassipes.
12 Isa H. Egbuche K.C. Malgwi M.M. Tukur N.A. Cytological studies in Eichhornia crassipes (Mart.) solms Am. J. Plant Physiol. 8 2013 50 62 10.3923/AJPP.2013.50.62
13 Gopal B. Biocontrol with arthropods 1987 Water Hyacinth 208 230
14 Pegg J. South J. Hill J.E. Durland-Donahou A. Weyl O.L. Impacts of alien invasive species on large wetlands Fundamentals of Tropical Freshwater Wetlands: From Ecology to Conservation Management 2022 Elsevier 487 516 10.1016/B978-0-12-822362-8.00018-9
15 Mahmood Q. Siddiqi M.R. Islam E.u. Azim M.R. Zheng P. Hayat Y. Anatomical studies on water hyacinth (Eichhornia crassipes (Mart.) Solms) under the influence of textile wastewater J. Zhejiang Univ. - Sci. B 6 2005 991 998 10.1631/JZUS.2005.B0991 16187412
16 Bote M.A. Naik V.R. Jagadeeshgouda K.B. Review on water hyacinth weed as a potential bio fuel crop to meet collective energy needs Mater. Sci. Energy Technol. 3 2020 397 406 10.1016/J.MSET.2020.02.003
17 Ganesh P.S. Ramasamy E.V. Gajalakshmi S. Abbasi S.A. Extraction of volatile fatty acids (VFAs) from water hyacinth using inexpensive contraptions, and the use of the VFAs as feed supplement in conventional biogas digesters with concomitant final disposal of water hyacinth as vermicompost Biochem. Eng. J. 27 2005 17 23 10.1016/J.BEJ.2005.06.010
18 Shanab S.M.M. Hanafy E.A. Shalaby E.A. Water Hyacinth as Non-edible Source for Biofuel Production Waste Biomass Valorization 9 2018 255 264 10.1007/S12649-016-9816-6
19 Priya P. Nikhitha S.O. Anand C. Dipin Nath R.S. Krishnakumar B. Biomethanation of water hyacinth biomass Bioresour. Technol. 255 2018 288 292 10.1016/j.biortech.2018.01.119 29428784
20 Vymazal J. Constructed Wetlands, Surface Flow Encyclopedia of Ecology, Five-Volume Set 2008 Elsevier 765 776 10.1016/B978-008045405-4.00079-3
21 Lahon D. Sahariah D. Debnath J. Nath N. Meraj G. Farooq M. Kanga S. Singh S.K. Chand K. Growth of water hyacinth biomass and its impact on the floristic composition of aquatic plants in a wetland ecosystem of the Brahmaputra floodplain of Assam, India PeerJ 11 2023 e14811 10.7717/PEERJ.14811/SUPP-5
22 Jones J.L. Jenkins R.O. Haris P.I. Extending the geographic reach of the water hyacinth plant in removal of heavy metals from a temperate Northern Hemisphere river OPEN Sci. Rep. 8 2018 11071 10.1038/s41598-018-29387-6
23 Yang J.S. Qian Z.H. Shi T. Li Z.Z. Chen J.M. Chromosome-level genome assembly of the aquatic plant Nymphoides indica reveals transposable element bursts and NBS-LRR gene family expansion shedding light on its invasiveness DNA Res. 29 2022 dsac022 10.1093/dnares/dsac022
24 Simpson J.T. Exploring genome characteristics and sequence quality without a reference Bioinformatics 30 2014 1228 1235 10.1093/BIOINFORMATICS/BTU023 24443382
25 Ranallo-Benavidez T.R. Jaron K.S. Schatz M.C. GenomeScope 2.0 and Smudgeplot for reference-free profiling of polyploid genomes Nat. Commun. 11 2020 1432 10.1038/s41467-020-14998-3 32188846
26 Ou S. Jiang N. LTR_retriever: A Highly Accurate and Sensitive Program for Identification of Long Terminal Repeat Retrotransposons Plant Physiol. 176 2018 1410 1422 10.1104/PP.17.01310 29233850
27 Kim N.-S. The genomes and transposable elements in plants: are they friends or foes? Genes Genomics 39 2017 359 370 10.1007/s13258-017-0522-y
28 Qiao X. Li Q. Yin H. Qi K. Li L. Wang R. Zhang S. Paterson A.H. Gene duplication and evolution in recurring polyploidization–diploidization cycles in plants Genome Biol. 20 2019 1 23 10.1186/S13059-019-1650-2 30606230
29 Liu Y. Zhang X. Han K. Li R. Xu G. Han Y. Cui F. Fan S. Seim I. Fan G. Insights into amphicarpy from the compact genome of the legume Amphicarpaea edgeworthii Plant Biotechnol. J. 19 2021 952 965 10.1111/PBI.13520 33236503
30 Gray W.M. Hormonal Regulation of Plant Growth and Development PLoS Biol. 2 2004 E311 10.1371/JOURNAL.PBIO.0020311
31 Pal P. Ansari S.A. Jalil S.U. Ansari M.I. Regulatory role of phytohormones in plant growth and development Plant Horm. Crop Improv. 2023 1 13 10.1016/B978-0-323-91886-2.00016-1
32 Xiao L. Ding J. Zhang J. Huang W. Siemann E. Chemical responses of an invasive plant to herbivory and abiotic environments reveal a novel invasion mechanism Sci. Total Environ. 741 2020 140452 10.1016/J.SCITOTENV.2020.140452
33 Manoharan B. Qi S.S. Dhandapani V. Chen Q. Rutherford S. Wan J.S. Jegadeesan S. Yang H.Y. Li Q. Li J. Gene Expression Profiling Reveals Enhanced Defense Responses in an Invasive Weed Compared to Its Native Congener During Pathogenesis Int. J. Mol. Sci. 20 2019 4916 10.3390/IJMS20194916
34 Malar S. Vikram S.S. Favas P.J.C. Perumal V. Lead heavy metal toxicity induced changes on growth and antioxidative enzymes level in water hyacinths [Eichhornia crassipes (Mart.)] Bot. Stud. 55 2016 1 11 10.1186/S40529-014-0054-6
35 Peixoto R.B. Marotta H. Bastviken D. Enrich-Prast A. Floating Aquatic Macrophytes Can Substantially Offset Open Water CO2 Emissions from Tropical Floodplain Lake Ecosystems Ecosystems 19 2016 724 736 10.1007/S10021-016-9964-3
36 Montesinos D. Fast invasives fastly become faster: Invasive plants align largely with the fast side of the plant economics spectrum J. Ecol. 110 2022 1010 1014 10.1111/1365-2745.13616
37 Seo Y.S. Lee S.K. Song M.Y. Suh J.P. Hahn T.R. Ronald P. Jeon J.S. The HSP90-SGT1-RAR1 molecular chaperone complex: A core modulator in plant immunity J. Plant Biol. 51 2008 1 10 10.1007/BF03030734/METRICS
38 Zhao Y. Auxin biosynthesis and its role in plant development Annu. Rev. Plant Biol. 61 2010 49 64 10.1146/ANNUREV-ARPLANT-042809-112308 20192736
39 Vandenbrink J.P. Kiss J.Z. Herranz R. Medina F.J. Light and gravity signals synergize in modulating plant development Front. Plant Sci. 5 2014 563 10.3389/FPLS.2014.00563
40 Hall J.L. Williams L.E. Transition metal transporters in plants J. Exp. Bot. 54 2003 2601 2613 10.1093/JXB/ERG303 14585824
41 Malar S. Shivendra Vikram S. Favas P.J. Perumal V. Lead heavy metal toxicity induced changes on growth and antioxidative enzymes level in water hyacinths [Eichhornia crassipes (Mart.)] Botanical studies 55 2014 1 11 10.1186/s40529-014-0054-6 28510906
42 Chase M.W. Christenhusz M.J.M. Fay M.F. Byng J.W. Judd W.S. Soltis D.E. Mabberley D.J. Sennikov A.N. Soltis P.S. Stevens P.F. An update of the Angiosperm Phylogeny Group classification for the orders and families of flowering plants: APG IV Bot. J. Linn. Soc. 181 2016 1 20 10.1111/BOJ.12385
43 Turcotte M.M. Kaufmann N. Wagner K.L. Zallek T.A. Ashman T.L. Neopolyploidy increases stress tolerance and reduces fitness plasticity across multiple urban pollutants: support for the “general-purpose” genotype hypothesis Evol. Lett. 8 2024 416 426 10.1093/EVLETT/QRAD072 38818423
44 Blanc G. Wolfe K.H. Functional Divergence of Duplicated Genes Formed by Polyploidy during Arabidopsis Evolution Plant Cell 16 2004 1679 1691 10.1105/TPC.021410 15208398
45 Qiao X. Yin H. Li L. Wang R. Wu J. Wu J. Zhang S. Different modes of gene duplication show divergent evolutionary patterns and contribute differently to the expansion of gene families involved in important fruit traits in pear (Pyrus bretschneideri) Front. Plant Sci. 9 2018 161 10.3389/FPLS.2018.00161/FULL
46 Baker H.G. Stebbins G.L. Genetics of Colonizing Species, Proceedings 172 1966 Stony Brook Foundation, Inc. S1 S3 10.1086/405172
47 Dudeque Zenni R. Lacerda da Cunha W. Sena G. Rapid increase in growth and productivity can aid invasions by a non-native tree AoB Plants 8 2016 48 10.1093/AOBPLA/PLW048
48 Xu J. Li X. Gao T. The Multifaceted Function of Water Hyacinth in Maintaining Environmental Sustainability and the Underlying Mechanisms: A Mini Review Int. J. Environ. Res. Publ. Health 19 2022 16725 10.3390/IJERPH192416725
49 Raines C.A. Increasing Photosynthetic Carbon Assimilation in C3 Plants to Improve Crop Yield: Current and Future Strategies Plant Physiol. 155 2011 36 42 10.1104/PP.110.168559 21071599
50 Weatherby K. Carter D. Chromera velia: The Missing Link in the Evolution of Parasitism Adv. Appl. Microbiol. 85 2013 119 144 10.1016/B978-0-12-407672-3.00004-6 23942150
51 Zhao L. Lü Y. Chen W. Yao J. Li Y. Li Q. Pan J. Fang S. Sun J. Zhang Y. Genome-wide identification and analyses of the AHL gene family in cotton (Gossypium) BMC Genom. 21 2020 1 14 10.1186/S12864-019-6406-6
52 Cosgrove D.J. Re-constructing our models of cellulose and primary cell wall assembly Curr. Opin. Plant Biol. 22 2014 122 131 10.1016/J.PBI.2014.11.001 25460077
53 Rongpipi S. Ye D. Gomez E.D. Gomez E.W. Progress and opportunities in the characterization of cellulose – an important regulator of cell wall growth and mechanics Front. Plant Sci. 9 2019 410940 10.3389/FPLS.2018.01894
54 Asaeda T. Fujino T. Manatunge J. Morphological adaptations of emergent plants to water flow: a case study with Typha angustifolia, Zizania latifolia and Phragmites australis Freshw. Biol. 50 2005 1991 2001 10.1111/J.1365-2427.2005.01445.X
55 Rabemanolontsoa H. Saka S. Comparative study on chemical composition of various biomass species RSC Adv. 3 2013 3946 3956 10.1039/C3RA22958K
56 Alagu K. Venu H. Jayaraman J. Raju V.D. Subramani L. Appavu P. S D. Novel water hyacinth biodiesel as a potential alternative fuel for existing unmodified diesel engine: Performance, combustion and emission characteristics Energy 179 2019 295 305 10.1016/J.ENERGY.2019.04.207
57 Bolger A.M. Lohse M. Usadel B. Trimmomatic: a flexible trimmer for Illumina sequence data Bioinformatics 30 2014 2114 2120 10.1093/BIOINFORMATICS/BTU170 24695404
58 Marçais G. Kingsford C. A fast, lock-free approach for efficient parallel counting of occurrences of k-mers Bioinformatics 27 2011 764 770 10.1093/BIOINFORMATICS/BTR011 21217122
59 Koren S. Walenz B.P. Berlin K. Miller J.R. Bergman N.H. Phillippy A.M. Canu: Scalable and accurate long-read assembly via adaptive κ-mer weighting and repeat separation Genome Res. 27 2017 722 736 10.1101/GR.215087.116 28298431
60 Kolmogorov M. Yuan J. Lin Y. Pevzner P.A. Assembly of long, error-prone reads using repeat graphs Nat. Biotechnol. 37 2019 540 546 10.1038/s41587-019-0072-8 30936562
61 Gurevich A. Saveliev V. Vyahhi N. Tesler G. Genome analysis QUAST: quality assessment tool for genome assemblies Bioinformatics 29 2013 1072 1075 10.1093/bioinformatics/btt086 23422339
62 Walker B.J. Abeel T. Shea T. Priest M. Abouelliel A. Sakthikumar S. Cuomo C.A. Zeng Q. Wortman J. Young S.K. Earl A.M. Pilon: An Integrated Tool for Comprehensive Microbial Variant Detection and Genome Assembly Improvement PLoS One 9 2014 e112963 10.1371/JOURNAL.PONE.0112963
63 Kajitani R. Yoshimura D. Okuno M. Minakuchi Y. Kagoshima H. Fujiyama A. Kubokawa K. Kohara Y. Toyoda A. Itoh T. Platanus-allee is a de novo haplotype assembler enabling a comprehensive access to divergent heterozygous regions Nat. Commun. 10 2019 1702 10.1038/s41467-019-09575-2
64 Warren R.L. Yang C. Vandervalk B.P. Behsaz B. Lagman A. Jones S.J.M. Birol I. LINKS: Scalable, alignment-free scaffolding of draft genomes with long reads GigaScience 4 2015 35 10.1186/S13742-015-0076-3/2707579
65 Zhang S.V. Zhuo L. Hahn M.W. AGOUTI: Improving genome assembly and annotation using transcriptome data GigaScience 5 2016 31 10.1186/S13742-016-0136-3/2558793
66 Prjibelski A. Antipov D. Meleshko D. Lapidus A. Korobeynikov A. Using SPAdes De Novo Assembler Curr. Protoc. Bioinformatics 70 2020 e102 10.1002/CPBI.102 32559359
67 Chakraborty M. Baldwin-Brown J.G. Long A.D. Emerson J.J. Contiguous and accurate de novo assembly of metazoan genomes with modest long read coverage Nucleic Acids Res. 44 2016 e147 10.1093/NAR/GKW654 27458204
68 Xu G.C. Xu T.J. Zhu R. Zhang Y. Li S.Q. Wang H.W. Li J.T. LR Gapcloser: a tiling path-based gap closer that uses long reads to complete genome assembly GigaScience 8 2018 1 14 10.1093/gigascience/giy157
69 Xu M. Guo L. Gu S. Wang O. Zhang R. Peters B.A. Fan G. Liu X. Xu X. Deng L. Zhang Y. TGS-GapCloser: A fast and accurate gap closer for large genomes with low coverage of error-prone long reads GigaScience 9 2020 giaa094 10.1093/gigascience/giaa094 32893860
70 Roach M.J. Schmidt S.A. Borneman A.R. Purge Haplotigs: Allelic contig reassignment for third-gen diploid genome assemblies BMC Bioinf. 19 2018 460 10.1186/S12859-018-2485-7
71 Manni M. Berkeley M.R. Seppey M. Simão F.A. Zdobnov E.M. BUSCO Update: Novel and Streamlined Workflows along with Broader and Deeper Phylogenetic Coverage for Scoring of Eukaryotic, Prokaryotic, and Viral Genomes Mol. Biol. Evol. 38 2021 4647 4654 10.1093/MOLBEV/MSAB199 34320186
72 Li H. Minimap2: pairwise alignment for nucleotide sequences Bioinformatics 34 2018 3094 3100 10.1093/BIOINFORMATICS/BTY191 29750242
73 Li H. Durbin R. Fast and accurate short read alignment with Burrows–Wheeler transform Bioinformatics 25 2009 1754 1760 10.1093/BIOINFORMATICS/BTP324 19451168
74 Li H. Handsaker B. Wysoker A. Fennell T. Ruan J. Homer N. Marth G. Abecasis G. Durbin R. 1000 Genome Project Data Processing Subgroup The Sequence Alignment/Map format and SAMtools Bioinformatics 25 2009 2078 2079 10.1093/BIOINFORMATICS/BTP352 19505943
75 Flynn J.M. Hubley R. Goubert C. Rosen J. Clark A.G. Feschotte C. Smit A.F. RepeatModeler2 for automated genomic discovery of transposable element families Proc. Natl. Acad. Sci. USA 117 2020 9451 9457 10.1073/PNAS.1921046117 32300014
76 Bickmann L. Rodriguez M. Jiang X. Makalowski W. TEclass2: Classification of transposable elements using Transformers Preprint at bioRxiv 2023 10.1101/2023.10.13.562246
77 Cantarel B.L. Korf I. Robb S.M.C. Parra G. Ross E. Moore B. Holt C. Sánchez Alvarado A. Yandell M. MAKER: An easy-to-use annotation pipeline designed for emerging model organism genomes Genome Res. 18 2008 188 196 10.1101/GR.6743907 18025269
78 Haas B.J. Papanicolaou A. Yassour M. Grabherr M. Blood P.D. Bowden J. Couger M.B. Eccles D. Li B. Lieber M. De novo transcript sequence reconstruction from RNA-seq using the Trinity platform for reference generation and analysis Nat. Protoc. 8 2013 1494 1512 10.1038/nprot.2013.084 23845962
79 Altschul S.F. Gish W. Miller W. Myers E.W. Lipman D.J. Basic local alignment search tool J. Mol. Biol. 215 1990 403 410 10.1016/S0022-2836(05)80360-2 2231712
80 Stanke M. Keller O. Gunduz I. Hayes A. Waack S. Morgenstern B. AUGUSTUS: ab initio prediction of alternative transcripts Nucleic Acids Res. 34 2006 W435 W439 10.1093/NAR/GKL200 16845043
81 Korf I. Gene finding in novel genomes BMC Bioinf. 5 2004 59 10.1186/1471-2105-5-59
82 Bray N.L. Pimentel H. Melsted P. Pachter L. Near-optimal probabilistic RNA-seq quantification Nat. Biotechnol. 34 2016 525 527 10.1038/nbt.3519 27043002
83 Chan P.P. Lowe T.M. tRNAscan-SE: Searching for tRNA genes in genomic sequences Methods Mol. Biol. 1962 2019 1 14 10.1007/978-1-4939-9173-0_1 31020551
84 Kozomara A. Birgaoanu M. Griffiths-Jones S. miRBase: from microRNA sequences to function Nucleic Acids Res. 47 2019 D155 D162 10.1093/NAR/GKY1141 30423142
85 Stamatakis A. RAxML version 8: a tool for phylogenetic analysis and post-analysis of large phylogenies Bioinformatics 30 2014 1312 1313 10.1093/BIOINFORMATICS/BTU033 24451623
86 Yang Z. PAML 4: Phylogenetic Analysis by Maximum Likelihood Mol. Biol. Evol. 24 2007 1586 1591 10.1093/MOLBEV/MSM088 17483113
87 Wang Y. Tang H. Debarry J.D. Tan X. Li J. Wang X. Lee T.H. Jin H. Marler B. Guo H. MCScanX: a toolkit for detection and evolutionary analysis of gene synteny and collinearity Nucleic Acids Res. 40 2012 e49 10.1093/NAR/GKR1293 22217600
88 Chen H. Zwaenepoel A. Van De Peer Y. wgd v2: a suite of tools to uncover and date ancient polyploidy and whole-genome duplication Bioinformatics 40 2024 btae272 10.1093/BIOINFORMATICS/BTAE272
89 Sun J. Lu F. Luo Y. Bie L. Xu L. Wang Y. OrthoVenn3: an integrated platform for exploring and visualizing orthologous data across genomes Nucleic Acids Res. 51 2023 W397 W403 10.1093/nar/gkad313 37114999
90 Finn R.D. Clements J. Eddy S.R. HMMER web server: interactive sequence similarity searching Nucleic Acids Res. 39 2011 W29 W37 10.1093/NAR/GKR367 21593126
91 Moriya Y. Itoh M. Okuda S. Yoshizawa A.C. Kanehisa M. KAAS: an automatic genome annotation and pathway reconstruction server Nucleic Acids Res. 35 2007 W182 W185 10.1093/NAR/GKM321 17526522
92 Cantalapiedra C.P. Hernández-Plaza A. Letunic I. Bork P. Huerta-Cepas J. eggNOG-mapper v2: Functional Annotation, Orthology Assignments, and Domain Prediction at the Metagenomic Scale Mol. Biol. Evol. 38 2021 5825 5829 10.1093/MOLBEV/MSAB293 34597405
93 Steinegger M. Söding J. MMseqs2 enables sensitive protein sequence searching for the analysis of massive data sets Nat. Biotechnol. 35 2017 1026 1028 10.1038/nbt.3988 29035372
94 Ruan J. Li H. Fast and accurate long-read assembly with wtdbg2 Nat. Methods 17 2019 155 158 10.1038/s41592-019-0669-3 31819265
95 Li W. Godzik A. Cd-hit: a fast program for clustering and comparing large sets of protein or nucleotide sequences Bioinformatics 22 2006 1658 1659 10.1093/BIOINFORMATICS/BTL158 16731699
96 Ness R.W. Wright S.I. Barrett S.C.H. Mating-System Variation, Demographic History and Patterns of Nucleotide Diversity in the Tristylous Plant Eichhornia paniculata Genetics 184 2010 381 392 10.1534/GENETICS.109.110130 19917767
97 Gaut B.S. Morton B.R. Mccaig B.C. Clegg M.T. Substitution rate comparisons between grasses and palms: synonymous rate differences at the nuclear gene Adh parallel rate differences at the plastid gene rbcL Proc. Natl. Acad. Sci. USA 93 1996 10274 10279 10.1073/PNAS.93.19.10274 8816790
98 Emms D.M. Kelly S. OrthoFinder: Phylogenetic orthology inference for comparative genomics Genome Biol. 20 2019 238 314 10.1186/S13059-019-1832-Y 31727128
99 Laetsch D.R. Blaxter M.L. KinFin: Software for taxon-aware analysis of clustered protein sequences G3 (Bethesda) 7 2017 3349 3357 10.1534/G3.117.300233 28866640
100 Katoh K. Standley D.M. MAFFT Multiple Sequence Alignment Software Version 7: Improvements in Performance and Usability Mol. Biol. Evol. 30 2013 772 780 10.1093/MOLBEV/MST010 23329690
101 Mendes F.K. Vanderpool D. Fulton B. Hahn M.W. CAFE 5 models variation in evolutionary rates among gene families Bioinformatics 36 2021 5516 5518 10.1093/BIOINFORMATICS/BTAA1022 33325502
102 Enright A.J. Van Dongen S. Ouzounis C.A. An efficient algorithm for large-scale detection of protein families Nucleic Acids Res. 30 2002 1575 1584 10.1093/NAR/30.7.1575 11917018
103 Jombart T. Balloux F. Dray S. adephylo: new tools for investigating the phylogenetic signal in biological traits Bioinformatics 26 2010 1907 1909 10.1093/BIOINFORMATICS/BTQ292 20525823
104 Ng P.C. Henikoff S. SIFT: predicting amino acid changes that affect protein function Nucleic Acids Res. 31 2003 3812 3814 10.1093/NAR/GKG509 12824425
105 Mahajan S. Bisht M.S. Chakraborty A. Sharma V.K. Genome of Phyllanthus emblica: the medicinal plant Amla with super antioxidant properties Front. Plant Sci. 14 2023 1210078 10.3389/FPLS.2023.1210078
106 Chakraborty A. Mahajan S. Bisht M.S. Sharma V.K. Genome sequencing and comparative analysis of Ficus benghalensis and Ficus religiosa species reveal evolutionary mechanisms of longevity iScience 25 2022 105100 10.1016/J.ISCI.2022.105100
107 Finn R.D. Bateman A. Clements J. Coggill P. Eberhardt R.Y. Eddy S.R. Heger A. Hetherington K. Holm L. Mistry J. Pfam: the protein families database Nucleic Acids Res. 42 2014 D222 D230 10.1093/NAR/GKT1223 24288371
108 Bairoch A. Apweiler R. The SWISS-PROT protein sequence database and its supplement TrEMBL in 2000 Nucleic Acids Res. 28 2000 45 48 10.1093/NAR/28.1.45 10592178
109 Ge S.X. Jung D. Yao R. ShinyGO: a graphical gene-set enrichment tool for animals and plants Bioinformatics 36 2020 2628 2629 10.1093/BIOINFORMATICS/BTZ931 31882993
