
==== Front
Sci Data
Sci Data
Scientific Data
2052-4463
Nature Publishing Group UK London

39294137
3861
10.1038/s41597-024-03861-5
Data Descriptor
A haplotype-resolved genome assembly of Coptis teeta, an endangered plant of significant medicinal value
Wang Ya 12
Liu Yan 12
Miao Ke 12
Hou Luxiao 12
Guo Xiaorong xrguo@ynu.edu.cn

3
http://orcid.org/0000-0002-1116-7485
Ji Yunheng jiyh@mail.kib.ac.cn

14
1 grid.9227.e 0000000119573309 Key Laboratory of Phytochemistry and Natural Medicines, Kunming Institute of Botany, Chinese Academy of Sciences, Kunming, 650201 China
2 https://ror.org/05qbk4x57 grid.410726.6 0000 0004 1797 8419 Kunming College of Life Science, University of Chinese Academy of Sciences, Kunming, 650201 China
3 https://ror.org/0040axw97 grid.440773.3 0000 0000 9342 2456 School of Ecology and Environmental Science, Yunnan Key Laboratory of Plant Reproductive Adaptation and Evolutionary Ecology and Institute of Biodiversity, Yunnan University, Kunming, 650201 China
4 grid.9227.e 0000000119573309 Yunnan Key Laboratory for Integrative Conservation of Plant Species with Extremely Small Population, Kunming Institute of Botany, Chinese Academy of Sciences, Kunming, 650201 China
18 9 2024
18 9 2024
2024
11 101221 6 2024
4 9 2024
© The Author(s) 2024
2024
https://creativecommons.org/licenses/by-nc-nd/4.0/ Open Access This article is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License, which permits any non-commercial use, sharing, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if you modified the licensed material. You do not have permission under this licence to share adapted material derived from this article or parts of it. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by-nc-nd/4.0/.
Coptis teeta Wall. (Ranunculaceae), an endangered plant species of significant medicinal value, predominantly undergoes clonal propagation, potentially compromising the species’ evolutionary potential and ultimately increase its risk of extinction. In this study, we successfully assembled two sets of haploid genomes (Hap1 and Hap2) for C. teeta, comprising nine homologous chromosome pairs, by employing Illumina and PacBio sequencing technologies. The genome annotation identified a total of 43,979 and 46,311 protein-coding genes in Hap1 and in Hap2, and most of them were functionally annotated. The high-quality reference genome will serve as an indispensable genomic resource for conservation and comprehensive exploitation of this endangered species. Between the two haploid genomes, numerous structural alterations were detected within the nine homologous chromosome pairs, potentially resulting in aberrant synapsis and irregular chromosomal segregation and thus contributing to the sustained preservation of clonal propagation in C. teeta. The findings offer new perspective for elucidating the genetic mechanism underlying the compromised sexual reproductive capacity of C. teeta, thereby facilitating its enhancement though molecular breeding and genetic improvement.

Subject terms

Plant evolution
Fertilization
https://doi.org/10.13039/501100001809 National Natural Science Foundation of China (National Science Foundation of China) 32370395 Ji Yunheng Yunnan Revitalization Talent Support Program “Top Team” Project (202305AT35), and the Key R & D program of Yunnan Province (202103AC100003).issue-copyright-statement© Springer Nature Limited 2024
==== Body
pmcBackground & Summary

Coptis teeta Wall. (Fig. 1), a rhizomatous herbaceous perennial belonging to the Ranunculaceae family, is sporadically distributed in the mountainous regions of Southwest China and the Eastern Himalaya, primarily occurring within the understory of broad-leaved forests, with a chromosome number 2n = 181–3. C. teeta is renowned for its medicinal significance in both China and the Himalayas, owing to the remarkable anti-inflammatory, antipyretic, dehumidifying, and detoxification properties of its subterranean rhizomes, which have traditionally been employed to effectively treat a wide range of ailments, such as infections, fever, acute diarrhea, vomiting, gastro-enteritis, eye diseases, skin diseases, and intoxications4,5. Modern researches have revealed that benzylisoquinoline alkaloids, such as berberine, coptisine, jatrorrhizine, palmatine, columbamine, and epiberberine derived from the rhizomes of C. teeta, are the primary bioactive components responsible for the therapeutic properties exhibited by this plant6, given that these compounds (particularly berberine) demonstrate significant pharmacological effects including antimicrobial, antiviral, anti-liver steatosis, anti-atherosclerosis, anti-myocardial ischemia/reperfusion injury, anti-diabetic, anti-arrhythmic, anti-hypertensive, anti-inflammatory, antioxidant, and anticancer bioactivities5,7–9. Notably, C. teeta exhibits a significantly higher abundance of berberine in its rhizomes compared with other congeneric species distributed in East Asia, thereby establishing its widespread recognition as the most efficacious medicinal Coptis species5. This advantage renders it highly promising for the drug development purposes.Fig. 1 Plant morphology of Coptis teeta. Whole plant (A), aerial shoots (B), and multibranched subterranean rhizome and the remarkable development of adventitious roots (C).

The natural propagation of C. teeta continues to present unresolved challenges despite its considerable medicinal significance, thereby hindering both the conservation efforts and comprehensive exploitation of this species. Specifically, previous studies have demonstrated that the majority of natural populations of C. teeta exhibit varying degrees of pollen sterility, resulting reduced or even absent seed production2,10 and an extremely low rate of seed germination11. The presence of these symptoms results in a diminished sexual reproductive capacity, constrained geographic distribution, and a limited population size in this species2. Considering this, the overexploitation of C. teeta for medicinal purposes has inevitably resulted in a rapid declination of its natural populations, thereby warranting its classification as an endangered species in the IUCN Red list of threatened species9,10. Furthermore, due to its reduced sexual reproductive capacity, C. teeta predominantly undergoes clonal propagation via stoloniferous rhizomes, characterized by the remarkable development of adventitious roots and apical/axillary buds that assimilate a significant quantity of nutrients1, consequently leading reduced rhizome biomass of C. teeta12. Theoretically, the absence of homologous recombination and segregation associated with sexual reproduction is expected to make prolonged asexual reproduction more prone to the generation and accumulation of deleterious mutations13–18, thereby potentially compromising the evolutionary potential and ultimately augmenting the risk of extinction for C. teeta. Accordingly, restoring its sexual reproduction and enhancing its sexual reproductive capacity are imperative for the conservation of this endangered species with significant medicinal value, as well as for facilitating its comprehensive exploitation.

A previous study has demonstrated that individuals of C. teeta, which exclusively undergo clonal propagation, exhibit aberrant synapsis of homologous chromosomes during meiotic prophase I and the irregular chromosomal segregation during meiotic anaphase I in their microspore mother cells19. The occurrence of such synaptic abnormality potentially results in the formation of sterile pollen grains, consequently contributing to their diminished sexual reproductive capacity19. Due to the lack of a reference genome, the elucidation of the genetic mechanisms underlying the observed synaptic mutation in C. teeta remains elusive, creating a crucial knowledge gap in terms of restoring and enhancing its sexual reproduction.

In this study, we assembled and annotated a haplotype-resolved reference genome of C. teeta, using PacBio high-fidelity (HiFi) sequencing, high-throughput chromosome conformation capture (Hi-C) sequencing, Illumina short-read sequencing, and RNA-seq data. The de novo genome assembly anchored approximately 97.59% of the assembled data to 18 pseudochromosomes and assigned them to two sets of haploid genomes, with a scaffold N50 length of more than 94 Mb. The C. teeta genome contains 46,311 protein-coding genes, with 92% of them being functionally annotated. The high-quality reference genome of C. teeta will serve as an unprecedented genetic resource for elucidating the underlying genetic mechanisms involved in aberrant synapsis of homologous chromosomes during the meiosis of microspore mother cells, thereby facilitating molecular breeding and genetic improvement for the conservation and sustainable utilization of this medicinally important endangered plant species.

Methods

Plant materials

The living plant individuals of C. teeta were collected from a natural population in Gaoligong Mountain, Fugong County, Yunnan Province, southwestern China, and subsequently cultivated in a greenhouse at the Kunming Institute of Botany, Chinese Academy of Sciences. Fresh young leaves were harvested from an individual plant to extract genomic DNA for Illumina short-read sequencing, HiFi sequencing, and Hi-C sequencing. Additionally, four types of tissues (leaf, stem, rhizome, and bud) were collected from the same individual plant and mixed for RNA extraction. The harvested materials were rapidly frozen in liquid nitrogen and then cryogenically preserved (−80 °C) until DNA and RNA extraction.

Genome sequencing

Genomic DNA was extracted from the leaf tissues using a modified CTAB method20. A short-read library with an average insert size of 400 bp was constructed using the TruSeq DNA PCR-free preparation kit (Illumina, MA, USA) following the manufacturer’s instructions. The constructed library was subjected to paired-end sequencing on the Illumina NovaSeq6000 platform (Illumina, MA, USA). The Illumina short-read sequencing generated approximately 66.424 Gb of raw data, consisting of 439,897,362 paired-end (2 × 150 bp) reads.

The PacBio HiFi library was prepared using the standard Pacbio Template Prep Kit 1.0 experimental protocol, specifically employing the Template Preparation Using Blue Pippin Size Selection method. Subsequently, the third-generation single-molecule sequencing of the HiFi library was conducted on the PacBio Sequel II sequencing platform, using the CCS (Circular Consensus Sequencing) mode21. In total, approximately 31.027 Gb of HiFi data were obtained.

For Hi-C library preparation, young and tender leaves were fixed in a 2% formaldehyde solution, and the DpnII restriction enzymes (NEB, UK) were then use to create sticky ends on both sides of the cross-linked DNA. Next, biotin-labeled bases were attached to the sticky ends through end-repair mechanism for subsequent DNA purification and capture. Subsequently, biotin-labeled DNA fragments (300−700 bp) with streptavidin beads were specifically enriched for constructing the Hi-C library, following the standard protocol of TruSeq DNA PCR-free prep kit (Illumina, MA, USA). The qualified Hi-C library was subjected to paired-end (2 × 150 bp) sequencing on the Illumina NovaSeq6000 platform (Illumina, MA, USA). In total, approximately 84.434 Gb of Hi-C data were generated.

Full-length transcriptome sequencing

The leaves, stems, rhizomes, and buds of a single plant individual were sampled to extract high-quality total RNA. The SMARTer PCR cDNA Synthesis Kit (Clontech) was employed for reverse transcription of the RNA into cDNA. Subsequently, the SMRTbell™ Template Prep Kit 1.0 (Clontech) was utilized to construct a full-length RNA-seq library, which was then sequenced on the PacBio Sequel II sequencing platform using the Sequel™ Sequencing Kit 2.0. A total of approximately 26.12 Gb of full-length RNA-seq data were generated for genome annotation.

Genome size and heterozygosity estimation

The genome traits (i.e., genome size, heterozygosity, and the proportion of repetitive sequences) of C. teeta were estimated using clean Illumina short-read sequencing data via a K-mer-based approach. Illumina raw reads were filtered using FastQC (v0.11.9,) and FASTP (v0.20.0)22. The frequency distribution of 19-mer indicators was calculated using the Jellyfish algorithm (2.3.0)23. Subsequently, the results were then imported to GenomeScope (v1.0)24 to infer the basic features of the genome. The haploid genome size C. teeta was estimated to be approximately 958.95 Mb, with a heterozygosity rate of approximately 0.22% (Table S1).

Haplotype-resolved genome assembly

PacBio HiFi reads were initially assembled into haplotype-resolved scaffolds using the software Hifiasm (0.16.1-r375)25 with default parameters. Next, the 3D-DNA pipeline (180922)26 was employed to anchor Hi-C reads onto each scaffold, thereby integrating these scaffolds into independent pseudochromosomes. The preliminary haplotype-resolved genome assembly was improved using purge-dups (1.2.5)27 to eliminate redundant sequences. Subsequently, the haplotype-resolved genome assembly was further improved using RagTag (v2.1.0)28 to rectify errors in the genome assembly and fill gaps in each pseudochromosome.

The de novo genome assembly anchored approximately 97.59% of the assembled data to 18 pseudochromosomes, which were categorized into two sets of haploid genomes referred to as haplotype 1 (Hap1) and haplotype 2 (Hap 2). Each individual haploid genome comprises a total of nine pseudochromosomes (Fig. 2). The result is in accordance with previous cytological investigations of C. teeta, which demonstrated that it possesses a diploid chromosome number of 2n = 181. The genome size of the two haplotypes were approximately 931.7 Mb and 898.9 Mb (Table 1), respectively, which is close to the estimated haploid genome size of this species (approximately 958.95 Mb) through K-mer analysis (Table S1). The scaffold N50 length of each haploid genome exceeded 94 Mb (Table 1).Fig. 2 The landscape of genome assembly and annotation of the Coptis teeta. Outer to inner tracks: (a) Length of chromosome, with each tick indicating a span of 5 Mb; (b) GC content distribution; (c) Coding gene density; (d) Repetitive sequence density; (e) Long terminal repeat (LTR) element density; (f) Density of LTR/Gypsy elements; (g) Density of LTR/Copia elements; (h) DNA transposon density; (i) The interconnections within the circle represent syntenic gene pairs.

Table 1 Summary of the genome assembly of Coptis teeta.

Statistic	Hap1	Hap2	
Min sequence length (bp)	1,002	1,005	
Max sequence length (bp)	110,731,202	113,735,707	
Total sequence number	1,458	1,474	
N20 (bp)	102,721,411	98,722,088	
N50 (bp)	98,128,514	94,220,073	
N90 (bp)	89,379,719	86,609,443	
Total sequence length (bp)	931,700,640	898,887,047	
GC content (%)	37.78	37.9	
Sequences greater than 1 kb	1,458	1,474	

Genome annotation

The identification and annotation of repetitive sequences in C. teeta genome were performed through a combination of ab initio and homology-based approaches. The programs RepeatModler (2.0.4)29 and RepeatScout (1.0.5)30 were employed to predict repetitive sequences to construct a reference library of C. teeta genome regarding repetitive sequences. The homology-based annotation was performed using RepeatMasker (4.1.4)31 to identify repetitive regions in the two haploid genomes and by comparing them with the reference library and the RepeatScou (1.0.5)30 database. In total, 996,075 and 989,247 repetitive sequences were identified in Hap1 and Hap2, respectively; the combined length of the repetitive sequences identified in Hap1 and Hap2 was 615,272,110 bp and 582,452,510, respectively, which account for 66.04% and 64.80% of corresponding haploid genome size (Table S2).

The majority of these repetitive sequences identified in the C. teeta genome are transposable elements (TEs), including long terminal repeat (LTR) retroelements, long interspersed nuclear element (LINE) retroelements, short interspersed nuclear element (SINE) Retroelements, and DNA transposons. Among them, the LTR retroelements exhibited the highest proportion, comprising a total length of 334,373,833 bp in Hap1 and 309,665,818 bp in Hap2, accounting for 35.89% and 34.45% of corresponding haploid genome size (Table S2).

TRNAscan-SE (1.3.1)32 was used to predict the presence of tRNA genes in the entire genome, and RNAmmer (1.2)33 was employed for rRNA gene prediction. The prediction of other non-coding RNAs was primarily obtained through comparison with the Rfam database (1.0)34. The total number of non-coding RNA identified in Hap 1 and Hap2 was 6,802 (with the combined sequence length 1,629,812 bp) and 7,600 copies (with the combined sequence length 4,903,501 bp), respectively; the sequence length account for approximately 0.17% and 0.55% of the corresponding haploid genomes size (Table S3).

The prediction of protein-coding genes within the assembled C. teeta genome was conducted using three distinct methods, namely de novo gene predictions, homology-based predictions, and transcriptome-based predictions. The de novo predictions of protein-coding genes was performed using four software tools, AUGUSTUS (3.3.2)35, SNAP (6.0, https://step.esa.int/main/snap-6-0-released/), GlimmerHMM (3.0.4)36, and GeneMark (4.35)37. The amino acid sequences derived from the publicly available genomes of Akebia trifoliata, Aquilegia coerulea, Arabidopsis thaliana, Copteis chinensis, Epimedium pubescens, and Vitis vinifera genomes were utilized as a reference database for homology-based prediction using AUGUSTUS (3.3.2)35. For transcriptome-based prediction, the software Trinity (r20140717)38 was employed to de novo assemble full-length RNA-Seq data. The software PASA (2.5.2)39 was employed to predict the gene structure of the assembled transcripts. Finally, EvidenceModeler (r2012-06-25)40 was utilized to conduct gene annotation based on the results of de novo gene predictions, homology-based predictions, and transcriptome-based predictions. Within the assembled C. teeta genome, a total of 43,979 and 46,311 protein-coding genes were identified in Hap1 and in Hap2, respectively (Table 2).Table 2 Summary of the structure annotation of protein-coding genes in the genome of Coptis teeta.

Statistic	Hap1	Hap2	
Total genes length (bp)	173,519,366	172,116,066	
Genes percentage of genome (%)	18.62	19.15	
Total genes number	43,979	46,311	
Average gene length (bp)	3,945.5	3,716.5	
Total coding sequence (CDS) number	205,970	203,156	
Average coding sequence (CDS) length (bp)	225.6	230.9	
Average introns length (bp)	784.1	798.1	
Total coding sequences (CDSs) length (bp)	46,487,106	46,928,784	
Coding sequences (CDSs) percentage of genome (%)	4.99	5.22	
Average transcription length (bp)	1,057	1,013.3	

The functional annotation of these predicted protein-coding genes within the C. teeta genome was accomplished through alignment with publicly accessible functional databases. Specifically, DIAMOND (2.0.14.152)41 was utilized to perform sequence similarity search (with an E-value threshold 1e-5) against the NCBI NR (Non-Redundant Protein Sequence) and SWISS-Prot databases. The GO (Gene Ontology) and Pfam (Protein Family Database) annotations were perform using InterProScan (5.61-93.0)42 to conduct domain similarity searches against respective database. Additionally, the KEGG Automatic Annotation Server (KAAS)43 was utilized to conduct KEGG (Kyoto Encyclopedia of Genes and Genomes) annotation for assigning these predicted protein-coding genes to their corresponding KEGG pathways. Finally, a total 41,831 and 42,605 genes were functionally annotated in at least one of the aforementioned databases (Table S4), accounting for 95.12% and 92.00%, respectively, of the predicted protein-coding genes within Hap 1 and Hap 2 of the C. teeta genome.

Genomic structural variation analysis

The chromosome-level sequences of C. teeta genome were aligned using Minimap2 (2.17-r941)44, and SyRI (1.6)38 was employed to identify collinearity and structural alterations between the homologous chromosomes of the diploid genome. Subsequently, PLSTSR (0.5.4)45 was employed to visualize the results. Although the two haploid genomes displayed a high level of collinearity, numerous structural alterations were detected within the nine homologous chromosome pairs (Fig. 3), including 106 inversion events, 1,488 translocation events, and 5,151 duplication loss/gain events. Consequently, the C. teeta genome exhibited varying degrees of differences in the lengths of homologous chromosomes (Table 3) and a substantial number of (11,285) non-aligned regions with remarkable sequence length (approximately 221.234 Mb in total). The remarkable heterogeneity of homologous chromosomes between the two haploid genomes of C. teeta provides empirical evidence supporting the theoretical expectation that the genomes of asexual taxa are more prone to generate and accumulate structural variations and sequence mutations, consequently leading to the divergence of intra-individual haploid genomes over time13–18.Fig. 3 Structural variation between the two haploid genomes of Coptis teeta.

Table 3 Comparison of the homologous chromosome lengths between the two haploid genomes of Coptis teeta.

Chromosome	Length in Hap1 (bp)	Length in Hap2 (bp)	
Chr1	110,731,202	113,735,707	
Chr2	99,748,014	98,722,088	
Chr3	102,721,411	98,045,796	
Chr4	96,796,468	97,641,301	
Chr5	98,887,026	94,220,073	
Chr6	93,232,271	90,922,627	
Chr7	98,128,514	89,664,123	
Chr8	89,379,719	89,593,540	
Chr9	92,659,245	86,609,443	

Notably, previous studies have demonstrated that the presence of complex structural variations between homologous chromosomes may impede homologous chromosome synapsis during meiosis, thereby diminishing the sexual fertility of Pogostemon cablin (Lamiaceae)46 and Actinidia zhejiangensis (Actinidiaceae)47. Therefore, the extensive structural alterations within the nine homoeologous chromosome pairs between the two haploid genomes of C. teeta can also potentially impede regular homoeologous chromosome synapsis and lead aberrant chromosomal segregation during meiosis, thus contributing to the previously observed meiotic abnormalities in microspore mother cells of this species19. Consequently, we hypothesized that the prolonged clonal propagation could lead to the divergence of intra-individual haploid genomes in C. teeta, subsequently reducing its sexual reproductive capacity. The validation of this hypothesis necessitates further empirical evidence to be generated in future studies.

Data Records

All sequencing data (including Illumina sequencing, PacBio HiFi sequencing, Hi-C sequencing, and full-length RNA-seq data) has been deposited in Genome Sequence Archive (GSA) of the National Genomics Data Center (NGDC) database under the accession number CRA01563648. The genome assembly and annotation files are available in the FigShare database49. The genome assembly files are also available in National Center for Biotechnology Information (NCBI) GenBank database under the accession number JBEGIT00000000050 and JBEGIU00000000051. The outputs of genomic structural variation analysis are available in the FigShare database52.

Technical Validation

The Hi-C data was aligned to the genome assembled at the chromosome level using the 3D-DNA pipeline (180922)26. The strength of the interaction signal between any two blocks (500 Kb in length) was quantified based on the number of Hi-C read pairs, and a heatmap was generated to visually represent this information. The heatmap clearly distinguished the nine chromosomes, and within each chromosome, the strongest interaction signal was observed along the diagonal, indicating a high level of interaction strength between adjacent sequences (Fig. 4), thereby demonstrating a high degree of accuracy in genome assembly. Additionally, the quality of genome assembly was further assessed using the Benchmarking Universal Single-Copy Orthology (BUSCO) software (5.4.5)53, based on the embryophyta_odb10 lineage dataset. The proportion of complete core genes detected in Hap1 and Hap2 was 96.6% (including 90.0% single-copy and 6.6% multi-copy genes) and 95.4% (including 90.8% single-copy and 4.6% multi-copy genes) (Table S5), respectively, indicating a high level of genome assembly completeness.Fig. 4 Hi-C interaction heatmap of the reference genome of Coptis teeta.

The BUSCO analysis was also conducted to assess the annotation of protein-coding genes. The coverage percentages of the 1,614 BUSCO groups were 96.3% and 94.8% in Hap1 and Hap2 (Table S5), respectively. The results indicate the high quality of gene annotation.

Supplementary information

Supplementary tables

Supplementary information

The online version contains supplementary material available at 10.1038/s41597-024-03861-5.

Acknowledgements

The study is financially supported by the National Natural Science Foundation of China (32370395), the Yunnan Revitalization Talent Support Program “Top Team” Project (202305AT35), and the Key R & D program of Yunnan Province (202103AC100003).

Author contributions

Y.J. and X.G. conceived the framework of this study. Y.W., Y.L., K.M. and L.H. collected and analyzed the data. Y.W., Y.L. and Y.J. wrote the draft manuscript. Y.J. and X.G. revised the manuscript. All authors contributed to the article and approved the submitted version.

Code availability

The manuals and protocols of the published bioinformatics tools were strictly adhered to for the execution of all software and pipelines. The Methods section provides a comprehensive description of the software version, code, and parameters employed. No custom code was utilized in this study for dataset curation and validation.

Competing interests

The authors declare no competing interests.

Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

These authors contributed equally: Ya Wang, Yan Liu.
==== Refs
References

1. Pandit M Cytology and taxonomy of Coptis teeta Wall. (Ranunculaceae) Bot J Linn Soc. 1993 111 371 378 10.1006/bojl.1993.1026
Pandit, M. Cytology and taxonomy of Coptis teeta Wall. (Ranunculaceae). Bot J Linn Soc. 111, 371–378 (1993).
2. Pandit MK Babu CR Biology and conservation of Coptis teeta Wall. – an endemic and endangered medicinal herb of Eastern Himalaya Envir Conserv. 1998 25 262 272 10.1017/S0376892998000320
Pandit, M. K. & Babu, C. R. Biology and conservation of Coptis teeta Wall. – an endemic and endangered medicinal herb of Eastern Himalaya. Envir Conserv. 25, 262–272 (1998).
3. Zhang J Variation patterns of Coptis teeta biomass and its major active compounds along an altitude gradient J Appl Ecol. 2008 19 1455 1461
Zhang, J. et al. Variation patterns of Coptis teeta biomass and its major active compounds along an altitude gradient. J Appl Ecol. 19, 1455–1461 (2008).
4. Moon M Coptidis rhizoma prevents heat stress-induced brain damage and cognitive impairment in mice Nutrients. 2017 9 1057 10.3390/nu9101057 28946610
Moon, M. et al. Coptidis rhizoma prevents heat stress-induced brain damage and cognitive impairment in mice. Nutrients. 9, 1057 (2017).28946610
5. Wang J Coptidis Rhizoma: a comprehensive review of its traditional uses, botany, phytochemistry, pharmacology and toxicology Pharm Biol. 2019 57 193 225 10.1080/13880209.2019.1577466 30963783
Wang, J. et al. Coptidis Rhizoma: a comprehensive review of its traditional uses, botany, phytochemistry, pharmacology and toxicology. Pharm Biol. 57, 193–225 (2019).30963783
6. Meng FC Coptidis rhizoma and its main bioactive components: recent advances in chemical investigation, quality evaluation and pharmacological activity Chin Med. 2018 13 13 10.1186/s13020-018-0171-3 29541156
Meng, F. C. et al. Coptidis rhizoma and its main bioactive components: recent advances in chemical investigation, quality evaluation and pharmacological activity. Chin Med. 13, 13 (2018).29541156
7. Yu HH Antimicrobial activity of berberine alone and in combination with ampicillin or oxacillin against methicillin-resistant Staphylococcus aureus J Med Food. 2005 8 454 461 10.1089/jmf.2005.8.454 16379555
Yu, H. H. et al. Antimicrobial activity of berberine alone and in combination with ampicillin or oxacillin against methicillin-resistant Staphylococcus aureus. J Med Food. 8, 454–461 (2005).16379555
8. Tan, H. L. et al. Rhizoma Coptidis: a potential cardiovascular protective agent. Front. Pharmacol. 7 (2016).
9. Plants for Immunity and Conservation Strategies. (Springer, Singapore, 2023).
10. Pandit MK Babu CR The effects of loss of sex in clonal populations of an endangered perennial Coptis teeta (Ranunculaceae) Bot J Linn Soc. 2003 143 47 54 10.1046/j.1095-8339.2003.00192.x
Pandit, M. K. & Babu, C. R. The effects of loss of sex in clonal populations of an endangered perennial Coptis teeta (Ranunculaceae). Bot J Linn Soc. 143, 47–54 (2003).
11. Yang YJ Xie SQ Meng ZG Chen GM Liang YL Pollination ecology of Coptis teeta Wall. an endangered medicinal plant Acta Bot. Boreal.-Occident. Sin. 2012 32 1372 1376
Yang, Y. J., Xie, S. Q., Meng, Z. G., Chen, G. M. & Liang, Y. L. Pollination ecology of Coptis teeta Wall. an endangered medicinal plant. Acta Bot. Boreal.-Occident. Sin. 32, 1372–1376 (2012).
12. Liu TH Zhang XM Tian SZ Chen LG Yuan JL Bioinformatics analysis of endophytic bacteria related to berberine in the Chinese medicinal plant Coptis teeta Wall 3 Biotech. 2020 10 96 10.1007/s13205-020-2084-y 32099737
Liu, T. H., Zhang, X. M., Tian, S. Z., Chen, L. G. & Yuan, J. L. Bioinformatics analysis of endophytic bacteria related to berberine in the Chinese medicinal plant Coptis teeta Wall. 3 Biotech. 10, 96 (2020).32099737
13. Butlin R The costs and benefits of sex: new insights from old asexual lineages Nat Rev Genet. 2002 3 311 317 10.1038/nrg749 11967555
Butlin, R. The costs and benefits of sex: new insights from old asexual lineages. Nat Rev Genet. 3, 311–317 (2002).11967555
14. Bast J Consequences of asexuality in natural populations: insights from stick insects Mol Biol and Evol. 2018 35 1668 1677 10.1093/molbev/msy058 29659991
Bast, J. et al. Consequences of asexuality in natural populations: insights from stick insects. Mol Biol and Evol. 35, 1668–1677 (2018).29659991
15. Brandt A Haplotype divergence supports long-term asexuality in the oribatid mite Oppiella nova Proc. Natl. Acad. Sci. USA 2021 118 e2101485118 10.1073/pnas.2101485118 34535550
Brandt, A. et al. Haplotype divergence supports long-term asexuality in the oribatid mite Oppiella nova. Proc. Natl. Acad. Sci. USA 118, e2101485118 (2021).34535550
16. Neiman M Lively CM Meirmans S Why sex? A pluralist approach revisited Trends Ecol Evol. 2017 32 589 600 10.1016/j.tree.2017.05.004 28606425
Neiman, M., Lively, C. M. & Meirmans, S. Why sex? A pluralist approach revisited. Trends Ecol Evol. 32, 589–600 (2017).28606425
17. Otto SP Selective interference and the evolution of sex J Hered. 2021 112 9 18 10.1093/jhered/esaa026 33047117
Otto, S. P. Selective interference and the evolution of sex. J Hered. 112, 9–18 (2021).33047117
18. Sharbrough J Luse M Boore JL Logsdon JM Neiman M Radical amino acid mutations persist longer in the absence of sex Evolution 2018 72 808 824 10.1111/evo.13465 29520921
Sharbrough, J., Luse, M., Boore, J. L., Logsdon, J. M. & Neiman, M. Radical amino acid mutations persist longer in the absence of sex. Evolution 72, 808–824 (2018).29520921
19. Pandit MK Babu CR Synaptic mutation associated with gametic sterility and population divergence in Coptis teeta (Ranunculaceae) Bot J Linn Soc. 2000 133 525 533 10.1111/j.1095-8339.2000.tb01594.x
Pandit, M. K. & Babu, C. R. Synaptic mutation associated with gametic sterility and population divergence in Coptis teeta (Ranunculaceae). Bot J Linn Soc. 133, 525–533 (2000).
20. Allen GC Flores-Vergara MA Krasynanski S Kumar S Thompson WF A modified protocol for rapid DNA isolation from plant tissues using cetyltrimethylammonium bromide Nat Protoc. 2006 1 2320 2325 10.1038/nprot.2006.384 17406474
Allen, G. C., Flores-Vergara, M. A., Krasynanski, S., Kumar, S. & Thompson, W. F. A modified protocol for rapid DNA isolation from plant tissues using cetyltrimethylammonium bromide. Nat Protoc. 1, 2320–2325 (2006).17406474
21. Wenger AM Accurate circular consensus long-read sequencing improves variant detection and assembly of a human genome Nat Biotechnol. 2019 37 1155 1162 10.1038/s41587-019-0217-9 31406327
Wenger, A. M. et al. Accurate circular consensus long-read sequencing improves variant detection and assembly of a human genome. Nat Biotechnol. 37, 1155–1162 (2019).31406327
22. Chen S Zhou Y Chen Y Gu J fastp: an ultra-fast all-in-one FASTQ preprocessor Bioinformatics. 2018 34 i884 i890 10.1093/bioinformatics/bty560 30423086
Chen, S., Zhou, Y., Chen, Y. & Gu, J. fastp: an ultra-fast all-in-one FASTQ preprocessor. Bioinformatics. 34, i884–i890 (2018).30423086
23. Marçais G Kingsford C A fast, lock-free approach for efficient parallel counting of occurrences of k -mers Bioinformatics. 2011 27 764 770 10.1093/bioinformatics/btr011 21217122
Marçais, G. & Kingsford, C. A fast, lock-free approach for efficient parallel counting of occurrences of k -mers. Bioinformatics. 27, 764–770 (2011).21217122
24. Vurture GW GenomeScope: fast reference-free genome profiling from short reads Bioinformatics. 2017 33 2202 2204 10.1093/bioinformatics/btx153 28369201
Vurture, G. W. et al. GenomeScope: fast reference-free genome profiling from short reads. Bioinformatics. 33, 2202–2204 (2017).28369201
25. Cheng HY Concepcion GT Feng XW Zhang HW Li H Haplotype-resolved de novo assembly using phased assembly graphs with hifiasm Nat Methods. 2021 18 170 175 10.1038/s41592-020-01056-5 33526886
Cheng, H. Y., Concepcion, G. T., Feng, X. W., Zhang, H. W. & Li, H. Haplotype-resolved de novo assembly using phased assembly graphs with hifiasm. Nat Methods. 18, 170–175 (2021).33526886
26. Dudchenko O De novo assembly of the Aedes aegypti genome using Hi-C yields chromosome-length scaffolds Science. 2017 356 92 95 10.1126/science.aal3327 28336562
Dudchenko, O. et al. De novo assembly of the Aedes aegypti genome using Hi-C yields chromosome-length scaffolds. Science. 356, 92–95 (2017).28336562
27. Guan DF Identifying and removing haplotypic duplication in primary genome assemblies Bioinformatics. 2020 36 2896 2898 10.1093/bioinformatics/btaa025 31971576
Guan, D. F. et al. Identifying and removing haplotypic duplication in primary genome assemblies. Bioinformatics. 36, 2896–2898 (2020).31971576
28. Alonge M Automated assembly scaffolding using RagTag elevates a new tomato system for high-throughput genome editing Genome Biol. 2022 23 258 10.1186/s13059-022-02823-7 36522651
Alonge, M. et al. Automated assembly scaffolding using RagTag elevates a new tomato system for high-throughput genome editing. Genome Biol. 23, 258 (2022).36522651
29. Flynn JM RepeatModeler2 for automated genomic discovery of transposable element families Proc. Natl. Acad. Sci. USA 2020 117 9451 9457 10.1073/pnas.1921046117 32300014
Flynn, J. M. et al. RepeatModeler2 for automated genomic discovery of transposable element families. Proc. Natl. Acad. Sci. USA 117, 9451–9457 (2020).32300014
30. Bao WD Kojima KK Kohany O Repbase Update, a database of repetitive elements in eukaryotic genomes Mobile DNA. 2015 6 11 10.1186/s13100-015-0041-9 26045719
Bao, W. D., Kojima, K. K. & Kohany, O. Repbase Update, a database of repetitive elements in eukaryotic genomes. Mobile DNA. 6, 11 (2015).26045719
31. Tempel, S. Using and understanding RepeatMasker. in Mobile Genetic Elements (ed. Bigot, Y.) vol. 859 29–51 (Humana Press, Totowa, NJ, 2012).
32. Lowe TM Eddy SR tRNAscan-SE: a program for improved detection of transfer RNA genes in genomic sequence Nucl Acids Res. 1997 25 955 964 10.1093/nar/25.5.955 9023104
Lowe, T. M. & Eddy, S. R. tRNAscan-SE: a program for improved detection of transfer RNA genes in genomic sequence. Nucl Acids Res. 25, 955–964 (1997).9023104
33. Lagesen K RNAmmer: consistent and rapid annotation of ribosomal RNA genes Nucl Acids Res. 2007 35 3100 3108 10.1093/nar/gkm160 17452365
Lagesen, K. et al. RNAmmer: consistent and rapid annotation of ribosomal RNA genes. Nucl Acids Res. 35, 3100–3108 (2007).17452365
34. Griffiths‐Jones, S. Annotating Non‐Coding RNAs with Rfam. CP in Bioinformatics. 9 (2005).
35. Hoff KJ Stanke M Predicting genes in single genomes with AUGUSTUS CP in Bioinformatics. 2019 65 e57 10.1002/cpbi.57
Hoff, K. J. & Stanke, M. Predicting genes in single genomes with AUGUSTUS. CP in Bioinformatics. 65, e57 (2019).
36. Majoros WH Pertea M Salzberg SL TigrScan and GlimmerHMM: two open source ab initio eukaryotic gene-finders Bioinformatics. 2004 20 2878 2879 10.1093/bioinformatics/bth315 15145805
Majoros, W. H., Pertea, M. & Salzberg, S. L. TigrScan and GlimmerHMM: two open source ab initio eukaryotic gene-finders. Bioinformatics. 20, 2878–2879 (2004).15145805
37. Besemer J Borodovsky M GeneMark: web software for gene finding in prokaryotes, eukaryotes and viruses Nucl Acids Res. 2005 33 W451 W454 10.1093/nar/gki487 15980510
Besemer, J. & Borodovsky, M. GeneMark: web software for gene finding in prokaryotes, eukaryotes and viruses. Nucl Acids Res. 33, W451–W454 (2005).15980510
38. Haas BJ De novo transcript sequence reconstruction from RNA-seq using the Trinity platform for reference generation and analysis Nat Protoc. 2013 8 1494 1512 10.1038/nprot.2013.084 23845962
Haas, B. J. et al. De novo transcript sequence reconstruction from RNA-seq using the Trinity platform for reference generation and analysis. Nat Protoc. 8, 1494–1512 (2013).23845962
39. Haas BJ Improving the Arabidopsis genome annotation using maximal transcript alignment assemblies Nucl Acids Res. 2003 31 5654 5666 10.1093/nar/gkg770 14500829
Haas, B. J. Improving the Arabidopsis genome annotation using maximal transcript alignment assemblies. Nucl Acids Res. 31, 5654–5666 (2003).14500829
40. Haas BJ Automated eukaryotic gene structure annotation using EVidenceModeler and the Program to Assemble Spliced Alignments Genome Biol. 2008 9 R7 10.1186/gb-2008-9-1-r7 18190707
Haas, B. J. et al. Automated eukaryotic gene structure annotation using EVidenceModeler and the Program to Assemble Spliced Alignments. Genome Biol. 9, R7 (2008).18190707
41. Li CX Gao MX Yang WX Zhong CQ Yu RS Diamond: a multi-modal DIA mass spectrometry data processing pipeline Bioinformatics 2021 37 265 267 10.1093/bioinformatics/btaa1093 33416868
Li, C. X., Gao, M. X., Yang, W. X., Zhong, C. Q. & Yu, R. S. Diamond: a multi-modal DIA mass spectrometry data processing pipeline. Bioinformatics 37, 265–267 (2021).33416868
42. Jones P InterProScan 5: genome-scale protein function classification Bioinformatics 2014 30 1236 1240 10.1093/bioinformatics/btu031 24451626
Jones, P. et al. InterProScan 5: genome-scale protein function classification. Bioinformatics 30, 1236–1240 (2014).24451626
43. Moriya Y Itoh M Okuda S Yoshizawa AC Kanehisa M KAAS: an automatic genome annotation and pathway reconstruction server Nucl Acids Res. 2007 35 W182 W185 10.1093/nar/gkm321 17526522
Moriya, Y., Itoh, M., Okuda, S., Yoshizawa, A. C. & Kanehisa, M. KAAS: an automatic genome annotation and pathway reconstruction server. Nucl Acids Res. 35, W182–W185 (2007).17526522
44. Li H Minimap2: pairwise alignment for nucleotide sequences Bioinformatics. 2018 34 3094 3100 10.1093/bioinformatics/bty191 29750242
Li, H. Minimap2: pairwise alignment for nucleotide sequences. Bioinformatics. 34, 3094–3100 (2018).29750242
45. Goel M Schneeberger K plotsr: visualizing structural similarities and rearrangements between multiple genomes Bioinformatics. 2022 38 5328 5328 10.1093/bioinformatics/btac196 36228123
Goel, M. & Schneeberger, K. plotsr: visualizing structural similarities and rearrangements between multiple genomes. Bioinformatics. 38, 5328–5328 (2022).36228123
46. Shen Y Chromosome-level and haplotype-resolved genome provides insight into the tetraploid hybrid origin of patchouli Nat Commun. 2022 13 3511 10.1038/s41467-022-31121-w 35717499
Shen, Y. et al. Chromosome-level and haplotype-resolved genome provides insight into the tetraploid hybrid origin of patchouli. Nat Commun. 13, 3511, 10.1038/s41467-022-31121-w (2022).35717499
47. Yu X Genomic analyses reveal dead-end hybridization between two deeply divergent kiwifruit species rather than homoploid hybrid speciation Plant J. 2023 115 1528 1543 10.1111/tpj.16336 37258460
Yu, X. et al. Genomic analyses reveal dead-end hybridization between two deeply divergent kiwifruit species rather than homoploid hybrid speciation. Plant J. 115, 1528–1543 (2023).37258460
48. 2024 NGDC Genome Sequence Archive https://ngdc.cncb.ac.cn/gsa/browse/CRA015636
NGDC Genome Sequence Archive https://ngdc.cncb.ac.cn/gsa/browse/CRA015636 (2024).
49. Wang Y 2024 Genome assembly and annotation for C. teeta figshare 10.6084/m9.figshare.25956079.v1
Wang, Y. Genome assembly and annotation for C. teeta. figshare10.6084/m9.figshare.25956079.v1 (2024).
50. 2024 NCBI GenBank GCA_040256895.1
NCBI GenBank https://identifiers.org/ncbi/insdc.gca:GCA_040256895.1 (2024).
51. 2024 NCBI GenBank GCA_040256905.1
NCBI GenBank https://identifiers.org/ncbi/insdc.gca:GCA_040256905.1 (2024).
52. Wang Y 2024 SVs in the two haplotypes of C. teeta figshare 10.6084/m9.figshare.25495405.v3
Wang, Y. SVs in the two haplotypes of C. teeta. figshare10.6084/m9.figshare.25495405.v3 (2024).
53. Manni M Berkeley MR Seppey M Simão FA Zdobnov EM BUSCO update: novel and streamlined workflows along with broader and deeper phylogenetic coverage for scoring of eukaryotic, prokaryotic, and viral genomes Mol Biol Evol. 2021 38 4647 4654 10.1093/molbev/msab199 34320186
Manni, M., Berkeley, M. R., Seppey, M., Simão, F. A. & Zdobnov, E. M. BUSCO update: novel and streamlined workflows along with broader and deeper phylogenetic coverage for scoring of eukaryotic, prokaryotic, and viral genomes. Mol Biol Evol. 38, 4647–4654 (2021).34320186
