
==== Front
Sci Data
Sci Data
Scientific Data
2052-4463
Nature Publishing Group UK London

3798
10.1038/s41597-024-03798-9
Data Descriptor
The first Chromosomal-level genome assembly of Sageretia thea using Nanopore long reads and Pore-C technology
Jo Jihoon 1
Park Jong-Soo 2
Won Hari 1
Jeong Jun Seong 1
Jung Tae Won 3
Lee Kyung Jun 1
Lee Shin Ae shinaelee@hnibr.re.kr

1
1 Division of Genetic Resources, Honam National Institute of Biological Resources (HNIBR), Mokpo, Republic of Korea
2 Division of Botany, Honam National Institute of Biological Resources (HNIBR), Mokpo, Republic of Korea
3 Division of Exhibition, Honam National Institute of Biological Resources (HNIBR), Mokpo, Republic of Korea
4 9 2024
4 9 2024
2024
11 95913 6 2024
19 8 2024
© The Author(s) 2024
2024
https://creativecommons.org/licenses/by-nc-nd/4.0/ Open Access This article is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License, which permits any non-commercial use, sharing, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if you modified the licensed material. You do not have permission under this licence to share adapted material derived from this article or parts of it. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by-nc-nd/4.0/.
Sageretia thea, a notable species within the mock buckthorn genus, is recognized for its intriguing biogeographical distribution and diverse medicinal properties. Despite this significance, genomic studies on S. thea are still in the nascent stages. We present the first chromosome-level genome assembly of S. thea that was generated using a combination of Oxford Nanopore long-read and Illumina short-read sequencing technologies complemented by Pore-C chromatin conformation capture. The genome assembly had a size of 197.8 Mb with 12 chromosomal scaffolds and a scaffold N50 length of 15.9 Mb. A total of 25,434 protein-coding genes were identified and functionally annotated, and the gene model indicated 96.5% complete eukaryotic BUSCOs. Additionally, orthologous gene profiling and synteny analysis were performed to elucidate the evolutionary relationships within the Rhamnaceae family and Rosales. This high-quality chromosomal genome is the first genomic view of S. thea, which will serve as the basis for future studies on its biological and medicinal properties, and evolutionary history.

Subject terms

Comparative genomics
Plant evolution
Honam National Institute of Biological Resources, HNIBR202302118 Honam National Institute of Biological Resources, HNIBR202401104issue-copyright-statement© Springer Nature Limited 2024
==== Body
pmcBackground & Summary

Sageretia thea, a representative species of the mock buckthorn genus, is an semi-evergreen shrub belonging to the family Rhamnaceae of Rosales. The species is widely distributed in tropical and temperate habitats, including, the Arabian Peninsula and East Asia. It is characterized by spicate (or spicate-panicle) inflorescences and persistent sepals at the base of its fruits1. Mock buckthorn trees, including S. thea are intriguing on account of their status as a biogeographical study model to demonstrate Old–New World disjunction1. Further, S. thea is a renowned nutrient-rich and medicinal herb in South Korea and China. Similar to the jujube Ziziphus jujuba, is a representative of the Rhamnaceae species, and has high flavonoid and phenolic contents2. While its fruit extract exhibits anti-melanogenic activity3, the leaf and branch extracts possess anticancer activity4.

Despite these valuable characteristics, the plant’s genetic background is fairly unknown due to the lack of comprehensive genomic studies. We generated long-read Oxford Nanopore Technologies (ONT) and Illumina short-read sequences to assemble the genome of S. thea de novo, and successfully constructed the first chromosomal-level genome of this plant using Pore-C chromatin conformation capture technology5. This genome assembly size was 197.8 Mb with 12 chromosomal scaffolds, indicating a scaffold N50 length of 15.9 Mb. We identified 25,434 protein-coding genes that encompassed 96.5% of complete BUSCOs (Benchmarking Universal Single-Copy Orthologs), of which 94.6% were annotated. Additionally, syntenic alignments of these genes showed highly conserved gene content with Z. jujuba with several rearrangements of sub-telomeric genes. Gain/loss profiles of orthologous gene families were explored by comparing 11 other Viridiplantae species. This high-quality genome for S. thea will facilitate understanding their evolution and developing useful genetic resources from their physiological characteristics.

Methods

Sample collection, DNA extraction, library construction, and sequencing

The S. thea plant was collected in Wando, Republic of Korea (34°21′47.02″ N, 126°38′31.52″ E) in September 2023. The botanical specimen was made from the branch with leaves and was deposited at the Honam National Institute of Biological Resources (http://en.hnibr.re.kr/) under voucher number HNIBRVP41778. The HMW genomic DNA of S. thea was extracted from fresh leaves using cetyltrimethylammonium bromide (CTAB) lysis buffer (2% CTAB, 0.1 M Tris-HCl pH 8.0, 20 mM EDTA pH 8.0, 1.4 M sodium chloride, and 1% beta-mercaptoethanol), phenol/chloroform/isoamylalcohol, and ethanol precipitation. The library for Illumina short-read sequencing was constructed using the QIAseq FX DNA library kit (Qiagen, Hilden, Germany), and sequencing was performed on the Illumina NovaSeq 6000 platform (Illumina, San Diego, CA, USA) with 151 bp paired ends. The Nanopore sequencing library was prepared using a Ligation Sequencing Kit (SQK-LSK114, Oxford Nanopore Technologies, Oxford, UK), and sequencing was performed on the PromethION platform. A total of 50.61 Gb of Illumina short-read data (256.3× coverage) and a total of 77.55 Gb of Nanopore long-read data (392.7× coverage) were generated. All Illumina raw reads were subsequently trimmed using Trimmomatic v0.366 with the ‘MINLEN:30’ option, and Nanopore reads were pre-processed by Guppy v6.5.77 and Porechop v0.2.48. Post quality control, 49.56 Gb and 73.13 Gb clean sequences were retained from the Illumina and Nanopore reads, respectively.

The transcriptome profile of S. thea was obtained by RNA sequencing to provide further evidence for gene prediction. Towards this end, total RNA was extracted from fresh leaves using a Qiagen RNeasy Mini Kit. The RNA-seq library was prepared using the TruSeq Total RNA Library Kit, and sequenced using the Illumina NovaSeq platform with 151 bp paired ends. A Total of 7.61 Gb of sequences was generated as a result of transcriptome sequencing, and a total of 7.39 Gb (97.1% of raw reads) was retained post-trimming using the above described Trimmomatic run.

Pore-C sequencing

Pore-C sequencing based on the Oxford Nanopore sequencing platform was employed to acquire chromosomal contact information for the S. thea genome5. Approximately 3 g of fresh leaves was chopped and fixed with 1% formaldehyde under 16 kPa vacuum conditions. The cross-linked material was crushed into a fine powder using a mortar and pestle in liquid nitrogen. Nuclei were extracted from the powder using a 40μm cell strainer in the nuclei extraction buffer comprising 1% PVP-40, 0.5% ECOSURF™ EH-9 surfactant, and 0.25% beta-mercaptoethanol. The isolated nuclei were digested using the NlaIII restriction enzyme (New England Biolabs, Ipswich, MA, USA) for 18 hours at 37 °C. The digested material was de-crosslinked using Proteinase K, and chromatin DNA was extracted using a phenol/chloroform solution and ethanol precipitation. Nanopore sequencing library construction and sequencing were conducted as described previously. A total of 24.18 Gb of raw data was generated, and 21.42 Gb (108.5× coverage) of clean reads were retained after pre-processing using Guppy and Porechop runs (Table 1).Table 1 General information on sequencing raw data of this study.

Read type (platform)	Molecule type	Raw-data size (Gb)	Cleaned-data size (Gb)	Use	NCBI SRA accession	
Nanopore long-reads (PromethION)	DNA	77.55	73.13	Assembly	SRR28713546	
Illumina short-reads (NovaSeq 6000)	DNA	50.61	49.56	Polishing	SRR28713547	
Pore-C reads (PromethION)	DNA	24.18	21.42	Scaffolding	SRR28713544	
Illumina RNA-seq reads (NovaSeq 6000)	RNA	7.61	7.39	Gene prediction	SRR28713545	

de novo genome assembly and scaffolding with Pore-C reads

The genome size of S. thea was determined using the k-mer-based estimation method prior to genome assembly. All k-mer (k = 21) frequencies of the trimmed Illumina reads were analyzed using Jellyfish v2.3.09, and genome size information was profiled using GenomeScope v1.010. The S. thea genome was thus found to have a genome size of 197.5 Mb with 0.92% heterozygosity (Fig. 1A).Fig. 1 De novo genome assembly of S. thea. (A) Genome size estimation summary using k-mer (21-mer) analysis. (B) Chromosomal contacts map for the scaffolding of S. thea genome using the Pore-C sequencing technology. Blue lines indicate chromosomal scaffold boundaries, and green lines indicate contig boundaries. (C) The chromosomal genome map of S. thea. The inner layers sequentially mean gene density (a), repeat sequence density (b), DNA element density (c), retroelement density (d), and interspecific co-linearity of protein-coding genes (e). Dark colors indicate high density and light colors indicate low density in all heatmap graphs (a–d).

The draft genome was initially assembled from Nanopore long reads using NextDenovo v2.5.211 with default parameters. The resultant assembly and Pore-C sequencing raw reads were pre-processed (e.g., filtering and alignment) using Pore-C-Snakemake v0.412 for subsequent scaffolding analysis. The scaffolding was subsequently performed using the ‘run-asm-pipeline.sh’ script from the 3D-DNA pipeline v19071613 with ‘-i 10000 -q 0 --editor-coarse-resolution 250000 --editor-coarse-region 1250000 --polisher-coarse-resolution 1000000 --polisher-coarse-region 30000000’ modified parameters. The scaffolding results were visualized using JuiceBox v1.1114 and manually curated based on a pairwise contact heatmap (Fig. 1B). The final scaffolded assembly was produced by the ‘run-asm-pipeline-post-review.sh’ script of 3D-DNA and polished using NextPolish v1.4.115 with trimmed Illumina reads. To filter out non-chromosomal scaffolds, we removed haplotig, junk, and too-short (<10,000 bp) scaffolds using the Purge Haplotigs pipeline v1.1.216. Small-organellar sequences were also removed based on a BLAST search against the NCBI organelle genomes of Rhamnaceae species (chloroplast genome accession; OR039202.117, and mitochondrial genome accession; KU187967.118)19,20. We finally assembled a total scaffold length of 197.8 Mb representing 12 scaffolds, with 15.9 Mb of the scaffold N50 and 34.49% GC content. Our assembly size was identical to the estimated genome size (100.002%, assembled size/estimated size = 197.8/197.5 Mb), and the BUSCO score21 revealed 98.4% of complete eukaryotic-BUSCOs and 97.9% of complete eudicots-BUSCOs. The assembly results are summarized in Table 2, and the chromosomal genome characteristics and self-collinearities profiled by reciprocal BLAST searches were visualized using Circos v0.6922 (Fig. 1C).Table 2 Genome assembly and scaffolding statistics.

Assembly statistics	Initial assembly (NextDenovo run only)	After Pore-C scaffolding	Final chromosomal genome	
Assembly size	200,225,435	201,068,823	197,778,613	
Number of contigs or scaffolds (>10 Kbp)	29 contigs	25 scaffolds	12 scaffolds	
N50 length	13,876,666	15,864,814	15,864,814	
Largest contig or scaffold length	23,456,224	27,570,280	27,570,280	
GC contents (%)	34.56	34.55	34.49	

Gene prediction and functional annotation

Prior to protein-coding gene prediction, transposable elements (TEs) in the S. thea genome were profiled. Genome-wide TEs were predicted using RepeatModeler v2.0.123, which integrates various ab initio repeat prediction tools, including, RepeatScout v1.0.624, RECON v1.0825, and TRF v4.1026, as well as the long-terminal repeat identifiers LtrHarvest v1.6.127 and Ltr_retriever v2.928. The predicted TEs were subsequently classified using DeepTE29, a machine learning-based de novo repeat classifier. All identified TE regions were masked using RepeatMasker v4.1.1 with a soft-masking option (Table 3). To predict protein-coding gene regions in the S. thea genome, three prediction methods were employed, namely, ab initio gene prediction, protein homology-based, and transcriptome-based gene prediction. First, the BRAKER3 pipeline30 was run for ab initio gene prediction. To this end, de novo gene prediction was conducted using GeneMark-ETP v1.0231 and AUGUSTUS v3.5.032 tools, and gene model accuracy was further improved using TSEBRA v1.1.233. Moreover, RNA-seq alignment data aligned using the STAR v2.7.11 aligner34 were also employed in the prediction procedure to provide accurate splice alignment information. Second, gene prediction with protein homology information was conducted using the GeMoMa v1.9 pipeline35. A total of 11 reference genomes, including, those of Chlamydomonas reinhardtii, Physcomitrella patens, Oryza sativa, Arabidopsis thaliana, Coffea canephora, Medicago truncatula, Rosa chinensis, Malus domestica, Prunus persica, Ochetophila trinervis, and Ziziphus jujuba that were downloaded from the NCBI for Biotechnology Information genome (https://www.ncbi.nlm.nih.gov/genome/), Ensembl Plants (https://plants.ensembl.org/index.html)36, or previous publications were used for the reference proteome37,38. The protein sequences were aligned to our genome using MMseq 2 v13.45139, and the final gene model was produced using the annotation finalizer script in GeMoMa. Finally, transcriptome-based prediction was performed using the alignment assembly script in the PASA pipeline v2.5.340 with de novo transcriptome assembly using Trinity v2.1541. All prediction results were integrated using EvidenceModeler (EVM) v2.0.542 with multiple rounds of weight adjustments. The annotation comparison and update script was run in the PASA with two rounds of repetition to explore the isoform profiles of the EVM gene model. The gene model was manually curated by editing abnormally concatenated gene features. The final gene model revealed 25,434 protein-coding genes, with 96.5% eukaryotes and 92.4% eudicots of complete BUSCOs (Table 4).Table 3 Repetitive sequences of S. thea genome.

Repeat classification	Length (proportion)	
Total repetitive sequences	65,959,671 bp (33.4%)	
Class I: Retroelements	21,947,516 bp (11.1%)	
Class II: DNA transposons	32,974,232 bp (16.7%)	
Other minor repeats	5,183,859 bp (2.6%)	
Unclassified repeats	5,854,064 bp (3.0%)	

Table 4 Detailed information for the S. thea gene model.

Gene model statistics	Value	
Number of genes	25,434	
Number of proteins (including isoforms)	34,645	
Complete BUSCOs, eukaryotes	96.5% (S: 62.0%, D: 34.5%)	
Complete BUSCOs, eudicots	92.4% (S: 70.1%, D: 22.3%)	
Number of database-hit genes	InterPro	23,410	
NCBI nr	23,265	
EggNog	22,050	
UniProtKB/Swiss-Prot	19,262	
KEGG	7,715	

Gene models were further annotated to profile their functions using five databases, namely, NCBI nr, UniProtKB/Swiss-Prot, KEGG, EggNOG, and InterPro (Table 4). We searched the homology of S. thea’s proteome against the NCBI nr and UniProtKB/Swiss-Prot databases43 using Diamond v2.1.8 with an e-value of 1e-10 in ultra-sensitive mode. The KEGG terms of the gene model were searched using the KEGG Automatic Annotation Server (KAAS) with a bidirectional best hit (BBH) option44. Annotation information from the eggNOG database was assigned using the eggNOG-mapper v2.1.1245. The protein domain information from the InterPro database46 was profiled using InterProScan v5.6447. A total of 24,063 genes (i.e., 94.6% of all predicted genes) were found to have at least one match across the above five databases. The functional annotation results are summarized in the upset plot made using the UpsetR R package48 in Fig. 2A.Fig. 2 Genome annotation summary and the orthologous gene profile of S. thea. (A) The Upset plot for functional annotation results of S. thea genes. (B) Orthologous gene gain and loss information of S. thea genome as compared to other reference plants. The numbers at each node or tip indicate the number of orthogroups (black), gained orthogroups (red), and lost orthogroups (blue). (C) Synteny alignments with the closest chromosomal-reference genome, Ziziphus jujuba.

Orthologous gene profiling and synteny analysis

The gene family repertoire of the S. thea genome was compared with that of other Rhamnaceae and Rosales members by profiling orthologous genes using 11 reference genomes downloaded from public genome databases. The Orthofinder v2.5.549 analysis with all-by-all ultrasensitive diamond searches was subsequently launched. The backbone species tree was inferred by aligning the protein sequences of 195 single-copy orthologs using MAFFT v7.45350, and constructing the maximum-likelihood tree using IQ-TREE v1.6.1251 with concatenated alignment. Further gene gain or loss profiles were estimated using the dollo-parsimony method in Count v10.04 (Fig. 2B)52. Additionally, we compared the genome-wide syntenies between the S. thea genome and that of Z. jujuba38,53, the closest reference genome within Rhamnaceae. MCscanX54 was employed to explore and visualize syntenic blocks with a cut-off of ≥ 10 colinear genes. This syntenic comparison showed highly conserved synteny between S. thea and Z. jujuba, and there were several rearrangements of sub-telomeric genes such as on S. thea’s chr2 and chr4 (Fig. 2C).

Data Records

Illumina sequencing, Nanopore long-read sequencing, Pore-C sequencing, and transcriptome sequencing raw reads were deposited in the NCBI Sequence Read Archive (SRA) database with accession numbers SRR2871354455, SRR2871354556, SRR2871354657, and SRR2871354758. The chromosomal genome assembly and genome annotation of S. thea are deposited in Figshare (10.6084/m9.figshare.25877698)59, and the assembly was also submitted to the NCBI GenBank with accession number JBGEWM00000000060.

Technical Validation

Quality control of nucleic acids and sequencing libraries

The DNA quality and purity were confirmed by 1% agarose gel electrophoresis and Nanodrop, and the quantity was measured using a Qubit 3.0 fluorometer with a Qubit dsDNA BR Assay Kit. The RNA integrity was qualified using an Agilent Bioanalyzer 2100 (Agilent Technologies, Santa Clara, CA, USA). The quality of all Illumina sequencing libraries was checked using the Bioanalyzer 2100.

Validation of the de novo genome assembly and genome annotation

Detailed results of each genome assembly were confirmed using QUAST v5.261. To verify the contiguity of the final genome assembly, we aligned all raw reads (i.e., Illumina short reads, and Nanopore long reads) to the assembly using Minimap2 v2.2662. These alignments were further visualized by the ‘wgscoverageplotter.py’ script in jvarkit v2023.0963. The resultant plots showed at least 45× (Illumina reads), and 115× (Nanopore reads) coverages across all chromosomal contigs, and the coverages were evenly distributed (Supplementary Figure S1). Assembly and gene model completeness were validated by BUSCO (Benchmarking Universal Single-Copy Orthologs) in the genome (-m genome) and proteins modes (-m prot). Two odb10 datasets, namely, eukaryota and eudicots, were used for all the BUSCO runs. The quality of the final gene model was corroborated using syntenic alignment and orthologous gene profiles (Fig. 2).

Supplementary information

Supplementary Figure S1

Supplementary information

The online version contains supplementary material available at 10.1038/s41597-024-03798-9.

Acknowledgements

This work was supported by grants from the Honam National Institute of Biological Resources (HNIBR), funded by the Ministry of Environment (MOE) of the Republic of Korea (Grant No. HNIBR202302118 and HNIBR202401104).

Author contributions

J.J., T.W.J., K.J.L. and S.A.L. conceived the project. J.S.P. and.S.A.L. performed the field collection. H.W., J.S.J. and S.A.L. performed the molecular experiments. J.J. performed all genome assembly and annotation analyses, and wrote the manuscript. All the authors have read and approved the final version of the manuscript.

Code availability

All software and/or pipelines used in this genome study were executed according to the developer’s manual from the corresponding bioinformatics publication or software webpage such as GitHub. Each software version and the detailed parameters are described in the Methods section, and the default parameters were employed unless otherwise stated.

Competing interests

The authors declare no competing interests.

Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
==== Refs
References

1. Yang Y Chen Y-S Zhang J-W Sun L Sun H Phylogenetics and historical biogeography of the mock buckthorn genus (Sageretia, Rhamnaceae) Botanical Journal of the Linnean Society 2019 189 244 261 10.1093/botlinnean/boy077
Yang, Y., Chen, Y.-S., Zhang, J.-W., Sun, L. & Sun, H. Phylogenetics and historical biogeography of the mock buckthorn genus (Sageretia, Rhamnaceae). Botanical Journal of the Linnean Society 189, 244–261, 10.1093/botlinnean/boy077 (2019).10.1093/botlinnean/boy077
2. Chung SK Chen CY Blumberg JB Flavonoid-rich fraction from Sageretia theezans leaves scavenges reactive oxygen radical species and increases the resistance of low-density lipoprotein to oxidation J Med Food 2009 12 1310 1315 10.1089/jmf.2008.1309 20041786
Chung, S. K., Chen, C. Y. & Blumberg, J. B. Flavonoid-rich fraction from Sageretia theezans leaves scavenges reactive oxygen radical species and increases the resistance of low-density lipoprotein to oxidation. J Med Food 12, 1310–1315, 10.1089/jmf.2008.1309 (2009).20041786 10.1089/jmf.2008.1309
3. Ko GA Shrestha S Kim Cho S Sageretia thea fruit extracts rich in methyl linoleate and methyl linolenate downregulate melanogenesis via the Akt/GSK3beta signaling pathway Nutr Res Pract 2018 12 3 12, 10.4162/nrp.2018.12.1.3 29399291
Ko, G. A., Shrestha, S. & Kim Cho, S. Sageretia thea fruit extracts rich in methyl linoleate and methyl linolenate downregulate melanogenesis via the Akt/GSK3beta signaling pathway. Nutr Res Pract 12, 3–12, 10.4162/nrp.2018.12.1.3 (2018).29399291 10.4162/nrp.2018.12.1.3
4. Kim HN Extracts from Sageretia thea reduce cell viability through inducing cyclin D1 proteasomal degradation and HO-1 expression in human colorectal cancer cells BMC Complement Altern Med 2019 19 43 10.1186/s12906-019-2453-4 30736789
Kim, H. N. et al. Extracts from Sageretia thea reduce cell viability through inducing cyclin D1 proteasomal degradation and HO-1 expression in human colorectal cancer cells. BMC Complement Altern Med 19, 43, 10.1186/s12906-019-2453-4 (2019).30736789 10.1186/s12906-019-2453-4
5. Deshpande AS Identifying synergistic high-order 3D chromatin conformations from genome-scale nanopore concatemer sequencing Nat Biotechnol 2022 40 1488 1499 10.1038/s41587-022-01289-z 35637420
Deshpande, A. S. et al. Identifying synergistic high-order 3D chromatin conformations from genome-scale nanopore concatemer sequencing. Nat Biotechnol 40, 1488–1499, 10.1038/s41587-022-01289-z (2022).35637420 10.1038/s41587-022-01289-z
6. Bolger AM Lohse M Usadel B Trimmomatic: a flexible trimmer for Illumina sequence data Bioinformatics 2014 30 2114 2120 10.1093/bioinformatics/btu170 24695404
Bolger, A. M., Lohse, M. & Usadel, B. Trimmomatic: a flexible trimmer for Illumina sequence data. Bioinformatics 30, 2114–2120, 10.1093/bioinformatics/btu170 (2014).24695404 10.1093/bioinformatics/btu170
7. Wick RR Judd LM Holt KE Performance of neural network basecalling tools for Oxford Nanopore sequencing Genome Biol 2019 20 129 10.1186/s13059-019-1727-y 31234903
Wick, R. R., Judd, L. M. & Holt, K. E. Performance of neural network basecalling tools for Oxford Nanopore sequencing. Genome Biol 20, 129, 10.1186/s13059-019-1727-y (2019).31234903 10.1186/s13059-019-1727-y
8. Wick, R. Porechop: adaptor trimmer for Oxford Nanopore reads., <https://github.com/rrwick/Porechop> (2017).
9. Marcais G Kingsford C A fast, lock-free approach for efficient parallel counting of occurrences of k-mers Bioinformatics 2011 27 764 770 10.1093/bioinformatics/btr011 21217122
Marcais, G. & Kingsford, C. A fast, lock-free approach for efficient parallel counting of occurrences of k-mers. Bioinformatics 27, 764–770, 10.1093/bioinformatics/btr011 (2011).21217122 10.1093/bioinformatics/btr011
10. Vurture GW GenomeScope: fast reference-free genome profiling from short reads Bioinformatics 2017 33 2202 2204 10.1093/bioinformatics/btx153 28369201
Vurture, G. W. et al. GenomeScope: fast reference-free genome profiling from short reads. Bioinformatics 33, 2202–2204, 10.1093/bioinformatics/btx153 (2017).28369201 10.1093/bioinformatics/btx153
11. Jiang, H. et al. An efficient error correction and accurate assembly tool for noisy long reads. bioRxiv, 2023.2003.2009.531669, 10.1101/2023.03.09.531669 (2023).
12. Technologies, O. N. Pore-C snakemake https://github.com/nanoporetech/Pore-C-Snakemake (2019).
13. Dudchenko O De novo assembly of the Aedes aegypti genome using Hi-C yields chromosome-length scaffolds Science 2017 356 92 95 10.1126/science.aal3327 28336562
Dudchenko, O. et al. De novo assembly of the Aedes aegypti genome using Hi-C yields chromosome-length scaffolds. Science 356, 92–95, 10.1126/science.aal3327 (2017).28336562 10.1126/science.aal3327
14. Durand NC Juicebox Provides a Visualization System for Hi-C Contact Maps with Unlimited Zoom Cell Syst 2016 3 99 101 10.1016/j.cels.2015.07.012 27467250
Durand, N. C. et al. Juicebox Provides a Visualization System for Hi-C Contact Maps with Unlimited Zoom. Cell Syst 3, 99–101, 10.1016/j.cels.2015.07.012 (2016).27467250 10.1016/j.cels.2015.07.012
15. Hu J Fan J Sun Z Liu S NextPolish: a fast and efficient genome polishing tool for long-read assembly Bioinformatics 2020 36 2253 2255 10.1093/bioinformatics/btz891 31778144
Hu, J., Fan, J., Sun, Z. & Liu, S. NextPolish: a fast and efficient genome polishing tool for long-read assembly. Bioinformatics 36, 2253–2255, 10.1093/bioinformatics/btz891 (2020).31778144 10.1093/bioinformatics/btz891
16. Roach MJ Schmidt SA Borneman AR Purge Haplotigs: allelic contig reassignment for third-gen diploid genome assemblies BMC Bioinformatics 2018 19 460 10.1186/s12859-018-2485-7 30497373
Roach, M. J., Schmidt, S. A. & Borneman, A. R. Purge Haplotigs: allelic contig reassignment for third-gen diploid genome assemblies. BMC Bioinformatics 19, 460, 10.1186/s12859-018-2485-7 (2018).30497373 10.1186/s12859-018-2485-7
17. 2023 Sageretia thea chloroplast, complete genome NCBI GenBank OR039202.1
Sageretia thea chloroplast, complete genome. NCBI GenBank. https://identifiers.org/ncbi/insdc:OR039202.1 (2023).
18. 2016 Ziziphus jujuba mitochondrion, complete genome NCBI GenBank KU187967.1
Ziziphus jujuba mitochondrion, complete genome. NCBI GenBank. https://identifiers.org/ncbi/insdc:KU187967.1 (2016).
19. Zhan M Wang X Chen W Huang X Complete chloroplast genome of sageretia thea (rhamnaceae), an ornamental fruit and medicinal tree Mitochondrial DNA B Resour 2024 9 376 380 10.1080/23802359.2024.2329667 38545567
Zhan, M., Wang, X., Chen, W. & Huang, X. Complete chloroplast genome of sageretia thea (rhamnaceae), an ornamental fruit and medicinal tree. Mitochondrial DNA B Resour 9, 376–380, 10.1080/23802359.2024.2329667 (2024).38545567 10.1080/23802359.2024.2329667
20. Wang X Organellar genome assembly methods and comparative analysis of horticultural plants Hortic Res 2018 5 3 10.1038/s41438-017-0002-1 29423233
Wang, X. et al. Organellar genome assembly methods and comparative analysis of horticultural plants. Hortic Res 5, 3, 10.1038/s41438-017-0002-1 (2018).29423233 10.1038/s41438-017-0002-1
21. Manni M Berkeley MR Seppey M Simao FA Zdobnov EM BUSCO Update: Novel and Streamlined Workflows along with Broader and Deeper Phylogenetic Coverage for Scoring of Eukaryotic, Prokaryotic, and Viral Genomes Mol Biol Evol 2021 38 4647 4654 10.1093/molbev/msab199 34320186
Manni, M., Berkeley, M. R., Seppey, M., Simao, F. A. & Zdobnov, E. M. BUSCO Update: Novel and Streamlined Workflows along with Broader and Deeper Phylogenetic Coverage for Scoring of Eukaryotic, Prokaryotic, and Viral Genomes. Mol Biol Evol 38, 4647–4654, 10.1093/molbev/msab199 (2021).34320186 10.1093/molbev/msab199
22. Krzywinski M Circos: an information aesthetic for comparative genomics Genome Res 2009 19 1639 1645 10.1101/gr.092759.109 19541911
Krzywinski, M. et al. Circos: an information aesthetic for comparative genomics. Genome Res 19, 1639–1645, 10.1101/gr.092759.109 (2009).19541911 10.1101/gr.092759.109
23. Flynn JM RepeatModeler2 for automated genomic discovery of transposable element families Proc Natl Acad Sci USA 2020 117 9451 9457 10.1073/pnas.1921046117 32300014
Flynn, J. M. et al. RepeatModeler2 for automated genomic discovery of transposable element families. Proc Natl Acad Sci USA 117, 9451–9457, 10.1073/pnas.1921046117 (2020).32300014 10.1073/pnas.1921046117
24. Price AL Jones NC Pevzner PA De novo identification of repeat families in large genomes Bioinformatics 2005 21 Suppl 1 i351 358, 10.1093/bioinformatics/bti1018 15961478
Price, A. L., Jones, N. C. & Pevzner, P. A. De novo identification of repeat families in large genomes. Bioinformatics 21(Suppl 1), i351–358, 10.1093/bioinformatics/bti1018 (2005).15961478 10.1093/bioinformatics/bti1018
25. Bao Z Eddy SR Automated de novo identification of repeat sequence families in sequenced genomes Genome Res 2002 12 1269 1276 10.1101/gr.88502 12176934
Bao, Z. & Eddy, S. R. Automated de novo identification of repeat sequence families in sequenced genomes. Genome Res 12, 1269–1276, 10.1101/gr.88502 (2002).12176934 10.1101/gr.88502
26. Benson G Tandem repeats finder: a program to analyze DNA sequences Nucleic Acids Res 1999 27 573 580 10.1093/nar/27.2.573 9862982
Benson, G. Tandem repeats finder: a program to analyze DNA sequences. Nucleic Acids Res 27, 573–580, 10.1093/nar/27.2.573 (1999).9862982 10.1093/nar/27.2.573
27. Ellinghaus D Kurtz S Willhoeft U LTRharvest, an efficient and flexible software for De novo detection of LTR retrotransposons BMC Bioinformatics 2008 9 18 10.1186/1471-2105-9-18 18194517
Ellinghaus, D., Kurtz, S. & Willhoeft, U. LTRharvest, an efficient and flexible software for De novo detection of LTR retrotransposons. BMC Bioinformatics 9, 18, 10.1186/1471-2105-9-18 (2008).18194517 10.1186/1471-2105-9-18
28. Ou S Jiang N LTR_retriever: A Highly Accurate and Sensitive Program for Identification of Long Terminal Repeat Retrotransposons Plant Physiol 2018 176 1410 1422 10.1104/pp.17.01310 29233850
Ou, S. & Jiang, N. LTR_retriever: A Highly Accurate and Sensitive Program for Identification of Long Terminal Repeat Retrotransposons. Plant Physiol 176, 1410–1422, 10.1104/pp.17.01310 (2018).29233850 10.1104/pp.17.01310
29. Yan H Bombarely A Li S DeepTE: a computational method for De novo classification of transposons with convolutional neural network Bioinformatics 2020 36 4269 4275 10.1093/bioinformatics/btaa519 32415954
Yan, H., Bombarely, A. & Li, S. DeepTE: a computational method for De novo classification of transposons with convolutional neural network. Bioinformatics 36, 4269–4275, 10.1093/bioinformatics/btaa519 (2020).32415954 10.1093/bioinformatics/btaa519
30. Gabriel L BRAKER3: Fully automated genome annotation using RNA-seq and protein evidence with GeneMark-ETP, AUGUSTUS and TSEBRA bioRxiv 2024 10.1101/2023.06.10.544449 38895358
Gabriel, L. et al. BRAKER3: Fully automated genome annotation using RNA-seq and protein evidence with GeneMark-ETP, AUGUSTUS and TSEBRA. bioRxiv10.1101/2023.06.10.544449 (2024).38895358 10.1101/2023.06.10.544449
31. Bruna, T., Lomsadze, A. & Borodovsky, M. GeneMark-ETP: Automatic Gene Finding in Eukaryotic Genomes in Consistency with Extrinsic Data. bioRxiv, 10.1101/2023.01.13.524024 (2024).
32. Stanke M AUGUSTUS: ab initio prediction of alternative transcripts Nucleic Acids Res 2006 34 W435 439, 10.1093/nar/gkl200 16845043
Stanke, M. et al. AUGUSTUS: ab initio prediction of alternative transcripts. Nucleic Acids Res 34, W435–439, 10.1093/nar/gkl200 (2006).16845043 10.1093/nar/gkl200
33. Gabriel L Hoff KJ Bruna T Borodovsky M Stanke M TSEBRA: transcript selector for BRAKER BMC Bioinformatics 2021 22 566 10.1186/s12859-021-04482-0 34823473
Gabriel, L., Hoff, K. J., Bruna, T., Borodovsky, M. & Stanke, M. TSEBRA: transcript selector for BRAKER. BMC Bioinformatics 22, 566, 10.1186/s12859-021-04482-0 (2021).34823473 10.1186/s12859-021-04482-0
34. Dobin A STAR: ultrafast universal RNA-seq aligner Bioinformatics 2013 29 15 21 10.1093/bioinformatics/bts635 23104886
Dobin, A. et al. STAR: ultrafast universal RNA-seq aligner. Bioinformatics 29, 15–21, 10.1093/bioinformatics/bts635 (2013).23104886 10.1093/bioinformatics/bts635
35. Keilwagen J Hartung F Paulini M Twardziok SO Grau J Combining RNA-seq data and homology-based gene prediction for plants, animals and fungi BMC Bioinformatics 2018 19 189 10.1186/s12859-018-2203-5 29843602
Keilwagen, J., Hartung, F., Paulini, M., Twardziok, S. O. & Grau, J. Combining RNA-seq data and homology-based gene prediction for plants, animals and fungi. BMC Bioinformatics 19, 189, 10.1186/s12859-018-2203-5 (2018).29843602 10.1186/s12859-018-2203-5
36. Yates AD Ensembl Genomes 2022: an expanding genome resource for non-vertebrates Nucleic Acids Res 2022 50 D996 D1003 10.1093/nar/gkab1007 34791415
Yates, A. D. et al. Ensembl Genomes 2022: an expanding genome resource for non-vertebrates. Nucleic Acids Res 50, D996–D1003, 10.1093/nar/gkab1007 (2022).34791415 10.1093/nar/gkab1007
37. Griesmann, M. et al. Phylogenomics reveals multiple losses of nitrogen-fixing root nodule symbiosis. Science 361, 10.1126/science.aat1743 (2018).
38. Shen LY Chromosome-Scale Genome Assembly for Chinese Sour Jujube and Insights Into Its Genome Evolution and Domestication Signature Front Plant Sci 2021 12 773090 10.3389/fpls.2021.773090 34899800
Shen, L. Y. et al. Chromosome-Scale Genome Assembly for Chinese Sour Jujube and Insights Into Its Genome Evolution and Domestication Signature. Front Plant Sci 12, 773090, 10.3389/fpls.2021.773090 (2021).34899800 10.3389/fpls.2021.773090
39. Steinegger M Soding J MMseqs 2 enables sensitive protein sequence searching for the analysis of massive data sets Nat Biotechnol 2017 35 1026 1028 10.1038/nbt.3988 29035372
Steinegger, M. & Soding, J. MMseqs 2 enables sensitive protein sequence searching for the analysis of massive data sets. Nat Biotechnol 35, 1026–1028, 10.1038/nbt.3988 (2017).29035372 10.1038/nbt.3988
40. Haas BJ Improving the Arabidopsis genome annotation using maximal transcript alignment assemblies Nucleic Acids Res 2003 31 5654 5666 10.1093/nar/gkg770 14500829
Haas, B. J. et al. Improving the Arabidopsis genome annotation using maximal transcript alignment assemblies. Nucleic Acids Res 31, 5654–5666, 10.1093/nar/gkg770 (2003).14500829 10.1093/nar/gkg770
41. Grabherr MG Full-length transcriptome assembly from RNA-Seq data without a reference genome Nat Biotechnol 2011 29 644 652 10.1038/nbt.1883 21572440
Grabherr, M. G. et al. Full-length transcriptome assembly from RNA-Seq data without a reference genome. Nat Biotechnol 29, 644–652, 10.1038/nbt.1883 (2011).21572440 10.1038/nbt.1883
42. Haas BJ Automated eukaryotic gene structure annotation using EVidenceModeler and the Program to Assemble Spliced Alignments Genome Biol 2008 9 R7 10.1186/gb-2008-9-1-r7 18190707
Haas, B. J. et al. Automated eukaryotic gene structure annotation using EVidenceModeler and the Program to Assemble Spliced Alignments. Genome Biol 9, R7, 10.1186/gb-2008-9-1-r7 (2008).18190707 10.1186/gb-2008-9-1-r7
43. UniProt C UniProt: the Universal Protein Knowledgebase in 2023 Nucleic Acids Res 2023 51 D523 D531 10.1093/nar/gkac1052 36408920
UniProt, C. UniProt: the Universal Protein Knowledgebase in 2023. Nucleic Acids Res 51, D523–D531, 10.1093/nar/gkac1052 (2023).36408920 10.1093/nar/gkac1052
44. Moriya Y Itoh M Okuda S Yoshizawa AC Kanehisa M KAAS: an automatic genome annotation and pathway reconstruction server Nucleic Acids Res 2007 35 W182 185, 10.1093/nar/gkm321 17526522
Moriya, Y., Itoh, M., Okuda, S., Yoshizawa, A. C. & Kanehisa, M. KAAS: an automatic genome annotation and pathway reconstruction server. Nucleic Acids Res 35, W182–185, 10.1093/nar/gkm321 (2007).17526522 10.1093/nar/gkm321
45. Cantalapiedra CP Hernandez-Plaza A Letunic I Bork P Huerta-Cepas J eggNOG-mapper v2: Functional Annotation, Orthology Assignments, and Domain Prediction at the Metagenomic Scale Mol Biol Evol 2021 38 5825 5829 10.1093/molbev/msab293 34597405
Cantalapiedra, C. P., Hernandez-Plaza, A., Letunic, I., Bork, P. & Huerta-Cepas, J. eggNOG-mapper v2: Functional Annotation, Orthology Assignments, and Domain Prediction at the Metagenomic Scale. Mol Biol Evol 38, 5825–5829, 10.1093/molbev/msab293 (2021).34597405 10.1093/molbev/msab293
46. Blum M The InterPro protein families and domains database: 20 years on Nucleic Acids Res 2021 49 D344 D354 10.1093/nar/gkaa977 33156333
Blum, M. et al. The InterPro protein families and domains database: 20 years on. Nucleic Acids Res 49, D344–D354, 10.1093/nar/gkaa977 (2021).33156333 10.1093/nar/gkaa977
47. Jones P InterProScan 5: genome-scale protein function classification Bioinformatics 2014 30 1236 1240 10.1093/bioinformatics/btu031 24451626
Jones, P. et al. InterProScan 5: genome-scale protein function classification. Bioinformatics 30, 1236–1240, 10.1093/bioinformatics/btu031 (2014).24451626 10.1093/bioinformatics/btu031
48. Conway JR Lex A Gehlenborg N UpSetR: an R package for the visualization of intersecting sets and their properties Bioinformatics 2017 33 2938 2940 10.1093/bioinformatics/btx364 28645171
Conway, J. R., Lex, A. & Gehlenborg, N. UpSetR: an R package for the visualization of intersecting sets and their properties. Bioinformatics 33, 2938–2940, 10.1093/bioinformatics/btx364 (2017).28645171 10.1093/bioinformatics/btx364
49. Emms DM Kelly S OrthoFinder: phylogenetic orthology inference for comparative genomics Genome Biol 2019 20 238 10.1186/s13059-019-1832-y 31727128
Emms, D. M. & Kelly, S. OrthoFinder: phylogenetic orthology inference for comparative genomics. Genome Biol 20, 238, 10.1186/s13059-019-1832-y (2019).31727128 10.1186/s13059-019-1832-y
50. Katoh K Standley DM MAFFT multiple sequence alignment software version 7: improvements in performance and usability Mol Biol Evol 2013 30 772 780 10.1093/molbev/mst010 23329690
Katoh, K. & Standley, D. M. MAFFT multiple sequence alignment software version 7: improvements in performance and usability. Mol Biol Evol 30, 772–780, 10.1093/molbev/mst010 (2013).23329690 10.1093/molbev/mst010
51. Nguyen LT Schmidt HA von Haeseler A Minh BQ IQ-TREE: a fast and effective stochastic algorithm for estimating maximum-likelihood phylogenies Mol Biol Evol 2015 32 268 274 10.1093/molbev/msu300 25371430
Nguyen, L. T., Schmidt, H. A., von Haeseler, A. & Minh, B. Q. IQ-TREE: a fast and effective stochastic algorithm for estimating maximum-likelihood phylogenies. Mol Biol Evol 32, 268–274, 10.1093/molbev/msu300 (2015).25371430 10.1093/molbev/msu300
52. Csuros M Count: evolutionary analysis of phylogenetic profiles with parsimony and likelihood Bioinformatics 2010 26 1910 1912 10.1093/bioinformatics/btq315 20551134
Csuros, M. Count: evolutionary analysis of phylogenetic profiles with parsimony and likelihood. Bioinformatics 26, 1910–1912, 10.1093/bioinformatics/btq315 (2010).20551134 10.1093/bioinformatics/btq315
53. Liu MJ The complex jujube genome provides insights into fruit tree biology Nat Commun 2014 5 5315 10.1038/ncomms6315 25350882
Liu, M. J. et al. The complex jujube genome provides insights into fruit tree biology. Nat Commun 5, 5315, 10.1038/ncomms6315 (2014).25350882 10.1038/ncomms6315
54. Wang Y MCScanX: a toolkit for detection and evolutionary analysis of gene synteny and collinearity Nucleic Acids Res 2012 40 e49 10.1093/nar/gkr1293 22217600
Wang, Y. et al. MCScanX: a toolkit for detection and evolutionary analysis of gene synteny and collinearity. Nucleic Acids Res 40, e49, 10.1093/nar/gkr1293 (2012).22217600 10.1093/nar/gkr1293
55. 2024 NCBI Sequence Read Archive SRR28713544
NCBI Sequence Read Archive https://identifiers.org/ncbi/insdc.sra:SRR28713544 (2024).
56. 2024 NCBI Sequence Read Archive SRR28713545
NCBI Sequence Read Archive https://identifiers.org/ncbi/insdc.sra:SRR28713545 (2024).
57. 2024 NCBI Sequence Read Archive SRR28713546
NCBI Sequence Read Archive https://identifiers.org/ncbi/insdc.sra:SRR28713546 (2024).
58. 2024 NCBI Sequence Read Archive SRR28713547
NCBI Sequence Read Archive https://identifiers.org/ncbi/insdc.sra:SRR28713547 (2024).
59. Jo J 2024 The chromosome-level genome assembly and annotation of mock buckthorn, Sageretia thea FigShare 10.6084/m9.figshare.25877698
Jo, J. et al. The chromosome-level genome assembly and annotation of mock buckthorn, Sageretia thea. FigShare10.6084/m9.figshare.25877698 (2024).10.6084/m9.figshare.25877698
60. Jo J 2024 The chromosome-level genome assembly of mock buckthorn, Sageretia thea GenBank JBGEWM000000000
Jo, J. et al. The chromosome-level genome assembly of mock buckthorn, Sageretia thea. GenBank https://identifiers.org/ncbi/insdc:JBGEWM000000000 (2024).
61. Gurevich A Saveliev V Vyahhi N Tesler G QUAST: quality assessment tool for genome assemblies Bioinformatics 2013 29 1072 1075 10.1093/bioinformatics/btt086 23422339
Gurevich, A., Saveliev, V., Vyahhi, N. & Tesler, G. QUAST: quality assessment tool for genome assemblies. Bioinformatics 29, 1072–1075, 10.1093/bioinformatics/btt086 (2013).23422339 10.1093/bioinformatics/btt086
62. Li H Minimap2: pairwise alignment for nucleotide sequences Bioinformatics 2018 34 3094 3100 10.1093/bioinformatics/bty191 29750242
Li, H. Minimap2: pairwise alignment for nucleotide sequences. Bioinformatics 34, 3094–3100, 10.1093/bioinformatics/bty191 (2018).29750242 10.1093/bioinformatics/bty191
63. Lindenbaum, P. JVarkit: java-based utilities for Bioinformatics, <https://github.com/lindenb/jvarkit> (2015).
