
==== Front
Sci Data
Sci Data
Scientific Data
2052-4463
Nature Publishing Group UK London

39294147
3860
10.1038/s41597-024-03860-6
Data Descriptor
Haplotype-resolved genome assembly of the upas tree (Antiaris toxicaria)
Miao Ke 12
Wang Ya 1
Hou Luxiao 12
Liu Yan 12
Liu Haiyang haiyangliu@mail.kib.ac.cn

3
http://orcid.org/0000-0002-1116-7485
Ji Yunheng jiyh@mail.kib.ac.cn

134
1 grid.9227.e 0000000119573309 CAS Key Laboratory for Plant Diversity and Biogeography of East Asia, Kunming Institute of Botany, Chinese Academy of Sciences, Kunming, 650201 China
2 https://ror.org/05qbk4x57 grid.410726.6 0000 0004 1797 8419 Kunming College of Life Science, University of Chinese Academy of Sciences, Kunming, 650201 China
3 grid.9227.e 0000000119573309 State Key Laboratory of Phytochemistry and Natural Medicine, Kunming Institute of Botany, Chinese Academy of Sciences, Kunming, 650201, China
4 grid.9227.e 0000000119573309 Yunnan Key Laboratory for Integrative Conservation of Plant Species with Extremely Small Population, Kunming Institute of Botany, Chinese Academy of Sciences, Kunming, 650201 China
18 9 2024
18 9 2024
2024
11 101115 3 2024
4 9 2024
© The Author(s) 2024
2024
https://creativecommons.org/licenses/by-nc-nd/4.0/ Open Access This article is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License, which permits any non-commercial use, sharing, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if you modified the licensed material. You do not have permission under this licence to share adapted material derived from this article or parts of it. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by-nc-nd/4.0/.
The upas tree (Antiaris toxicaria Lesch.) is a medically important plant that contains various specialized metabolites with significant bioactivity. The lack of a reference genome hinders the in-depth study as well as rational exploitation and conservation of this plant. Here, we present the first holotype-resolved chromosome-scale genome of the upas tree. The assembled genome consisted of 26 chromosomes that contain 1.34 Gb of sequencing data with a contig N50 length of 60 Mb. Genome annotation identified 43,500 protein-coding genes in the upas tree genome, of which 98.75% were functionally annotated. This high-quality reference genome will lay the foundation for further studies on the evolution and functional genomics of the upas tree.

Subject terms

Conservation genomics
Genome evolution
Medical genomics
Yunnan Revitalization Talent Support Program "Top Team&quot; Project (202305AT35)Yunnan Revitalization Talent Support Program "Top Team&quot; Project (202305AT35)issue-copyright-statement© Springer Nature Limited 2024
==== Body
pmcBackground & Summary

The upas tree (Antiaris toxicaria Lesch.; Fig. 1) is a deciduous woody plant within the Moraceae family, which is widely distributed throughout the Old World tropics1. It is a medicinally important plant with various therapeutic properties. Its seeds, latex, bark, and leaves have been traditionally employed as analgesics, anticonvulsants, antipyretics, anesthetics, cardiotonics, and emetics in the tropical areas of Africa and Asia2. Latex of upas tree is extremely toxic and has been used as traditional arrow and dart poison for hunting wild animals in China and Indochina3,4. Modern phytochemical and pharmacological studies indicate that the plant is rich in cardenolides, flavonoids, and lignan, of which the cardenolides are the main bioactive and toxic ingredients5–11, which are involved in chemical defense against herbivores.Fig. 1 Plant morphology and habitat of the upas tree (A. toxicaria).

Plant-derived cardenolides have been isolated from approximately 60 genera and 12 families of angiosperms belonging to core eudicots, monocots, and early-diverging eudicots12,13. The dispersal distribution of cardenolides in the phylogenetic tree of angiosperms implies that cardenolide biosynthesis represents an ancient pathway with multiple evolutionary losses and reversals14,15. Therefore, exploring the evolutionary trajectories of cardenolide biosynthesis in angiosperms may provide new insights into the evolution of plant defense and secondary metabolism.

Within the family Moraceae, Antiaris Lesch. is the only genus that has been found to contain cardenolides to date. The upas tree contains a complex mixture of cardenolides, and a series of cardenolides have been isolated from its latex16,17, seeds18,19, stems20, and bark21. These compounds exhibit a wide spectrum of biological activities, such as increasing the cardiac contractile force and inhibiting the proliferation of human gastric cancer, hepatocellular carcinoma, and lung cancer cells20,22,23. Due to the absence of a reference genome, the genetic basis of cardenolide and flavonoid enrichment in the upas tree remains to be explored, and the gene pathways encoding the biosynthesis of plant-derived cardenolides have not been deciphered14.

Here, we present a high-quality reference genome of the upas tree (Fig. 2) by combining PacBio high-fidelity (HiFi) sequencing, Illumina short-read sequencing, and high-throughput chromosome conformation capture (Hi-C) sequencing data. To the best of our knowledge, the upas tree genome is the first holotype-resolved genome of the Moraceae family. The reference genome will serve as a valuable genetic resource for investigating evolutionary biology and functional genomics of the upas tree, as well as for exploring the gene pathways involved in plant-derived cardenolide biosynthesis.Fig. 2 Characterization of the haplotype-resolved genome assembly of the upas tree (A. toxicaria). From the outer ring to the inner ring are (a) the distributions of pseudochromosomes, (b) class I TE density, (c) class II TE density, (d) protein-coding gene density, (e) proportion of tandem repeat, (f) GC content, and g: syntenic blocks.

Methods

Plant materials

Plant individuals of A. toxicaria were originally collected from a natural population in Mengla, Yunnan Province, southwest China, and cultivated in a greenhouse at Kunming Institute of Botany, Chinese Academy of Science. Fresh young leaves were collected for genomic DNA extraction. Two types of tissues (stems and leaves) were collected from the same individual plant for RNA extraction. These materials were frozen immediately in liquid nitrogen and then cryogenically preserved (−80 °C) until DNA and RNA extraction. The voucher specimen (Gao & Ji 002) was deposited at the herbarium of Kunming Institute of Botany, Chinese Academy of Science (KUN).

Illumina short-read sequencing

Genomic DNA was extracted from the leaf tissues using a modified CTAB method24. A PCR-free library for Illumina short-read with an average insert size of 300 bp was constructed using the Truseq nano DNA HT library preparation kit (Illumina, MA, USA) following the manufacturer’s instructions. The constructed library was subjected to paired-end sequencing on the Illumina NovaSeq 6000 platform (Illumina, MA, USA). The Illumina short-read sequencing generated approximately 184 Gb of raw data (approximately 615 million 2 × 150 bp reads), resulting in a coverage of approximately 269 times. The fastp program (v. 0.19.3)25 was used to trim adapters and filter low-quality reads from the Illumina raw data.

PacBio high-fidelity (HiFi) sequencing

High-quality genomic DNA was obtained from young leaves of the upas tree using the QIAGEN Plant Genomic DNA Extraction Kit (Germantown, MD, USA) and randomly fragmented using a Megaruptor (Diagenode, NJ, USA). The SMRTbell Express Template Prep Kit 2.0 (Pacific Biosciences, CA, USA) and the PacBio Sequel II Binding Kit 2.2 (Pacific Biosciences) were employed to construct a PacBio single-molecule real-time (SMRT) library, following the manufacturer’s instruction. Insert sizes over 500 bp were selected for HiFi sequencing, using the BluePippin system (Sage Science, Beverly, MA, USA). The library was sequenced on the PacBio Sequel II platform (Pacific Biosciences, CA, USA), and raw reads were filtered using the CCS program (v. 6.3.0)26, generating approximately 35 Gb (around 1.8 million reads) of high-quality HiFi data, with an average read length and N50 read length of approximately 20 kb. The HiFi data yielded a coverage of approximately 51 times.

High-throughput chromosome conformation capture (Hi-C) sequencing

Young and tender leaves were fixed in a 2% formaldehyde solution, and the cross-linked DNA was digested with DpnII (NEB, UK). The sticky ends of the digested DNA fragments were attached to biotin-labeled adapters, which were specifically enriched to obtain fragments of approximately 500 bp for constructing the Hi-C library. The Hi-C library was sequenced using the Illumina NovaSeq 6000 platform. In total, approximately 76 Gb (approximately 508 million reads) of high-quality Hi-C data were generated after the adapters were trimmed and the low-quality reads were filtered. The filtered Hi-C data yielded a coverage of approximately 111 times.

Transcriptome sequencing

A U-mRNAseq Library Prep Kit (Tecan, CA, USA) was used to extract RNA from separate tissues (stems and leaves). RNA-seq libraries were prepared separately for each tissue type, with one library constructed for stems and one library constructed for leaves. The libraries were sequenced using the Illumina NovaSeq 6000 platform, generating approximately 6.69 Gb of data from stems and 7.23 Gb of data from leaves. In total, approximately 14 Gb (approximately 47 million 2 × 150 bp reads) of RNA-seq data were acquired for genome annotation.

Genome size and heterozygosity estimation

Based on the clean Illumina sequencing data, the size and heterozygosity of the upas tree genome were evaluated using K-mer analysis22. Inferred from the frequency distribution of the K-mers (Supplementary Fig. 1), the genomic traits (i.e., genome size, heterozygosity, and the proportion of repetitive sequences) of the upas tree were evaluated using GCE (v. 1.0.0)27 and findGSE (v. 1.94)28. The haploid genome size of the upas tree was estimated to be approximately 711 Mb (Supplementary Fig. 2), with a heterozygosity of approximately 0.6% and a repetitiveness rate of approximately 58.52% (Supplementary Table 1).

Haplotype-resolved genome assembly

The program Hifiasm (v. 0.16.1-r375)29 was used to assemble PacBio HiFi reads into haplotype-resolved contigs with default parameters. Next, Hi-C reads were mapped to these contigs using Juicer (v. 1.5.6)30; the 3D-DNA pipeline (v. 180922)31 was used to integrate these contigs into separate pseudochromosomes. The de novo assembled chromosome-level genome was improved using the manually operated Juicebox (v. 1.11.08)32. For further optimization, LR_GapCloser33 was used to fill gaps in the genome based on the HiFi reads. Although most assembled pseudochromosomes possess telomere sequence features, some have shorter or absent telomere sequences owing to incomplete assembly or insufficient extension. To further improve the assembly, HiFi reads were aligned back to these incompletely assembled pseudochromosomes, and the reads near the telomeres were assembled into contigs using Hifiasm (v. 0.16.1-r375)29. These contigs were then mapped to the incompletely assembled pseudochromosomes to recover the telomere sequences as completely as possible. Then, the GetOrganelle (v. 1.7.7.0)34 was utilized to assemble the chloroplast and mitochondrial genomes. Finally, NextPolish (v. 1.3.1)35 was used to run two rounds of polishing based on the Illumina short-read sequencing data. Redundans (v. 0.13c)36 was used to compare the scattered contig sequences with each pseudochromosome and remove any redundancies (such as haplotigs and rDNA fragments) following the manual inspection. The genome assembly anchored approximately 99.96% of the assembled data to 26 pseudochromosomes (Supplementary Table 2) and assigned them to two haploid genomes (Fig. 2), which were labeled Chr 01a–13a and Chr 01b–13b. The assembled organelle genomes contained a chloroplast genome and a mitochondrial genome, with length of 161,699 bp and 341,913 bp respectively. The assembled nuclear genome of upas tree (approximately 1.34 Gb) comprised two complete haplotypes, with a genome size of approximately 682.4 Mb and 661.4 Mb, respectively (Table 1). The assembly sizes are similar to the estimated haploid genome size of the upas tree. The contig N50 length of each haploid genome reached 60 Mb, and the number of gaps was less than 10 in both haploid genomes (Table 1).Table 1 Summary of the A. toxicaria genome assembly data.

Statistic	Haplotype A	Haplotype B	
Total size (bp)	682,498,812	661,432,422	
Number of gaps	4	7	
GC content (%)	35.64	35.54	
Characteristic	Contig	Scaffold	Contig	Scaffold	
Number	19	15	20	13	
Max. (bp)	114,031,235	114,031,235	99,273,849	103,092,922	
Mean (bp)	35,920,969	45,499,920	33,071,586	50,879,417	
Min. (bp)	161,699	161,699	291,585	18,536,726	
Median (bp)	22,828,143	27,970,581	20,607,115	36,734,386	
N10 (bp)	114,031,235	114,031,235	99,273,849	103,092,922	
N50 (bp)	60,259,137	89,883,092	79,464,572	87,196,807	
N90 (bp)	21,245,517	21,245,517	18,108,114	21,048,986	
L10	1	1	1	1	
L50	4	4	4	4	
L90	12	10	13	10	

Genome annotations

Repetitive sequences in the upas tree genome were identified and annotated using a combination of de novo and homology-based approaches. The software EDTA (v. 1.9.9; parameters:–sensitive 1–anno 1)37 was employed to identify the transposable elements (TE) and construct a TE library of the upas tree genome. Repetitive sequences in the upas tree genome were identified using RepeatMasker (version 4.0.7)38. Based on the combination of the two library files, a RepeatMasker search was performed to annotate repetitive sequences. In total, 2,205,812 repetitive sequences were identified in the upas tree genome, with a total length of 1,030,162,123 bp, accounting for 76.65% of the genome size. Among them, long terminal repeats (LTRs) possessed the highest proportion, with a total of 945,677 count numbers and a total length of 831,823,764 bp, accounting for 61.9% of the upas tree genome (Supplementary Table 3). Additionally, we conducted separate statistical analyses for the repeat sequences in the two haplotypes. The results indicate that the types and proportions of repeat sequences are largely consistent between the two haploid genomes (Supplementary Table 3). The presence of a significant proportion of repetitive sequences within the upas tree genome suggests that a chromosome-level genome assembly may potentially result in redundancy and an inflated size, when compared to our haplotype-resolved genome assembly.

The annotation of protein-coding genes in the upas tree genome was based on de novo annotations, transcriptome analyses, and homology blasts. A total of 290,667 non-redundant protein-encoding gene sequences were obtained from the publicly accessible genomes of Fragaria vesca39, Prunus persica40, Gillenia trifoliata41, Artocarpus nanchuanensis42, Morus notabilis43, Ficus microcarpa44, Hippophae rhamnoides45, Cannabis sativa46, Ziziphus jujuba47, Boehmeria nivea48, Vitis vinifera49, and Arabidopsis thaliana50, which were used as references for the homology-based gene annotation.

RNA-seq reads were assembled using de novo and reference-guided approaches. The de novo transcriptome assembly was performed using Trinity software (v. 2.0.6)51. Additionally, the RNA-seq reads were aligned to the reference genome using HISAT (v. 2.1.0)52 and assembled using StringTie (v.1.3.5)53. Inferred from transcriptome data, the gene structure annotation was performed using PASA (v. 2.4.1)54, and full-length genes were detected by comparing them with reference protein sequences using PASA. The AUGUSTUS parametric model (v. 3.4.0)55 was trained using the full-length gene set for five rounds of optimization (Supplementary Table 4).

Subsequently, the MAKER2 (v. 2.31.9)56 pipeline was used for annotation based on ab initio predictions, transcript annotations, and homologous protein evidence. Given the relatively low accuracy of the MAKER2 annotation process, further integration of the MAKER2 and PASA annotations was performed using EVidenceModeler (EVM, v. 1.1.1)57 to generate consistent gene annotations. To avoid introducing TE-coding regions, TEsorter (v. 1.4.1)58 was used to identify the TE protein domains in the genome, and EVM was used to shield them. In addition, PASA was used to optimize the EVM annotation by adding UTR and alternative splicing, and the gene annotations with abnormal coding frames (containing internal termination codons or ambiguous bases, lacking start or termination codons) and those that were too short (<50 aa) were removed. Additionally, tRNAs were annotated using tRNAScan-SE (v. 2.0.7)59, rRNAs were annotated using Barrnap (v. 0.9)60 with partial results filtered out, and various non-coding RNAs were aligned and annotated using RfamScan (v. 14.2)61.

In total, 43,500 coding genes were identified in the upas tree genome, with 21,817 and 21,622 being predicted in Haplotypes A and B, respectively, while the remainders are organelle genes, among which (Table 2). The coding genes in Haplotypes A and B encode 31,612 and 28,507 CDS, respectively, while those identified in the mitochondrial and chloroplast genome encode 24 and 37 CDS, respectively. In addition, a total of 4,222 non-coding RNAs were identified, comprising 1,765 rRNAs, 804 tRNAs, and 1,653 other non-coding RNAs (ncRNA, as referred to in Table 2). Among them, Haplotypes A and B contain a total of 819 and 802 ncRNAs, 940 and 825 rRNAs, as well as 381 and 385 tRNAs, respectively. Additionally, a total of 14 and 18 ncRNAs, as well as 13 and 25 tRNAs were identified in the mitochondrial and chloroplast genome, respectively (Table 2).Table 2 Gene annotation statistics.

Category	Total	Haplotype A	Haplotype B	Mitochondrial genome	Chloroplast genome	
All genes	47,722	23,957	23,634	51	80	
Coding genes	43,500	21,817	21,622	24	37	
ncRNA	1,653	819	802	14	18	
rRNA	1,765	940	825	0	0	
tRNA	804	381	385	13	25	
CDS	60,180	31,612	28,507	24	37	

The functional annotation of protein-coding genes was based on three strategies: (i) eggNOG-mapper (v. 2.0.1)62 annotation, which compares the gene with the eggNOG ortholog database to annotate its function, including Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG); (ii) Sequence similarity search, using DIAMOND (v. 2.0.4)63 to align protein sequences with protein databases (such as Swiss Prot, TrEMBL, and NR), to identify the best alignments with a similarity percentage greater than 30% and an E-value less than 1e-5; (iii) Domain similarity search, using InterProScan (v. 5.14–53.0)64 to align the sub-databases PRINTS, Pfam, SMART, PANTHER, CDD, and others in the InterPro database to obtain the conserved amino acid sequences, motifs, and domains of the predicted protein. The annotation results for 18 databases, including GO, EC, BRITE, were statistically analyzed. Overall, 43,008 genes were functionally annotated in at least one of these databases, accounting for 98.86% of the predicted protein-coding genes (Supplementary Table 5). The percentage of unannotated genes in Haplotypes A and B was only 1.19% and 1.07%, respectively, across all databases.

Data Records

The genome sequencing data (Illumina, PacBio HiFi, Hi-C), genome assembly, genome annotation, and RNA-Seq data are available in National Genomics Data Center (NGDC)65, Beijing Institute of Genomics, Chinese Academy of Sciences/China National Center for Bioinformation, under the BioProject ID PRJCA020380. The genome sequencing raw data and RNA-Seq data have been deposited in Genome Sequence Archive (GSA) of NGDC under the accession number CRA01296066 and CRA01295267. The genome assembly and annotation data have been deposited in Genome Warehouse (GWH) of NGDC under the accession number GWHEQBD0000000068. They can also be found in National Center for Biotechnology Information (NCBI) GenBank database under accession number GCA_035233585.169 and GCA_035234565.170. The genome assembly and detailed gene structure annotation files also can be accessed on FigShare71.

Technical Validation

Evaluation of the assembled genome and gene annotation

To assess the quality of the genome assembly, Illumina short-read sequencing data were mapped to the genome using BWA (v. 0.7.17-r1188)72, long reads were mapped to the genome using Minimap2 (v. 2.24-r1122)73, and RNA-Seq reads were mapped to the genome using HISAT2. Non-primary alignments were filtered, and the mapping ratio and coverage percentage were calculated. The sequencing data showed high coverage of the genome (Table 3). BUSCO (v. 5.3.0)74 was used to evaluate the assembled genome and identified 98.4% complete BUSCOs and 0.6% fragmented BUSCOs (Table 4), indicating that the genome assembly exhibited a good completeness. To assess gene annotation quality, BUSCO also was used to evaluate the proteins with integrated annotations. The proportion of complete core gene coverage was 98.3%, including 3.6% single-copy and 94.7% duplicated genes. Additionally, there were only 0.5% fragmented genes and 1.2% missing genes, indicating a high-quality annotation (Table 4).Table 3 Statistics of map rate and coverage of different types of sequencing data.

Data set	reads mapped	properly paired	bases mapped	> = 1×	> = 5×	> = 10×	> = 20×	
HiFi	99.57%	—	99.52%	99.99%	99.93%	99.42%	80.89%	
RNA-Seq	91.95%	88.27%	92.00%	10.52%	6.56%	5.23%	4.11%	
Illumina short-read sequencing	99.86%	98.01%	99.87%	99.96%	99.92%	99.87%	99.78%	

Table 4 Statistics of BUSCO evaluation of the genome and proteins.

Type	BUSCO groups	
Haplotype A	Haplotype B	Proteins (HapA and B)	
Complete BUSCOs (C)	1,589 (98.4%)	1,587 (98.4%)	1,586 (98.3%)	
Complete and single-copy BUSCOs (S)	1,576 (97.6%)	1,570 (97.3%)	58 (3.6%)	
Complete and duplicated BUSCOs (D)	13 (0.8%)	17 (1.1%)	1,528 (94.7%)	
Fragmented BUSCOs (F)	9 (0.6%)	10 (0.6%)	8 (0.5%)	
Missing BUSCOs (M)	16 (1.0%)	17 (1.0%)	40 (1.2%)	
Total BUSCO groups searched	1,614	1,614	1,614	

We also evaluated the genome assembly quality using the Merqury (v. v1.3)75 software, employing second-generation DNA sequencing data for the assessment. The results indicate that both haplotypes exhibit a QV (Quality Value) greater than 54 and a completeness rate exceeding 86%. The combined haplotype shows a completeness rate of over 98% (Supplementary Table 6).

To detect redundancies in the genome assembly, the sequencing data were re-aligned to the assembled genome using BWA (filtering out any non-primary alignments to ensure that each read aligned at most once), and the coverage depth at all sites was examined. Based on the core genes obtained from BUSCO, the analysis indicated that the coverage depth of single-copy and multi-copy core genes matched the Poisson distribution, and there were no significant heterozygous peaks in the BUSCO core single-copy and multiple-copy gene regions (Supplementary Fig. 3). These results indicate that the assembled genome had no redundancy. Mapping the Hi-C data to the final genome assembly using Juicer showed that the chromosome clustering effect was acceptable, with no evident chromosome assembly errors (Fig. 3a). The whole-genome alignment revealed good collinearity between the two haplotypes (Fig. 3b). Based on the Illumina and HiFi data, the coverage depth analysis detected no evident guanine-cytosine (GC) bias at different GC contents (Supplementary Fig. 4).Fig. 3 Hi-C interaction heatmap (a) and syntenic plot (b) between the two haplotypes of the upas tree (A. toxicaria).

To identify the distribution of characteristic sequences on the chromosomes, such as telomere sequences, rDNA, and tandem repeats, the sequences were aligned to the genome. Most of the chromosomes (21 out of 26) contained telomere sequences (TTTAGGG) at both ends (Fig. 4a). Most of the chromosomes also contained a high number of tandem repeats, likely in the centromere (Fig. 4b). In addition, the 18-5.8-28S rDNA arrays were located only on chromosome 3 (Fig. 4c), whereas the 5S rDNA arrays were identified on chromosomes 1, 6, 8, and 11 (Fig. 4d).Fig. 4 The distribution of telomere sequences TTTAGGG (a), high tandem repeat likely representing centromere (b), 18-5.8-28S rDNA (c), and 5S rDNA (d) on the chromosomes of upas tree (A. toxicaria).

The chromosome-level sequences of the two haploid genomes were aligned using Minimap2, and SyRI (v. 1.6)76 was used to identify synteny and structural rearrangements between the homologous chromosomes of the two haploid genomes. In total, 2,236 syntenic regions (approximately total 897 Mb) were detected, indicating good collinearity between the two haplotypes. Notably, consistent with the estimated heterozygosity of the upas tree genome, many sequence and structural variations were discovered between the homologous chromosomes of the two haploid genomes, including 2,651,915 SNPs, 182,004 insertions, 183,169 deletions, 1,201 translocations, and 65 inversions. At the chromosomal level, 11 inversions of more than 5 Mb were identified in chromosomes 1, 2, 3, and 4. In addition, 1,612 and 3,188 duplications were detected in haplotypes A and B, respectively (Fig. 5, Supplementary Table 7). Additionally, we utilized SNPs to calculate the heterozygosity of two sets of haplotypes, employing a window size of 1000 bp and a step size of 500 bp to count the number of SNPs. The results indicate an average of 3.85 SNPs per 1000 bases in the two haplotype genomes (Supplementary Table 8).Fig. 5 Structural variation between the two haploid genomes of the upas tree (A. toxicaria).

Supplementary information

Supplementary Figures 1–4

Supplementary Tables 1–8

Supplementary information

The online version contains supplementary material available at 10.1038/s41597-024-03860-6.

Acknowledgements

This study was supported by National Natural Science Foundation of China (31872673, 32370395), and Yunnan Revitalization Talent Support Program “Top Team” Project (202305AT35). The authors are grateful to Dr. Fu Gao for his help in plant sample collection.

Author contributions

Y.J. conceived the framework of this study. K.M., Y.W., L.H., Y.L. and L.T. collected and analyzed the data. K.M., Y.W. and Y.J. wrote the draft manuscript. H.L. and Y.J. discussed the results and revised the manuscript. All authors contributed to the article and approved the submitted version.

Code availability

The manuals and protocols of the published bioinformatics tools were followed for the execution of all the software and pipelines. The Methods section provides a description of the software version and code/parameters used. No custom code was utilized in this study for the curation and validation of the dataset.

Competing interests

The authors declare no competing interests.

Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

These authors contributed equally: Ke Miao, Ya Wang.
==== Refs
References

1. Wu, Z., Raven, P. H. & Hong, D. In Flora of China. vol. 5: Ulmaceae through Basellaceae (eds Wu, Z. Y., Raven, P. H. & Hong, D. Y.) (Science Press & Missouri Botanical Garden Press, 2003).
2. Mante PK Adongo DW Woode E Kukuia KKE Ameyaw EO Anticonvulsant Effect of Antiaris toxicaria (Pers.) Lesch. (Moraceae) Aqueous Extract in Rodents ISRN Pharmacology 2013 2013 1 9 10.1155/2013/519208
Mante, P. K., Adongo, D. W., Woode, E., Kukuia, K. K. E. & Ameyaw, E. O. Anticonvulsant Effect of Antiaris toxicaria (Pers.) Lesch. (Moraceae) Aqueous Extract in Rodents. ISRN Pharmacology 2013, 1–9 (2013).
3. Bisset NG Arrow poisons in China. Part 1 J Ethnopharmacol 1979 1 325 384 10.1016/S0378-8741(79)80002-1 397373
Bisset, N. G. Arrow poisons in China. Part 1. J Ethnopharmacol 1, 325–384 (1979).397373
4. Kopp B Bauer WP Bernkop Schmürch A Analysis of some Malaysian dart poisons J Ethnopharmacol 1992 36 57 62 10.1016/0378-8741(92)90061-U 1501494
Kopp, B., Bauer, W. P. & Bernkop Schmürch, A. Analysis of some Malaysian dart poisons. J Ethnopharmacol 36, 57–62 (1992).1501494
5. Li XS Three new compounds from the bark of Antiaris toxicaria Phytochem Lett 2015 13 182 186 10.1016/j.phytol.2015.06.006
Li, X. S. et al. Three new compounds from the bark of Antiaris toxicaria. Phytochem Lett 13, 182–186 (2015).
6. Jiang MM Phenylpropanoid and lignan derivatives from Antiaris toxicaria and their effects on proliferation and differentiation of an osteoblast-like cell line Planta Med 2009 75 340 345 10.1055/s-0028-1112212 19184966
Jiang, M. M. et al. Phenylpropanoid and lignan derivatives from Antiaris toxicaria and their effects on proliferation and differentiation of an osteoblast-like cell line. Planta Med 75, 340–345 (2009).19184966
7. Carter CA Toxicarioside A. A new cardenolide isolated from Antiaris toxicaria latex-derived dart poison. Assignment of the 1H- and 13C-NMR shifts for an antiarigenin aglycone Tetrahedron 1997 53 13557 13566 10.1016/S0040-4020(97)00895-8
Carter, C. A. et al. Toxicarioside A. A new cardenolide isolated from Antiaris toxicaria latex-derived dart poison. Assignment of the 1H- and 13C-NMR shifts for an antiarigenin aglycone. Tetrahedron 53, 13557–13566 (1997).
8. Gan YJ Mei WL Zhao YX Dai HF A new cytotoxic cardenolide from the latex of Antiaris toxicaria Chinese Chem Lett 2009 20 450 452 10.1016/j.cclet.2008.12.043
Gan, Y. J., Mei, W. L., Zhao, Y. X. & Dai, H. F. A new cytotoxic cardenolide from the latex of Antiaris toxicaria. Chinese Chem Lett 20, 450–452 (2009).
9. Li XS Cardiac glycosides from the bark of Antiaris toxicaria Fitoterapia 2014 97 71 77 10.1016/j.fitote.2014.05.013 24879902
Li, X. S. et al. Cardiac glycosides from the bark of Antiaris toxicaria. Fitoterapia 97, 71–77 (2014).24879902
10. Hano Y Mitsui P Nomura T Kawai T Yoshida Y Two new dihydrochalcone derivatives, antiarones J and K, from the root bark of Antiaris toxicaria J Nat Prod 1991 54 1049 1055 10.1021/np50076a020
Hano, Y., Mitsui, P., Nomura, T., Kawai, T. & Yoshida, Y. Two new dihydrochalcone derivatives, antiarones J and K, from the root bark of Antiaris toxicaria. J Nat Prod 54, 1049–1055 (1991).
11. Shi LS Cytotoxic cardiac glycosides and coumarins from Antiaris toxicaria Bioorgan Med Chem 2014 22 1889 1898 10.1016/j.bmc.2014.01.052
Shi, L. S. et al. Cytotoxic cardiac glycosides and coumarins from Antiaris toxicaria. Bioorgan Med Chem 22, 1889–1898 (2014).
12. Wink, M. Biochemistry of Plant Secondary Metabolism. 10.1002/9781444320503 (Wiley-Blackwell, Oxford, UK, 2010).
13. Kreis W Hensel A Stuhlemmer U Cardenolide Biosynthesis in Foxglove 1 Planta Med 1998 64 491 499 10.1055/s-2006-957500
Kreis, W., Hensel, A. & Stuhlemmer, U. Cardenolide Biosynthesis in Foxglove 1. Planta Med 64, 491–499 (1998).
14. Kunert M Promiscuous CYP87A enzyme activity initiates cardenolide biosynthesis in plants Nat Plants 2023 9 1607 1617 10.1038/s41477-023-01515-9 37723202
Kunert, M. et al. Promiscuous CYP87A enzyme activity initiates cardenolide biosynthesis in plants. Nat Plants 9, 1607–1617 (2023).37723202
15. Agrawal AA Petschenka G Bingham RA Weber MG Rasmann S Toxic cardenolides: chemical ecology and coevolution of specialized plant–herbivore interactions New Phytol 2012 194 28 45 10.1111/j.1469-8137.2011.04049.x 22292897
Agrawal, A. A., Petschenka, G., Bingham, R. A., Weber, M. G. & Rasmann, S. Toxic cardenolides: chemical ecology and coevolution of specialized plant–herbivore interactions. New Phytol 194, 28–45 (2012).22292897
16. Carter CA Toxicarioside B and toxicarioside C. New cardenolides isolated from Antiaris toxicaria latex-derived dart poison Tetrahedron 1997 53 16959 16968 10.1016/S0040-4020(97)10174-0
Carter, C. A. et al. Toxicarioside B and toxicarioside C. New cardenolides isolated from Antiaris toxicaria latex-derived dart poison. Tetrahedron 53, 16959–16968 (1997).
17. Dai HF Two new cytotoxic cardenolides from the latex of Antiaris toxicaria J Asian Nat Prod Res 2009 11 832 837 10.1080/10286020903164285 20183332
Dai, H. F. et al. Two new cytotoxic cardenolides from the latex of Antiaris toxicaria. J Asian Nat Prod Res 11, 832–837 (2009).20183332
18. Zuo WJ Two new strophanthidol cardenolides from the seeds of Antiaris toxicaria Phytochem Lett 2013 6 1 4 10.1016/j.phytol.2012.10.001
Zuo, W. J. et al. Two new strophanthidol cardenolides from the seeds of Antiaris toxicaria. Phytochem Lett 6, 1–4 (2013).
19. Wu XL A new periplogenin cardenolide from the seeds of Antiaris toxicaria J Asian Nat Prod Res 2014 16 418 421 10.1080/10286020.2014.885506 24597720
Wu, X. L. et al. A new periplogenin cardenolide from the seeds of Antiaris toxicaria. J Asian Nat Prod Res 16, 418–421 (2014).24597720
20. Jiang MM Cardenolides from Antiaris toxicariaas Potent Selective Nur77 Modulators Chem Pharm Bull 2008 56 1005 1008 10.1248/cpb.56.1005
Jiang, M. M. et al. Cardenolides from Antiaris toxicariaas Potent Selective Nur77 Modulators. Chem Pharm Bull 56, 1005–1008 (2008).
21. Levrier C Kiremire B Guéritte F Litaudon M Toxicarioside M, a new cytotoxic 10β-hydroxy-19-nor-cardenolide from Antiaris toxicaria Fitoterapia 2012 83 660 664 10.1016/j.fitote.2012.02.001 22348979
Levrier, C., Kiremire, B., Guéritte, F. & Litaudon, M. Toxicarioside M, a new cytotoxic 10β-hydroxy-19-nor-cardenolide from Antiaris toxicaria. Fitoterapia 83, 660–664 (2012).22348979
22. El-Seedi HR Cardenolides: Insights from chemical structure and pharmacological utility Pharmacol Res 2019 141 123 175 10.1016/j.phrs.2018.12.015 30579976
El-Seedi, H. R. et al. Cardenolides: Insights from chemical structure and pharmacological utility. Pharmacol Res 141, 123–175 (2019).30579976
23. Li YN Toxicarioside A, isolated from tropical Antiaris toxicaria, blocks endoglin/TGF-β signaling in a bone marrow stromal cell line Asian Pac J Trop Med 2012 5 91 97 10.1016/S1995-7645(12)60002-9 22221748
Li, Y. N. et al. Toxicarioside A, isolated from tropical Antiaris toxicaria, blocks endoglin/TGF-β signaling in a bone marrow stromal cell line. Asian Pac J Trop Med 5, 91–97 (2012).22221748
24. Doyle, J. A rapid DNA isolation procedure for small quantities of fresh leaf tissue. Phytochem Bull 19 (1987).
25. Chen S Zhou Y Chen Y Gu J fastp: an ultra-fast all-in-one FASTQ preprocessor Bioinformatics 2018 34 i884 i890 10.1093/bioinformatics/bty560 30423086
Chen, S., Zhou, Y., Chen, Y. & Gu, J. fastp: an ultra-fast all-in-one FASTQ preprocessor. Bioinformatics 34, i884–i890 (2018).30423086
26. Wenger AM Accurate circular consensus long-read sequencing improves variant detection and assembly of a human genome Nat Biotechnol 2019 37 1155 1162 10.1038/s41587-019-0217-9 31406327
Wenger, A. M. et al. Accurate circular consensus long-read sequencing improves variant detection and assembly of a human genome. Nat Biotechnol 37, 1155–1162 (2019).31406327
27. Liu, B. et al. Estimation of genomic characteristics by analyzing k-mer frequency in de novo genome projects. arXiv: Genomics (2013).
28. Sun H Ding J Piednoël M Schneeberger K findGSE: estimating genome size variation within human and Arabidopsis using k -mer frequencies Bioinformatics 2018 34 550 557 10.1093/bioinformatics/btx637 29444236
Sun, H., Ding, J., Piednoël, M. & Schneeberger, K. findGSE: estimating genome size variation within human and Arabidopsis using k -mer frequencies. Bioinformatics 34, 550–557 (2018).29444236
29. Cheng HY Concepcion GT Feng XW Zhang HW Li H Haplotype-resolved de novo assembly using phased assembly graphs with hifiasm Nat Methods 2021 18 170 175 10.1038/s41592-020-01056-5 33526886
Cheng, H. Y., Concepcion, G. T., Feng, X. W., Zhang, H. W. & Li, H. Haplotype-resolved de novo assembly using phased assembly graphs with hifiasm. Nat Methods 18, 170–175 (2021).33526886
30. Durand NC Juicer provides a one-click system for analyzing loop-resolution Hi-C experiments Cell Syst 2016 3 95 98 10.1016/j.cels.2016.07.002 27467249
Durand, N. C. et al. Juicer provides a one-click system for analyzing loop-resolution Hi-C experiments. Cell Syst 3, 95–98 (2016).27467249
31. Dudchenko O De novo assembly of the Aedes aegypti genome using Hi-C yields chromosome-length scaffolds Science 2017 356 92 95 10.1126/science.aal3327 28336562
Dudchenko, O. et al. De novo assembly of the Aedes aegypti genome using Hi-C yields chromosome-length scaffolds. Science 356, 92–95 (2017).28336562
32. Durand NC Juicebox provides a visualization system for Hi-C contact maps with unlimited zoom Cell Syst 2016 3 99 101 10.1016/j.cels.2015.07.012 27467250
Durand, N. C. et al. Juicebox provides a visualization system for Hi-C contact maps with unlimited zoom. Cell Syst 3, 99–101 (2016).27467250
33. Xu, G. C. et al. LR_Gapcloser: a tiling path-based gap closer that uses long reads to complete genome assembly. Gigascience 8 (2019).
34. Jin JJ GetOrganelle: a fast and versatile toolkit for accurate de novo assembly of organelle genomes Genome Biol 2020 21 241 10.1186/s13059-020-02154-5 32912315
Jin, J. J. et al. GetOrganelle: a fast and versatile toolkit for accurate de novo assembly of organelle genomes. Genome Biol 21, 241 (2020).32912315
35. Hu J Fan J Sun Z Liu S NextPolish: a fast and efficient genome polishing tool for long-read assembly Bioinformatics 2020 36 2253 2255 10.1093/bioinformatics/btz891 31778144
Hu, J., Fan, J., Sun, Z. & Liu, S. NextPolish: a fast and efficient genome polishing tool for long-read assembly. Bioinformatics 36, 2253–2255 (2020).31778144
36. Pryszcz LP Gabaldón T Redundans: an assembly pipeline for highly heterozygous genomes Nucleic Acids Res 2016 44 e113 e113 10.1093/nar/gkw294 27131372
Pryszcz, L. P. & Gabaldón, T. Redundans: an assembly pipeline for highly heterozygous genomes. Nucleic Acids Res 44, e113–e113 (2016).27131372
37. Ou, S. et al. Benchmarking Transposable Element Annotation Methods for Creation of a Streamlined, Comprehensive Pipeline. Genome Biol, 20, 275 (2019).
38. Tempel, S. Using and understanding RepeatMasker. in Mobile Genetic Elements (ed. Bigot, Y.) vol. 859 29–51 (Humana Press, Totowa, NJ, 2012).
39. Li Y Pi M Gao Q Liu Z Kang C Updated annotation of the wild strawberry Fragaria vesca V4 genome Hortic Res 2019 6 61 10.1038/s41438-019-0142-6 31069085
Li, Y., Pi, M., Gao, Q., Liu, Z. & Kang, C. Updated annotation of the wild strawberry Fragaria vesca V4 genome. Hortic Res 6, 61 (2019).31069085
40. The International Peach Genome Initiative The high-quality draft genome of peach (Prunus persica) identifies unique patterns of genetic diversity, domestication and genome evolution Nat Genet 2013 45 487 494 10.1038/ng.2586 23525075
The International Peach Genome Initiative. et al. The high-quality draft genome of peach (Prunus persica) identifies unique patterns of genetic diversity, domestication and genome evolution. Nat Genet 45, 487–494 (2013).23525075
41. Ireland HS The Gillenia trifoliata genome reveals dynamics correlated with growth and reproduction in Rosaceae Hortic Res 2021 8 233 10.1038/s41438-021-00662-4 34719690
Ireland, H. S. et al. The Gillenia trifoliata genome reveals dynamics correlated with growth and reproduction in Rosaceae. Hortic Res 8, 233 (2021).34719690
42. He J A chromosome-level genome assembly of Artocarpus nanchuanensis (Moraceae), an extremely endangered fruit tree Gigascience 2022 11 giac042 10.1093/gigascience/giac042 35701376
He, J. et al. A chromosome-level genome assembly of Artocarpus nanchuanensis (Moraceae), an extremely endangered fruit tree. Gigascience 11, giac042 (2022).35701376
43. Xia Z Chromosome-Level genomes reveal the genetic basis of descending dysploidy and sex determination in Morus Plants Genomics, Proteomics & Bioinformatics 2022 20 1119 1137 10.1016/j.gpb.2022.08.005
Xia, Z. et al. Chromosome-Level genomes reveal the genetic basis of descending dysploidy and sex determination in Morus Plants. Genomics, Proteomics & Bioinformatics 20, 1119–1137 (2022).
44. Zhang X Genomes of the banyan tree and pollinator aasp provide Insights into fig-wasp coevolution Cell 2020 183 875 889.e17 10.1016/j.cell.2020.09.043 33035453
Zhang, X. et al. Genomes of the banyan tree and pollinator aasp provide Insights into fig-wasp coevolution. Cell 183, 875–889.e17 (2020).33035453
45. Wu Z Genome of Hippophae rhamnoides provides insights into a conserved molecular mechanism in actinorhizal and rhizobial symbioses New Phytol 2022 235 276 291 10.1111/nph.18017 35118662
Wu, Z. et al. Genome of Hippophae rhamnoides provides insights into a conserved molecular mechanism in actinorhizal and rhizobial symbioses. New Phytol 235, 276–291 (2022).35118662
46. Gao S A high-quality reference genome of wild Cannabis sativa Hortic Res 2020 7 73 10.1038/s41438-020-0295-3 32377363
Gao, S. et al. A high-quality reference genome of wild Cannabis sativa. Hortic Res 7, 73 (2020).32377363
47. Shen LY Chromosome-scale genome assembly for chinese sour jujube and insights Into its genome evolution and domestication signature Front Plant Sci 2021 12 773090 10.3389/fpls.2021.773090 34899800
Shen, L. Y. et al. Chromosome-scale genome assembly for chinese sour jujube and insights Into its genome evolution and domestication signature. Front Plant Sci 12, 773090 (2021).34899800
48. Wang Y Genomic analyses provide comprehensive insights into the domestication of bast fiber crop ramie (Boehmeria nivea) Plant J 2021 107 787 800 10.1111/tpj.15346 33993558
Wang, Y. et al. Genomic analyses provide comprehensive insights into the domestication of bast fiber crop ramie (Boehmeria nivea). Plant J 107, 787–800 (2021).33993558
49. The French–Italian Public Consortium for Grapevine Genome Characterization The grapevine genome sequence suggests ancestral hexaploidization in major angiosperm phyla Nature 2007 449 463 467 10.1038/nature06148 17721507
The French–Italian Public Consortium for Grapevine Genome Characterization. The grapevine genome sequence suggests ancestral hexaploidization in major angiosperm phyla. Nature 449, 463–467 (2007).17721507
50. Wang B High-Quality Arabidopsis Thaliana Genome Assembly with Nanopore and HiFi Long Reads Genomics, Proteomics & Bioinformatics 2022 20 4 13 10.1016/j.gpb.2021.08.003
Wang, B. et al. High-Quality Arabidopsis Thaliana Genome Assembly with Nanopore and HiFi Long Reads. Genomics, Proteomics & Bioinformatics 20, 4–13 (2022).
51. Grabherr MG Full-length transcriptome assembly from RNA-Seq data without a reference genome Nat Biotechnol 2011 29 644 652 10.1038/nbt.1883 21572440
Grabherr, M. G. et al. Full-length transcriptome assembly from RNA-Seq data without a reference genome. Nat Biotechnol 29, 644–652 (2011).21572440
52. Kim D Langmead B Salzberg SL HISAT: a fast spliced aligner with low memory requirements Nat Methods 2015 12 357 360 10.1038/nmeth.3317 25751142
Kim, D., Langmead, B. & Salzberg, S. L. HISAT: a fast spliced aligner with low memory requirements. Nat Methods 12, 357–360 (2015).25751142
53. Pertea M StringTie enables improved reconstruction of a transcriptome from RNA-seq reads Nat Biotechnol 2015 33 290 295 10.1038/nbt.3122 25690850
Pertea, M. et al. StringTie enables improved reconstruction of a transcriptome from RNA-seq reads. Nat Biotechnol 33, 290–295 (2015).25690850
54. Haas BJ Improving the Arabidopsis genome annotation using maximal transcript alignment assemblies Nucleic Acids Research 2003 31 5654 5666 10.1093/nar/gkg770 14500829
Haas, B. J. Improving the Arabidopsis genome annotation using maximal transcript alignment assemblies. Nucleic Acids Research 31, 5654–5666 (2003).14500829
55. Stanke M Diekhans M Baertsch R Haussler D Using native and syntenically mapped cDNA alignments to improve de novo gene finding Bioinformatics 2008 24 637 644 10.1093/bioinformatics/btn013 18218656
Stanke, M., Diekhans, M., Baertsch, R. & Haussler, D. Using native and syntenically mapped cDNA alignments to improve de novo gene finding. Bioinformatics 24, 637–644 (2008).18218656
56. Cantarel BL MAKER: An easy-to-use annotation pipeline designed for emerging model organism genomes Genome Res 2008 18 188 196 10.1101/gr.6743907 18025269
Cantarel, B. L. et al. MAKER: An easy-to-use annotation pipeline designed for emerging model organism genomes. Genome Res 18, 188–196 (2008).18025269
57. Haas BJ Automated eukaryotic gene structure annotation using EVidenceModeler and the Program to Assemble Spliced Alignments Genome Biol 2008 9 R7 10.1186/gb-2008-9-1-r7 18190707
Haas, B. J. et al. Automated eukaryotic gene structure annotation using EVidenceModeler and the Program to Assemble Spliced Alignments. Genome Biol 9, R7 (2008).18190707
58. Zhang RG TEsorter: An accurate and fast method to classify LTR-retrotransposons in plant genomes Hortic Res 2022 9 uhac017 10.1093/hr/uhac017 35184178
Zhang, R. G. et al. TEsorter: An accurate and fast method to classify LTR-retrotransposons in plant genomes. Hortic Res 9, uhac017 (2022).35184178
59. Chan PP Lin BY Mak AJ Lowe TM tRNAscan-SE 2.0: improved detection and functional classification of transfer RNA genes Nucleic Acids Res 2021 49 9077 9096 10.1093/nar/gkab688 34417604
Chan, P. P., Lin, B. Y., Mak, A. J. & Lowe, T. M. tRNAscan-SE 2.0: improved detection and functional classification of transfer RNA genes. Nucleic Acids Res 49, 9077–9096 (2021).34417604
60. Seemann, T. BAsic Rapid Ribosomal RNA Predictor. https://github.com/tseemann/barrnap (2024).
61. Nawrocki EP Rfam 12.0: updates to the RNA families database Nucleic Acids Res 2015 43 D130 D137 10.1093/nar/gku1063 25392425
Nawrocki, E. P. et al. Rfam 12.0: updates to the RNA families database. Nucleic Acids Res 43, D130–D137 (2015).25392425
62. Huerta-Cepas J Fast genome-wide functional annotation through orthology assignment by eggNOG-Mapper Mol Biol Evol 2017 34 2115 2122 10.1093/molbev/msx148 28460117
Huerta-Cepas, J. et al. Fast genome-wide functional annotation through orthology assignment by eggNOG-Mapper. Mol Biol Evol 34, 2115–2122 (2017).28460117
63. Buchfink B Xie C Huson DH Fast and sensitive protein alignment using DIAMOND Nat Methods 2015 12 59 60 10.1038/nmeth.3176 25402007
Buchfink, B., Xie, C. & Huson, D. H. Fast and sensitive protein alignment using DIAMOND. Nat Methods 12, 59–60 (2015).25402007
64. Jones P InterProScan 5: genome-scale protein function classification Bioinformatics 2014 30 1236 1240 10.1093/bioinformatics/btu031 24451626
Jones, P. et al. InterProScan 5: genome-scale protein function classification. Bioinformatics 30, 1236–1240 (2014).24451626
65. CNCB-NGDC Members and Partners Database resources of the national genomics data center, China National Center for Bioinformation in 2022 Nucleic Acids Res 2022 50 D27 D38 10.1093/nar/gkab951 34718731
CNCB-NGDC Members and Partners. et al. Database resources of the national genomics data center, China National Center for Bioinformation in 2022. Nucleic Acids Res 50, D27–D38 (2022).34718731
66. 2023 NGDC Genome Sequence Archive https://ngdc.cncb.ac.cn/gsa/browse/CRA012960
NGDC Genome Sequence Archive https://ngdc.cncb.ac.cn/gsa/browse/CRA012960 (2023).
67. 2023 NGDC Genome Sequence Archive https://ngdc.cncb.ac.cn/gsa/browse/CRA012952
NGDC Genome Sequence Archive https://ngdc.cncb.ac.cn/gsa/browse/CRA012952 (2023).
68. 2023 NGDC Genome Warehouse https://ngdc.cncb.ac.cn/gwh/Assembly/82938/show
NGDC Genome Warehouse https://ngdc.cncb.ac.cn/gwh/Assembly/82938/show (2023).
69. 2024 NCBI Assembly GCA_035233585.1
NCBI Assembly https://identifiers.org/insdc.gca:GCA_035233585.1 (2024).
70. 2024 NCBI Assembly GCA_035234565.1
NCBI Assembly https://identifiers.org/insdc.gca:GCA_035234565.1 (2024).
71. Wang Y 2024 Genome assembly and annotation files of Antiaris toxicaria FigShare Dataset. 10.6084/m9.figshare.26315620.v1
Wang, Y. Genome assembly and annotation files of Antiaris toxicaria. FigShare Dataset.10.6084/m9.figshare.26315620.v1 (2024).
72. Li, H. Aligning sequence reads, clone sequences and assembly contigs with BWA-MEM. ArXiv 1303 (2013).
73. Li H Minimap2: pairwise alignment for nucleotide sequences Bioinformatics 2018 34 3094 3100 10.1093/bioinformatics/bty191 29750242
Li, H. Minimap2: pairwise alignment for nucleotide sequences. Bioinformatics 34, 3094–3100 (2018).29750242
74. Simão FA Waterhouse RM Ioannidis P Kriventseva EV Zdobnov EM BUSCO: assessing genome assembly and annotation completeness with single-copy orthologs Bioinformatics 2015 31 3210 3212 10.1093/bioinformatics/btv351 26059717
Simão, F. A., Waterhouse, R. M., Ioannidis, P., Kriventseva, E. V. & Zdobnov, E. M. BUSCO: assessing genome assembly and annotation completeness with single-copy orthologs. Bioinformatics 31, 3210–3212 (2015).26059717
75. Rhie A Walenz BP Koren S Phillippy AM Merqury: reference-free quality, completeness, and phasing assessment for genome assemblies Genome Biol 2020 21 245 10.1186/s13059-020-02134-9 32928274
Rhie, A., Walenz, B. P., Koren, S. & Phillippy, A. M. Merqury: reference-free quality, completeness, and phasing assessment for genome assemblies. Genome Biol 21, 245 (2020).32928274
76. Goel M Sun H Jiao WB Schneeberger K SyRI: finding genomic rearrangements and local sequence differences from whole-genome assemblies Genome Biol 2019 20 277 10.1186/s13059-019-1911-0 31842948
Goel, M., Sun, H., Jiao, W. B. & Schneeberger, K. SyRI: finding genomic rearrangements and local sequence differences from whole-genome assemblies. Genome Biol 20, 277 (2019).31842948
