
==== Front
Sci Data
Sci Data
Scientific Data
2052-4463
Nature Publishing Group UK London

3846
10.1038/s41597-024-03846-4
Data Descriptor
A chromosome-level genome assembly of the Brontispa longissima
Li Zaiyuan 12
Ma Guangchang 1
Tang Chao 1
Wen Haibo 1
Liu Conghui 2
http://orcid.org/0000-0002-7840-9450
Liu Bo 2
Qiao Xi 2
Jin Tao 1
Qian Wanqiang 2
Wan Fanghao wanfanghao@caas.cn

2
Peng Zhengqiang lypzhq@163.com

1
Gong Zhi zhigong11@163.com

1
1 grid.453499.6 0000 0000 9835 1415 Environment and Plant Protection Institute, Chinese Academy of Tropical Agricultural Sciences, Haikou, 571101 China
2 grid.410727.7 0000 0001 0526 1937 Shenzhen Branch, Guangdong Laboratory for Lingnan Modern Agriculture, Genome Analysis Laboratory of the Ministry of Agriculture and Rural Affairs, Agricultural Genomics Institute at Shenzhen, Chinese Academy of Agricultural Sciences, Shenzhen, 518120 China
14 9 2024
14 9 2024
2024
11 100224 5 2024
2 9 2024
© The Author(s) 2024
2024
https://creativecommons.org/licenses/by-nc-nd/4.0/ Open Access This article is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License, which permits any non-commercial use, sharing, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if you modified the licensed material. You do not have permission under this licence to share adapted material derived from this article or parts of it. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by-nc-nd/4.0/.
Brontispa longissima is a highly destructive pest that affects coconut and ornamental palm plants. It is widely distributed across Southeast and East Asia and the Pacific region, causing production losses of up to 50–70%. While control methods and ecological phenomena have been the primary focus of research, there is a significant lack of studies on the molecular mechanisms underlying these ecological phenomena. The absence of a reference genome has also hindered the development of new molecular-targeted control technologies. In this study, we conducted a karyotype analysis of B. longissima and assembled the first high-quality chromosome-level genome. The assembled genome is 582.24 Mb in size, with a scaffold N50 size of 63.81 Mb, consisting of 10 chromosomes and a GC content of 33.71%. The BUSCO assessment indicated a completeness estimate of 98.1%. A total of 23,051 protein-coding genes were predicted. Our study provides a valuable genomic resource for understanding the mechanisms of adaptive evolution and facilitates the development of new molecular-targeted control methods for B. longissima.

Subject terms

Zoology
Entomology
the National Key R&D Program of China (2021YFC2600400 & 2021YFC2600402); the Central Public-interest Scientific Institution Basal Research Fund for Chinese Academy of Tropical Agricultural Sciences (No. 1630042019029); the Hainan Province Science and Technology Special Fund (ZDKJ2021016); the Shenzhen Science and Technology Program (KQTD20180411143628272); and the agricultural science and technology innovation program.issue-copyright-statement© Springer Nature Limited 2024
==== Body
pmcBackground & Summary

Brontispa longissima Gestro, a destructive pest of coconut and ornamental palm plants, originates from Indonesia and New Guinea and has spread extensively across Southeast and East Asia and the Pacific region1. The potential for its range to expand further poses a significant risk, threatening coconut production, which is vital for many developing countries2,3. In invaded areas, B. longissima can rapidly reach high incidence rates, causing severe agricultural and economic damage, with production losses reaching as high as 50–70%4. The larvae and adults feed on the soft leaf tissues of coconut palms, resulting in brown leaves, reduced photosynthetic ability, stunted growth, and even death of the palms4. Given its destructive impact, effective management strategies are crucial to mitigate the threat posed by B. longissima.

Currently, the primary method for urgently controlling B. longissima in most countries is through the use of chemical insecticides. However, the effectiveness of pesticide sprays is limited as B. longissima spends most of its stages hidden within unopened buds of coconut trees. Furthermore, populations collected in Southeast Asia have demonstrated varying levels of resistance to certain insecticides, such as β-cypermethrin and avermectin5. Despite this, the specific molecular mechanisms underlying insecticide resistance in B. longissima remain unknown. Additionally, the feeding experience of B. longissima adults influences their subsequent host preferences, facilitating the beetle’s successful establishment in newly invaded habitats where the original host plant is scarce6. Furthermore, chemosensory gene families crucial for insect olfaction were identified based on antennal and abdominal transcriptomes of males and females using RNA-Seq7. However, RNA-Seq only provides limited information on the expression of chemosensory genes in specific tissues and time frames. The sequencing of the B. longissima genome will enable a more comprehensive identification of genes involved in key biological processes such as digestion, insecticide detoxification, and chemosensory perception. This comprehensive genomic dataset will also facilitate the development of targeted pest control methods, including RNA interference (RNAi) and gene editing techniques.

In this study, we utilized PacBio long-read sequencing and Hi-C sequencing technologies to construct the first high-quality chromosome-level reference genome of B. longissima. The final genome size was 582.24 Mb, organized into 10 chromosomes with N50 sizes of 63.81 Mb. Within this genome, we predicted a total of 23,051 protein-coding genes. This high-quality genome serves as a crucial genetic resource for investigating the molecular mechanisms underlying ecological phenomena in B. longissima, such as insecticide resistance and host selection.

Methods

Insect rare and sample collection

The adult B. longissima (Fig. 1A) specimens were collected from a coconut plantation (Fig. 1B) in Haikou, Hainan Province, in 2003. These insects were reared in the laboratory under controlled conditions: 26 ± 1 °C temperature, 14:10 (L:D) photoperiod, and 80 ± 5% relative humidity. To improve the accuracy and continuity of genome assembly, an inbred B. longissima laboratory population was obtained through sibling mating. The surface-sterilized male adults from this inbred laboratory population were used for Illumina, PacBio, and High-throughput chromosome conformation capture (Hi-C) sequencing, with sample sizes of 10, 10, and 15, respectively. Additionally, eggs, larvae at the 1st, 4th, and 5th instar stages, pupae, and male adults were collected from the laboratory-reared population for transcriptome sequencing. The experiment required 80 eggs, 50 1st instar larvae, 15 individuals each of 4th and 5th instar larvae and pupae, and 10 adults at different stages.Fig. 1 Karyotype analysis of the B. longissima. (A) Photo of B. longissima. (B) Photos of coconut palm (Cocos nucifera L.) damage caused by B. longissima. (C,D) karyotype analysis of B. longissima male chromosome number.

Genome sequencing and assembly

Genomic DNA was extracted from male adults within one day of emergence using the QIAamp DNA Mini Kit (Qiagen, Hilden, Germany) for Illumina, PacBio, and Hi-C sequencing. The purity and integrity of the DNA were verified with a NanoDrop 2000C spectrophotometer (Thermo Fisher Scientific, Wilmington, DE, USA) and 1.5% agarose gel electrophoresis. Approximately 0.5 μg of genomic DNA was used to create a PCR-free Illumina genomic library with the TruSeq Nano DNA HT Sample Preparation Kit (Illumina), targeting a 350-bp insert size. The library was sequenced in a 2 × 150-bp format on the Illumina NovaSeq 6000 platform, generating 39.06 Gb of raw data, achieving a sequence coverage of 67.08× (Table 1). The quality control of the raw Illumina reads was performed using fastp v0.20.18. Clean reads were then used to construct a 17-mer frequency distribution map with Jellyfish v2.3.19. The genome size of B. longissima was further estimated to be 553 Mb (Supplementary Figure 1) using GCE v1.0.210.Table 1 Sequencing data statistics for genome assembly.

Genomic libraries	Clean data (Gb)	Mean length (bp)	Sequencing coverage (X)	
Illumina	39.06	150	67.08	
PacBio	154.76	21198.5	265.8	
Hi-C	73.48	150	126.2	
RNA-Seq	111.85	150	—	
Iso-Seq	68.27	1568.2	—	

For PacBio sequencing, 5 μg of genomic DNA was used to generate ~20-kb insert libraries. These libraries were sequenced on the PacBio Sequel (Pacific Biosciences) platform. A total of 154.76 Gb of clean data was generated, providing a sequence coverage of 265.8× (Table 1). The PacBio clean reads were then subjected to de novo assembly using Canu v1.911 following parameters were used for de novo assembly: correctedErrorRate = 0.035 utgOvlErrorRate = 0.065 trimReadsCoverage = 2 trimReadsOverlap = 500. The parameter correctedErrorRate = 0.035 sets the maximum allowed error rate for corrected reads, ensuring high accuracy of the reads used in the assembly process. The parameter utgOvlErrorRate = 0.065 specifies the maximum error rate allowed for overlaps between unitigs, which helps in correctly assembling contiguous sequences. The trimReadsCoverage = 2 parameter ensures that only reads with at least two times coverage are used, removing low-coverage reads that might introduce errors. Finally, trimReadsOverlap = 500 sets the minimum overlap length required between reads, ensuring that only significant overlaps are considered during assembly. These settings help to improve the accuracy and continuity of the assembled genome. To further generate non-redundant genome sequences, haplotigs and contig overlaps in the initial assemblies were removed using purge_dups v1.2.612. The genome assembly was further refined for errors using both PacBio and Illumina reads with NextPolish v1.2.413 with default parameters.

Hi-C scafolding

Firstly, we conducted cytogenetic analysis using testicular material from sexually mature male coconut leaf beetles. The analysis revealed that the male diploid complement of coconut leaf beetles consists of 18 autosomes along with X and Y sex chromosomes (Fig. 1C,D). To achieve the contig to chromosome-level assembly, the Hi-C technique was used to identify contacts between different regions of chromatin filaments. Male adults within 1 day of emergence were selected for Hi-C library construction. The Hi-C library was constructed following the standard library preparation protocol, where nuclear DNA was cross-linked in situ, extracted, and then digested with the restriction enzyme DpnII. The Hi-C libraries were sequenced on the Illumina NovaSeq 6000 platform with 2 × 150-bp reads. Low-quality reads and adapters from the Hi-C library were filtered using fastp v0.20.18. Under the default parameters, a total of 73.48 Gb of clean data was generated (Table 1) and then mapped to the assembled contigs using Chromap v0.2.514 with the parameter: “--remove-pcr-duplicates”. YaHS v1.2a.115 was applied to perform clustering, ordering, and orientation based on the agglomerative hierarchical clustering algorithm. The chromosome interaction matrix was further manually adjusted using JuicerBox v1.11.0816. Finally, the genome had a total size of 582.24 Mb and an N50 size of 63.81 Mb (Table 2). The genome contained 10 chromosomes (Fig. 2A), which included 98.01% of the assembled contigs (Table 2).Table 2 B. longissima genome assembly and annotation.

Genomic features	Brontispa longissima	
Estimated male genome size (Mb)	553	
Total assembly size (Mb)	582.24	
Assembly level	Chromosome	
Number of chromosomes	10	
Sequences anchored to chromosomes (%)	98.01	
Scaffold N50 size (Mb)	63.81	
BUSCO complete rate of the genome (%)	98.1	
GC content (%)	33.71	
Number of genes BUSCO complete rate of the gene sets (%)	23051 95.4	
Total CDS size (Mb) and % in the genome	35.75 (6.14%)	

Fig. 2 Assembly and quality evaluation of the chromosome genome of B. longissima. (A) Circos plot of distribution of genomic elements in B. longissima. (B) Comparison of BUSCO completeness of the genome among 14 coleoptera species, as a percentage of 1367 insect genes from insecta_odb10. (C) Comparison of contig contiguity among 8 Coleoptera species, including Rhagonycha fulva, H. axyridis, B. longissima, Leptinotarsa decemlineata, C. septempunctata, P. serraticornis, Dendroctonus valens, Tribolium castaneum. N (x)% graphs show contig sizes (y-axis), in which x percent of the assembly consists of contigs of at least that size.

Following Hi-C scaffolding, the genome integrity of B. longissima was assessed using Benchmarking Universal Single-Copy Orthologs (BUSCO v5.4.3)17. The assessment showed that the B. longissima chromosome-level assembly achieved the following BUSCO scores: C: 98.1% [S: 97.8%, D: 0.3%], F: 0.4%, M: 1.5%, n: 1367. These results confirm that the B. longissima genome assembly is both complete and accurate. Furthermore, B. longissima demonstrates relatively greater completeness compared to other Coleoptera insects (Fig. 2B). Remarkably, the assembly continuity of B. longissima surpasses that of other high-quality Coleoptera genome assemblies (Fig. 2C).

Repetitive elements and non-coding RNA annotations

The tandem repeats were firstly annotated using the software GMATA v2.318 and Tandem Repeats Finder (TRF) v4.0919 with default parameters. The simple sequence repeats (SSRs) and all tandem repeat elements across the entire genome were identified using GMATA and TRF, respectively. The transposable elements (TEs) sequences, including SINEs, Penelope, LINEs, LTR elements, DNA transposons, and Rolling-circles were annotated using de novo approaches. We first created a de novo repeat library using RepeatModeler v2.0.420 based on the assembly sequences with default parameters. Transposable element (TE) sequences were then identified through homology searches against this library using RepeatMasker v4.1.521. As a result, 288.54 Mb of repetitive element sequences were identified, representing 49.55% of the genome assembly (Table 3).Table 3 Summary statistics of repetitive elements in the assembled B. longissima genome.

Class	Number of elements	Length of sequence (bp)	Percentage of sequence (%)	
Tandem repeat	77,279	17,632,009	3.0282	
SSR	97,471	2,240,809	0.3848	
SINEs	47	4,570	0.0008	
Penelope	0	0	0.0000	
LINEs	142,367	49,861,125	8.5636	
LTR elements	8,651	2,832,547	0.4865	
DNA transposons	626,192	151,643,605	26.0446	
Rolling-circles	38,584	16,195,098	2.7815	
Unclassified:	428,674	48,132,223	8.2667	
Low complexity	0	0	0.0000	
Simple repeats	174	4214	0.0007	
Satellites	0	0	0.0000	
Total Repeats	1,419,439	288,546,200	49.5575	

To obtain non-coding RNAs (ncRNAs) in the B. longissima genome, two strategies were employed: database searches and model predictions. Transfer RNAs (tRNAs) were identified using tRNAscan-SE v2.0.922 with eukaryote-specific parameters, while microRNAs (miRNAs), ribosomal RNAs (rRNAs), small nuclear RNAs (snRNAs), and small nucleolar RNAs (snoRNAs) were detected by searching the Rfam database v14.1023 with Infernal cmscan v1.1.4. Additionally, rRNAs and their subunits were predicted using Barrnap v0.9 with the parameter --kingdom euk. This comprehensive approach identified 459 tRNAs, 83 rRNAs, 68 miRNAs, and 84 other ncRNAs (Table 4).Table 4 Summary statistics of Non-coding RNA annotation results.

Type	Copy Number	Average Length(bp)	Total Length(bp)	
18S rRNA	26	1040.08	27042	
28S rRNA	36	1421.50	51174	
5S rRNA	13	109.23	1420	
5.8S rRNA	8	137.13	1097	
tRNA	459	75.09	34465	
miRNA	68	77.52	5271	
snRNA	61	136.61	8333	
sRNA	23	137.74	3168	

Protein-coding gene prediction and function annotation

To annotate protein-coding genes (PCGs) in the genome, we utilized a comprehensive approach combining ab initio prediction, homology-based prediction, and transcriptome-based evidence. Ab initio predictions were performed using Augustus v3.2.324 with default parameters. For homology-based prediction, we aligned and homology annotation the genome sequence against non-overlapping protein sequences from closely related seven species, including Drosophila melanogaster, T. castaneum, Coccinella septempunctata, Harmonia axyridis, Octodonta nipae, Propylea japonica, and Pyrochroa serraticornis, using Miniprot v0.1225 with default parameters. Transcriptome-based predictions were divided into two approaches: RNA-Seq prediction and Iso-Seq prediction. For RNA-Seq prediction, RNA-seq data were mapped to the genome using HISAT2 v2.1.026 and gene models were predicted with Cufflinks v2.2.127 with default parameters. For Iso-Seq prediction, we mapped the Iso-Seq data using GMAP v2023-12-0128 and Minimap2 v2.2829 with default parameters, followed by gene prediction with PASA v2.5.330. The RNA-Seq samples were collected from the egg, 1st, 4th, and 5th instars, pupa, and male adults, generating 111.85 Gb of data (Table 1). The Iso-Seq samples were mixed samples including all developmental stages, generating 68.27 Gb of data (Table 1). Finally, we integrated the gene prediction results from these methods using EVidenceModeler v1.1.131 with default parameters to generate unified consensus gene models. This analysis predicted a total of 23,051 genes in B. longissima, with the total length of coding sequences (CDSs) reaching 35.75 Mb, representing 6.14% of the genome (Table 1). The predicted protein gene sequences assessed for BUSCO completeness were 95.4% (n: 1367), including 94.4% single-copy, 1.0% duplicated, 2.9% fragmented and 1.7% missing BUSCOs.

To perform functional annotation of the protein-coding genes, we employed multiple approaches. First, the predicted genes were aligned against the NR and UniProtKB databases using Diamond v0.9.30.13132 with a threshold of 1e-5, resulting in 21,142 genes (91.71%) showing hits in the NR database and 11,952 genes (51.85%) in the UniProtKB database (Table 5). Additionally, to annotate Gene Ontology (GO) terms, KEGG and Reactome pathways, and identify protein domains, we used HMMER v3.1b233 for the Pfam database, KofamKOALA v1.3.034 for KEGG, and eggNOG-mapper v2.1.53235 for eggNOG. This resulted in 14,752 genes (63.99%) matching the Pfam database and 20,205 genes (87.63%) matching the eggNOG database. Furthermore, 11,755 genes (50.99%) were annotated with GO terms and 13,373 genes (58.01%) with KEGG Orthology (KO) terms. In total, 21,368 of the 23,051 predicted genes (92.69%) were annotated by at least one public database, demonstrating substantial functional annotation coverage (Table 5).Table 5 Statistics of functional annotation based on various databases.

Annotation in database	Number of genes	Percentage %	
EggNOG	20205	87.65346406	
NR	21142	91.71836363	
Pfam	14752	63.99722355	
KEGG	13373	58.01483667	
GO	11755	50.99561841	
Uniport	11952	51.85024511	
At least one database	21368	92.69879832	

Data Records

All data were associated with the BioProject PRJNA1085852. The reference genome of B. longissima was deposited in GenBank (JBBPDZ000000000.136). Raw data from Pacbio (SRR2827742437), Hi-C (SRR2827742338) and Illumina (SRR2827742539) genome sequencing and RNA-seq (SRR28277416-SRR2827742240–46) were deposited in the NCBI SRA database with the accession number SRP49423347. The annotation of the B. longissima genome have been deposited at figshare2634425248.

Technical Validation

Genome assembly quality assessment

To assess the quality of the genome, the Illumina genomic reads, Pacbio genomic reads and RNA-Seq reads were mapped to B. longissima genome using the BWA v0.7.17-r118849, Minimap2 v2.22-r110129, and HISAT2 v2.1.026, respectively. The results showed that 98.04% of the Illumina genomic reads mapped back to the assembly, achieving a genome coverage rate of 98.20% (Table 6). Additionally, 94.42% of Pacbio genomic reads were mapped back to the assembly, with a genome coverage rate of 99.79%. More than 86.79% of all RNA-Seq reads were also successfully recovered in the genome, indicating that most reads were successfully assembled (Table 6). The completeness of the B. longissima genome assembly was further evaluated using BUSCO (v5.3.2), which revealed that 98.1% of Insecta dataset core genes (odb_10, released on 2024-01-08) from OrthoDB (http://www.orthodb.org) were identified in the genome assembly (Table 2). The predicted protein gene sequences assessed for BUSCO completeness were 95.4% (n: 1367), including 94.4% single-copy, 1.0% duplicated, 2.9% fragmented and 1.7% missing BUSCOs. A total of 23,051 genes were predicted, and the predicted protein gene sequences showed a BUSCO completeness of 95.4%. Together, these assessments strongly suggest that the assembled B. longissima genome was complete and high quality.Table 6 Summary of Illumina and Pacbio reads aligned to the B. longissima genome assembly.

Reads type	Reads length (bp)	No. of reads	Mapped read	Percentage of mapped reads (%)	Percentage of genome coverage (%)	Sequencing depth	
Genome Illumina reads	150	153,930,157	150,912,605	98.04	98.20	35	
Genome Pacbio reads	21198	15,499,342	14,634,933	94.42	99.79	207	
RNA-Seq of egg	150	131,676,945	114,279,620	86.79	—	—	
RNA-Seq of L1	150	140,079,852	124,200,706	88.66	—	—	
RNA-Seq of L4	150	130,294,350	117,905,694	90.49	—	—	
RNA-Seq of L5	150	137,936,102	126,662,588	91.83	—	—	
RNA-Seq of pupa	150	125,268,489	113,242,464	90.40	—	—	
RNA-Seq of male	150	127,790,441	116,927,613	91.50	—	—	

Supplementary information

Supplementary Figure 1

Supplementary information

The online version contains supplementary material available at 10.1038/s41597-024-03846-4.

Acknowledgements

This work was supported by the National Key R&D Program of China (2021YFC2600400 & 2021YFC2600402); the Central Public-interest Scientific Institution Basal Research Fund for Chinese Academy of Tropical Agricultural Sciences (No. 1630042019029); the Hainan Province Science and Technology Special Fund (ZDKJ2021016); the Shenzhen Science and Technology Program (KQTD20180411143628272); and the agricultural science and technology innovation program.

Author contributions

Z.G., Z.P. and F.W. conceived the study. C.T., T.J. and C.L. prepared the samples. Z.L. and Z.G. performed the experiments, analyzed the data, and wrote the manuscript. H.W. and X.Q. conducted the field investigation and took photos. Z.P., F.W., B.L. and W.Q. evaluated the results. G.M. provided funding support. All authors reviewed the manuscript.

Code availability

This study followed the protocols and manuals provided by the bioinformatics software developers for data processing and analysis, as described in the Methods section. No custom scripts were used.

Competing interests

The authors declare no competing interests.

Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

These authors contributed equally: Zaiyuan Li, Guangchang Ma.
==== Refs
References

1. Chen Z Development of Single Nucleotide Polymorphism (SNP) Markers for Analysis of Population Structure and Invasion Pathway in the Coconut Leaf Beetle Brontispa longissima (Gestro) Using Restriction Site-Associated DNA (RAD) Genotyping in Southern China Insects. 2020 11 230 10.3390/insects11040230 32272596
Chen, Z. et al. Development of Single Nucleotide Polymorphism (SNP) Markers for Analysis of Population Structure and Invasion Pathway in the Coconut Leaf Beetle Brontispa longissima (Gestro) Using Restriction Site-Associated DNA (RAD) Genotyping in Southern China. Insects. 11, 230 (2020).32272596 10.3390/insects11040230
2. Lu B Inter-country trade, genetic diversity and bio-ecological parameters upgrade pest risk maps for the coconut hispid Brontispa longissima Pest Manag. Sci. 2020 76 1483 1491 10.1002/ps.5663 31659862
Lu, B. et al. Inter-country trade, genetic diversity and bio-ecological parameters upgrade pest risk maps for the coconut hispid Brontispa longissima. Pest Manag. Sci. 76, 1483–1491 (2020).31659862 10.1002/ps.5663
3. Zhang X Tang B Hou Y A Rapid Diagnostic Technique to Discriminate between Two Pests of Palms, Brontispa longissima and Octodonta nipae (Coleoptera: Chrysomelidae), for Quarantine Applications J. Econ. Entomol. 2015 108 95 99 10.1093/jee/tou025 26470108
Zhang, X., Tang, B. & Hou, Y. A Rapid Diagnostic Technique to Discriminate between Two Pests of Palms, Brontispa longissima and Octodonta nipae (Coleoptera: Chrysomelidae), for Quarantine Applications. J. Econ. Entomol. 108, 95–99 (2015).26470108 10.1093/jee/tou025
4. Voegele J Biological control of Brontispa longissima in Western Samoa: An ecological and economic evaluation Agric. Ecosyst. Environ. 1989 27 315 329 10.1016/0167-8809(89)90095-9
Voegele, J. Biological control of Brontispa longissima in Western Samoa: An ecological and economic evaluation. Agric. Ecosyst. Environ. 27, 315–329 (1989).10.1016/0167-8809(89)90095-9
5. Lin YY Jin T Jin QA Wen HB Peng ZQ Differential susceptibilities of Brontispa longissima (Coleoptera: Hispidae) to insecticides in Southeast Asia J. Econ. Entomol. 2012 105 988 93 10.1603/EC11387 22812140
Lin, Y. Y., Jin, T., Jin, Q. A., Wen, H. B. & Peng, Z. Q. Differential susceptibilities of Brontispa longissima (Coleoptera: Hispidae) to insecticides in Southeast Asia. J. Econ. Entomol. 105, 988–93 (2012).22812140 10.1603/EC11387
6. Takano S Takasu K Ichiki R Fushimi T Nakamura S Induction of host‐plant preference in Brontispa longissima (Gestro) (Coleoptera: Chrysomelidae) J. Appl. Entomol. 2011 135 634 640 10.1111/j.1439-0418.2010.01591.x
Takano, S., Takasu, K., Ichiki, R., Fushimi, T. & Nakamura, S. Induction of host‐plant preference in Brontispa longissima (Gestro) (Coleoptera: Chrysomelidae). J. Appl. Entomol. 135, 634–640 (2011).10.1111/j.1439-0418.2010.01591.x
7. Bin SY Antennal and abdominal transcriptomes reveal chemosensory gene families in the coconut hispine beetle, Brontispa longissima Sci. Rep. 2017 7 2809 10.1038/s41598-017-03263-1 28584273
Bin, S. Y. et al. Antennal and abdominal transcriptomes reveal chemosensory gene families in the coconut hispine beetle, Brontispa longissima. Sci. Rep. 7, 2809 (2017).28584273 10.1038/s41598-017-03263-1
8. Chen S Zhou Y Chen Y Gu J fastp: an ultra-fast all-in-one FASTQ preprocessor Bioinformatics. 2018 34 i884 i890 10.1093/bioinformatics/bty560 30423086
Chen, S., Zhou, Y., Chen, Y. & Gu, J. fastp: an ultra-fast all-in-one FASTQ preprocessor. Bioinformatics. 34, i884–i890 (2018).30423086 10.1093/bioinformatics/bty560
9. Marçais G Kingsford C A fast, lock-free approach for efficient parallel counting of occurrences of k-mers Bioinformatics. 2011 27 6 764 770 10.1093/bioinformatics/btr011 21217122
Marçais, G. & Kingsford, C. A fast, lock-free approach for efficient parallel counting of occurrences of k-mers. Bioinformatics. 27(6), 764–770 (2011).21217122 10.1093/bioinformatics/btr011
10. Liu, B. H. et al. Estimation of genomic characteristics by analyzing k-mer frequency in de novo genome projects. arXiv preprint arXiv:1308.2012 (2013).
11. Koren S Canu: scalable and accurate long-read assembly via adaptive k-mer weighting and repeat separation Genome Res. 2017 27 722 736 10.1101/gr.215087.116 28298431
Koren, S. et al. Canu: scalable and accurate long-read assembly via adaptive k-mer weighting and repeat separation. Genome Res. 27, 722–736 (2017).28298431 10.1101/gr.215087.116
12. Guan D Identifying and removing haplotypic duplication in primary genome assemblies Bioinformatics. 2020 36 2896 2898 10.1093/bioinformatics/btaa025 31971576
Guan, D. et al. Identifying and removing haplotypic duplication in primary genome assemblies. Bioinformatics. 36, 2896–2898 (2020).31971576 10.1093/bioinformatics/btaa025
13. Hu J Fan J Sun Z Liu S NextPolish: a fast and efficient genome polishing tool for long-read assembly Bioinformatics. 2020 36 2253 2255 10.1093/bioinformatics/btz891 31778144
Hu, J., Fan, J., Sun, Z. & Liu, S. NextPolish: a fast and efficient genome polishing tool for long-read assembly. Bioinformatics. 36, 2253–2255 (2020).31778144 10.1093/bioinformatics/btz891
14. Zhang H Fast alignment and preprocessing of chromatin profiles with Chromap Nat. Commun. 2021 12 6566 10.1038/s41467-021-26865-w 34772935
Zhang, H. et al. Fast alignment and preprocessing of chromatin profiles with Chromap. Nat. Commun. 12, 6566 (2021).34772935 10.1038/s41467-021-26865-w
15. Zhou C McCarthy SA Durbin R YaHS: yet another Hi-C scaffolding tool Bioinformatics. 2023 39 btac808 10.1093/bioinformatics/btac808 36525368
Zhou, C., McCarthy, S. A. & Durbin, R. YaHS: yet another Hi-C scaffolding tool. Bioinformatics. 39, btac808 (2023).36525368 10.1093/bioinformatics/btac808
16. Durand NC Juicebox Provides a Visualization System for Hi-C Contact Maps with Unlimited Zoom Cell Syst. 2016 3 99 101 10.1016/j.cels.2015.07.012 27467250
Durand, N. C. et al. Juicebox Provides a Visualization System for Hi-C Contact Maps with Unlimited Zoom. Cell Syst. 3, 99–101 (2016).27467250 10.1016/j.cels.2015.07.012
17. Simao FA Waterhouse RM Ioannidis P Kriventseva EV Zdobnov EM BUSCO: assessing genome assembly and annotation completeness with single-copy orthologs Bioinformatics. 2015 31 3210 3212 10.1093/bioinformatics/btv351 26059717
Simao, F. A., Waterhouse, R. M., Ioannidis, P., Kriventseva, E. V. & Zdobnov, E. M. BUSCO: assessing genome assembly and annotation completeness with single-copy orthologs. Bioinformatics. 31, 3210–3212 (2015).26059717 10.1093/bioinformatics/btv351
18. Wang X Wang L GMATA An Integrated Sofware Package for Genome-Scale SSR Mining, Marker Development and Viewing Front. Plant Sci. 2016 7 1350 27679641
Wang, X., Wang, L. & GMATA An Integrated Sofware Package for Genome-Scale SSR Mining, Marker Development and Viewing. Front. Plant Sci. 7, 1350 (2016).27679641
19. Benson G Tandem repeats finder: a program to analyze DNA sequences Nucleic Acids Res. 1999 27 573 580 10.1093/nar/27.2.573 9862982
Benson, G. Tandem repeats finder: a program to analyze DNA sequences. Nucleic Acids Res. 27, 573–580 (1999).9862982 10.1093/nar/27.2.573
20. Flynn JM RepeatModeler2 for automated genomic discovery of transposable element families Proc. Natl. Acad. Sci. USA 2020 117 9451 9457 10.1073/pnas.1921046117 32300014
Flynn, J. M. et al. RepeatModeler2 for automated genomic discovery of transposable element families. Proc. Natl. Acad. Sci. USA 117, 9451–9457 (2020).32300014 10.1073/pnas.1921046117
21. Tarailo-Graovac M Chen N Using RepeatMasker to identify repetitive elements in genomic sequences Curr. Protoc. Bioinformatics. 2009 4 4101 41014
Tarailo-Graovac, M. & Chen, N. Using RepeatMasker to identify repetitive elements in genomic sequences. Curr. Protoc. Bioinformatics. 4, 4101–41014 (2009).
22. Lowe TM Eddy SR tRNAscan-SE: a program for improved detection of transfer RNA genes in genomic sequence Nucleic. Acids. Res. 1997 25 955 64 10.1093/nar/25.5.955 9023104
Lowe, T. M. & Eddy, S. R. tRNAscan-SE: a program for improved detection of transfer RNA genes in genomic sequence. Nucleic. Acids. Res. 25, 955–64 (1997).9023104 10.1093/nar/25.5.955
23. Kalvari I Rfam 14: expanded coverage of metagenomic, viral and microRNA families Nucleic Acids Res. 2021 49 D192 D200 10.1093/nar/gkaa1047 33211869
Kalvari, I. et al. Rfam 14: expanded coverage of metagenomic, viral and microRNA families. Nucleic Acids Res. 49, D192–D200 (2021).33211869 10.1093/nar/gkaa1047
24. Stanke M AUGUSTUS: ab initio prediction of alternative transcripts Nucleic. Acids. Res. 2006 34 W435 9 10.1093/nar/gkl200 16845043
Stanke, M. et al. AUGUSTUS: ab initio prediction of alternative transcripts. Nucleic. Acids. Res. 34, W435–9 (2006).16845043 10.1093/nar/gkl200
25. Li H Protein-to-genome alignment with miniprot Bioinformatics. 2023 39 btad014 10.1093/bioinformatics/btad014 36648328
Li, H. Protein-to-genome alignment with miniprot. Bioinformatics. 39, btad014 (2023).36648328 10.1093/bioinformatics/btad014
26. Kim D Paggi JM Park C Bennett C Salzberg SL Graph-based genome alignment and genotyping with HISAT2 and HISAT-genotype Nat. Biotechnol. 2019 37 907 915 10.1038/s41587-019-0201-4 31375807
Kim, D., Paggi, J. M., Park, C., Bennett, C. & Salzberg, S. L. Graph-based genome alignment and genotyping with HISAT2 and HISAT-genotype. Nat. Biotechnol. 37, 907–915 (2019).31375807 10.1038/s41587-019-0201-4
27. Trapnell C Transcript assembly and quantification by RNA-Seq reveals unannotated transcripts and isoform switching during cell differentiation Nat. Biotechnol. 2010 28 511 515 10.1038/nbt.1621 20436464
Trapnell, C. et al. Transcript assembly and quantification by RNA-Seq reveals unannotated transcripts and isoform switching during cell differentiation. Nat. Biotechnol. 28, 511–515 (2010).20436464 10.1038/nbt.1621
28. Wu TD Watanabe CK GMAP: a genomic mapping and alignment program for mRNA and EST sequences Bioinformatics. 2005 21 1859 75 10.1093/bioinformatics/bti310 15728110
Wu, T. D. & Watanabe, C. K. GMAP: a genomic mapping and alignment program for mRNA and EST sequences. Bioinformatics. 21, 1859–75 (2005).15728110 10.1093/bioinformatics/bti310
29. Li H Minimap2: pairwise alignment for nucleotide sequences Bioinformatics. 2018 34 3094 3100 10.1093/bioinformatics/bty191 29750242
Li, H. Minimap2: pairwise alignment for nucleotide sequences. Bioinformatics. 34, 3094–3100 (2018).29750242 10.1093/bioinformatics/bty191
30. Haas BJ Improving the Arabidopsis genome annotation using maximal transcript alignment assemblies Nucleic Acids Res. 2003 31 5654 66 10.1093/nar/gkg770 14500829
Haas, B. J. et al. Improving the Arabidopsis genome annotation using maximal transcript alignment assemblies. Nucleic Acids Res. 31, 5654–66 (2003).14500829 10.1093/nar/gkg770
31. Haas BJ Automated eukaryotic gene structure annotation using EVidenceModeler and the Program to Assemble Spliced Alignments Genome Biol. 2008 9 R7 10.1186/gb-2008-9-1-r7 18190707
Haas, B. J. et al. Automated eukaryotic gene structure annotation using EVidenceModeler and the Program to Assemble Spliced Alignments. Genome Biol. 9, R7 (2008).18190707 10.1186/gb-2008-9-1-r7
32. Buchfink B Reuter K Drost HG Sensitive protein alignments at tree-of-life scale using DIAMOND Nat. Methods. 2021 18 366 368 10.1038/s41592-021-01101-x 33828273
Buchfink, B., Reuter, K. & Drost, H. G. Sensitive protein alignments at tree-of-life scale using DIAMOND. Nat. Methods. 18, 366–368 (2021).33828273 10.1038/s41592-021-01101-x
33. Eddy SR Accelerated Profile HMM Searches PLoS Comput. Biol. 2011 7 e1002195 10.1371/journal.pcbi.1002195 22039361
Eddy, S. R. Accelerated Profile HMM Searches. PLoS Comput. Biol. 7, e1002195 (2011).22039361 10.1371/journal.pcbi.1002195
34. Aramaki T KofamKOALA: KEGG Ortholog assignment based on profile HMM and adaptive score threshold Bioinformatics. 2020 36 2251 2252 10.1093/bioinformatics/btz859 31742321
Aramaki, T. et al. KofamKOALA: KEGG Ortholog assignment based on profile HMM and adaptive score threshold. Bioinformatics. 36, 2251–2252 (2020).31742321 10.1093/bioinformatics/btz859
35. Cantalapiedra CP Hernandez-Plaza A Letunic I Bork P Huerta-Cepas J eggNOG-mapper v2: Functional Annotation, Orthology Assignments, and Domain Prediction at the Metagenomic Scale Mol. Biol. Evol. 2021 38 5825 5829 10.1093/molbev/msab293 34597405
Cantalapiedra, C. P., Hernandez-Plaza, A., Letunic, I., Bork, P. & Huerta-Cepas, J. eggNOG-mapper v2: Functional Annotation, Orthology Assignments, and Domain Prediction at the Metagenomic Scale. Mol. Biol. Evol. 38, 5825–5829 (2021).34597405 10.1093/molbev/msab293
36. Li ZY 2024 Brontispa longissima isolate ZL-2024, whole genome shotgun sequencing project GenBank JBBPDZ000000000.1
Li, Z. Y. Brontispa longissima isolate ZL-2024, whole genome shotgun sequencing project. GenBank https://identifiers.org/ncbi/insdc:JBBPDZ000000000.1 (2024).
37. 2024 NCBI Sequence Read Archive SRR28277424
NCBI Sequence Read Archive https://identifiers.org/ncbi/insdc.sra:SRR28277424 (2024).
38. 2024 NCBI Sequence Read Archive SRR28277423
NCBI Sequence Read Archive https://identifiers.org/ncbi/insdc.sra:SRR28277423 (2024).
39. 2024 NCBI Sequence Read Archive SRR28277425
NCBI Sequence Read Archive https://identifiers.org/ncbi/insdc.sra:SRR28277425 (2024).
40. 2024 NCBI Sequence Read Archive SRR28277416
NCBI Sequence Read Archive https://identifiers.org/ncbi/insdc.sra:SRR28277416 (2024).
41. 2024 NCBI Sequence Read Archive SRR28277417
NCBI Sequence Read Archive https://identifiers.org/ncbi/insdc.sra:SRR28277417 (2024).
42. 2024 NCBI Sequence Read Archive SRR28277418
NCBI Sequence Read Archive https://identifiers.org/ncbi/insdc.sra:SRR28277418 (2024).
43. 2024 NCBI Sequence Read Archive SRR28277419
NCBI Sequence Read Archive https://identifiers.org/ncbi/insdc.sra:SRR28277419 (2024).
44. 2024 NCBI Sequence Read Archive SRR28277420
NCBI Sequence Read Archive https://identifiers.org/ncbi/insdc.sra:SRR28277420 (2024).
45. 2024 NCBI Sequence Read Archive SRR28277421
NCBI Sequence Read Archive https://identifiers.org/ncbi/insdc.sra:SRR28277421 (2024).
46. 2024 NCBI Sequence Read Archive SRR28277422
NCBI Sequence Read Archive https://identifiers.org/ncbi/insdc.sra:SRR28277422 (2024).
47. 2024 NCBI Sequence Read Archive SRP494233
NCBI Sequence Read Archive https://identifiers.org/ncbi/insdc.sra:SRP494233 (2024).
48. Li ZY 2024 Genome information files for Brontispa longissima figshare. Dataset. 10.6084/m9.figshare.26344252.v1
Li, Z. Y. Genome information files for Brontispa longissima. figshare. Dataset.10.6084/m9.figshare.26344252.v1 (2024).10.6084/m9.figshare.26344252.v1
49. Li H Durbin R Fast and accurate short read alignment with Burrows-Wheeler transform Bioinformatics. 2009 25 1754 60 10.1093/bioinformatics/btp324 19451168
Li, H. & Durbin, R. Fast and accurate short read alignment with Burrows-Wheeler transform. Bioinformatics. 25, 1754–60 (2009).19451168 10.1093/bioinformatics/btp324
