
==== Front
Sci Data
Sci Data
Scientific Data
2052-4463
Nature Publishing Group UK London

3856
10.1038/s41597-024-03856-2
Data Descriptor
Chromosome-level genome assembly of Chinese water Scorpion Ranatra chinensis (Heteroptera: Nepidae)
Liu Xinzhi 1
Ma Ling 1
http://orcid.org/0000-0002-7288-9676
Tian Li 1
Song Fan 1
Xie Tongyin 2
Wu Yunfei 3
Li Hu 1
Cai Wanzhi caiwz@cau.edu.cn

1
http://orcid.org/0000-0003-2311-9859
Duan Yuange duanyuange@cau.edu.cn

1
1 https://ror.org/04v3ywz14 grid.22935.3f 0000 0004 0530 8290 Department of Entomology and MOA Key Lab of Pest Monitoring and Green Management, College of Plant Protection, China Agricultural University, Beijing, 100193 China
2 https://ror.org/0515nd386 grid.412243.2 0000 0004 1760 1136 College of Plant Protection, Northeast Agricultural University, Harbin, 150030 China
3 https://ror.org/037663q52 grid.411671.4 0000 0004 1757 5070 School of Biology Science and Food Engineering, Chuzhou University, Anhui, 293000 China
18 9 2024
18 9 2024
2024
11 101622 7 2024
3 9 2024
© The Author(s) 2024
2024
https://creativecommons.org/licenses/by-nc-nd/4.0/ Open Access This article is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License, which permits any non-commercial use, sharing, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if you modified the licensed material. You do not have permission under this licence to share adapted material derived from this article or parts of it. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by-nc-nd/4.0/.
Heteroptera (the true bugs), one of the most diverse lineages of insects, diversified in feeding strategies and living habitats, and thus become an ideal lineage for studies on adaptive evolution. Chinese water scorpion Ranatra chinensis (Heteroptera: Nepidae) is a predaceous bug living in lentic water systems, representing an ideal model for studying habitat transition and adaptation to water environment. However, genetic studies on this water bug remain limited. Here, we obtained a chromosome-level genome of R. chinensis using PacBio HiFi long reads and Hi-C sequencing reads. The total assembly size of genome is 867.89 Mb, with a scaffold N50 length of 26.48 Mb and the GC content of 39.50%. All contigs were assembled into 23 pseudo-chromosomes (N = 19 A + X1X2X3X4), and we predicted 18,424 protein-coding genes in this genome. This study will provide valuable genomic resources for future studies on the biology, water adaptation, and genome evolution of water bugs.

Subject terms

Genome
Genomics
https://doi.org/10.13039/501100001809 National Natural Science Foundation of China (National Science Foundation of China) 32120103006 31922012 32120103006 31922012 32120103006 31922012 32120103006 31922012 32120103006 31922012 32120103006 31922012 32120103006 31922012 32120103006 31922012 32120103006 31922012 Liu Xinzhi Ma Ling Tian Li Song Fan Xie Tongyin Wu Yunfei Li Hu Cai Wanzhi Duan Yuange the 2115 Talent Development Program of China Agricultural Universityissue-copyright-statement© Springer Nature Limited 2024
==== Body
pmcBackground & Summary

Biodiversity is the backbone of earth’s life support system, it plays a vital role in supporting and sustaining life on earth and maintaining the ecosystems functioning, it also underpins numerous essential benefits from nature that are vital for human well-being1–3. The ecosystem involved enormous organisms, raising an interesting question that how animals adapt into different habitats. The ability of flight and water adaptation, for example, enables animals gain novel niches and escape predators. For a number of animals, the molecular bases underlying the adaptive evolution have been uncovered. These animals include marine mammals4, Asian honeybee5, the water strider6 and the poultry shaft louse7. However, the adaptive mechanisms of the majority of extant species remain unknown.

Insects are a crucial component of biodiversity and play a vital role in maintenance of ecological balance, which are estimated to be as many as 5.5 million species on earth8. They make up around 80% of animals and around half of all living species, and occupy nearly every terrestrial habitat on the planet8,9. The true bugs (Hemiptera: Heteroptera) is one of the most diverse lineages of insects and the most diversified lineage among hemimetabola, which possesses a diversity of feeding strategies, ranging from predation on other arthropods and hematophagy on vertebrates, to mycetophagy and phytophagy. Hemipterans also vary in living habitats, including various terrestrial, aquatic and even marine habitats10,11. The enormous diversity makes Hemiptera an ideal lineage for exploration of adaptive evolution. In previous studies, the common ancestor of Heteroptera was considered as terrestrial and experienced multiple independent evolutionary events on habitat transitions, from terrestrial to water surface, aquatic and other habitation, and from aquatic to shoreline11,12. The infraorder Nepomorpha, specifically, containing many species living in water, is helpful for understanding on water adaptation and habitats transition. Exploring the water adaptation of water bugs can pave the way for the following studies related to living habitat adaptation. However, the lack of chromosome-level genomes prevents us from a deep genomic analysis of aquatic true bugs and also impedes the discovery of the adaptive mechanism.

The family Nepidae (Heteroptera: Nepomorpha), or water scorpions, can be recognized by the characteristics of the antennae hidden under the head and the long tail-like siphon (the respiratory siphon) on the rear end of their body that cannot be retracted into the apex of the abdomen, which is unique within the aquatic insects13–15. Normally, they can be found in ponds, lakes, marshes and riversides and often move far from its aquatic habitat at night16. As predatory insects, they usually reside in the shade of plants in water or between stones, waiting for its preys, including small fish and aquatic insects and other invertebrates in water16. This means, Nepidae have evolved multiple derived traits compared to ancestors and are also potential natural enemies of health pests for biocontrol, because they are found to be able to control mosquito population by preying the larvae17,18. Ranatra chinensis, also known as Chinese water scorpion, is a representative species of family Nepidae with a long caudally breathing tube. Its body size ranges from 41 mm to 48 mm, and the length of its respiratory siphon can be equal to that of torso18. Ranatra chinensis is widely distributed in China, from the northernmost Heilongjiang Province to southernmost Guangdong Province, making it a relatively accessible specimen for scientific studies.

There is no chromosome-level genome available within family Nepidae yet. Analysis on R. chinensis genome will benefit our interpretation about the molecular mechanisms of water bugs’ habitat transition and adaptative evolution. With the development of sequencing technologies and bioinformatic tools, obtaining a high-quality reference genome becomes feasible for most organisms, enabling us to investigate the genome evolution and the underlying molecular mechanisms. In this study, we assembled and annotated the chromosome-level genome of R. chinensis, combining of PacBio long-read, Illumina short-read sequencing, and chromosome conformation capture (Hi-C) technologies. Our data will provide genomic resources for future exploration on the biology and evolution of water bugs, and will facilitate the understanding of habitat transition and water adaptation.

Methods

Sample preparation

The R. chinensis specimen were collected from Xichuan County, Henan Province, China (33.24° N, 111.02° E), and were put into dry ice for transportation. All samples were stored at −80°C until further usage. The female adult was used for genomic DNA extraction based on the CTAB method, and extracted DNA were purified using a Blood and Cell Culture DNA Midi Kit (QIAGEN, Germany). The DNA degradation and contamination of the extracted DNA was monitored on 1% agarose gels. The purity of DNA samples was then detected using NanoDrop™. One UV-Vis spectrophotometer (Thermo Fisher Scientific, USA) and the Qubit® 4.0 Fluorometer (Invitrogen, USA) were used to measure the DNA concentration.

Genomic DNA sequencing

Illumina short read library was firstly constructed with an insertion size of about 350 bp, and then sequenced using the Illumina Novaseq6000 to generate 150 bp paired-end short reads. We obtained 59.89 Gb raw data of Illumina short reads and finally got 57.55 Gb clean data after removing adapters and low-quality short reads using Fastp version 0.21.019 with default parameters (Table 1).Table 1 Library sequencing data and methods used in this study to assemble the Ranatra chinensis genome.

Sequencing strategy	Platform	Usage	Insertion size	Raw data (Gb)	Clean data (Gb)	Coverage (×)	
Short-reads	Illumina	Genome survey	350 bp	59.89	57.55	66.30	
Long-reads	PacBio-sequel II	Genome assembly	15 kb	872.49	60.41	69.60	
Hi-C	Illumina	Hi-C assembly	350 bp	59.18	57.87	66.67	
RNA-seq	Illumina	Anno-evidence	350 bp	5.26	4.86	44.71	
The coverages of DNA-Seq data were calculated using clean data size divided by the genome size 867.89 Mb = 0.86789 Gb. The coverage of RNA-Seq data was calculated using clean data size divided by the total length of merged gene region 108.69 Mb = 0.10869 Gb.

For long-read sequencing, a PacBio HiFi-read library with insertion sizes of 15 kb was generated, and sequenced the long DNA fragments using a SMRT cell on PacBio Sequel II sequencing platform (Pacific Biosciences, Menlo Park, USA). A total of 60.41 Gb clean data were obtained from the raw long reads generated using Circular Consensus Sequencing (CCS) model (Table 1).

A single female adult was used for chromosome conformation capture (Hi-C) sequencing and the library was prepared according to the standard protocol described by Belton with minor modifications20. The sample was cut into pieces and mixed with 2% formaldehyde solution for cross-linking, and then treated with New England Biolabs (NEB) buffer to digest nuclei. Biotinylated nucleotides were used to fill the cohesive ends and purified DNA was sheared to fragments of 350 bp in length after ligation. DNA purification was achieved using QIAamp DNA Mini Kits (Qiagen). The final generated Hi-C library was sequenced on Illumina Novoseq6000 platform with paired-end 150 bp. The sequencing yields 59.18 Gb raw data and 57.87 Gb clean data obtained after applying the same filter criteria for short reads (Table 1).

Transcriptome sequencing

Total RNA was extracted from two adults (one female and one male) with the TRIzol reagent (Thermo Fisher Scientific, USA) for transcriptome sequencing. The construction of a paired-end library was obtained by using the TruSeq RNA Library Preparation Kit (Illumina, USA). The transcriptome sequencing was finished on an Illumina Novoseq6000 platform, resulting in a total of 4.86 Gb RNA-seq clean data (Table 1).

Genome size estimation

Genomic characteristics including genome size, heterozygosity, and duplication were estimated using 57.55 Gb clean Illumina short-reads. The distribution of k-mer copy number was calculated to perform this estimation in JELLYFISH version 2.1.321. Genome size and genome heterozygosity were estimated based on 17-mer depth analysis in GenomeScope version 2.022 with default parameters, and the results were 605.54 Mb and 2.27%, respectively (Fig. 1).Fig. 1 The K-mer distribution of Illumina paired-end reads using GenomeScope version 2.0 based on a k value of 17. The K-mer distributions showed double peaks: the first peak with a coverage of ~40 indicates genome duplication, and the second peak with a coverage of ~60 represents a genome size peak.

Chromosome-level genome assembly

The initial de novo assembly of R. chinensis genome was performed based on PacBio sequence data using Hifiasm version 0.13 with default parameters23. After assembly, the genome was then polished by the Purge_Dups pipeline24 to remove alternative haplotype and redundant fragments from the genome. Then, the subsequent polishing was performed using Illumina sequencing data to enhance the quality of the contigs. Finally, an 867.89 Mb contig-level genome assembly of R. chinensis was obtained based on PacBio sequencing data, containing 689 contigs with contig N50 and N90 sizes of 26.48 Mb and 3.80 Mb, respectively, and the GC content of 39.50% (Table 2).Table 2 Statistics of the Ranatra chinensis genome assembly.

Contig-level		
Assembly length (bp)	867,893,761	
N50 (bp)	26,478,726	
N90 (bp)	3,803,519	
Longest reads length (bp)	49,587,281	
Contig number	689	
GC content %	39.50	
Chromosome-level	
Number of chromosomes	23	
Chromosome length (bp)	600,782,313	
N50 (bp)	29,801,000	

The high-quality chromosome-scale genome was generated using a scaffolding pipeline based on Durand25. Initially, BWA-MEM version 0.7.1726 with the parameters: ‘mem -SP5M’ was used for mapping Hi-C data to the contig assembly genome. The DpnII sites were generated using the ‘generate_site_positions.py’ script in Juicer version 1.525. Subsequently, contigs were assembled into the chromosome-level scaffolds using the 3D-DNA pipeline with the parameter “-r 2”27. After the confirmation by Hi-C contact maps, chromosome interaction matrix was manually adjusted and corrected using Juicebox version 1.11.0828. Ultimately, we anchored and generated 23 pseudo-chromosomes, and the final chromosome-level genome assembly of R. chinensis was obtained with a scaffold N50 of 29.80 Mb (Fig. 2; Table 2).Fig. 2 The Circos atlas of the Ranatra chinensis chromosome-level genome. Tracks represent (a) the distribution of chromosome karyotypes, (b) gene density, (c) transposable element content, (d) DNA transposon and (e) GC density. Densities were calculated in 100-kb windows. Chr8, Chr21, Chr22 and Chr23 are predicted to be the sex chromosomes and the remaining are autosomes.

Synteny analysis and the determination of sex chromosome in R. chinensis

Based on JCVI v1.1.17 with default parameters29, we performed synteny analysis to confirm the sex chromosome in R. chinensis using public chromosome-level genomes whose sex chromosomes have been verified, including Rhynocoris fuscipes (Hemiptera: Reduviidae)30, Triatoma rubrofasciata (Hemiptera: Reduviidae)31 and Riptortus pedestris (Hemiptera: Alydidae)32. The result showed that the Chr8, Chr21, Chr22 and Chr23 of R. chinensis exhibited high homology with Chr12 and Chr13 of T. rubrofasciata, Chr13, Chr14 and Chr15 of R. fuscipes, and the ChrX of R. pedestris (Fig. S1). This result indicates that R. chinensis has the same sex chromosome system as that of Nepa cinerea (Heteroptera: Nepidae, N = 14 A + X1X2X3X4) and Ranatra linearis (Heteroptera: Nepidae, N = 19 A + X1X2X3X4) reported in previous study33, and Chr8, Chr21, Chr22 and Chr23 correspond to the X1, X2, X3, and X4 chromosomes.

Prediction of repetitive elements

Repeat sequences of R. chinensis genome were detected in Extensive de novo TE Annotator (EDTA) version 1.9.434. LTR retrotransposons were determined in LTR FINDER version 1.0735, LTRharvest36, and LTR retriever version 2.9.037 with default parameters. DNA transposons were classified utilizing TIR Learner38 and HelitronScanner39 with default parameters. RepeatMasker version 4.0.7 (parameters: -gff -xsmall -no_is)40 and RepeatProteinMask version 4.0.7 (parameters: -engine wublast) were used to find the interspersed repeats against the RepBase database41 (http://www.girinst.org/repbase). In addition, Tandem Repeats Finder version 4.0942 was used to classify tandem repeats with parameters ‘2 7 7 80 10 50 500 -f -d -m’ based on the de novo prediction. RepeatModeler version 2.0.443 (parameters: ‘-engine ncbi -pa 4’) was utilized to construct a repetitive sequence library and RepeatMasker version 4.0.740 was used for annotation of the repeat element against this repeat library. In the genomic sequences, a total of 391.73 Mb (45.13%) repetitive elements were identified, mainly including 22.96% retrotransposon, 3.35% DNA transposons and 11.26% tandem repeat (Table 3). Retrotransposons include LTR, SINE, and LINE; and LTR is further classified in to Copia, Gypsy, and other LTR.Table 3 Repeats elements statistics in genome of Ranatra chinensis.

Item	Number of elements	Percentage (%)	
Retrotransposon		22.96	
LTR Copia	385,921	0.04	
LTR Gypsy	42,399,582	4.88	
LTR-other	9,846,383	1.13	
SINE	68,061	0.01	
LINE	146,571,201	16.88	
DNA transposons		3.35	
CACTA	227,978	0.03	
Harbinger	45,923	0.01	
Helitron	215,872	0.02	
Tcl/Mariner	19,216,292	2.21	
hAT	344,217	0.04	
Mutator	85,513	0.01	
DNA-other	8,910,395	1.03	
Low complexity	724,366	0.08	
Tandem repeat	97,740,021	11.26	
Unclassified	64,748,693	7.46	
Total content	391,728,587	45.13	
SINEs: short interspersed nuclear elements; LINEs: long interspersed nuclear elements; LTR: long terminal repeat.

Protein-coding gene prediction and functional annotation

Protein-coding genes (PCGs) within the genome were predicted by a combined method of homology-based prediction, ab initio prediction and transcriptome-based prediction. HISAT2 version 2.2.144 was utilized to map RNA-seq short data to the genome with the parameter ‘-k 2’. Then the StringTie version 2.4.045 was used to assemble the mapped reads into transcripts with default parameters. For the homology-based prediction, the protein sequences of eight representative insect species were downloaded from the NCBI GenBank database (Table S1). Homologous proteins were aligned against R. chinensis genome using Exonerate version 2.4.0 with default parameters to train the gene sets. Additionally, the bam2hints program (parameter: -intronsonly) in AUGUSTUS version 3.2.346 was employed to transfer the sorted and mapped bam file of RNA-seq data into a hints file. To predict coding genes from the assembled genome, AUGUSTUS version 3.2.346 with default parameters was performed for prediction, in which the combination of trained gene sets and hint files was the input. In the end, MAKER version 2.31.1047 was utilized to merge and generate a consensus high-confidence gene set on the basis of homology-based, de novo-derived and transcript genes. We predicted a total of 18,424 genes in R. chinensis genome with an average gene length of 5,900.48 bp (Table 4). The average length of coding sequences (CDS) and protein sequence were 1,189.79 bp and 396.60 AA, respectively (Table 4). The above statistics on sex chromosomes and autosomes were provided in Table S2.Table 4 Statistics of predicted protein-coding genes of Ranatra chinensis genome assembly.

Class	Result	
Number of genes	18,424	
Maximum gene length (bp)	99,814	
Minimum gene length (bp)	198	
Average gene length (bp)	5,900.48	
Number of mRNAs	18,565	
Average mRNA length (bp)	1,189.79	
Average mRNA count per gene	1.01	
Average CDS length (bp)	1,189.79	
Average protein sequence length (AA)	396.60	
Average exon length (bp)	239.85	
Average exon counts per gene	4.96	

The predicted genes were functionally annotated using multiple methods, includeincludeing eggnog-mapper48 (parameter: -m diamond–tax_scope auto–go_evidence experimental–target_orthologs all–seed_ortholog_evalue 0.001–seed_ortholog_score 60–query-cover 20–subject-cover 0 –override), InterProscan version 5.049 (parameter: -iprlookup -goterms -appl Pfam -f TSV), BLAST version 2.2.2850 (parameter: -evalue 1e-5), and HMMER version 3.3.251 (parameter: –noali–cut_ga Pfam-A.hmm). All these approaches were performed to search against several public databases: Gene Ontology (GO), Clusters of Orthologous Groups of Proteins (COG), Kyoto Encyclopedia of Genes and Genomes (KEGG), NCBI non-redundant protein (Nr), Swiss-Prot, and Pfam. Overall, 16,262 genes were functionally annotated with at least one public database (Table 5).Table 5 Number of functionally annotated protein-coding gene of Ranatra chinensis genome.

Database name	Annotated number	Percent (%)	
GO	5,712	30.77	
KEGG	9,988	53.80	
PFAM	12,933	69.66	
NR	16,095	86.70	
SW	13,578	73.14	
COG	14,087	75.88	
Total	16,262	87.59	

Data Records

Genomic Illumina short-reads data were deposited at the NCBI Sequence Read Archive database under accession number SRR2878529252. Genomic PacBio HiFi sequencing data were deposited at the NCBI Sequence Read Archive database under accession number SRR2878938053. RNA-seq data was deposited at the NCBI Sequence Read Archive database under accession number SRR2878888054. The Hi-C sequencing data were deposited at the NCBI Sequence Read Archive database under accession number SRR2878753855.

The final chromosome assembly was submitted to GenBank at NCBI under accession number JBFDAA00000000056. The genome sequence and raw reads have been deposited in GenBank and Sequence Read Archive at NCBI under BioProject PRJNA110371857.

Technical Validation

To assess the accuracy of the final genome assembly, we mapped the Illumina short-reads to the R. chinensis genome with BWA-MEM version 0.7.1726, and the result showed 97.08% of short reads were successfully mapped to the genome. Benchmarking Universal Single-Copy Orthologs (BUSCO version 3.0.2)58 was used to evaluate the genome completeness based on the insecta_odb10 database, revealing the completeness was 95.7%. Among 1,309 orthologous, 1,299 genes were classified as complete single-copy genes and 10 genes were complete duplicated genes, eight genes were fragmented and 48 genes were missing (Table 6).Table 6 BUSCO evaluation for the final genome assembly of Ranatra chinensis.

Contig level	Number	Percent (%)	
Complete BUSCOs (C)	1,309	95.7	
Complete and single-copy BUSCOs (S)	1,299	95.0	
Complete and duplicated BUSCOs (D)	10	0.7	
Fragmented BUSCOs (F)	10	0.7	
Missing BUSCOs (M)	48	3.6	
Total BUSCO groups searched	1,367	/	

Supplementary information

supplementary_clean version

Supplementary information

The online version contains supplementary material available at 10.1038/s41597-024-03856-2.

Acknowledgements

This work was supported by the National Natural Science Foundation of China (Nos. 32120103006, 31922012) and the 2115 Talent Development Program of China Agricultural University. We thank Shuxian Chen for the help in statistics.

Author contributions

Y.D. and W.C. conceived the project. L.M., X.L., Y.W. and T.X. collected samples and extracted genomic DNA. L.M., X.L. and Y.W. performed data analysis and wrote the manuscript. L.T., F.S., T.X. and H.L. contributed to data analyses. All authors contributed to revising the manuscript. All authors have read and approved the final version.

Code availability

The bioinformatic analyses were performed using the manuals and protocols by the software developers, with the manually adjusted parameters clearly described in the Methods. No custom script or code was used in this study.

Competing interests

The authors declare no competing interests.

Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
==== Refs
References

1. Cardinale B Biodiversity loss and its impact on humanity Nature 2012 486 59 67 10.1038/nature11148 22678280
Cardinale, B. et al. Biodiversity loss and its impact on humanity. Nature 486, 59–67, 10.1038/nature11148 (2012).22678280
2. Brauman KA Global trends in nature’s contributions to people Proc. Natl. Acad. Sci. USA 2020 117 32799 32805 10.1073/pnas.2010473117 33288690
Brauman, K. A. et al. Global trends in nature’s contributions to people. Proc. Natl. Acad. Sci. USA 117, 32799–32805, 10.1073/pnas.2010473117 (2020).33288690
3. Kim HJ Understanding the role of biodiversity in the climate, food, water, energy, transport and health nexus in Europe Sci. Total Environ. 2024 925 171692 10.1016/j.scitotenv.2024.171692 38485013
Kim, H. J. et al. Understanding the role of biodiversity in the climate, food, water, energy, transport and health nexus in Europe. Sci. Total Environ. 925, 171692, 10.1016/j.scitotenv.2024.171692 (2024).38485013
4. Yuan Y Comparative genomics provides insights into the aquatic adaptations of mammals Proc. Natl. Acad. Sci. USA 2021 118 e2106080118 10.1073/pnas.2106080118 34503999
Yuan, Y. et al. Comparative genomics provides insights into the aquatic adaptations of mammals. Proc. Natl. Acad. Sci. USA 118, e2106080118, 10.1073/pnas.2106080118 (2021).34503999
5. Ji Y Gene reuse facilitates rapid radiation and independent adaptation to diverse habitats in the Asian honeybee Sci. Adv. 2020 6 eabd3590 10.1126/sciadv.abd3590 33355133
Ji, Y. et al. Gene reuse facilitates rapid radiation and independent adaptation to diverse habitats in the Asian honeybee. Sci. Adv. 6, eabd3590 https://www.science.org/doi/10.1126/sciadv.abd3590 (2020).33355133
6. Santos M Le Bouquin A Crumière A Khila A Taxon-restricted genes at the origin of a novel trait allowing access to a new environment Science 2017 358 386 390 10.1126/science.aan2748 29051384
Santos, M., Le Bouquin, A., Crumière, A. & Khila, A. Taxon-restricted genes at the origin of a novel trait allowing access to a new environment. Science 358, 386–390, https://www.science.org/doi/10.1126/science.aan2748 (2017).29051384
7. Xu Y Chromosome-level genome of the poultry shaft louse Menopon gallinae provides insight into the host-switching and adaptive evolution of parasitic lice GigaScience 2024 13 giae004 10.1093/gigascience/giae004 38372702
Xu, Y. et al. Chromosome-level genome of the poultry shaft louse Menopon gallinae provides insight into the host-switching and adaptive evolution of parasitic lice. GigaScience 13, giae004, 10.1093/gigascience/giae004 (2024).38372702
8. Stork NE How many species of insects and other terrestrial arthropods are there on earth? Annu. Rev. Entomol. 2018 63 31 45 10.1146/annurev-ento-020117-043348 28938083
Stork, N. E. How many species of insects and other terrestrial arthropods are there on earth? Annu. Rev. Entomol. 63, 31–45, 10.1146/annurev-ento-020117-043348 (2018).28938083
9. Slade EM Ong XR The future of tropical insect diversity: strategies to fill data and knowledge gaps Curr. Opin. Insect Sci. 2023 58 101063 10.1016/j.cois.2023.101063 37247774
Slade, E. M. & Ong, X. R. The future of tropical insect diversity: strategies to fill data and knowledge gaps. Curr. Opin. Insect Sci. 58, 101063, 10.1016/j.cois.2023.101063 (2023).37247774
10. Schuh, R. T. & Slater, J. A. True bugs of the world (Hemiptera: Heteroptera) classification and natural history. Cornell University Press, Ithaca, NY. (1995).
11. Weirauch C Schuh RT Cassis G Wheeler WC Revisiting habitat and lifestyle transitions in Heteroptera (Insecta: Hemiptera): insights from a combined morphological and molecular phylogeny Cladistics 2019 35 67 105 10.1111/cla.12233 34622978
Weirauch, C., Schuh, R. T., Cassis, G. & Wheeler, W. C. Revisiting habitat and lifestyle transitions in Heteroptera (Insecta: Hemiptera): insights from a combined morphological and molecular phylogeny. Cladistics 35, 67–105, 10.1111/cla.12233 (2019).34622978
12. Li H Mitochondrial phylogenomics of Hemiptera reveals adaptive innovations driving the diversification of true bugs Proc. R. Soc. B 2017 284 20171223 10.1098/rspb.2017.1223 28878063
Li, H. et al. Mitochondrial phylogenomics of Hemiptera reveals adaptive innovations driving the diversification of true bugs. Proc. R. Soc. B 284, 20171223, 10.1098/rspb.2017.1223 (2017).28878063
13. Chen P Nieser N Ho J Review of Chinese Ranatrinae (Hemiptera: Nepidae), with descriptions of four new species of Ranatra Fabricius Tijd. Entomol. 2004 147 81 102 10.1163/22119434-900000142
Chen, P., Nieser, N. & Ho, J. Review of Chinese Ranatrinae (Hemiptera: Nepidae), with descriptions of four new species of Ranatra Fabricius. Tijd. Entomol. 147, 81–102, https://api.semanticscholar.org/CorpusID:84596642 (2004).
14. Chen, P., Nieser, N. & Zettel, H. The aquatic and semi-aquatic bugs (Heteroptera: Nepomorpha & Gerromorpha) of Malesia. Fauna Malesiana Handbooks 5. Leiden and Boston, Brill (546 pp), https://api.semanticscholar.org/CorpusID:82862404 (2005).
15. Polhemus DA Polhemus JT Guide to the aquatic Heteroptera of Singapore and Peninsular Malaysia. X. Infraorder Nepomorpha – Families Belostomatidae and Nepidae Raffles Bull. Zool. 2013 61 25 45
Polhemus, D. A. & Polhemus, J. T. Guide to the aquatic Heteroptera of Singapore and Peninsular Malaysia. X. Infraorder Nepomorpha – Families Belostomatidae and Nepidae. Raffles Bull. Zool. 61, 25–45, https://api.semanticscholar.org/CorpusID:87760711 (2013).
16. Ueno, M. Hemiptera. In: Freshwater Biology of Japan. (Edited by UCno M.), Hokuryukan Pub. Co. Ltd, Tokyo. pp. 567–575, (1973).
17. Shaalan EA Canyon DV Aquatic insect predators and mosquito control Trop. Biomed 2009 26 223 261 20237438
Shaalan, E. A. & Canyon, D. V. Aquatic insect predators and mosquito control. Trop. Biomed. 26, 223–261, https://pubmed.ncbi.nlm.nih.gov/20237438/ (2009).20237438
18. Cui J Cai W Taxonomic note on Ranatra Fabricius from Henan of China Journal of Henan Agricultural Sciences 2015 44 87 90
Cui, J. & Cai, W. Taxonomic note on Ranatra Fabricius from Henan of China. Journal of Henan Agricultural Sciences 44, 87–90, https://www.hnnykx.org.cn/EN/Y2015/V44/I2/87 (2015).
19. Chen S Zhou Y Chen Y Gu J Fastp: an ultra-fast all-in-one FASTQ preprocessor Bioinformatics 2018 34 884 890 10.1093/bioinformatics/bty560 29126246
Chen, S., Zhou, Y., Chen, Y. & Gu, J. Fastp: an ultra-fast all-in-one FASTQ preprocessor. Bioinformatics 34, 884–890, 10.1093/bioinformatics/bty560 (2018).29126246
20. Belton JM Hi-C: a comprehensive technique to capture the conformation of genomes Methods 2012 58 268 276 10.1016/j.ymeth.2012.05.001 22652625
Belton, J. M. et al. Hi-C: a comprehensive technique to capture the conformation of genomes. Methods 58, 268–276 (2012).22652625
21. Marcais G Kingsford C A fast, lock-free approach for efficient parallel counting of occurrences of k-mers Bioinformatics 2011 27 764 770 10.1093/bioinformatics/btr011 21217122
Marcais, G. & Kingsford, C. A fast, lock-free approach for efficient parallel counting of occurrences of k-mers. Bioinformatics 27, 764–770, 10.1093/bioinformatics/btr011 (2011).21217122
22. Vurture GW GenomeScope: fast reference-free genome profiling from short reads Bioinformatics 2017 33 2202 2204 10.1093/bioinformatics/btx153 28369201
Vurture, G. W. et al. GenomeScope: fast reference-free genome profiling from short reads. Bioinformatics 33, 2202–2204, 10.1093/bioinformatics/btx153 (2017).28369201
23. Cheng H Haplotype-resolved de novo assembly using phased assembly graphs with hifiasm Nat. Methods 2021 18 170 5, 10.1038/s41592-020-01056-5 33526886
Cheng, H. et al. Haplotype-resolved de novo assembly using phased assembly graphs with hifiasm. Nat. Methods 18, 170–5, 10.1038/s41592-020-01056-5 (2021).33526886
24. Guan D Identifying and removing haplotypic duplication in primary genome assemblies Bioinformatics 2020 36 2896 2898 10.1093/bioinformatics/btaa025 31971576
Guan, D. et al. Identifying and removing haplotypic duplication in primary genome assemblies. Bioinformatics 36, 2896–2898, 10.1093/bioinformatics/btaa025 (2020).31971576
25. Durand NC Juicer provides a one-click system for analyzing loop-resolution Hi-C experiments Cell Syst 2016 3 95 98 10.1016/j.cels.2016.07.002 27467249
Durand, N. C. et al. Juicer provides a one-click system for analyzing loop-resolution Hi-C experiments. Cell Syst. 3, 95–98, 10.1016/j.cels.2016.07.002 (2016).27467249
26. Li H Durbin R Fast and accurate short read alignment with Burrows-Wheeler transform Bioinformatics 2009 25 1754 1760 10.1093/bioinformatics/btp324 19451168
Li, H. & Durbin, R. Fast and accurate short read alignment with Burrows-Wheeler transform. Bioinformatics 25, 1754–1760, 10.1093/bioinformatics/btp324 (2009).19451168
27. Dudchenko O De novo assembly of the Aedes aegypti genome using Hi-C yields chromosome-length scaffolds Science 2017 356 92 95 10.1126/science.aal3327 28336562
Dudchenko, O. et al. De novo assembly of the Aedes aegypti genome using Hi-C yields chromosome-length scaffolds. Science 356, 92–95, 10.1126/science.aal3327 (2017).28336562
28. Durand NC Juicebox provides a visualization system for Hi-C contact maps with unlimited zoom Cell Syst 2016 3 99 101 10.1016/j.cels.2015.07.012 27467250
Durand, N. C. et al. Juicebox provides a visualization system for Hi-C contact maps with unlimited zoom. Cell Syst. 3, 99–101, 10.1016/j.cels.2015.07.012 (2016).27467250
29. Tang H Synteny and collinearity in plant genomes Science 2008 320 486 488 10.1126/science.1153917 18436778
Tang, H. et al. Synteny and collinearity in plant genomes. Science 320, 486–488, https://www.science.org/doi/10.1126/science.1153917 (2008).18436778
30. Ma L Comparative genomic analyses on assassin bug Rhynocoris fuscipes (Hemiptera: Reduviidae) reveal genetic bases governing the diet-shift iScience 2024 27 110411 10.1016/j.isci.2024.110411 39108731
Ma, L. et al. Comparative genomic analyses on assassin bug Rhynocoris fuscipes (Hemiptera: Reduviidae) reveal genetic bases governing the diet-shift. iScience 27, 110411, 10.1016/j.isci.2024.110411 (2024).39108731
31. Liu Q A chromosomal-level genome assembly for the insect vector for Chagas disease, Triatoma rubrofasciata GigaScience 2019 8 giz089 10.1093/gigascience/giz089 31425588
Liu, Q. et al. A chromosomal-level genome assembly for the insect vector for Chagas disease, Triatoma rubrofasciata. GigaScience 8, giz089, 10.1093/gigascience/giz089 (2019).31425588
32. Huang HJ Chromosome-level genome assembly of the bean bug Riptortus pedestris Mol. Ecol. Resour. 2021 21 2423 2436 10.1111/1755-0998.13434 34038033
Huang, H. J. et al. Chromosome-level genome assembly of the bean bug Riptortus pedestris. Mol. Ecol. Resour. 21, 2423–2436, 10.1111/1755-0998.13434 (2021).34038033
33. Angus RB Jeangirard C Stoianova D Grozeva S Kuznetsova VG A chromosomal analysis of Nepa cinerea Linnaeus, 1758 and Ranatra linearis (Linnaeus, 1758) (Heteroptera, Nepidae) Comp. Cytogenet 2017 11 641 657 10.3897/CompCytogen.v11i4.14928 29114353
Angus, R. B., Jeangirard, C., Stoianova, D., Grozeva, S. & Kuznetsova, V. G. A chromosomal analysis of Nepa cinerea Linnaeus, 1758 and Ranatra linearis (Linnaeus, 1758) (Heteroptera, Nepidae). Comp. Cytogenet. 11, 641–657, https://pubmed.ncbi.nlm.nih.gov/29114353/ (2017).29114353
34. Ou S Benchmarking transposable element annotation methods for creation of a streamlined, comprehensive pipeline Genome Biol. 2019 20 1 18 10.1186/s13059-019-1905-y 30606230
Ou, S. et al. Benchmarking transposable element annotation methods for creation of a streamlined, comprehensive pipeline. Genome Biol. 20, 1–18, 10.1186/s13059-019-1905-y (2019).30606230
35. Ou S Jiang N LTR_FINDER_parallel: parallelization of LTR_FINDER enabling rapid identification of long terminal repeat retrotransposons Mobile DNA 2019 10 48 48 10.1186/s13100-019-0193-0 31857828
Ou, S. & Jiang, N. LTR_FINDER_parallel: parallelization of LTR_FINDER enabling rapid identification of long terminal repeat retrotransposons. Mobile DNA 10, 48–48, 10.1186/s13100-019-0193-0 (2019).31857828
36. Ellinghaus D Kurtz S Willhoeft U LTRharvest, an efficient and flexible software for de novo detection of LTR retrotransposons BMC Bioinform. 2008 9 1 14 10.1186/1471-2105-9-18
Ellinghaus, D., Kurtz, S. & Willhoeft, U. LTRharvest, an efficient and flexible software for de novo detection of LTR retrotransposons. BMC Bioinform. 9, 1–14, 10.1186/1471-2105-9-18 (2008).
37. Ou S Jiang N LTR_retriever: a highly accurate and sensitive program for identification of long terminal repeat retrotransposons Plant Physiol. 2017 176 1410 1422 10.1104/pp.17.01310 29233850
Ou, S. & Jiang, N. LTR_retriever: a highly accurate and sensitive program for identification of long terminal repeat retrotransposons. Plant Physiol. 176, 1410–1422, 10.1104/pp.17.01310 (2017).29233850
38. Su W Gu X Peterson T TIR-Learner, a new ensemble method for TIR transposable element annotation, provides evidence for abundant new transposable elements in the maize genome Mol. Plant 2019 12 447 460 10.1016/j.molp.2019.02.008 30802553
Su, W., Gu, X. & Peterson, T. TIR-Learner, a new ensemble method for TIR transposable element annotation, provides evidence for abundant new transposable elements in the maize genome. Mol. Plant 12, 447–460, 10.1016/j.molp.2019.02.008 (2019).30802553
39. Xiong W He L Lai J Dooner HK Du C HelitronScanner uncovers a large overlooked cache of Helitron transposons in many plant genomes Proc. Natl. Acad. Sci. USA 2014 111 10263 10268 10.1073/pnas.141006811 24982153
Xiong, W., He, L., Lai, J., Dooner, H. K. & Du, C. HelitronScanner uncovers a large overlooked cache of Helitron transposons in many plant genomes. Proc. Natl. Acad. Sci. USA 111, 10263–10268, 10.1073/pnas.141006811 (2014).24982153
40. Chen N Using RepeatMasker to identify repetitive elements in genomic sequences Curr. Protoc. Bioinformatics 2004 5 1 14 10.1002/0471250953.bi0410s25
Chen, N. Using RepeatMasker to identify repetitive elements in genomic sequences. Curr. Protoc. Bioinformatics 5, 1–14, 10.1002/0471250953.bi0410s25 (2004).
41. Jurka J Repbase update, a database of eukaryotic repetitive elements Cytogenet. Genome Res 2005 110 462 467 10.1186/s13100-015-0041-9 16093699
Jurka, J. et al. Repbase update, a database of eukaryotic repetitive elements. Cytogenet. Genome Res. 110, 462–467, 10.1186/s13100-015-0041-9 (2005).16093699
42. Benso G Tandem repeats finder: a program to analyze DNA sequences Nucleic Acids Res. 1999 27 573 580 10.1093/nar/27.2.573 9862982
Benso, G. Tandem repeats finder: a program to analyze DNA sequences. Nucleic Acids Res. 27, 573–580, 10.1093/nar/27.2.573 (1999).9862982
43. Flynn JM RepeatModeler2 for automated genomic discovery of transposable element families Proc. Natl. Acad. Sci. USA 2020 117 9451 9457 10.1073/pnas.1921046117 32300014
Flynn, J. M. et al. RepeatModeler2 for automated genomic discovery of transposable element families. Proc. Natl. Acad. Sci. USA 117, 9451–9457 (2020).32300014
44. Kim D Langmead B Salzberg SL HISAT: a fast spliced aligner with low memory requirements Nat. Methods 2015 12 357 360 10.1038/nmeth.3317 25751142
Kim, D., Langmead, B. & Salzberg, S. L. HISAT: a fast spliced aligner with low memory requirements. Nat. Methods 12, 357–360, 10.1038/nmeth.3317 (2015).25751142
45. Kovaka S Transcriptome assembly from long-read RNA-seq alignments with StringTie2 Genome Biol. 2019 20 1 13 10.1186/s13059-019-1910-1 30606230
Kovaka, S. et al. Transcriptome assembly from long-read RNA-seq alignments with StringTie2. Genome Biol. 20, 1–13, 10.1186/s13059-019-1910-1 (2019).30606230
46. Stanke M AUGUSTUS: ab initio prediction of alternative transcripts Nucleic Acids Res. 2006 34 435 439 10.1093/nar/gkl200
Stanke, M. et al. AUGUSTUS: ab initio prediction of alternative transcripts. Nucleic Acids Res. 34, 435–439, 10.1093/nar/gkl200 (2006).
47. Cantarel BL MAKER: an easy-to-use annotation pipeline designed for emerging model organism genomes Genome Res. 2008 18 188 196 10.1101/gr.6743907 18025269
Cantarel, B. L. et al. MAKER: an easy-to-use annotation pipeline designed for emerging model organism genomes. Genome Res. 18, 188–196, http://www.genome.org/cgi/doi/10.1101/gr.6743907 (2008).18025269
48. Cantalapiedra CP Hernández-Plaza A Letunic I Bork P Huerta-Cepas J eggNOG-mapper v2: functional annotation, orthology assignments, and domain prediction at the metagenomic scale Mol. Biol. Evol. 2021 38 5825 5829 10.1093/molbev/msab293 34597405
Cantalapiedra, C. P., Hernández-Plaza, A., Letunic, I., Bork, P. & Huerta-Cepas, J. eggNOG-mapper v2: functional annotation, orthology assignments, and domain prediction at the metagenomic scale. Mol. Biol. Evol. 38, 5825–5829, 10.1093/molbev/msab293 (2021).34597405
49. Jones P InterProScan 5: genome-scale protein function classification Bioinformatics 2014 30 1236 1240 10.1093/bioinformatics/btu031 24451626
Jones, P. et al. InterProScan 5: genome-scale protein function classification. Bioinformatics 30, 1236–1240, 10.1093/bioinformatics/btu031 (2014).24451626
50. Camacho C BLAST+: architecture and applications BMC Bioinform. 2009 10 421 429, 10.1186/1471-2105-10-421
Camacho, C. et al. BLAST+: architecture and applications. BMC Bioinform. 10, 421–429, 10.1186/1471-2105-10-421 (2009).
51. Finn RD Clements J Eddy SR HMMER web server: interactive sequence similarity searching Nucleic Acids Res. 2011 39 29 37 10.1093/nar/gkr367
Finn, R. D., Clements, J. & Eddy, S. R. HMMER web server: interactive sequence similarity searching. Nucleic Acids Res. 39, 29–37, 10.1093/nar/gkr367 (2011).
52. 2024 NCBI Sequence Read Archive SRR28785292
NCBI Sequence Read Archive. https://identifiers.org/ncbi/insdc.sra:SRR28785292 (2024).
53. 2024 NCBI Sequence Read Archive SRR28789380
NCBI Sequence Read Archive https://identifiers.org/ncbi/insdc.sra:SRR28789380 (2024).
54. 2024 NCBI Sequence Read Archive SRR28788880
NCBI Sequence Read Archive https://identifiers.org/ncbi/insdc.sra:SRR28788880 (2024).
55. 2024 NCBI Sequence Read Archive SRR28787538
NCBI Sequence Read Archive https://identifiers.org/ncbi/insdc.sra:SRR28787538 (2024).
56. 2024 NCBI GenBank JBFDAA000000000
NCBI GenBank https://identifiers.org/ncbi/insdc:JBFDAA000000000 (2024).
57. 2024 NCBI BioProject https://www.ncbi.nlm.nih.gov/bioproject/PRJNA1103718
NCBI BioProject https://www.ncbi.nlm.nih.gov/bioproject/PRJNA1103718 (2024).
58. Simao FA Waterhouse RM Ioannidis P Kriventseva EV Zdobnov EM BUSCO: assessing genome assembly and annotation completeness with single-copy orthologs Bioinformatics 2015 31 3210 3212 10.1093/bioinformatics/btv351 26059717
Simao, F. A., Waterhouse, R. M., Ioannidis, P., Kriventseva, E. V. & Zdobnov, E. M. BUSCO: assessing genome assembly and annotation completeness with single-copy orthologs. Bioinformatics 31, 3210–3212, 10.1093/bioinformatics/btv351 (2015).26059717
