
==== Front
Sci Data
Sci Data
Scientific Data
2052-4463
Nature Publishing Group UK London

3854
10.1038/s41597-024-03854-4
Data Descriptor
A chromosome-level genome assembly of the legume pod borer, Maruca vitrata Fabricius (Lepidoptera: Crambidae)
Wang Jiale 123
Liu Shuai 123
Tang Xuemei 123
Huang Congling 123
Wan Kai wankai202308@163.com

123
1 grid.135769.f 0000 0001 0561 6611 Institute of Quality Standard and Monitoring Technology for Agro- Products of Guangdong Academy of Agricultural Sciences, Guangzhou, Guangdong China
2 grid.484195.5 Guangdong Provincial Key Laboratory of Quality & Safety Risk Assessment for Agro-Products, Guangzhou, Guangdong China
3 grid.418524.e 0000 0004 0369 6250 Key Laboratory of Testing and Evaluation for Agro-Product Safety and Quality of Ministry of Agriculture and Rural Affairs of the People’s Republic of China, Guangzhou, Guangdong China
18 9 2024
18 9 2024
2024
11 101010 5 2024
4 9 2024
© The Author(s) 2024
2024
https://creativecommons.org/licenses/by-nc-nd/4.0/ Open Access This article is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License, which permits any non-commercial use, sharing, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if you modified the licensed material. You do not have permission under this licence to share adapted material derived from this article or parts of it. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by-nc-nd/4.0/.
Maruca vitrata, a significant pest of legumes, impacts food security in Asia and Africa. This study presents a high-quality genome assembly of M. vitrata, utilizing advanced sequencing technologies including Nanopore long-read, MGI short-read, and Hi-C. The genome, totaling 482.3 Mb with a contig N50 of 2.91 Mb, features 41.58% repetitive sequences and encompasses 13,320 protein-coding genes. We performed comparative genomic analyses to affirm the accuracy and completeness of the protein sequences assembled, ensuring the assembly’s integrity. Additionally, the annotation of 83 Cytochrome P450 (CYP) genes further confirms the comprehensive nature of the genome assembly and its annotations. This genome assembly not only deepens our understanding of M. vitrata biology but also supports the development of sustainable pest management strategies. This research highlights the importance of genomics in advancing sustainable agricultural solutions through innovative pest management approaches.

Subject terms

Entomology
Agricultural genetics
Project of Collaborative Innovation Center of GDAAS (XTXM202202).https://doi.org/10.13039/501100001809 National Natural Science Foundation of China (National Science Foundation of China) 32302444 Wang Jiale https://doi.org/10.13039/501100004000 Guangzhou Science and Technology Program key projects 2023A04J0783 Wang Jiale GuangDong Basic and Applied Basic Research Foundation (2022A1515010362, 2023A1515012720), Special Fund for Scientific Innovation Strategy- Construction of High- Level Academy of Agriculture Science (R2021YJ -YB3023).issue-copyright-statement© Springer Nature Limited 2024
==== Body
pmcBackground & Summary

The legume pod borer, Maruca vitrata, poses a critical threat to leguminous crop production worldwide, particularly in Asia and Africa where these crops are crucial dietary protein sources and economic commodities1,2. The lifecycle of this pantropical pest begins when eggs are deposited on the flowers of host plants3; the emerging larvae then penetrate the pods to consume the seeds within. Notably, M. vitrata damages over 70 species within the Fabaceae family1, including essential crops like cowpea (Vigna unguiculata), mung bean (V. radiata), soybean (Glycine max), common bean (Phaseolus vulgaris) and pea (Pisum sativum). The severity of this issue is exemplified by cowpea yield losses reaching up to 72%4, signifying a major threat to food legume production and underscoring the urgent need for effective pest management strategies.

The prevailing control strategies for M. vitrata primarily involve the use of chemical pesticides. However, the efficacy of these interventions is significantly compromised by the pod-penetrative behavior of M. vitrata larvae, which bore into pods to consume the seeds. This adaptation not only renders pest management efforts inadequate but also exacerbates food safety concerns by contributing to substantial pesticide residues. Moreover, the challenge is intensified by the routine application of pesticides close to the harvest period of cowpea, a crop characterized by the concurrent production of fruits and flowers, which further complicates adherence to the Preharvest Interval (PHI). This situation is thereby linked to the difficulty of scheduling pesticide applications effectively. Alarmingly, documented evidence of consistent pesticide misuse in the cultivation of various legumes amplifies these issues5. Data from the Taiwan Ministry of Agriculture’s Agricultural Chemicals Research Institute (2011–2021) indicates legumes had the highest incidence of substandard quality among inspected vegetables (https://www.acri.gov.tw). Given the rising resistance to pesticides in M. vitrata6,7 and significant ecological and health impacts from the overuse of pesticides, genomic research into M. vitrata becomes essential. It offers potential for new pest management strategies by revealing insights into its evolutionary biology, molecular mechanisms.

This study delivers a comprehensive chromosome-level genome assembly of M. vitrata, achieved through Nanopore long-read, MGI short-read sequencing, and Hi-C technologies. The assembly features a genome size of 482.3 Mb size, a contig N50 of 2.91 Mb, and a completeness of 98.83%, demonstrating the high quality of the dataset. Our analysis catalogues 41.58% repeat sequences and identifies 13,320 protein-coding genes, confirming the breadth and depth of genetic information covered. Additionally, 83 Cytochrome P450 (CYP) genes were annotated within the M. vitrata genome, providing essential data for further research into gene functions, particularly those potentially involved in pesticide resistance. This genome assembly not only enriches our understanding of M. vitrata biology but also supports the development of effective and sustainable pest management strategies, crucial for addressing the impact of this pest on legume agriculture worldwide.

Methods

Genome sequencing

M. vitrata specimens, initially collected from cowpea plants located in the experimental fields of the Guangdong Academy of Agricultural Sciences, Guangzhou, China (23°5′ N, 113°32′ E), were bred through successive cycles on a specialized artificial diet in a controlled laboratory environment. This procedure ensured a genetically uniform population for genome sequencing, enhancing data integrity. Genomic DNA was subsequently extracted from newly-pupated female pupae using the Sodium Dodecyl Sulfate (SDS) method, followed by purification with a commercially available genomic DNA extraction kit (QIAGEN®, Cat#13343). The standard operating procedure of manufacturer was strictly followed to ensure optimal DNA quality and yield.

For genome sequencing, two distinct libraries were prepared. The short-read library was constructed by fragmenting DNA to lengths of 200~400 bp using Covaris, and sequenced on the MGISEQ 2000 platform. The long-read library, on the other hand, involved size selection via the BluePippin system and utilized the NEBNext Ultra II kit (New England Biolabs, USA) along with the LSK109 adapter (Oxford Nanopore Technologies, UK) for preparation, followed by sequencing with the Nanopore PromethION system. These processes were conducted at GrandOmics Biosciences Co., Ltd. (Wuhan, China), yielding 30.93 Gb and 64.69 Gb of clean data from the MGI (short-read) and ONT (long-read) libraries, respectively. The Hi-C sequencing library was generated according to the protocol outlined by Belton et al.8. This method involved cross-linking ground female M. vitrata pupal samples, isolating nuclei, and digesting with MboI. DNA fragments were biotin-labeled, sheared to 300~600 bp, blunt-end repaired, A-tailed, and then selected via biotin-streptavidin pull-down. After 12~14 cycles of PCR amplification, the library was sequenced on the MGISEQ 2000 platform.

For transcriptomic analysis, RNA from female M. vitrata pupae was extracted in triplicate using TRIzol® Reagent (TIANGEN, China). This RNA was enriched and fragmented with the Dynabeads mRNA Purification Kit (Cat#61006, Invitrogen) and the MGIEasy RNA Library Prep Kit (Cat# 1000005276, MGI), respectively, to prepare a cDNA library. Quality assessment of the library was performed using the Agilent 2100 Bioanalyzer before circularization with the MGIEasy Circularization Module (CAT#1000005260, MGI). Sequencing was carried out on the DNBSEQ-T7RS platform at GrandOmics Biosciences Co., Ltd. (Wuhan, China).A comprehensive overview of the sequencing data used in this study is provided (Supplementary Table S1).

Genome assembly

The estimated genome size of M. vitrata was determined through 17-mer frequency distribution analysis, utilizing the KMC software9. This analysis provided an estimated genome size of approximately 454.19 Mb. For de novo genome assembly, we constructed an assembly using Oxford Nanopore Technologies (ONT) long reads by employing a string graph method with NextDenovo (v 2.3.1). Considering the inherent error rate of ONT raw reads, we first self-corrected the original subreads using the NextCorrect module, resulting in 30 Gb of Corrected and Normalized Sequence (CNS reads). These CNS reads were then processed with the NextGraph module to analyze correlations among the sequences and produce a preliminary genome assembly. This assembly spanned 543.44 Mb with an N50 of 2.56 Mb. To enhance the accuracy of the assembly, we refined the contigs using a two-step polishing process. Initially, NextPolish (v 1.3.0)10 was employed with ONT long reads and MGI short reads. Specifically, the ONT long reads underwent three rounds of correction using minimap211 to align the reads back to the genome, followed by correction with NextPolish. Subsequently, MGI short reads, filtered using fastp, were used for four additional rounds of polishing with NextPolish, resulting in the final polished genome assembly which had a size of 542.67 Mb and an N50 of 2.55 Mb.

To evaluate the genome assembly, we performed a series of assessments. BUSCO (v3.0.1)12 analysis using the insecta_odb10 database13 indicated that approximately 98.76% of the complete orthologs were present, suggesting a high completeness of the conserved gene set. CEGMA (v 2.5)14 analysis showed that 236 cluster of essential genes (95.16%) were present, with 199 (80.24%) being complete, further supporting the integrity of the assembly. Sequence consistency was confirmed by aligning second-generation sequencing data to the genome using BWA (v 0.7.12-r1039)15, with an alignment rate of 99.34%. SNP and Indel analysis revealed a homozygous SNP count of 23,263 and an Indel count of 15,713, corresponding to a single-base accuracy of 99.99%. GC-Depth analysis using minimap2 indicated a read alignment rate of 99.56% and an average depth of 107.01X, with a genome coverage of 99.97% at 1X depth. To address genome redundancy, likely due to a 2.9% heterozygosity, we employed Purge_Dups to generate a non-redundant genome sequence, resulting in a genome size of 486.28 Mb with an N50 of 2.91 Mb. Post-redundancy BUSCO analysis revealed 98.83% completeness of complete orthologs, and GC-Depth analysis showed consistent GC content and sequencing depth distribution, with improved redundancy metrics.

Hi-C sequencing reads were aligned to the assembled genome employing bowtie2 (v2.3.2)16. The raw reads were filtered to remove low-quality sequences and adapters using fastp (v0.12.6)17, ensuring high-quality clean reads for subsequent analyses. Valid interaction pairs were identified using HiC-Pro (v2.7.8)18, and the scaffolds were organized and oriented into pseudochromosomes using LACHESIS19. This process involved clustering scaffolds based on their interaction frequencies using hierarchical clustering and assigning them to chromosome groups. Specifically, 66,986,033 paired-end reads uniquely mapped to the genome, accounting for 26.17% of the total clean reads. After filtering, 47,049,532 valid interaction pairs were identified, representing 70.24% of the uniquely mapped reads and 18.38% of the total clean reads. A total of 259 scaffolds were clustered into 32 chromosome groups based on interaction frequencies (Fig. 1), constituting 99.04% of the entire genome assembly. Within each group, scaffolds were ordered and oriented to form linear assemblies using a weighted directed acyclic graph (WDAG) approach, resulting in a chromosome-level assembly with a final size of 482.30 Mb, including 35 scaffolds with a scaffold N50 of 16.16 Mb and a contig N50 of 2.91 Mb (Table 1). While the comprehensive genome assembly has been released to NCBI GenBank with the accession number GCA_039566165.1, we have also provided a correlation table in Supplementary Table S2, which aligns the original chromosome names with those renamed in the NCBI database. A Hi-C interaction heatmap was generated to illustrate the interaction of each chromosome (Fig. 2), showing clear chromosome boundaries and strong interactions along the diagonal, indicative of high assembly accuracy. The absence of significant off-diagonal interactions confirmed the lack of contamination and high-quality assembly.Fig. 1 Circos Diagram of the Maruca vitrata Genome. This diagram provides a comprehensive visualization of the structure of M. vitrata genome. Starting from the outermost layer and moving inward, the layers represent: (I) Distribution of markers across 32 chromosomes, denoted at a megabase scale; (II) Gene density per chromosomal region; (III) ncRNA counts; and (IV-VII) Distribution of different repeat elements (DNA repeats, Long Terminal Repeats (LTR), Long Interspersed Nuclear Elements (LINE), and Short Interspersed Nuclear Elements (SINE)) across chromosomes. At the core of the diagram, an illustration of an adult female M. vitrata highlights the subject of this genomic study.

Table 1 Statistics of the Maruca vitrata genome assembly.

Genome assembly	Features	
Genome size	482.30 Mb	
Contig N50	2.91 Mb	
Number of Contigs	262	
Scaffold N50	16.16 Mb	
Number of Scaffolds	35	
Completed BUSCOs	98.76%	

Fig. 2 Heatmap of genome-wide Hi-C data and overview of the genomic landscape of Maruca vitrata. This map identifies 32 linkage groups, with color intensity (yellow for low, red for high) representing the frequency of Hi-C interactions.

Genome annotation

The M. vitrata genome was annotated for tandem repeats (TRs) and transposable elements (TEs) using various software tools. GMATA (v2.2)20 identified 19,803 simple sequence repeats (SSRs) covering 252,778 bp, accounting for 0.05% of the genome. Tandem Repeats Finder (TRF, v4.07b)21 detected 24,791 TRs, totaling 5,030,665 bp, representing 1.04% of the genome. These TRs were soft-masked to reduce conflicts during subsequent TE annotations. For TE identification, both ab initio and homology-based approaches were utilized. MITE-hunter22 identified miniature inverted-repeat transposable elements (MITEs), which were included in a comprehensive TE library used for hard-masking the genome. RepeatModeler (v2.0.1)23 performed de novo repeat identification, and the resulting library, enrich in unknown elements, was classified using TEclass24 against the RepBase database (v16.02)25. RepeatMasker (v4.0.7)26 further mapped sequences to both the de novo repeat library and the RepBase TE library, consolidating overlapping elements. This process identified 200.54 Mb of repetitive sequences, constituting 41.58% of the genome. Specifically, long terminal repeats (LTRs), long interspersed nuclear elements (LINEs), DNA elements (DNAs), and short interspersed nuclear elements (SINEs) accounted for 5.44%, 11.42%, 12.98%, and 3.18% of the genome, respectively (Table 2).Table 2 Annotation of repeat elements in the Maruca vitrata genome.

Type	Number of elements	Length (bp)	Percentage (%)	
LTRs	113,095	26,229,249	5.44	
LINEs	247,427	55,097,390	11.42	
DNAs	355,707	62,599,353	12.98	
SINEs	132,002	15,321,726	3.18	
Rolling-circles	142,681	15,239,798	3.16	
Microsatellites	19,803	232,975	0.05	
MITEs	45,066	8,249,432	1.71	
Simple repeats	2,937	648,985	0.13	
Low complexity	32	4,969	0.00	
Unclassified	110,647	11,778,005	2.44	
Total repeats	1,196,685	200,541,755	41.58	

Gene prediction on the repeat-masked genome utilized three methodologies: reference-guided transcriptome assembly, homology search, and ab initio prediction. RNA-seq-based prediction involved sequencing raw data, filtering using Fastp, and quality control with FastQC. Clean data was aligned to the reference genome using STAR (v2.7.3a)27, and the alignments were processed with StringTie (v1.3.4c)28 to extract transcript coordinates. PASA (v2.3.3)29 was used to predict UTRs and splicing regions, generating 15,785 gene predictions and forming a training set for subsequent predictions. Homolog prediction was performed using GeMoMa (v1.6.1)30, which integrated protein sequences from closely related species with transcriptome data when available, predicting 29,220 genes. De novo prediction was performed using Augustus (v3.1)31, trained with a set of 3,000 high-confidence genes selected from transcriptome predictions. This ab initio method identified 13,130 genes. An integrated gene set was compiled with EVidenceModeler (EVM, v1.1.1)29, combining results from PASA, GeMoMa, and ab initio predictions with default weight settings. TransposonPSI32 was used to remove genes associated with TEs and filter for miscoded genes. InterProScan (v5.32–71.0)33 identified putative domains and GO terms. EVM-integrated protein sequences were compared against four databases (SwissProt, NR, KEGG, KOG) using BLASTp (v2.7.1). This comprehensive approach led to the prediction of 14,481 protein-coding genes. Of these, 13,320 (91.98%) were successfully annotated in at least one database (Table 3).Table 3 Functional annotation of Maruca vitrata proteins.

Type	Number	Percentage (%)	
KOG	8,643	59.69	
KEGG	7,155	49.41	
NR	13,228	91.35	
Swiss-Prot	10,173	70.25	
GO	7,914	54.65	
Total Annotated	13,320	91.98	
Total Gene	14,481	—	

Comparative genomic analysis

Comparative genomic analysis was undertaken to validate the assembly quality of the M. vitrata genome. Protein sequences from 14 Lepidoptera species, including M. vitrata (this study), Amyelois transitella (RefSeq assembly accession: GCF_001186105.1), Bombyx mori (GCF_014905235.1), Chilo suppressalis (GeneBank assembly accession: GCA_902850365.2), Cnaphalocrocis medinalis (InsectBase assembly accession: IBG_00192), Diatraea saccharalis (GCA_918026875.4), Galleria mellonella (GCF_026898425.1), Helicoverpa armigera (GCF_023701775.1), Manduca secta (GCF_014839805.1), Ostrinia furnacalis (GCF_004193835.2), Ostrinia nubilalis (GCA_008921685.1), Pieris rapae (GCF_905147795.1), Plutella xylostella (GCF_932276165.1) and Spodoptera frugiperda (GCF_023101765.2), were analyzed to confirm orthologue and paralogue relationships, ensuring assembly accuracy. Orthologue and paralogue gene groups were identified using OrthoFinder (v2.5.4)34 under default parameters, resulting in a total of 14,657 orthogroups. Among these, 2,534 orthogroups were identified as single-copy gene families. MAFFT (v7.429)35 was employed for protein sequence alignment, with Gblock (v0.91b)36 used to refine these alignments. The aligned sequences were concatenated into a single supergene sequence. Phylogenetic trees were constructed using the Maximum Likelihood (ML) method via IQ-TREE (v1.6.12)37. ModelFinder determined the optimal model (JTT + F + I + G4) for tree construction, with P. xylostella serving as an outgroup to root the phylogenetic tree. The support values for the phylogenetic tree were calculated using 1,000 ultrafast bootstrap replicates, ensuring robust and reliable tree topology. Divergence times were estimated utilizing the MCMCTree program within the PAML package (v4.10)38. Calibration points for divergence time estimation were derived from fossil records, as obtained from the TimeTree website (http://www.timetree.org/). Specifically, these included divergences between S. frugiperda and H. armigera (38.0~60.3 mya)39,40, G. mellonella and A. transitella (71.6~76.4 mya)39,41, and O. nubilalis and P. xylostella (79.8~127.3 mya)39,42, covering a range of evolutionary periods. The evolution of gene families in M. vitrata was examined using Cafe (v4.2.1)43, suggesting significant expansions and contractions in 115 and 12 gene families (p < 0.01), respectively (Fig. 3). To identify potentially fast-evolving genes (FEGs) and positively selected genes (PSGs) in M. vitrata genome, investigations were conducted using the branch model and branch-site model by codeml in PAML (v4.10)38. Protein sequence alignments were initially converted to corresponding coding sequence alignments using pal2nal (v14)44. A dataset that includes 167 potential FEGs and 256 PSGs in M. vitrata is accessible on Zenodo.Fig. 3 Phylogenetic Tree and Comparative Genomics of Maruca vitrata. The left portion displays a ML phylogenetic tree, constructed from 2,534 concatenated single-copy orthologous groups across M. vitrata and 13 other lepidopteran species using IQ-TREE. Plutella xylostella, serving as the outgroup, anchors the base of the tree. Each node is supported by 100% bootstrap confidence. Branch annotations indicate the number of gene families that have expanded (green) or contracted (red). On the right, a summary of gene counts reveals the distribution of orthologous groups across these genomes: “1:1:1” denotes universal single-copy genes found in all species; “N:N:N” represents other universal genes; “Crambidae-specific” highlights genes unique to five Crambidae species (C. medinalis, O. nubilalis, O. furnacalis, D. saccharalis, C. suppressalis); “Species-specific” identifies genes exclusive to one species but present in multiple copies; “Unassigned genes” are unique genes found in only one copy per genome; “Others” encompasses the remaining genes not classified in the above categories.

Annotation of cytochrome P450 genes

The annotation of Cytochrome P450 (CYP) genes in the M. vitrata genome was performed to assess the comprehensiveness of the genome assembly and annotation. Well-annotated CYP protein sequences from species across four orders, including Aedes aegypti, Apis mellifera, Drosophila melanogaster, Solenopsis invicta, Tribolium castaneum, A. transitella, B. mori, M. sexta, O. furnacalis, P. rapae and S. frugiperda, were retrieved from the NCBI RefSeq database and compared against the M. vitrata genome using tBLASTn and BLASTp, ensuring a sensitive detection process. HMMER (v3.3.2)45 confirmed the presence of the CYP domain (PF00067.25), validating the annotation accuracy. Subsequent alignment was conducted using MAFFT (v7.429)46, with trimming performed by trimAl (v1.2.rev59)47 to guarantee the precision of our gene predictions, establishing a robust framework for further genomic studies. A total of 81 potential CYP genes were annotated within the M. vitrata genome.

To achieve a more accurate annotation of the CYP genes, we conducted additional analyses to clarify gene structures and domain architectures. We generated a figure illustrating the domain architectures and a phylogenetic tree of the annotated CYP proteins (Fig. 4), along with a table detailing domain completeness for each (Supplementary Table S4). These resources aid in assessing the presence of multiple identical domains, which might suggest structural complexities such as TRs. Our analysis identified four CYP genes with multiple identical domains, warranting further investigation. To clarify these annotations, we extracted sequences extending ± 5000 base pairs from the annotated genomic regions and re-annotated them using WebAugustus (https://bioinf.uni-greifswald.de/webaugustus). These refinements, thoroughly documented in Supplementary Table S5, provide a more accurate representation of the CYP family in M. vitrata, totaling 83 genes. Additionally, we constructed a phylogenetic tree for further categorize this gene family, using characterized CYP proteins from B. mori48,49 and employing ML methods with the LG + I + G4 model, as recommended by ModelFinder through IQ-TREE (Fig. 5).Fig. 4 Domain Architectures and Phylogenetic Relationships of Annotated Cytochrome P450 (CYP) Genes in Maruca vitrata. The phylogenetic tree (left panel) was constructed using the Maximum Likelihood (ML) method via IQ-TREE (v1.6.12), with ModelFinder determining the optimal model (JTT + F + I + G4) and 1000 bootstrap replicates for support. The middle panel shows motif detection using the MEME Suite, and the right panel presents domain architectures identified via the Conserved Domain Database (CDD). Black asterisks indicate the four potential CYP genes with multiple domains, which require further analysis for re-annotation.

Fig. 5 Phylogenetic Tree of Cytochrome P450 Proteins in Maruca vitrata and Bombyx mori. This ML tree suggests the potential relationships between CYP proteins, with M. vitrata proteins marked with red stars and B. mori with blue dots. CYPs are distributed across four clades: mitochondrial (purple), CYP2 (blue), CYP4 (green), and CYP3 (orange). Bootstrap values from 1,000 replicates are on branches. Branches with background colors indicate potential gene contraction (yellow) or expansion (blue) in M. vitrata genome.

Data Records

The M. vitrata genome sequencing data, utilizing MGI, Nanopore, and Hi-C technologies, are hosted under BioProject PRJNA1092972, with the corresponding SRA accessions SRR2857612150, SRR2857612251, and SRR2857612352. A comprehensive archive including the detailed genome assembly53, alongside gene coding sequences (CDS)54 and protein sequences55, alongside additional datasets such as annotations for TRs and TEs56, functional predictions for 14,481 protein-coding genes57, a dataset of potential FEGs and PSGs58, is available in Zenodo, accessible via DOI 10.5281/zenodo.12747486. Moreover, MGI transcriptome data for M. vitrata female pupae, including three biological replicates, are available under BioProject PRJNA1092342, with SRA accessions SRR2847983659, SRR2847983760, and SRR2847983861. Additionally, the comprehensive genome assembly has been released to NCBI GenBank with the accession number GCA_039566165.162.

Technical Validation

Library quality was assessed using established methods. DNA concentration and fragment size were quantified with the Qubit® 3.0 Fluorometer (Invitrogen, USA) and pulse-field gel electrophoresis (0.7%), respectively. RNA integrity and quantity were evaluated using the Agilent 2100 Bioanalyzer (Agilent, USA). Only high-quality nucleic acids were used for library preparation and sequencing. MGI paired-end sequencing data underwent quality filtering using Fastp (v.0.20.0) with default parameters to remove adapter sequences, low-quality reads (containing Ns, low Phred scores), and potential PCR duplicates. Additionally, 100,000 random reads were screened against the NT database to identify potential contamination. For Nanopore sequencing data, basecalling was performed using Guppy to convert FAST5 signal data to FASTQ format. Reads with a mean quality score below 7 were excluded, resulting in high-quality filtered reads for downstream analyses.

Supplementary information

Supplementary Tables

Supplementary information

The online version contains supplementary material available at 10.1038/s41597-024-03854-4.

Acknowledgements

We are particularly grateful to Dr. Jin Kyo Jung (National Institute of Crop Science, Republic of Korea) and Professor Yonggyun Kim (Andong National University) for their guidance in improving our Maruca vitrata laboratory rearing. This research was supported by the National Natural Science Foundation of China (32302444), GuangDong Basic and Applied Basic Research Foundation (2022A1515010362, 2023A1515012720), Funding by Science and Technology Projects in Guangzhou (2023A04J0783), Special Fund for Scientific Innovation Strategy- Construction of High- Level Academy of Agriculture Science (R2021YJ -YB3023) and Project of Collaborative Innovation Center of GDAAS (XTXM202202).

Author contributions

J.W. and K.W. conceived of this project. J.W. and S.L. participated in the data analysis. C.H. collected the samples. J.W. wrote the manuscript. K.W. and X.T. revised the manuscript. All authors have read, revised, and approved the final manuscript for submission.

Code availability

Genome Assembly Parameters:

KMC: -k17 -ci1 -cs1000000; NextDenovo (v2.3.1): default; NextCorrect module: default; NextGraph module: default; NextPolish (v1.3.0): default; fastp: -n 0 -f 5 -F 5 -t 5 -T 5 –q 20; Guppy: -x map-ont; Minimap2: -x map-ont; CEGMA (v1.0): default; BWA (v0.7.12-r1039): default; Purge_Dups: -f 0.9; bowtie2 (v2.3.2): -end-to-end–very-sensitive -L 30; HiC-Pro (v2.7.8): default; LACHESIS: CLUSTER MIN RE SITES = 100, CLUSTER MAX LINK DENSITY = 2.5, CLUSTER NONINFORMATIVE RATIO = 1.4, ORDER MIN N RES IN TRUNK = 60, ORDER MIN N RES IN SHREDS = 60; BUSCO: -l insecta_odb10 -g genome.

Genome Annotation Parameters:

TRF (v4.07b): 2 7 7 80 10 50 500 -f -d -h –r; MITE-Hunter: -n 20 -P 0.2 -c 3; RepeatModeler (v2.0.1): -engine wublast; TEclass: default; RepeatMasker (v4.0.7): nolow -no_is -gff -norna -engineabblast -lib lib; fastp (v0.12.6): default; STAR (v2.7.3a): --outWigType bedGraph --outSAMtype BAM SortedByCoordinate --outSAMstrandField intronMotif; StringTie (v1.3.4c): default; PASA (v2.3.3): -c alignAssembly.config -C -R –g genome.fasta -T -u trans.fasta –t trans.clean.fasta -f fl.acc --ALIGNERS gmap; GeMoMa: default; Augustus (v3.1): --gff3 = on --hintsfile = hints.gff --extrinsicCfgFile = extrinsic.cfg --allow_hinted_splicesites = gcag,atac–min_intron_len = 30 --softmasking = 1; STAR: --outWigType bedGraph --outSAMtype BAM SortedByCoordinate --outSAMstrandField intronMotif; EVM (v1.1.1): --segmentSize 1000000 --overlapSize 100000; TransposonPSI: prot; InterProScan (v5.32–71.0): default; BUSCO: -l BUSCO_db -m prot; Blastp (v2.7.1): -e 1e-5.

Orthologue and Phylogenetic Analyses Parameters:

MAFFT: --localpair --maxiterate 1000; Gblocks: -b5 = h; IQ-TREE: -bb 1000; PAML: seqtype = 2, burnin = 40000, nsample = 2000000; Cafe: -p 0.01.

FEGs and PSGs Identification Parameters:

Branch Model for FEGs: Model = 0, NSite = 0 (null model); Model = 2, NSite = 0 (alternative model). Branch-Site Model for PSGs: NSsites = 2, Model = 2, Fix_omega = 1, Omega = 1 (null model); NSsites = 2, Model = 2, Fix_omega = 0, Omega > 1 (alternative model).

Gene Family Annotation and Phylogenetic Analysis Parameters:

HMMER: -E 1e-5; MAFFT: --localpair --maxiterate 1000; trimAl: -automated1; IQ-TREE: -bb 1000.

Transcriptome Reanalysis Parameters:

Trimmomatic: ILLUMINACLIP: adapters/TruSeq 3-PE.fa:2:30:10 SLIDINGWINDOW:5:20 LEADING:3 TRAILING:3 MINLEN:36 HEADCROP:13; Cufflinks: --library-type fr-unstranded.

All software parameters not explicitly specified were set to default. Custom scripts used in our study are available in a publicly accessible GitHub repository (http://github.com/jialegz/Scripts_Comparative_Genomics). The scripts, ‘get_longest_pep_for_gffread0.9.9.pl’ and ‘get_super_gene_for_peptree.pl’, are used to extract longest protein sequences from various specie genomes and concatenate aligned protein sequences into a supergene sequence for phylogenetic tree construction, respectively. For identifying FEGs and PSGs, ‘calculate_FEGs.pl’ and ‘calculate_PSG.pl’ generate configuration files and execute the codeml program from the PAML software suite. The Python script ‘SelectFEGs_using_mclResults_FDR_pvalue0.05.py’ identifies FEGs using the likelihood ratio test (LRT) with significance assessed by false discovery rate (FDR) adjusted p-values < 0.05. The detection of PSGs is automated through the Bayes Empirical Bayes (BEB) method, also using FDR-adjusted p-values < 0.05.

Competing interests

The authors declare no competing interests.

Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
==== Refs
References

1. Srinivasan R Tamo M Malini P Emergence of Maruca vitrata as a major pest of food legumes and evolution of management practices in Asia and Africa Annu Rev Entomol 2021 66 141 161 10.1146/annurev-ento-021220-084539 33417822
Srinivasan, R., Tamo, M. & Malini, P. Emergence of Maruca vitrata as a major pest of food legumes and evolution of management practices in Asia and Africa. Annu Rev Entomol 66, 141–161, 10.1146/annurev-ento-021220-084539 (2021).33417822
2. Ba NM The legume pod borer, Maruca vitrata Fabricius (Lepidoptera: Crambidae), an important insect pest of cowpea: a review emphasizing West Africa Int J Trop Insect Sci 2019 39 93 106 10.1007/s42690-019-00024-7
Ba, N. M. et al. The legume pod borer, Maruca vitrata Fabricius (Lepidoptera: Crambidae), an important insect pest of cowpea: a review emphasizing West Africa. Int J Trop Insect Sci 39, 93–106, 10.1007/s42690-019-00024-7 (2019).
3. Ahraed M Miah M Amin M Hossain M Evaluation of some plant materials against pod borer infestation in country bean with reference to flower production Ann Bangladesh Agric 2015 19 71 78
Ahraed, M., Miah, M., Amin, M. & Hossain, M. Evaluation of some plant materials against pod borer infestation in country bean with reference to flower production. Ann Bangladesh Agric 19, 71–78 (2015).
4. Sharma H Bionomics, host plant resistance, and management of the legume pod borer, Maruca vitrata - a review Crop Prot 1998 17 373 386 10.1016/S0261-2194(98)00045-3
Sharma, H. Bionomics, host plant resistance, and management of the legume pod borer, Maruca vitrata - a review. Crop Prot 17, 373–386 (1998).
5. Schreinemachers P Too much to handle? Pesticide dependence of smallholder vegetable farmers in Southeast Asia Sci Total Environ 2017 593 470 477 10.1016/j.scitotenv.2017.03.181 28359998
Schreinemachers, P. et al. Too much to handle? Pesticide dependence of smallholder vegetable farmers in Southeast Asia. Sci Total Environ 593, 470–477 (2017).28359998
6. Sreelakshmi P Paul A Antu M Sheela M George T Assessment of insecticide resistance in field populations of spotted pod borer, Maruca vitrata Fabricius on vegetable cowpea Int J Farm Sci 2015 5 159 164
Sreelakshmi, P., Paul, A., Antu, M., Sheela, M. & George, T. Assessment of insecticide resistance in field populations of spotted pod borer, Maruca vitrata Fabricius on vegetable cowpea. Int J Farm Sci 5, 159–164 (2015).
7. Ekesi S Insecticide resistance in field populations of the legume pod-borer, Maruca vitrata Fabricius (Lepidoptera: Pyralidae), on cowpea, Vigna unguiculata (L.), Walp in Nigeria Int J Pest Manag 1999 45 57 59 10.1080/096708799228058
Ekesi, S. Insecticide resistance in field populations of the legume pod-borer, Maruca vitrata Fabricius (Lepidoptera: Pyralidae), on cowpea, Vigna unguiculata (L.), Walp in Nigeria. Int J Pest Manag 45, 57–59 (1999).
8. Belton J-M Hi-C: A comprehensive technique to capture the conformation of genomes Methods 2012 58 268 276 10.1016/j.ymeth.2012.05.001 22652625
Belton, J.-M. et al. Hi-C: A comprehensive technique to capture the conformation of genomes. Methods 58, 268–276, 10.1016/j.ymeth.2012.05.001 (2012).22652625
9. Deorowicz S Kokot M Grabowski S Debudaj-Grabysz A KMC 2: fast and resource-frugal k-mer counting Bioinformatics 2015 31 1569 1576 10.1093/bioinformatics/btv022 25609798
Deorowicz, S., Kokot, M., Grabowski, S. & Debudaj-Grabysz, A. KMC 2: fast and resource-frugal k-mer counting. Bioinformatics 31, 1569–1576 (2015).25609798
10. Hu J Fan J Sun Z Liu S NextPolish: a fast and efficient genome polishing tool for long-read assembly Bioinformatics 2020 36 2253 2255 10.1093/bioinformatics/btz891 31778144
Hu, J., Fan, J., Sun, Z. & Liu, S. NextPolish: a fast and efficient genome polishing tool for long-read assembly. Bioinformatics 36, 2253–2255 (2020).31778144
11. Li H Minimap and miniasm: fast mapping and de novo assembly for noisy long sequences Bioinformatics 2016 32 2103 2110 10.1093/bioinformatics/btw152 27153593
Li, H. Minimap and miniasm: fast mapping and de novo assembly for noisy long sequences. Bioinformatics 32, 2103–2110 (2016).27153593
12. Simão FA BUSCO: assessing genome assembly and annotation completeness with single-copy orthologs Bioinformatics 2015 31 3210 3212 10.1093/bioinformatics/btv351 26059717
Simão, F. A. et al. BUSCO: assessing genome assembly and annotation completeness with single-copy orthologs. Bioinformatics 31, 3210–3212 (2015).26059717
13. Zdobnov EM OrthoDB in 2020: evolutionary and functional annotations of orthologs Nucleic Acids Res 2021 49 D389 D393 10.1093/nar/gkaa1009 33196836
Zdobnov, E. M. et al. OrthoDB in 2020: evolutionary and functional annotations of orthologs. Nucleic Acids Res 49, D389–D393 (2021).33196836
14. Parra G Bradnam K Korf I CEGMA: a pipeline to accurately annotate core genes in eukaryotic genomes Bioinformatics 2007 23 1061 1067 10.1093/bioinformatics/btm071 17332020
Parra, G., Bradnam, K. & Korf, I. CEGMA: a pipeline to accurately annotate core genes in eukaryotic genomes. Bioinformatics 23, 1061–1067 (2007).17332020
15. Li H Durbin R Fast and accurate long-read alignment with Burrows–Wheeler transform Bioinformatics 2010 26 589 595 10.1093/bioinformatics/btp698 20080505
Li, H. & Durbin, R. Fast and accurate long-read alignment with Burrows–Wheeler transform. Bioinformatics 26, 589–595 (2010).20080505
16. Langmead B Salzberg SL Fast gapped-read alignment with Bowtie 2 Nat Methods 2012 9 357 359 10.1038/nmeth.1923 22388286
Langmead, B. & Salzberg, S. L. Fast gapped-read alignment with Bowtie 2. Nat Methods 9, 357–359, 10.1038/nmeth.1923 (2012).22388286
17. Chen S Zhou Y Chen Y Gu J fastp: an ultra-fast all-in-one FASTQ preprocessor Bioinformatics 2018 34 i884 i890 10.1093/bioinformatics/bty560 30423086
Chen, S., Zhou, Y., Chen, Y. & Gu, J. fastp: an ultra-fast all-in-one FASTQ preprocessor. Bioinformatics 34, i884–i890 (2018).30423086
18. Servant N HiC-Pro: an optimized and flexible pipeline for Hi-C data processing Genome Biol 2015 16 259 10.1186/s13059-015-0831-x 26619908
Servant, N. et al. HiC-Pro: an optimized and flexible pipeline for Hi-C data processing. Genome Biol 16, 259, 10.1186/s13059-015-0831-x (2015).26619908
19. Burton JN Chromosome-scale scaffolding of de novo genome assemblies based on chromatin interactions Nat Biotechnol 2013 31 1119 1125 10.1038/nbt.2727 24185095
Burton, J. N. et al. Chromosome-scale scaffolding of de novo genome assemblies based on chromatin interactions. Nat Biotechnol 31, 1119–1125 (2013).24185095
20. Wang X Wang L GMATA: an integrated software package for genome-scale SSR mining, marker development and viewing Front Plant Sci 2016 7 215951
Wang, X. & Wang, L. GMATA: an integrated software package for genome-scale SSR mining, marker development and viewing. Front Plant Sci 7, 215951 (2016).
21. Benson G Tandem repeats finder: a program to analyze DNA sequences Nucleic Acids Res 1999 27 573 580 10.1093/nar/27.2.573 9862982
Benson, G. Tandem repeats finder: a program to analyze DNA sequences. Nucleic Acids Res 27, 573–580 (1999).9862982
22. Han Y Wessler SR MITE-Hunter: a program for discovering miniature inverted-repeat transposable elements from genomic sequences Nucleic Acids Res 2010 38 e199 e199 10.1093/nar/gkq862 20880995
Han, Y. & Wessler, S. R. MITE-Hunter: a program for discovering miniature inverted-repeat transposable elements from genomic sequences. Nucleic Acids Res 38, e199–e199 (2010).20880995
23. Flynn JM RepeatModeler2 for automated genomic discovery of transposable element families Proc Natl Acad Sci USA 2020 117 9451 9457 10.1073/pnas.1921046117 32300014
Flynn, J. M. et al. RepeatModeler2 for automated genomic discovery of transposable element families. Proc Natl Acad Sci USA 117, 9451–9457, 10.1073/pnas.1921046117 (2020).32300014
24. Abrusán G Grundmann N DeMester L Makalowski W TEclass-a tool for automated classification of unknown eukaryotic transposable elements Bioinformatics 2009 25 1329 1330 10.1093/bioinformatics/btp084 19349283
Abrusán, G., Grundmann, N., DeMester, L. & Makalowski, W. TEclass-a tool for automated classification of unknown eukaryotic transposable elements. Bioinformatics 25, 1329–1330 (2009).19349283
25. Bao W Kojima KK Kohany O Repbase Update, a database of repetitive elements in eukaryotic genomes Mob DNA 2015 6 11 10.1186/s13100-015-0041-9 26045719
Bao, W., Kojima, K. K. & Kohany, O. Repbase Update, a database of repetitive elements in eukaryotic genomes. Mob DNA 6, 11, 10.1186/s13100-015-0041-9 (2015).26045719
26. Tarailo-Graovac M Chen N Using RepeatMasker to identify repetitive elements in genomic sequences Curr Protoc Bioinformatics 2009 Chapter 4 Unit 4.10 10.1002/0471250953.bi0410s25
Tarailo-Graovac, M. & Chen, N. Using RepeatMasker to identify repetitive elements in genomic sequences. Curr Protoc Bioinformatics Chapter 4, Unit 4.10, 10.1002/0471250953.bi0410s25 (2009).
27. Dobin A STAR: ultrafast universal RNA-seq aligner Bioinformatics 2013 29 15 21 10.1093/bioinformatics/bts635 23104886
Dobin, A. et al. STAR: ultrafast universal RNA-seq aligner. Bioinformatics 29, 15–21 (2013).23104886
28. Pertea M StringTie enables improved reconstruction of a transcriptome from RNA-seq reads Nat Biotechnol 2015 33 290 295 10.1038/nbt.3122 25690850
Pertea, M. et al. StringTie enables improved reconstruction of a transcriptome from RNA-seq reads. Nat Biotechnol 33, 290–295, 10.1038/nbt.3122 (2015).25690850
29. Haas BJ Automated eukaryotic gene structure annotation using EVidenceModeler and the Program to Assemble Spliced Alignments Genome Biol 2008 9 1 22 10.1186/gb-2008-9-1-r7
Haas, B. J. et al. Automated eukaryotic gene structure annotation using EVidenceModeler and the Program to Assemble Spliced Alignments. Genome Biol 9, 1–22 (2008).
30. Keilwagen J Using intron position conservation for homology-based gene prediction Nucleic Acids Res 2016 44 e89 e89 10.1093/nar/gkw092 26893356
Keilwagen, J. et al. Using intron position conservation for homology-based gene prediction. Nucleic Acids Res 44, e89–e89 (2016).26893356
31. Stanke M Diekhans M Baertsch R Haussler D Using native and syntenically mapped cDNA alignments to improve de novo gene finding Bioinformatics 2008 24 637 644 10.1093/bioinformatics/btn013 18218656
Stanke, M., Diekhans, M., Baertsch, R. & Haussler, D. Using native and syntenically mapped cDNA alignments to improve de novo gene finding. Bioinformatics 24, 637–644 (2008).18218656
32. Urasaki N Draft genome sequence of bitter gourd (Momordica charantia), a vegetable and medicinal plant in tropical and subtropical regions DNA Res 2017 24 51 58 28028039
Urasaki, N. et al. Draft genome sequence of bitter gourd (Momordica charantia), a vegetable and medicinal plant in tropical and subtropical regions. DNA Res 24, 51–58 (2017).28028039
33. Zdobnov EM Apweiler R InterProScan–an integration platform for the signature-recognition methods in InterPro Bioinformatics 2001 17 847 848 10.1093/bioinformatics/17.9.847 11590104
Zdobnov, E. M. & Apweiler, R. InterProScan–an integration platform for the signature-recognition methods in InterPro. Bioinformatics 17, 847–848 (2001).11590104
34. Emms DM Kelly S OrthoFinder: phylogenetic orthology inference for comparative genomics Genome Biol 2019 20 238 10.1186/s13059-019-1832-y 31727128
Emms, D. M. & Kelly, S. OrthoFinder: phylogenetic orthology inference for comparative genomics. Genome Biol 20, 238, 10.1186/s13059-019-1832-y (2019).31727128
35. Katoh K Standley DM MAFFT multiple sequence alignment software version 7: improvements in performance and usability Mol Biol Evol 2013 30 772 780 10.1093/molbev/mst010 23329690
Katoh, K. & Standley, D. M. MAFFT multiple sequence alignment software version 7: improvements in performance and usability. Mol Biol Evol 30, 772–780 (2013).23329690
36. Castresana J Selection of conserved blocks from multiple alignments for their use in phylogenetic analysis Mol Biol Evol 2000 17 540 552 10.1093/oxfordjournals.molbev.a026334 10742046
Castresana, J. Selection of conserved blocks from multiple alignments for their use in phylogenetic analysis. Mol Biol Evol 17, 540–552 (2000).10742046
37. Nguyen LT Schmidt HA von Haeseler A Minh BQ IQ-TREE: a fast and effective stochastic algorithm for estimating maximum-likelihood phylogenies Mol Biol Evol 2015 32 268 274 10.1093/molbev/msu300 25371430
Nguyen, L. T., Schmidt, H. A., von Haeseler, A. & Minh, B. Q. IQ-TREE: a fast and effective stochastic algorithm for estimating maximum-likelihood phylogenies. Mol Biol Evol 32, 268–274, 10.1093/molbev/msu300 (2015).25371430
38. Yang Z PAML 4: phylogenetic analysis by maximum likelihood Mol Biol Evol 2007 24 1586 1591 10.1093/molbev/msm088 17483113
Yang, Z. PAML 4: phylogenetic analysis by maximum likelihood. Mol Biol Evol 24, 1586–1591, 10.1093/molbev/msm088 (2007).17483113
39. Kawahara AY Phylogenomics reveals the evolutionary timing and pattern of butterflies and moths Proc Natl Acad Sci USA 2019 116 22657 22663 10.1073/pnas.1907847116 31636187
Kawahara, A. Y. et al. Phylogenomics reveals the evolutionary timing and pattern of butterflies and moths. Proc Natl Acad Sci USA 116, 22657–22663 (2019).31636187
40. Wheat CW Wahlberg N Phylogenomic insights into the Cambrian explosion, the colonization of land and the evolution of flight in Arthropoda Syst Biol 2013 62 93 109 10.1093/sysbio/sys074 22949483
Wheat, C. W. & Wahlberg, N. Phylogenomic insights into the Cambrian explosion, the colonization of land and the evolution of flight in Arthropoda. Syst Biol 62, 93–109 (2013).22949483
41. Barber JR Anti-bat ultrasound production in moths is globally and phylogenetically widespread Proc Natl Acad Sci USA 2022 119 e2117485119 10.1073/pnas.2117485119 35704762
Barber, J. R. et al. Anti-bat ultrasound production in moths is globally and phylogenetically widespread. Proc Natl Acad Sci USA 119, e2117485119 (2022).35704762
42. Misof B Phylogenomics resolves the timing and pattern of insect evolution Science 2014 346 763 767 10.1126/science.1257570 25378627
Misof, B. et al. Phylogenomics resolves the timing and pattern of insect evolution. Science 346, 763–767 (2014).25378627
43. De Bie T Cristianini N Demuth JP Hahn MW CAFE: a computational tool for the study of gene family evolution Bioinformatics 2006 22 1269 1271 10.1093/bioinformatics/btl097 16543274
De Bie, T., Cristianini, N., Demuth, J. P. & Hahn, M. W. CAFE: a computational tool for the study of gene family evolution. Bioinformatics 22, 1269–1271 (2006).16543274
44. Zhang Z ParaAT: a parallel tool for constructing multiple protein-coding DNA alignments Biochem Biophys Res Commun 2012 419 779 781 10.1016/j.bbrc.2012.02.101 22390928
Zhang, Z. et al. ParaAT: a parallel tool for constructing multiple protein-coding DNA alignments. Biochem Biophys Res Commun 419, 779–781, 10.1016/j.bbrc.2012.02.101 (2012).22390928
45. Johnson LS Eddy SR Portugaly E Hidden Markov model speed heuristic and iterative HMM search procedure BMC Bioinformatics 2010 11 1 8 10.1186/1471-2105-11-431 20043860
Johnson, L. S., Eddy, S. R. & Portugaly, E. Hidden Markov model speed heuristic and iterative HMM search procedure. BMC Bioinformatics 11, 1–8 (2010).20043860
46. Katoh K Standley DM MAFFT multiple sequence alignment software version 7: improvements in performance and usability Mol Biol Evol 2013 30 772 780 10.1093/molbev/mst010 23329690
Katoh, K. & Standley, D. M. MAFFT multiple sequence alignment software version 7: improvements in performance and usability. Mol Biol Evol 30, 772–780, 10.1093/molbev/mst010 (2013).23329690
47. Capella-Gutierrez S Silla-Martinez JM Gabaldon T trimAl: a tool for automated alignment trimming in large-scale phylogenetic analyses Bioinformatics 2009 25 1972 1973 10.1093/bioinformatics/btp348 19505945
Capella-Gutierrez, S., Silla-Martinez, J. M. & Gabaldon, T. trimAl: a tool for automated alignment trimming in large-scale phylogenetic analyses. Bioinformatics 25, 1972–1973, 10.1093/bioinformatics/btp348 (2009).19505945
48. Ai J Genome-wide analysis of cytochrome P450 monooxygenase genes in the silkworm, Bombyx mori Gene 2011 480 42 50 10.1016/j.gene.2011.03.002 21440608
Ai, J. et al. Genome-wide analysis of cytochrome P450 monooxygenase genes in the silkworm, Bombyx mori. Gene 480, 42–50 (2011).21440608
49. Kawamoto M High-quality genome assembly of the silkworm, Bombyx mori Insect Biochem Mol Biol 2019 107 53 62 10.1016/j.ibmb.2019.02.002 30802494
Kawamoto, M. et al. High-quality genome assembly of the silkworm, Bombyx mori. Insect Biochem Mol Biol 107, 53–62 (2019).30802494
50. 2024 NCBI Sequence Read Archive SRR28576121
NCBI Sequence Read Archive https://identifiers.org/ncbi/insdc.sra:SRR28576121 (2024).
51. 2024 NCBI Sequence Read Archive SRR28576122
NCBI Sequence Read Archive https://identifiers.org/ncbi/insdc.sra:SRR28576122 (2024).
52. 2024 NCBI Sequence Read Archive SRR28576123
NCBI Sequence Read Archive https://identifiers.org/ncbi/insdc.sra:SRR28576123 (2024).
53. Wang JL 2024 Mvit.genome.gff Zenodo https://zenodo.org/records/12747486
Wang, J. L. Mvit.genome.gff. Zenodo https://zenodo.org/records/12747486 (2024).
54. Wang JL 2024 Mvit.genome.cds.fasta Zenodo https://zenodo.org/records/12747486
Wang, J. L. Mvit.genome.cds.fasta. Zenodo https://zenodo.org/records/12747486 (2024).
55. Wang JL 2024 Mvit.genome.pep.fasta Zenodo https://zenodo.org/records/12747486
Wang, J. L. Mvit.genome.pep.fasta. Zenodo https://zenodo.org/records/12747486 (2024).
56. Wang JL 2024 Mvit.genome.repeats.gff Zenodo https://zenodo.org/records/12747486
Wang, J. L. Mvit.genome.repeats.gff. Zenodo https://zenodo.org/records/12747486 (2024).
57. Wang JL 2024 Annotation_Summary.xlsx Zenodo https://zenodo.org/records/12747486
Wang, J. L. Annotation_Summary.xlsx. Zenodo https://zenodo.org/records/12747486 (2024).
58. Wang JL 2024 FEGs_PSGs_compared_Table.xlsx Zenodo https://zenodo.org/records/12747486
Wang, J. L. FEGs_PSGs_compared_Table.xlsx. Zenodo https://zenodo.org/records/12747486 (2024).
59. 2024 NCBI Sequence Read Archive SRR28479836
NCBI Sequence Read Archive https://identifiers.org/ncbi/insdc.sra:SRR28479836 (2024).
60. 2024 NCBI Sequence Read Archive SRR28479837
NCBI Sequence Read Archive https://identifiers.org/ncbi/insdc.sra:SRR28479837 (2024).
61. 2024 NCBI Sequence Read Archive SRR28479838
NCBI Sequence Read Archive https://identifiers.org/ncbi/insdc.sra:SRR28479838 (2024).
62. 2024 NCBI GenBank GCA_039566165.1
NCBI GenBank https://identifiers.org/ncbi/insdc.gca:GCA_039566165.1 (2024).
