
==== Front
Sci Data
Sci Data
Scientific Data
2052-4463
Nature Publishing Group UK London

39231989
3794
10.1038/s41597-024-03794-z
Data Descriptor
Two near-chromosomal-level genomes of globally-distributed Macroascomycete based on single-molecule fluorescence and Hi-C methods
Liu Wei 1
Shi Xiaofei 1
Cai Yingli 2
Sun Wenhua 3
He Peixin 3
Perez-Moreno Jesus 4
Liu Dong liudongc@mail.kib.ac.cn

1
Yu Fuqiang fqyu@mail.kib.ac.cn

1
1 grid.9227.e 0000000119573309 The Germplasm Bank of Wild Species & Yunnan Key Laboratory for Fungal Diversity and Green Development, Kunming Institute of Botany, Chinese Academy of Sciences, Kunming, 650201 China
2 https://ror.org/02z2d6373 grid.410732.3 0000 0004 1799 1111 Institute of Agro-products Processing, Yunnan Academy of Agricultural Sciences, Kunming, 650221 China
3 https://ror.org/05fwr8z16 grid.413080.e 0000 0001 0476 2801 College of Food and Biological Engineering, Zhengzhou University of Light Industry, Zhengzhou, 450002 China
4 https://ror.org/00qfnf017 grid.418752.d 0000 0004 1795 9752 Edafología, Campus Montecillo, Colegio de Postgraduados, Texcoco, 56230 Mexico
4 9 2024
4 9 2024
2024
11 96430 4 2024
16 8 2024
© The Author(s) 2024
2024
https://creativecommons.org/licenses/by-nc-nd/4.0/ Open Access This article is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License, which permits any non-commercial use, sharing, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if you modified the licensed material. You do not have permission under this licence to share adapted material derived from this article or parts of it. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by-nc-nd/4.0/.
Discinaceae holds significant importance within the Pezizales, representing a prominent group of macroascomycetes distributed globally. However, there is a dearth of genomic studies focusing on this family, resulting in gaps in our understanding of its evolution, development, and ecology. Here we utilized state-of-the-art genome assembly methodologies, incorporating third-generation single-molecule fluorescence and Hi-C-assisted methods, to elucidate the genomic landscapes of Gyromitra esculenta and Paragyromitra xinjiangensis. The genome sizes of two species were determined to be 47.10 Mb and 48.20 Mb, with 23 and 22 scaffolds, respectively. 10,438 and 11,469 coding proteins were identified, with functional annotations encompassing over 96.47% and 94.40%, respectively. Assessment of completeness using BUSCO revealed that 98.71% and 98.89% of the conserved proteins were identified. The application of comparative genomic technology has helped in identifying traits associated with of heterothallic life cycle traits and elucidating unique patterns of chromosomal evolution. Additionally, we identified potential saprotrophic nutritional modes and systematic phylogenetic relationships between the two species. Therefore, this study provides crucial genomic insights into the evolution, nutritional type, and ecological roles of species within the Pezizales.

Subject terms

Fungal genomics
Forestry
issue-copyright-statement© Springer Nature Limited 2024
==== Body
pmcBackground & Summary

Discinaceae comprises a widely distributed group of macroascomycetes, predominantly found in the Northern hemisphere1. Originally, Discinaceae was thought to encompass two genera: Gyromitra characterized by distinct discoid, cerebriform, or saddle-shaped ascomata, and Hydnotrya exhibiting globose or tuberous ascomata1. The higher-level taxonomic classification of Gyromitra and Hydnotrya was uncertain prior to the advent of molecular evidence1,2. Recent phylogenetic analyses, utilizing three gene partitions (internal transcribed spacer [ITS], large subunit ribosomal DNA [LSU], and translation elongation factor [TEF]), have refined the taxonomy of the Discinaceae, delineating it into eight genera, comprising seven epigeous genera and one hypogeous genus, encompassing a total of 43 species1. Gyromitra, the most prominent genus within the Discinaceae, exhibits a global distribution and contains species of significant commercial interest. For example, G. esculenta (Pers.) Fr. is commonly traded and consumed in countries such as Estonia, Finland, Poland, Sweden, Mexico, and North America1,3–5. Nonetheless, species like G. esculenta and G. gigas are often approached with caution due to their association with gyromitrin, a notorious toxin that metabolizes into the toxic monomethylhydrazine, leading to various adverse effects, including vomiting, bloody diarrhea, abdominal pain, dehydration, sweating, headaches, muscle pain, dizziness, delirium, hemolysis, liver damage, and kidney failure6–8. Given its heat-volatile nature, gyromitrin requires thorough cooking, repeated boiling, or sun-drying to render it safe for consumption8,9. Paragyromitra, a close relative of Gyromitra, was formerly classified as Gyromitra xinjiangensis and is widely distributed in mainland China. Despite its prevalence, research on its toxicity, physiological ecology, and evolution remains limited1,10.

Macroascomycetes within Pezizales, encompassing species from families such as Tuberaceae11, Morchellaceae12,13, Helvellaceae14, Pyronemataceae15, and Pezizaceae16, have been extensively investigated to elucidate their nutritional type, life cycle, population structure, evolution, and multicellular development. However, the family Discinaceae remains largely unexplored beyond basic taxonomic and toxicological studies. This lack of comprehensive research limits our understanding of the evolution of Pezizales fruiting body morphology, toxin formation, and differentiation of epigeous and hypogeous species. Genomic research is the most effective approach to addressing these deficiencies. Despite the preliminary draft genomes of G. infula (45.88 Mb, 51 contigs, and 51 scaffolds)7 submitted under the US Department of Energy Joint Genome Institute (JGI) initiative17, and G. esculenta (45.05 Mb, 774 contigs, and 271 scaffolds)18, high-quality genomes for species within Discinaceae have yet to be reported. This knowledge gap hinders our comprehensive understanding of the growth, development, evolution, nutrient requirements and phylogeny of Discinaceae. One plausible explanation for the limited research on Discinaceae may be the difficulties associated with acquiring pure culture strains. Investigations into pharmacological activity19, toxicity8,20, trophic21, developmental processes22,23, and phylogenetic relationships1 have predominantly relied on wild fruiting body specimens. This reliance may be attributed to the mycorrhizal symbiotic lifestyle of related species, which presents challenges in obtaining suitable research materials22.

In this study, we initially acquired pure culture materials of two Discinaceae species, G. esculenta and P. xinjiangensis, employing advanced tissue isolation techniques. Subsequently, we conducted de novo genome assembly for both species using single-molecule fluorescence sequencing. The high-error-rate bases were refined through two rounds of second-generation DNBSEQ sequencing, and with the aid of high-throughput Chromosome Conformation Capture Sequencing (Hi-C), we achieved pseudo-chromosomal level genome assemblies for both species. The genome sizes were determined to be 47.10 Mb and 48.20 Mb, with Scaffold N50 values of 1.97 Mb and 2.10 Mb, and scaffold numbers of 23 and 22, respectively. For gene prediction, we employed a comprehensive approach integrating RNA-seq alignment, homology prediction, and ab initio prediction, resulting in predicted gene counts of 10,438 and 11,469 for the two species, with 96.47% and 94.40% of them functionally annotated by public database, respectively. Evaluation of completeness using the Benchmarking Universal Single-Copy Orthology (BUSCO) revealed completeness scores of 98.7% and 98.9%, indicating the near-perfect reference genomes and gene structures obtained for both species. Collinearity analysis demonstrated high similarity between the genes of the two species, albeit each exhibited substantial rearrangement on two respective chromosomes. Notably, the chromosome 23 of G. esculenta displayed independent evolution and lacked similarity to P. xinjiangensis. Mating-type locus analysis revealed the presence of only the Mat1-2-1 gene in both genomes, suggesting a heterothallic life cycle. Characterization of Carbohydrate activity enzymes (CAZyme) and plant cell wall–degrading enzymes (PCWDE) families provided insights into the potential saprotrophic nutritional types of the two species. Based on these findings, evolutionary analysis at the family level was conducted for species within Pezizales (Fig. 1). This study provides valuable information that support novel research insights into the growth, development, molecular markers, and systematic evolution of the Discinaceae.Fig. 1 The flowchart of the technical operations and three sub-branching analyses relevant to near-chromosomal-level genome.

Methods

Collection, isolation, identification, and mycelial culture of samples

The samples of G. esculenta No.21701 and P. xinjiangensis No.20138 were collected from Yichun City, Heilongjiang Province (47°42.1917′N, 128°53.1081′E, 293.8 m a.s.l., August 2, 2022), and Shangri-La City, Yunnan Province (27°55.5770′N, 99°36.9540′E, 3562.2 m a.s.l., August 20, 2021), respectively (Fig. 2a). Upon collection, fresh specimens were air-dried in a well-ventilated lab and stored for long-term preservation at 4 °C. Specimen identification was performed utilizing a combination of morphological characteristics and multi-gene phylogenetic (ITS-lsu-tef) analysis1 as shown in Fig. 2a. Tissue isolation was conducted using modified PDA medium (200 g boiled potato extract, 15 g wheat bran extract, 20 g glucose, 20 g agar powder, distilled water 1 L, pH 7.0–7.5). Approximately 0.5 × 0.5 cm2 tissue sample from the middle part of the ascomata was used for isolation. The tissue was submerged in sterile distilled water for 3–5 minutes and vigorously shaken to remove surface contaminants or intercellular microorganisms. The cleaned tissue sections were then cut into 1–2 mm3 pieces, and placed on the PDA medium for incubation at 18 °C. After approximately 3–4 days, mycelium colonies became visible, initially displaying thin, sparse, creeping growth. Once the colony diameter reached 1–2 cm, the leading edge of the colony was carefully transferred to a new PDA plate for purification. Fungal mycelium cultivation was subsequently performed in liquid PDA medium (agar-free) and statically incubated at 18 °C for 10 days. Following incubation, the fungal mycelium was rinsed with sterile water 2–3 times, and excess moisture was removed from the mycelium using sterile absorbent paper. The mycelium was then stored at −80 °C for subsequent DNA and RNA extraction.Fig. 2 Multi-gene phylogenetic tree, morphological characteristics, genome size assessment and heat maps of the interaction between genome chromosome of G. esculenta and P. xinjiangensis. (a) Multi-gene phylogenetic tree for Gyromitra and Paragyromitra based on combined ITS, LSU, and TEF sequences; the right panel depicts the natural morphological characteristics of the two species. (b–e) Genome size, sequencing depth and heat map of the interaction between genome chromosome of G. esculenta and P. xinjiangensis (shown in b and c; and d and e, respectively).

DNA extraction, sequencing and data filter

Genomic DNA extraction was conducted using the CTAB method, followed by quality control assessment24. Subsequently, genomic sequencing libraries were prepared and sequenced utilizing the DNBSEQ-G400 platform (MGI Tech, Shenzhen, China). For each species, two paired-end (PE150) sequencing libraries with insert sizes of 270 and 350 bp were generated. After sequencing, G. esculenta produced approximately 19.72 M and 19.43 M reads, while P. xinjiangensis produced approximately 14.79 M and 14.88 M reads. Following trimmomatic (v0.33) processing to eliminate low-quality bases and quality control filtering25, G. esculenta provided approximately 6.28 Gb and 6.43 Gb of clean data, while P. xinjiangensis yielded approximately 4.82 Gb and 4.85 Gb of clean data. The second-generation DNA sequencing data were primarily utilized for the initial assessment of genome size and subsequent refinement the third-generation sequencing assembly. Genome size estimation and evaluation based on Kmer frequency analysis indicated genome sizes of approximately 50.83 Mb and 59.34 Mb for the two species (Fig. 2b,d), respectively, both being haploid genomes, with sequencing depths reaching approximately 159× and 113×.

Single Molecule Fluorescent Nanopore Sequencing technology was employed using the PromethION sequencing platform from Oxford Nanopore Technologies, primarily employed for de novo genome assembly. G. esculenta yielded approximately 0.78 M reads with a total base count of 5.61 Gb, N50 counts of 189.9 kb, and an N50 length of 9.64 kb, resulting in a sequencing depth of approximately 112 × relative to the genome size. P. xinjiangensis generated approximately 1.50 M reads with a total base count of 8.12 Gb, N50 counts of 133.1 kb, and an N50 length of 15.7 kb, resulting in a sequencing depth of approximately 158 × relative to the genome size. The sequencing data from Nanopore technology for both species met the standard requirements for conventional de novo genome assembly.

RNA extraction, sequencing and data filter

The juvenile mycelium, cultured in PDA liquid medium for 7 days, underwent formaldehyde fixation and cleavage with the Sau3AI enzyme. The resulting cleaved fragments underwent repair, biotin labeling, and ligation with T4 DNA ligase to facilitate the formation of cross-linked fragments. Following the library construction sequencing requirements for DNBSEQ’s, a 450 bp insertion fragment library was established for PE150 dual-end sequencing, generating a total of 30 Gb of raw data. The post-sequencing assessment using Hic-pro (3.1.0) software revealed that G. esculenta obtained a total of 22.23 M valid interaction reads, with 13.41 M unique paired alignments, accounting for 60.31% of the total reads. Similarly, P. xinjiangensis obtained a total of 17.76 M valid interaction reads, with 12.94 M unique paired alignments after deduplication, accounting for 72.86% of the total reads. These outcomes fulfill the prerequisites for subsequent chromatin conformation analyses.

RNA extraction and purification were conducted employing the TRIzol® Reagent (Invitrogen) extraction kit and the Plant RNA Purification Reagent (Invitrogen) purification kit, following the manufacturer’s protocols. Subsequently, RNA library preparation was performed using the TruSeq Stranded mRNA Library Prep Kit (Illumina, USA) according to the provided instructions. For both species, PE150 paired-end sequencing was conducted with an insertion fragment length of 270 bp. G. esculenta yielded 13.75 M reads, resulting in 4.35 Gb of clean data after Trimmomatic quality control25. Similarly, P. xinjiangensis yielded 23.73 M reads, resulting in 7.78 Gb of clean data after Trimmomatic quality control. The RNA sequencing data primarily serve for gene structure prediction and annotation purposes.

Genome assembly and comparative analysis

i) For Genomic assembly, Hi-C mounting

The schematic overview of the chromosome-level genome assembly and annotation pipelines is illustrated in Fig. 1. NextDenovo (v2.5.0) was employed for de novo genome assembly utilizing third-generation Nanopore data26. NextPolish (v1.4.1) was applied to refine the raw genome using quality-controlled second-generation sequencing data from each species, culminating in high-quality genome assemblies27. Alignment with conserved proteins from the mitochondrial genomes of Tuber calosporum and Morchella importuna facilitated the exclusion of mitochondrial sequences, yielding the final nuclear genome sequences. Hi-C genome-assisted assembly entailed aligning and establishing interaction relationships with sequencing data utilizing Juice (v1.6) software28. Chromosome scaffolding was accomplished via 3D-DNA29, with visualization facilitated by Juicebox and manual chromosome annotation30.

ii) For Gene structure and function annotation

Gene prediction and structural annotation were executed through a comprehensive approach integrating homologous alignment, RNA-seq data, and ab initio prediction methodologies. Initially, RepeatMask (v4.1.2) and RepeatModeler (v2.0.3) were enlisted to detect and mask repetitive sequences within the genome31,32. Subsequently, RNA-seq data underwent alignment to the genome utilizing Hisat233, followed by protein structure prediction encoded by the genes using TransDecoder (v5.5.0). GeneWise (v2.4.1)34 was then deployed to predict genes by leveraging homologous proteins from closely related species: T. melanosporum11, M. importuna35 and Pyronema confluens15. Based on the acquired results, an accurate and complete gene model was chosen, and gene prediction was executed using Augustus (v3.4.0)36. Finally, predictions from the three methodologies were integrated, and the PFAM database was utilized for filtering to yield the final prediction outcome37. Gene annotation was conducted utilizing a spectrum of databases, including Nr, eggNOG, InterPro, GO, SwissProt, KOG, KEGG, and Pfam.

iii) Genomic collinearity, Mating type gene structure and heterothallic life cycle

Utilizing JCVI software38, we conducted a thorough synteny analysis between the two species based on their protein sequences, revealing a notably high degree of collinearity (Fig. 3). We conducted blast comparisons utilizing conserved mating-type loci encoding proteins (APN2, Mat1-1-1, Mat1-2-1, SLA2, MBA1) from closely related species Verpa bohemia, V. conica39 and M. importuna35 against G. esculenta and P. xinjiangensis. By integrating these results with functional annotations, we delineated the mating-type loci of both species.Fig. 3 Global and local genomic collinearity of G. esculenta and P. xinjiangensis, and annotation and density analysis of chromosome 23 of G. esculenta. (a) Synteny analysis based on the protein sequence of G. esculenta and P. xinjiangensis. The chromosome diagram shown above in pink color (above) represents G. esculenta, while that in green color (below) represents P. xinjiangensis. (b–f) Corroborated interaction heatmap of complete chromosome structure of chromosomes 6, 12 and 23 (b–d) in G. esculenta (b–d); and chromosomes 16 and 15 (e,f) in P. xinjiangensis; the species annotated information for proteins in chromosome 23 of G. esculenta is shown in (f); (g,h) Display of the best species matching result of Nr functional annotation of proteins encoded on scaffold 23 of G. esculenta (g) and the protein density distribution on the syntenic chromosomes of the two species, with sorting P. xinjiangensis using the chromosome order of G. esculenta as a framework (h).

iv) Identification and characterization of repetitive sequences

A combination of ab initio and homology-based methodologies was employed to discern repetitive sequences within the genomes. Initially, LTR_FINDER (v1.05)40, RepeatScout (v1.05)41, and RepeatModeler (v2.0.1)32 were utilized to generate an ab initio library of repeat sequences. Subsequently, RepeatMasker (v4.1.0)31 was engaged to identify both known and novel repetitive elements by aligning sequences against the ab initio repeat library and the Repbase (v.19.06) database42.

v) Annotation of non-coding RNA

High-confidence tRNAs were predicted using tRNAscan-SE (v1.3.1)43, while rRNA prediction was conducted using rnammer-1.244. Using infernal-1.1.3 software45 and referencing the Rfam (v12.2) database46, miRNA and snRNA were sought.

vi) Carbohydrate activity enzymes (CAZymes)

Fungi possess a diverse repertoire of CAZymes essential for the degradation of plant polysaccharides, facilitating infection and nutrient acquisition47,48. Comparative genomic analysis of CAZymes serves as a fundamental approach to elucidate species-specific nutritional strategies48. Employing dbCAN v6.0 software49 and the CAZyme database50, we retrieved protein sequences of all CAZymes and extracted HMM structural domain information for each CAZyme family using HMM algorithms51. By applying the BLASTP method against the protein sequence database and the Hmmscan method matching against the HMM database in dbCAN, we obtained precise results by intersecting both approaches. Thresholds for each gene family were calculated from both methods to refine the outcomes49. The PCWDE is an important enzyme system secreted by fungi for degrading plant cell wall components, and is used to systematically analyze the nutritional types of fungal species. Relevant information was extracted from the CAZYme annotation results for downstream analysis.

vii) Comparative genomic and phylogenomic

We compiled genome sequences from 24 publicly available species within Pezizales, complemented by sequences from 3 Orbiliomycetes. Additionally, we included representatives from Lecanoromycetes (Graphis scripta)52, Eurotiomycetes (Aspergillus niger)53, Dothideomycetes (Alternaria alternata)54, and Leotiomycetes (Sclerotinia sclerotiorum)55 as outgroup taxa to construct a comprehensive species tree. Initially, we identified 1393 single-copy orthologous proteins using OrthoFinder (v2.5.4)56. Subsequently, multiple sequence alignment was conducted using MAFFT (v7.158b)57, followed by the extraction of conserved sites employing Gblocks (v0.91b)58. This process resulted in a total sequence alignment length of 1.38 Mb. For the construction of the phylogenetic tree, we employed the maximum likelihood (ML) method with the GAMMA model of rate heterogeneity and the GTR substitution matrix. RAxML (v8.1.24) and IQtree (v2.2.0)59 were both utilized, yielding similar tree topologies. The evolutionary relationships were visualized using FigTree (v1.4.3)60. Divergence time analysis was conducted using the mcmctree module in the PAML (v4.9i) program61, employing repetitive approximate likelihood analysis. Given the absence of fossil records for divergence time calibration within the studied population and the limited research on Pezizales, we adopted widely accepted divergence time nodes from existing literature. Specifically, we utilized divergence times of 233.8–367.0 million years ago (MYA) between S. sclerotiorum and A. niger, and 153.8–172.9 MYA between M. dunensis and T. melanosporum, obtained from Timetree62. To assess the contraction and expansion of gene families, CAFE software (v5.1.0)63 was utilized.

Genomic assembly and Hi-C mounting

The assembly results disclosed a genome size of 47.11 Mb with 24 contigs and 23 scaffolds for G. esculenta and 48.13 Mb with 22 contigs and 22 scaffolds for P. xinjiangensis (Table 1). Interaction heatmaps indicated that the de novo assembly had achieved near chromosome-level resolution, with Hi-C-assisted assembly primarily refining the ends of small contigs and aiding in contig merging (Fig. 2c,e). The scaffold N50 stands at 1.97 Mb and 2.10 Mb, with a GC content of 46.29% and 46.93% for G. esculenta and P. xinjiangensis, respectively.Table 1 Genomic and structural information of G. esculenta and P. xinjiangensis.

	G. esculenta No.21701	P. xinjiangensis No.21138	
the genome scaffolds number	23	22	
the genome contigs number	24	22	
the longest length (bp)	4207259	4206640	
the shortest length (bp)	904152	1271118	
the genome scaffolds size (bp)	47100447	48197954	
the genome contig size (bp)	47100247	48197954	
the rate of N	4.25E-06	0	
the percentage of GC (%)	46.39	46.93	
the scaffold N50 (bp)	1970509	2103843	
the scaffold L50	9	9	
tRNA numbers	253	255	
Av. tRNA length (bp)	85	87	
18/28 s tandem repetition times	8	12	
18/28 s tandem repetition length	69660	87009	

Gene structure and function annotation

For G. esculenta, gene prediction efforts yielded 10,438 coding proteins, with an average gene length of 2062.77 bp. Remarkably, approximately 96.47% of the coding proteins were successfully annotated in public databases. The completeness evaluation by BUSCO achieved a score of 98.7% (Table 2). Conversely, for P. xinjiangensis, 11,469 coding proteins were predicted, with an average gene length of 2027.62 bp. Moreover, approximately 94.40% of the coding proteins were annotated, and the BUSCO completeness score reached 98.9% (Table 2). These outcomes affirm the attainment of chromosome- or pseudochromosome-level genomes for both species, coupled with gene models of relatively high quality.Table 2 Encoding proteins, functional annotation, and completeness evaluation.

	G. esculenta No.21701	P. xinjiangensis No.21138	
Proteins Number (n)	10438	11469	
Nr	8943	9443	
InterPro5	9500	10142	
eggNOG	8041	8384	
swiss-prot	5864	6006	
Pfam	6876	7147	
KOG	3706	3827	
GO	6327	6551	
KEGG	3888	3903	
All annotated gene number	10070	10827	
All annotated gene ratio	96.5	94.4	
Busco completeness (%)	98.7	98.9	

Genomic collinearity

A substantial segmental rearrangement was evident at one terminus of chromosome 6 and 12 in G. esculenta and chromosome 16 and 15 in P. xinjiangensis, as corroborated by the chromosomal interaction heatmap (Fig. 3b–c, e–f). Notably, chromosome 23 in G. esculenta spanned 0.904 Mb and encompassed 71 coding proteins. This chromosome displayed a significantly elevated gene density compared to the mean of the preceding 22 chromosomes, reaching 12.7 kb per gene (Fig. 3h). Chromosome 23 exhibited no homology with the initial 22 chromosomes of G. esculenta or with the genome of P. xinjiangensis (Fig. 3a). Annotation of the 71 coding proteins revealed that 31 proteins could be accurately annotated, with the top matching species being closely related species within the Pezizales (Fig. 3g). Sequence analysis further confirmed the authenticity of this scaffold as a distinctive and newly evolved complete chromosome, evidenced by the presence of a (TTTAGGG)3 telomere repeat at both termini.

Mating type gene structure and heterothallic life cycle

Notably, substantial disparities in gene structure were observed compared to closely related species in Morchellaceae, such as V. bohemia, V. conica and M. importuna (Fig. 4). These differences suggest conspicuous gene rearrangements and inversions. Specifically, in G. esculenta and P. xinjiangensis, the APN2 and SLA2 genes exhibit congruent orientation. In contrast, in V. bohemia, V. conica and M. importuna, these genes are positioned in opposite or reversed directions. This trend also extends to Mat1-1-10 and Mat1-1-11. Unlike in V. bohemia and M. importuna, where the MAT genes are contiguous, in G. esculenta and P. xinjiangensis, they are conserved (UBX and SUPF1) and unknown (Hyp) proteins between Mat1-2-1 and Mat1-1-10/1135,39. Both G. esculenta and P. xinjiangensis exhibit similar gene architectures, encompassing only one independent Mat1-2-1 gene within this mating-type locus. Comprehensive alignment of conserved proteins and SPAdes (v3.15.3)64 assembly utilizing second-generation DNA short sequence data, confirms that both species lack the Mat1-1-1 gene. This absence, suggests a propensity towards heterothallic mating characteristics of G. esculenta and P. xinjiangensis.Fig. 4 Collinearity analysis of mating type genomic loci of G. esculenta and P. xinjiangensis.

Identification and characterization of repetitive sequences

In G. esculenta, a total of 7.15 Mb of repetitive sequences were identified, constituting 15.17% of the genome’s total length (Table 3). These repetitive sequences comprise various types, predominantly composed of transposable elements such as LTRs (3.46 Mb, 7.34%), LINEs (0.71 Mb, 1.49%), and SINEs (2.3 kb, 0.00%), collectively representing 8.85% of the genome sequence. Additionally, interspersed repeats of unknown function accounted for 4.69% of the genome (Table 3). In contrast, P. xinjiangensis exhibited different proportions and types of repetitive sequences, with a total length of 6.40 Mb, constituting 13.28% of the genome (Table 3). Within the interspersed repeats, while retroelements were predominant, the proportion of LINEs slightly exceeded that of LTRs, comprising 3.80% (1.83 Mb) and 2.59% (1.25 Mb), respectively. The proportions of non-interspersed repeats were similar between the two species, accounting for approximately 1.30% and 1.25% of the genomes, respectively (Table 3).Table 3 Repetitive sequences and transposon subtypes of G. esculenta and P. xinjiangensis.

	G. esculenta No.21701	P. xinjiangensis No.21138	
Length (bp)	ratio	Length (bp)	ratio	
Total Repeat	7149094	15.17%	6401846	13.28%	
Non-interspersed Repeats	616107	1.30%	603814	1.25%	
Simple_repeat	479237	1.01%	401791	0.83%	
Low_complexity	106834	0.22%	108068	0.22%	
rRNA	30903	0.06%	49668	0.10%	
Satellite	5078	1.00%	46381	0.09%	
Interspersed Repeats	6594811	14.00%	5890837	12.22%	
Retroelements	4169594	8.85%	3085184	6.40%	
LTR	3460696	7.34%	1253064	2.59%	
LINE	705785	1.49%	1834576	3.80%	
SINE	2304	0.00%	0	0.00%	
DNA transposons	65695	0.13%	154441	0.32%	
Rolling-circles	160054	0.33%	1548	0.00%	
Unknown	2213044	4.69%	2707385	5.61%	

Non-coding RNA

The annotation of non-coding RNA revealed that G. esculenta and P. xinjiangensis each harbored 253 and 255 predicted tRNAs, respectively, capable of decoding the standard 20 amino acids. In G. esculenta, the rRNA subunits were situated on chromosome 4, with the large subunits, 5.8S RNA and small subunits alternately repeated eight times, spanning 69.7 kb (scaffold04: 136308-205968). Additionally, 47 copies of 5S RNA were identified, with an average length of 119.98 bp. Similarly, in P. xinjiangensis, the large and small subunits were identified on chromosome 4, repeated twelve times and spanning 87.0 kb (scaffold04: 296256–383265). Moreover, 52 copies of 5S RNA were recognized, with an average length of 119.04 bp. Both species did not yield miRNA identification; however they were consistent in the quantity and type of snRNA identified. This included 8 C/D box snoRNA, 4 Sm-class snRNA, 3 H/ACA box snoRNA, and 1 Lsm-class snRNA (Table S1).

CAZymes

In both G. esculenta and P. xinjiangensis, a similar number of CAZymes were identified, totaling 356 and 352, respectively. Further categorization revealed comparable counts in glycoside hydrolases (GH), glycosyl transferases (GT), polysaccharide lyases (PL), carbohydrate esterases (CE), carbohydrate-binding modules (CBM), and auxiliary activities (AA), with figures of 164 and 167, 57 and 56, 22 and 21, 20 and 17, 33 and 36, 81 and 78, respectively (Fig. 5b). The CAZyme profiles of these species closely resemble those of closely related genera such as Verpa and Morchella within the Pezizales order. Additionally, similarities were observed with pyrophilous species like Geopyxis carbonaria (CAZyme count: 379)65. Furthermore, their enzyme profiles exhibited resemblance to those of more distantly related taxa such as Orbiliomycetes, Lecanoromycetes, Eurotiomycetes, Dothideomycetes, and Leotiomycetes (Fig. 5b). These fungi demonstrate robust cellulose degradation capabilities, indicative of a preference for saprotrophic or parasitic lifestyles, distinct from the mycorrhizal ecological types prevalent in most members of Pezizales. In the category of PCWDE, particularly cellulose-binding motif 1 (CMB1), copper-dependent lytic polysaccharide monooxygenases (AA9), cellobiohydrolase (GH6), and reducing end-acting cellobiohydrolase (GH7), the number of these genes is significantly higher in saprotrophic fungi compared to mycorrhizal fungi (Fig. 5b). This suggests that G. esculenta and P. xinjiangensis can be well regarded as potential saprotrophic fungi.Fig. 5 Species tree based on whole genome-encoded proteins of Pezizales (a) and their corresponding CAZYme (left size of b) and PCWDE categories (right size of b). The species order arrangement in a and b is consistent, with different colored boxes indicating the nutritional lifestyles of the species. GH = glycoside hydrolases; GT = glycosyl transferases; PL = polysaccharide lyases; CE = carbohydrate esterases; CBM = carbohydrate-binding modules; AA = auxiliary activities. For panel (a), blue numbers indicate the contraction of gene families, and red numbers indicate the expansion of gene families, applicable to both clades in the species tree and the species names. Black numbers represent the divergence time, with the 95% confidence interval value of the divergence time in the brackets. A time scale is provided at the bottom of panel (a) for reference, with all numbers expressed in million years. For panel (b), the numbers inside the bars denote the number of genes. The square colors, in panel (b), following the species names describe the lifestyle of the species, according to the colors illustrated at the bottom of the figure.

Comparative genomic and phylogenomic

Based on whole genome-encoded proteins, a robust species tree was constructed with 100% support for all branches. At the family level, the tree clearly delineated Ascobolaceae, Pezizaceae, Pyronemataceae, Tuberaceae, Morchellaceae, and Discinaceae. Notably, Morchellaceae and Discinaceae exhibited the closest genetic distance, followed by Tuberaceae and Pyronemataceae (Fig. 5). Divergence analysis indicated that the split between Morchellaceae and Discinaceae occurred approximately 107.42 MYA [95% highest density (HD) interval: 89.69–126.32], with their divergence from Tuberaceae occurring around 164.69 MYA [95% HD interval: 154.42–173.31] (Fig. 5). Under the ancestral node shared with Tuberaceae, both Morchellaceae and Discinaceae experienced extensive expansion of gene families, totaling 157, with 59 gene families undergoing contraction. This pattern likely reflects their saprophytic lifestyle, enabling rapid growth under laboratory conditions, in contrast to the subterranean lifestyle and mycorrhizal characteristics of Tuberaceae. Within Morchellaceae and Discinaceae, the number of gene families undergoing expansion or contraction significantly decreased, suggesting similar growth and metabolic traits despite their classification into different families (Fig. 5).

Data Records

The raw sequencing data and genome assembly of G. esculenta and P. xinjiangensis have been deposited at the National Center for Biotechnology Information (NCBI) with the BioProject accession number of PRJNA109698066. The assembled genome has been deposited in the NCBI assembly with the accession numbers JBCAML000000000 and JBCAMK000000000, respectively67,68. Additionally, the results of annotation for repeated sequences, gene structure, and functional prediction have been deposited in the Figshare database69.

Technical Validation

The third-generation genome assembly software was validated using five tools: NextDenovo26, Flye70, Minimap71, NECAT72, and Wtdbg273. These methods were employed for de novo genome assembly of G. esculenta and P. xinjiangensis, considering criteria such as genome size, contig count, and contig N50. The findings indicate that NextDenovo outperformed the other four software options overall in both species (refer to Table S2).

Using Merqury74, the filtered second-generation sequencing reads were aligned to the assembled genomes using a Kmer-based strategy to assess the completeness and error rate of the genome assembly. The evaluation results indicate that both G. esculenta and P. xinjiangensis achieved high completeness and quality assembly. Specifically, their QV values were 33.67 and 41.13, respectively, with error rates of 4.30E-04 and 7.71E-05 (Fig. 6j–k).Fig. 6 Response evaluation of chromatin interaction sequencing data, genomic integrity, and second-generation sequencing data. Reads alignment of the original HiC data and the DNA sequencing data for the genome of G. esculenta (a–d,j) and P. xinjiangensis (e–h,k), as well as the genome BUSCO completeness assessment for G. esculenta and P. xinjiangensis (i).

Employing HicPro software75, Hi-C chromatin interaction data were aligned to their respective genomes to evaluate the sequencing data volume and quality. The results indicate that G. esculenta and P. xinjiangensis generated 222 M and 178 M read pairs, respectively. After removing unmapped, low-quality, and singleton paired reads, they obtained 134 M and 129 M unique paired reads, respectively. Upon eliminating duplicated paired reads from valid interactions, they obtained 75.6 M and 28.2 M valid interaction paired reads, including both trans- and cis- interactions. The alignment of paired-end reads to the genome, including global and local mapping states, was consistent (see Fig. 6a–h), facilitating subsequent chromosomal conformation analysis.

Furthermore, the completeness of genome assembly was evaluated using BUSCO software (v5.2.2) based on the ascomycota_odb10 database76. The results reveal that G. esculenta and P. xinjiangensis species respectively predicted 98.71% and 98.89% of single-copy genes, of which 1.46% and 1.58% were duplicated, and only 0.88% and 0.82% were missing, respectively (Fig. 6i). This indicates that we have obtained high-quality genomes for both species.

Supplementary information

Supplementary Information

Supplementary Information

Supplementary information

The online version contains supplementary material available at 10.1038/s41597-024-03794-z.

Acknowledgements

This work was supported by the Yunnan Key Project of Science and Technology (202402AE090030), Yunnan Revitalization Talent Support Program to Jesús Pérez-Moreno and Xinhua He, and the Yunnan Technology Innovation Program (202205AD160036) to Fuqiang Yu.

Author contributions

W.L. data curation, methodology, resources, software, and writing – original draft; X.F.S., Y.L.C. and W.H.S.: methodology, review, and editing; P.X.H. and P.M.J.: conceptualization, resources, and supervision; D.L. and F.Q.Y.: conceptualization, writing – review, editing, and supervision.

Code availability

All pipelines and software utilized in this study were applied in accordance with their respective manuals and protocols. The parameters and software versions are specified in the Methods section. In cases where detailed parameters for a software are not provided, default settings were employed.

Competing interests

The authors declare no competing interests.

Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
==== Refs
References

1. Wang XC Phylogeny and taxonomic revision of the family Discinaceae (Pezizales, Ascomycota) Microbiol Spectr. 2023 11 e00207 00223 37102868
Wang, X. C. et al. Phylogeny and taxonomic revision of the family Discinaceae (Pezizales, Ascomycota). Microbiol Spectr. 11, e00207–00223 (2023).37102868
2. O’Donnell K Cigelnik E Weber NS Trappe JM Phylogenetic relationships among ascomycetous truffles and the true and false morels inferred from 18S and 28S ribosomal DNA sequence analysis Mycologia 1997 89 48 65 10.1080/00275514.1997.12026754
O’Donnell, K., Cigelnik, E., Weber, N. S. & Trappe, J. M. Phylogenetic relationships among ascomycetous truffles and the true and false morels inferred from 18S and 28S ribosomal DNA sequence analysis. Mycologia 89, 48–65, 10.1080/00275514.1997.12026754 (1997).10.1080/00275514.1997.12026754
3. Huhtinen S Ruotsalainen J Notes on the taxonomy and occurrence of some species of Gyromitra in Finland Karstenia 2004 44 25 34 10.29203/ka.2004.396
Huhtinen, S. & Ruotsalainen, J. Notes on the taxonomy and occurrence of some species of Gyromitra in Finland. Karstenia 44, 25–34 (2004).10.29203/ka.2004.396
4. Karelia S Notes on Gyromitra esculenta coll. and G. recurva, a noteworthy species of western North America Karstenia 1979 19 46 49 10.29203/ka.1979.185
Karelia, S. Notes on Gyromitra esculenta coll. and G. recurva, a noteworthy species of western North America. Karstenia 19, 46–49 (1979).10.29203/ka.1979.185
5. Medel R A review of the genus Gyromitra (Ascomycota, Pezizales, Discinaceae) in Mexico Mycotaxon 2005 94 103 110
Medel, R. A review of the genus Gyromitra (Ascomycota, Pezizales, Discinaceae) in Mexico. Mycotaxon 94, 103–110 (2005).
6. Patocka J Pita R Kuca K Gyromitrin mushroom toxin of Gyromitra spp Mil Med Sci Lett. 2012 81 61 67 10.31482/mmsl.2012.008
Patocka, J., Pita, R. & Kuca, K. Gyromitrin mushroom toxin of Gyromitra spp. Mil Med Sci Lett. 81, 61–67 (2012).10.31482/mmsl.2012.008
7. White J Mushroom poisoning: A proposed new clinical classification Toxicon 2019 157 53 65 10.1016/j.toxicon.2018.11.007 30439442
White, J. et al. Mushroom poisoning: A proposed new clinical classification. Toxicon 157, 53–65 (2019).30439442 10.1016/j.toxicon.2018.11.007
8. Dirks AC Not all bad: Gyromitrin has a limited distribution in the false morels as determined by a new ultra high-performance liquid chromatography method Mycologia 2023 115 1 15 10.1080/00275514.2022.2146473 36541902
Dirks, A. C. et al. Not all bad: Gyromitrin has a limited distribution in the false morels as determined by a new ultra high-performance liquid chromatography method. Mycologia 115, 1–15 (2023).36541902 10.1080/00275514.2022.2146473
9. Michelot D Toth B Poisoning by Gyromitra esculenta–a review J Appl Toxicol. 1991 11 235 243 10.1002/jat.2550110403 1939997
Michelot, D. & Toth, B. Poisoning by Gyromitra esculenta–a review. J Appl Toxicol. 11, 235–243 (1991).1939997 10.1002/jat.2550110403
10. Cao J Fan L Liu B Notes on the genus Gyromitra from China Mycosystema 1990 9 100 108
Cao, J., Fan, L. & Liu, B. Notes on the genus Gyromitra from China. Mycosystema 9, 100–108 (1990).
11. Martin F Périgord black truffle genome uncovers evolutionary origins and mechanisms of symbiosis Nature 2010 464 1033 10.1038/nature08867 20348908
Martin, F. et al. Périgord black truffle genome uncovers evolutionary origins and mechanisms of symbiosis. Nature 464, 1033 (2010).20348908 10.1038/nature08867
12. Cai Y Physiological characteristics and comparative secretome analysis of Morchella importuna grown on glucose, rice straw, sawdust, wheat grain, and MIX substrates Front Microbiol. 2021 12 636344 10.3389/fmicb.2021.636344 34113321
Cai, Y. et al. Physiological characteristics and comparative secretome analysis of Morchella importuna grown on glucose, rice straw, sawdust, wheat grain, and MIX substrates. Front Microbiol. 12, 636344 (2021).34113321 10.3389/fmicb.2021.636344
13. Liu W Subchromosome-scale nuclear and complete mitochondrial genome characteristics of Morchella crassipes Int J Mol Sci. 2020 21 483 10.3390/ijms21020483 31940908
Liu, W. et al. Subchromosome-scale nuclear and complete mitochondrial genome characteristics of Morchella crassipes. Int J Mol Sci. 21, 483 (2020).31940908 10.3390/ijms21020483
14. Hansen K Schumacher T Skrede I Huhtinen S Wang XH Pindara revisited–evolution and generic limits in Helvellaceae Persoonia 2019 42 186 204 10.3767/persoonia.2019.42.07 31551618
Hansen, K., Schumacher, T., Skrede, I., Huhtinen, S. & Wang, X. H. Pindara revisited–evolution and generic limits in Helvellaceae. Persoonia 42, 186–204 (2019).31551618 10.3767/persoonia.2019.42.07
15. Traeger S The genome and development-dependent transcriptomes of Pyronema confluens: A window into fungal evolution PLoS Genet. 2013 9 e1003820 10.1371/journal.pgen.1003820 24068976
Traeger, S. et al. The genome and development-dependent transcriptomes of Pyronema confluens: A window into fungal evolution. PLoS Genet. 9, e1003820 (2013).24068976 10.1371/journal.pgen.1003820
16. Steindorff AS Diversity of genomic adaptations to the post-fire environment in Pezizales fungi points to crosstalk between charcoal tolerance and sexual development New Phytol. 2022 236 1154 1167 10.1111/nph.18407 35898177
Steindorff, A. S. et al. Diversity of genomic adaptations to the post-fire environment in Pezizales fungi points to crosstalk between charcoal tolerance and sexual development. New Phytol. 236, 1154–1167 (2022).35898177 10.1111/nph.18407
17. JGI Mycocosm https://mycocosm.jgi.doe.gov/Gyrinf1/Gyrinf1.home.html (2024).
18. JGI Mycocosm https://mycocosm.jgi.doe.gov/Gyresc1/Gyresc1.home.html (2024).
19. Morel, S. et al. Antibacterial activity of wild mushrooms from France. Int J Med Mushrooms 23 (2021).
20. Suleimen YM Isolation, crystal structure, and in silico aromatase inhibition activity of ergosta-5, 22-dien-3β-ol from the fungus Gyromitra esculenta J Chem-Ny. 2021 2021 1 10 10.1155/2021/5529786
Suleimen, Y. M. et al. Isolation, crystal structure, and in silico aromatase inhibition activity of ergosta-5, 22-dien-3β-ol from the fungus Gyromitra esculenta. J Chem-Ny. 2021, 1–10 (2021).10.1155/2021/5529786
21. Hobbie, E. A., Weber, N. S. & Trappe, J. M. Mycorrhizal vs saprotrophic status of fungi: the isotopic evidence. New Phytol. 601-610 (2001).
22. Jalkanen R Jalkanen E Development of the fruit bodies of Gyromitra esculenta Karstenia 1981 21 50 52 10.29203/ka.1981.203
Jalkanen, R. & Jalkanen, E. Development of the fruit bodies of Gyromitra esculenta. Karstenia 21, 50–52 (1981).10.29203/ka.1981.203
23. Carris LM Peever TL McCotter SW Mitospore stages of Disciotis, Gyromitra and Morchella in the inland Pacific Northwest USA Mycologia 2015 107 729 744 10.3852/14-207 25911699
Carris, L. M., Peever, T. L. & McCotter, S. W. Mitospore stages of Disciotis, Gyromitra and Morchella in the inland Pacific Northwest USA. Mycologia 107, 729–744 (2015).25911699 10.3852/14-207
24. Porebski S Bailey LG Baum BR Modification of a CTAB DNA extraction protocol for plants containing high polysaccharide and polyphenol components Plant Mol Biol Rep. 1997 15 8 15 10.1007/BF02772108
Porebski, S., Bailey, L. G. & Baum, B. R. Modification of a CTAB DNA extraction protocol for plants containing high polysaccharide and polyphenol components. Plant Mol Biol Rep. 15, 8–15 (1997).10.1007/BF02772108
25. Bolger AM Lohse M Usadel B Trimmomatic: a flexible trimmer for Illumina sequence data Bioinformatics 2014 30 2114 2120 10.1093/bioinformatics/btu170 24695404
Bolger, A. M., Lohse, M. & Usadel, B. Trimmomatic: a flexible trimmer for Illumina sequence data. Bioinformatics 30, 2114–2120 (2014).24695404 10.1093/bioinformatics/btu170
26. Hu, J. et al. NextDenovo: an efficient error correction and accurate assembly tool for noisy long reads. Genome Biol. 25, 107 (2024).
27. Hu J Fan J Sun Z Liu S NextPolish: a fast and efficient genome polishing tool for long-read assembly Bioinformatics 2019 36 2253 2255 10.1093/bioinformatics/btz891
Hu, J., Fan, J., Sun, Z. & Liu, S. NextPolish: a fast and efficient genome polishing tool for long-read assembly. Bioinformatics 36, 2253–2255 (2019).10.1093/bioinformatics/btz891
28. Durand NC Juicer provides a one-click system for analyzing loop-resolution Hi-C experiments Cell Syst. 2016 3 95 98 10.1016/j.cels.2016.07.002 27467249
Durand, N. C. et al. Juicer provides a one-click system for analyzing loop-resolution Hi-C experiments. Cell Syst. 3, 95–98 (2016).27467249 10.1016/j.cels.2016.07.002
29. Dudchenko O De novo assembly of the Aedes aegypti genome using Hi-C yields chromosome-length scaffolds Science 2017 356 92 95 10.1126/science.aal3327 28336562
Dudchenko, O. et al. De novo assembly of the Aedes aegypti genome using Hi-C yields chromosome-length scaffolds. Science 356, 92–95 (2017).28336562 10.1126/science.aal3327
30. Durand NC Juicebox provides a visualization system for Hi-C contact maps with unlimited zoom Cell Syst. 2016 3 99 101 10.1016/j.cels.2015.07.012 27467250
Durand, N. C. et al. Juicebox provides a visualization system for Hi-C contact maps with unlimited zoom. Cell Syst. 3, 99–101 (2016).27467250 10.1016/j.cels.2015.07.012
31. Nansheng C Using repeatmasker to identify repetitive elements in genomic sequences Current Protocols in Bioinformatics 2004 5 4 10
Nansheng, C. Using repeatmasker to identify repetitive elements in genomic sequences. Current Protocols in Bioinformatics 5, 4–10 (2004).
32. Smit, A. F. & Hubley, R. RepeatModeler Open-1.0. Available fom http://www.repeatmasker.org (2008).
33. Kim D Langmead B Salzberg SL HISAT: a fast spliced aligner with low memory requirements Nat Methods. 2015 12 357 10.1038/nmeth.3317 25751142
Kim, D., Langmead, B. & Salzberg, S. L. HISAT: a fast spliced aligner with low memory requirements. Nat Methods. 12, 357 (2015).25751142 10.1038/nmeth.3317
34. Birney E Clamp M Durbin R GeneWise and genomewise Genome Res. 2004 14 988 995 10.1101/gr.1865504 15123596
Birney, E., Clamp, M. & Durbin, R. GeneWise and genomewise. Genome Res. 14, 988–995 (2004).15123596 10.1101/gr.1865504
35. Liu W Chen L Cai Y Zhang Q Bian Y Opposite polarity monospore genome de novo sequencing and comparative analysis reveal the possible heterothallic life cycle of Morchella importuna Int J Mol Sci. 2018 19 2525 10.3390/ijms19092525 30149649
Liu, W., Chen, L., Cai, Y., Zhang, Q. & Bian, Y. Opposite polarity monospore genome de novo sequencing and comparative analysis reveal the possible heterothallic life cycle of Morchella importuna. Int J Mol Sci. 19, 2525 (2018).30149649 10.3390/ijms19092525
36. Stanke M AUGUSTUS: ab initio prediction of alternative transcripts Nucleic Acids Res. 2006 34 W435 W439 10.1093/nar/gkl200 16845043
Stanke, M. et al. AUGUSTUS: ab initio prediction of alternative transcripts. Nucleic Acids Res. 34, W435–W439 (2006).16845043 10.1093/nar/gkl200
37. Punta M The Pfam protein families database Nucleic Acids Res. 2011 40 D290 D301 10.1093/nar/gkr1065 22127870
Punta, M. et al. The Pfam protein families database. Nucleic Acids Res. 40, D290–D301 (2011).22127870 10.1093/nar/gkr1065
38. Goll J METAREP: JCVI metagenomics reports—an open source tool for high-performance comparative metagenomics Bioinformatics 2010 26 2631 2632 10.1093/bioinformatics/btq455 20798169
Goll, J. et al. METAREP: JCVI metagenomics reports—an open source tool for high-performance comparative metagenomics. Bioinformatics 26, 2631–2632 (2010).20798169 10.1093/bioinformatics/btq455
39. Sun W Structure of the mating-type genes and mating systems of Verpa bohemica and Verpa conica (Ascomycota, Pezizomycotina) J Fungi. 2023 9 1202 10.3390/jof9121202
Sun, W. et al. Structure of the mating-type genes and mating systems of Verpa bohemica and Verpa conica (Ascomycota, Pezizomycotina). J Fungi. 9, 1202 (2023).10.3390/jof9121202
40. Xu Z Wang H LTR_FINDER: an efficient tool for the prediction of full-length LTR retrotransposons Nucleic Acids Res. 2007 35 W265 W268 10.1093/nar/gkm286 17485477
Xu, Z. & Wang, H. LTR_FINDER: an efficient tool for the prediction of full-length LTR retrotransposons. Nucleic Acids Res. 35, W265–W268 (2007).17485477 10.1093/nar/gkm286
41. Price AL Jones NC Pevzner PA De novo identification of repeat families in large genomes Bioinformatics 2005 21 i351 i358 10.1093/bioinformatics/bti1018 15961478
Price, A. L., Jones, N. C. & Pevzner, P. A. De novo identification of repeat families in large genomes. Bioinformatics 21, i351–i358 (2005).15961478 10.1093/bioinformatics/bti1018
42. Bao W Kojima KK Kohany O Repbase Update, a database of repetitive elements in eukaryotic genomes Mobile DNA 2015 6 1 6 10.1186/s13100-015-0041-9
Bao, W., Kojima, K. K. & Kohany, O. Repbase Update, a database of repetitive elements in eukaryotic genomes. Mobile DNA 6, 1–6 (2015).10.1186/s13100-015-0041-9
43. Lowe TM Eddy SR tRNAscan-SE: a program for improved detection of transfer RNA genes in genomic sequence Nucleic Acids Res. 1997 25 955 964 10.1093/nar/25.5.955 9023104
Lowe, T. M. & Eddy, S. R. tRNAscan-SE: a program for improved detection of transfer RNA genes in genomic sequence. Nucleic Acids Res. 25, 955–964 (1997).9023104 10.1093/nar/25.5.955
44. Lagesen K RNAmmer: consistent and rapid annotation of ribosomal RNA genes Nucleic Acids Res. 2007 35 3100 3108 10.1093/nar/gkm160 17452365
Lagesen, K. et al. RNAmmer: consistent and rapid annotation of ribosomal RNA genes. Nucleic Acids Res. 35, 3100–3108 (2007).17452365 10.1093/nar/gkm160
45. Nawrocki EP Eddy SR Infernal 1.1: 100-fold faster RNA homology searches Bioinformatics 2013 29 2933 2935 10.1093/bioinformatics/btt509 24008419
Nawrocki, E. P. & Eddy, S. R. Infernal 1.1: 100-fold faster RNA homology searches. Bioinformatics 29, 2933–2935 (2013).24008419 10.1093/bioinformatics/btt509
46. Nawrocki EP Rfam 12.0: updates to the RNA families database Nucleic Acids Res. 2015 43 D130 D137 10.1093/nar/gku1063 25392425
Nawrocki, E. P. et al. Rfam 12.0: updates to the RNA families database. Nucleic Acids Res. 43, D130–D137 (2015).25392425 10.1093/nar/gku1063
47. Zhao Z Liu H Wang C Xu J-R Erratum to: comparative analysis of fungal genomes reveals different plant cell wall degrading capacity in fungi BMC Genomics 2014 15 1 15 10.1186/1471-2164-15-6 24382143
Zhao, Z., Liu, H., Wang, C. & Xu, J.-R. Erratum to: comparative analysis of fungal genomes reveals different plant cell wall degrading capacity in fungi. BMC Genomics 15, 1–15 (2014).24382143 10.1186/1471-2164-15-6
48. Sista Kameshwar AK Qin W Comparative study of genome-wide plant biomass-degrading CAZymes in white rot, brown rot and soft rot fungi Mycology 2018 9 93 105 10.1080/21501203.2017.1419296 30123665
Sista Kameshwar, A. K. & Qin, W. Comparative study of genome-wide plant biomass-degrading CAZymes in white rot, brown rot and soft rot fungi. Mycology 9, 93–105 (2018).30123665 10.1080/21501203.2017.1419296
49. Yin Y dbCAN: a web resource for automated carbohydrate-active enzyme annotation Nucleic Acids Res. 2012 40 W445 W451 10.1093/nar/gks479 22645317
Yin, Y. et al. dbCAN: a web resource for automated carbohydrate-active enzyme annotation. Nucleic Acids Res. 40, W445–W451 (2012).22645317 10.1093/nar/gks479
50. Drula E The carbohydrate-active enzyme database: functions and literature Nucleic Acids Res. 2022 50 D571 D577 10.1093/nar/gkab1045 34850161
Drula, E. et al. The carbohydrate-active enzyme database: functions and literature. Nucleic Acids Res. 50, D571–D577 (2022).34850161 10.1093/nar/gkab1045
51. Söding J Protein homology detection by HMM–HMM comparison Bioinformatics 2005 21 951 960 10.1093/bioinformatics/bti125 15531603
Söding, J. Protein homology detection by HMM–HMM comparison. Bioinformatics 21, 951–960 (2005).15531603 10.1093/bioinformatics/bti125
52. McDonald TR Mueller O Dietrich FS Lutzoni F High-throughput genome sequencing of lichenizing fungi to assess gene loss in the ammonium transporter/ammonia permease gene family BMC Genomics 2013 14 1 14 10.1186/1471-2164-14-225 23323973
McDonald, T. R., Mueller, O., Dietrich, F. S. & Lutzoni, F. High-throughput genome sequencing of lichenizing fungi to assess gene loss in the ammonium transporter/ammonia permease gene family. BMC Genomics 14, 1–14 (2013).23323973 10.1186/1471-2164-14-225
53. Vesth TC Investigation of inter-and intraspecies variation through genome sequencing of Aspergillus section Nigri Nat Genet. 2018 50 1688 1695 10.1038/s41588-018-0246-1 30349117
Vesth, T. C. et al. Investigation of inter-and intraspecies variation through genome sequencing of Aspergillus section Nigri. Nat Genet. 50, 1688–1695 (2018).30349117 10.1038/s41588-018-0246-1
54. Mesny F Genetic determinants of endophytism in the Arabidopsis root mycobiome Nat commun. 2021 12 7227 10.1038/s41467-021-27479-y 34893598
Mesny, F. et al. Genetic determinants of endophytism in the Arabidopsis root mycobiome. Nat commun. 12, 7227 (2021).34893598 10.1038/s41467-021-27479-y
55. Amselem, J. et al. Genomic analysis of the necrotrophic fungal pathogens Sclerotinia sclerotiorum and Botrytis cinerea. PLoS Genet. 7 (2011).
56. Emms DM Kelly S OrthoFinder: phylogenetic orthology inference for comparative genomics Genome Biol. 2019 20 238 10.1186/s13059-019-1832-y 31727128
Emms, D. M. & Kelly, S. OrthoFinder: phylogenetic orthology inference for comparative genomics. Genome Biol. 20, 238 (2019).31727128 10.1186/s13059-019-1832-y
57. Katoh K Misawa K Kuma KI Miyata T MAFFT: a novel method for rapid multiple sequence alignment based on fast Fourier transform Nucleic Acids Res. 2002 30 3059 3066 10.1093/nar/gkf436 12136088
Katoh, K., Misawa, K., Kuma, K. I. & Miyata, T. MAFFT: a novel method for rapid multiple sequence alignment based on fast Fourier transform. Nucleic Acids Res. 30, 3059–3066 (2002).12136088 10.1093/nar/gkf436
58. Castresana J Selection of conserved blocks from multiple alignments for their use in phylogenetic analysis Mol Biol Evol. 2002 17 540 552 10.1093/oxfordjournals.molbev.a026334
Castresana, J. Selection of conserved blocks from multiple alignments for their use in phylogenetic analysis. Mol Biol Evol. 17, 540–552 (2002).10.1093/oxfordjournals.molbev.a026334
59. Minh BQ IQ-TREE 2: new models and efficient methods for phylogenetic inference in the genomic era Mol Biol Evol. 2020 37 1530 1534 10.1093/molbev/msaa015 32011700
Minh, B. Q. et al. IQ-TREE 2: new models and efficient methods for phylogenetic inference in the genomic era. Mol Biol Evol. 37, 1530–1534 (2020).32011700 10.1093/molbev/msaa015
60. FigTree. Available fom http://tree.bio.ed.ac.uk/software/figtree/ (2018).
61. Yang Z PAML 4: phylogenetic analysis by maximum likelihood Mol Biol Evol. 2007 24 1586 1591 10.1093/molbev/msm088 17483113
Yang, Z. PAML 4: phylogenetic analysis by maximum likelihood. Mol Biol Evol. 24, 1586–1591 (2007).17483113 10.1093/molbev/msm088
62. Kumar S Stecher G Suleski M Hedges SB TimeTree: a resource for timelines, timetrees, and divergence times Mol Biol Evol. 2017 34 1812 1819 10.1093/molbev/msx116 28387841
Kumar, S., Stecher, G., Suleski, M. & Hedges, S. B. TimeTree: a resource for timelines, timetrees, and divergence times. Mol Biol Evol. 34, 1812–1819 (2017).28387841 10.1093/molbev/msx116
63. De Bie T Cristianini N Demuth JP Hahn MW CAFE: a computational tool for the study of gene family evolution Bioinformatics 2006 22 1269 1271 10.1093/bioinformatics/btl097 16543274
De Bie, T., Cristianini, N., Demuth, J. P. & Hahn, M. W. CAFE: a computational tool for the study of gene family evolution. Bioinformatics 22, 1269–1271 (2006).16543274 10.1093/bioinformatics/btl097
64. Bankevich A SPAdes: a new genome assembly algorithm and its applications to single-cell sequencing J Comput Biol. 2012 19 455 477 10.1089/cmb.2012.0021 22506599
Bankevich, A. et al. SPAdes: a new genome assembly algorithm and its applications to single-cell sequencing. J Comput Biol. 19, 455–477 (2012).22506599 10.1089/cmb.2012.0021
65. Filialuna O Cripps C Evidence that pyrophilous fungi aggregate soil after forest fire Forest Ecol Manag. 2021 498 119579 10.1016/j.foreco.2021.119579
Filialuna, O. & Cripps, C. Evidence that pyrophilous fungi aggregate soil after forest fire. Forest Ecol Manag. 498, 119579 (2021).10.1016/j.foreco.2021.119579
66. 2024 NCBI Sequence Read Archive PRJNA1096980
NCBI Sequence Read Archive https://identifiers.org/ncbi/bioproject:PRJNA1096980 (2024).
67. 2024 NCBI Assembly JBCAML000000000
NCBI Assembly https://identifiers.org/ncbi/insdc:JBCAML000000000 (2024).
68. 2024 NCBI Assembly JBCAMK000000000
NCBI Assembly https://identifiers.org/ncbi/insdc:JBCAMK000000000 (2024).
69. Liu W 2024 Genome annotation of Gyromitra esculenta and Paragyromitra xinjiangensis figshare. 10.6084/m9.figshare.25593078
Liu, W. Genome annotation of Gyromitra esculenta and Paragyromitra xinjiangensis. figshare.10.6084/m9.figshare.25593078 (2024).10.6084/m9.figshare.25593078
70. Kolmogorov M Yuan J Lin Y Pevzner PA Assembly of long, error-prone reads using repeat graphs Nat Biotechnol. 2019 37 540 546 10.1038/s41587-019-0072-8 30936562
Kolmogorov, M., Yuan, J., Lin, Y. & Pevzner, P. A. Assembly of long, error-prone reads using repeat graphs. Nat Biotechnol. 37, 540–546 (2019).30936562 10.1038/s41587-019-0072-8
71. Li H Minimap and miniasm: fast mapping and de novo assembly for noisy long sequences Bioinformatics 2016 32 2103 2110 10.1093/bioinformatics/btw152 27153593
Li, H. Minimap and miniasm: fast mapping and de novo assembly for noisy long sequences. Bioinformatics 32, 2103–2110 (2016).27153593 10.1093/bioinformatics/btw152
72. Chen Y Efficient assembly of nanopore reads via highly accurate and intact error correction Nat Commun. 2021 12 60 10.1038/s41467-020-20236-7 33397900
Chen, Y. et al. Efficient assembly of nanopore reads via highly accurate and intact error correction. Nat Commun. 12, 60 (2021).33397900 10.1038/s41467-020-20236-7
73. Ruan J Li H Fast and accurate long-read assembly with wtdbg2 Nat Methods. 2020 17 155 158 10.1038/s41592-019-0669-3 31819265
Ruan, J. & Li, H. Fast and accurate long-read assembly with wtdbg2. Nat Methods. 17, 155–158 (2020).31819265 10.1038/s41592-019-0669-3
74. Rhie A Walenz BP Koren S Phillippy AM Merqury: reference-free quality, completeness, and phasing assessment for genome assemblies Genome Biol. 2020 21 1 27 10.1186/s13059-020-02134-9
Rhie, A., Walenz, B. P., Koren, S. & Phillippy, A. M. Merqury: reference-free quality, completeness, and phasing assessment for genome assemblies. Genome Biol. 21, 1–27 (2020).10.1186/s13059-020-02134-9
75. Servant N HiC-Pro: an optimized and flexible pipeline for Hi-C data processing Genome Biol. 2015 16 1 11 10.1186/s13059-015-0831-x 25583448
Servant, N. et al. HiC-Pro: an optimized and flexible pipeline for Hi-C data processing. Genome Biol. 16, 1–11 (2015).25583448 10.1186/s13059-015-0831-x
76. Simão FA Waterhouse RM Ioannidis P Kriventseva EV Zdobnov EM BUSCO: assessing genome assembly and annotation completeness with single-copy orthologs Bioinformatics 2015 31 3210 3212 10.1093/bioinformatics/btv351 26059717
Simão, F. A., Waterhouse, R. M., Ioannidis, P., Kriventseva, E. V. & Zdobnov, E. M. BUSCO: assessing genome assembly and annotation completeness with single-copy orthologs. Bioinformatics 31, 3210–3212 (2015).26059717 10.1093/bioinformatics/btv351
