
==== Front
Plant Commun
Plant Commun
Plant Communications
2590-3462
Elsevier

S2590-3462(24)00205-0
10.1016/j.xplc.2024.100935
100935
Correspondence
Complete genome assembly provides insights into the centromere architecture of pumpkin (Cucurbita maxima)
Zeng Qingguo 1
Wei Minghua 1
Li Shuai 1
Wang Haiyan 1
Mo Changjuan 1
Yang Li 1
Li Xinzheng 3
Bie Zhilong biezl@mail.hzau.edu.cn
12∗
Kong Qiusheng qskong@mail.hzau.edu.cn
1∗∗
1 National Key Laboratory for Germplasm Innovation & Utilization of Horticultural Crops, College of Horticulture and Forestry Sciences, Huazhong Agricultural University, Wuhan 430070, China
2 Hubei Hongshan Laboratory, Wuhan 430070, China
3 School of Horticulture and Landscape Architecture, Henan Institute of Science and Technology, Xinxiang 453003, China
∗ Corresponding author biezl@mail.hzau.edu.cn
∗∗ Corresponding author qskong@mail.hzau.edu.cn
30 4 2024
09 9 2024
30 4 2024
5 9 1009357 3 2024
23 4 2024
28 4 2024
© 2024 The Author(s)
2024
https://creativecommons.org/licenses/by-nc-nd/4.0/ This is an open access article under the CC BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/).
Published: April 30, 2024
==== Body
pmcDear Editor,

Pumpkin (Cucurbita maxima) belongs to the Cucurbita genus of the Cucurbitaceae family and is a globally cultivated vegetable crop with great economic significance. The global production and planting area of pumpkin, including C. maxima, C. moschata, and C. pepo, reached 7.38 million tons and more than 0.39 million hectares in 2022 (https://www.fao.org/). Pumpkin is also widely used as a rootstock for other cucurbit crops to overcome the obstacles caused by continuous cropping and enhance stress tolerance, yield, and fruit quality (Kong et al., 2014).

A high-quality complete genome is crucial for genetic and genomic research (Mo et al., 2024). The first pumpkin genome was assembled using short reads and released in 2017 (Sun et al., 2017), greatly facilitating research on the genetic diversity (Wei et al., 2023), genome evolution (Wang et al., 2022), and molecular breeding of pumpkin (Han et al., 2022). However, this genome is fragmented owing to the limitations of short reads, restricting its value for gene cloning and investigation of genome architecture.

To improve the completeness of the pumpkin reference genome, we selected and sequenced the highly salt-tolerant inbred line Huazhong Agricultural University (HZAU) (C. maxima) using the NovaSeq 6000, PromethION 48, and PacBio Sequel Ⅱ platforms (Supplemental Table 1). We assembled the genome de novo from High Fidelity and Oxford Nanopore Technologies data using multiple assemblers. With High Fidelity and Oxford Nanopore Technologies reads as the input data, Hifiasm generated the best assembly, with a contig N50 of 16.35 Mb, a Benchmarking Universal Single-copy Orthologs (BUSCO) value of 98.50%, a contig number of 519, and a total length of 377.15 Mb, which was then selected for subsequent analysis (Supplemental Table 2; Supplemental Figure 1). We also assembled the chloroplast and mitochondrial genomes of HZAU (Supplemental Table 3; Supplemental Figure 2). Contigs that originated from organelle genomes were removed, and the remaining 24 contigs with a cumulative size of 345.98 Mb were used for scaffolding (Supplemental Table 4). Using the 3D-DNA pipeline, 22 contigs were anchored onto 20 pseudo-chromosomes with an anchoring ratio of 99.75%; these included 18 pseudo-chromosomes, each consisting of a single contig, and two pseudo-chromosomes, each consisting of two contigs (Supplemental Figure 3; Supplemental Table 5). The two unanchored contigs were less than 1 Mb in length. We then manually filled the two gaps on chromosome 09 (chr09) and chr20 using contigs assembled by other assemblers (Supplemental Figure 4). The correctness of gap filling was confirmed by the uniform distribution and high coverage of the mapped reads (Supplemental Figure 5). A total of 40 telomere repeats were identified at both ends of the 20 chromosomes (Supplemental Figure 6; Supplemental Table 6). The final HZAU genome assembly was 345.14 Mb in length, with a scaffold N50 of 19.03 Mb, a BUSCO value of 98.50%, and a quality value (QV) score of 58.40; it was 133.64 Mb longer than the previously published genome of C. maxima (Rimu) (Sun et al., 2017) (Figure 1A; Supplemental Table 7). These results demonstrated the successful assembly of a high-quality telomere-to-telomere and gap-free genome for HZAU (Supplemental Table 8).Figure 1 Complete genome assembly and annotation of HZAU.

(A) Collinearity analysis between the genomes of HZAU and C. maxima (Rimu).

(B) Circos plot of the HZAU genome annotation. Quantitative tracks are aggregated in 10-kb windows. (a) Chromosomes, with orange bands representing centromeres. (b–g) Densities of Copia, Gypsy, DNA transposons, TEs, annotated genes, and GC content, respectively. (h) Ripe fruit of HZAU.

(C) StainedGlass sequence-identity heatmaps of HZAU centromeres. Repeat annotations are shown at the bottom of each centromere type.

A total of 227.43 Mb of repetitive sequences were identified in the HZAU genome, accounting for 65.89% of the assembly, a higher percentage than that reported for C. maxima (Rimu) (40.31%) (Supplemental Figure 7). The interspersed repeats, which consisted mainly of long terminal repeats (LTRs), terminal inverted repeats, and non-terminal inverted repeats, accounted for 41.70% (143.94 Mb) of the entire genome. A total of 170 279 LTR-retrotransposons (RTs) and 316 247 DNA transposons were identified among these transposable elements (Figure 1B; Supplemental Table 9). Copia and Gypsy were the most abundant LTR-RTs, accounting for 9.04% and 7.59% of the whole genome, respectively. Mutator was the most abundant DNA transposon, accounting for 4.06% of the whole genome. The 45S rDNA arrays were located mainly on chr01, chr08, chr09, and chr11, and the majority of 5S rDNA arrays were located on chr10 (Supplemental Figure 8). A total of 28 329 protein-coding genes were annotated, representing a complete BUSCO value of 98.50%. The average gene length was 3.10 kb, and the average exon number was 7.11. The ratio of single-exon genes to multiple-exon genes was 0.22, which is close to the ideal ratio of 0.20 (Jain et al., 2008) (Supplemental Table 10). A total of 27 717 (97.84%) genes were functionally annotated by searching multiple databases, including 214 resistance-gene analogs and 40 nucleotide-binding leucine-rich repeat receptors (Supplemental Tables 11, 12, and 13). In addition, 1984 segmental duplications comprising 139 duplicated genes were identified in the HZAU genome (Supplemental Table 14; Supplemental Figure 9A). Among C. maxima and other genera in the Cucurbitaceae family, only those genes within segmental duplications in C. maxima were significantly enriched in pathways related to stress tolerance, which is probably the reason for the higher tolerance of C. maxima to stress (Supplemental Figure 9B).

The centromere is an important genomic element that plays a crucial role in chromosomal segregation across eukaryotes (Naish et al., 2021). A complete centromere was found in the middle of each chromosome in the HZAU assembly. The lengths of the 20 centromeres varied from 0.54 Mb to 9.05 Mb, with an average of 4.13 Mb (Supplemental Figure 15). As the repeat units of the centromere, monomers are arranged into repeat arrays in a head-to-tail manner (Melters et al., 2013). A single centromere-specific monomer shared by all chromosomes has been observed in many plants, such as CentO in rice (Song et al., 2021) and CEN180 in Arabidopsis (Naish et al., 2021). Multiple centromeric monomers have also been reported in several plants. For example, five monomers, CentGm91, CentGm92, CentGm273, CentGm413, and CentGm444, were identified in soybean (Liu et al., 2023). In this study, six types of centromeric monomers with lengths of 169, 253, 315, 324, 327, and 654 bp were identified and referred to as CEN169, CEN253, CEN315, CEN324, CEN327, and CEN654 (Supplemental Figure 10; Supplemental Table 16). In the Cucurbitaceae family, the centromeric satellite monomer CEN354 has only been accurately identified in all chromosomes of melon (Mo et al., 2024). By contrast, no centromere-specific monomer was common to all chromosomes in the HZAU genome. In contrast to those of other species, particularly those of the Cucurbitaceae family, the HZAU centromeres could be classified into five types on the basis of their monomer compositions (Figure 1C). The type 1 centromere was represented by chr02 and was composed of CEN324, CEN327, and CEN654 with insertions of Copia and Gypsy. The type 2 centromere was represented by chr04 and comprised only CEN315 with the insertion of Copia. The type 3 centromere was represented by chr07, chr09, and chr11 and contained CEN169 and CEN253. The type 4 centromere was represented by chr01, chr05, chr08, and chr10 and consisted of only CEN253 with insertions of Gypsy and Copia. Gypsy and Copia insertions resulted in less dense tandem repeats in a heatmap compared with other centromeres. The type 5 centromere was represented by chr03, chr06, chr12, chr13, chr14, chr15, chr16, chr17, chr18, chr19, and chr20 and consisted of only CEN169 with the insertion of Gypsy.

A total of 872 intact LTR-RTs were identified in the assembly; 391 old LTR-RTs had insertion times before 0.5 million years ago (mya), and 481 young LTR-RTs had insertion times after 0.5 mya. In the centromeres, 138 intact LTR-RTs were identified, including 90 young LTR-RTs and 48 old LTR-RTs (Supplemental Table 17). The average insertion time of LTR-RTs in the centromeric regions was 0.55 mya, significantly later than that of non-centromeric regions (1.57 mya) (Supplemental Table 17; Supplemental Figure 11). These results suggest that rapid expansion of LTR-RTs in the centromeres may have driven their evolution.

Collectively, our results provide a valuable high-quality genome resource as well as important insights into the genome architecture of C. maxima.

Data and code availability

All genome assembly and annotation-associated raw sequencing data for C. maxima (HZAU) have been deposited in the National Genomics Data Center BioProject database under accession number PRJNA1060488. The assembled genomes and associated annotation files were submitted to the National Genomics Data Center (https://ngdc.cncb.ac.cn/) under accession number PRJCA022674 (SAM-C3303151).

Funding

This work was supported by the 10.13039/501100012166 National Key Research and Development Program of China (2023YFD2300703 ), the 10.13039/501100001809 National Natural Science Foundation of China (32372794 and 32072653 ), the Science and Technology Program of Xinjiang Production and Construction Corps (2023AB050 ), the Fundamental Research Funds for the Central Universities (2662023YLPY008 ), and the Ningbo Scientific and Technological Project (2021Z006 ).

Author contributions

Q.K. and Z.B. conceived and designed the work. Q.Z. performed data analyses and wrote the manuscript. S.L., L.Y., and X.L. contributed to the preparation of plant samples. M.W., H.W., and C.M. participated in the data analyses. Q.K. revised the manuscript. All authors have read and approved the final manuscript.

Supplemental information

Document S1. Figures S1–S11, Tables S1–S18, and Supplemental methods

Document S2. Article plus supplemental information

Acknowledgments

We are grateful to Associate Professor Yuan Huang for cultivating the experimental materials. No conflict of interest is declared.

Published by the Plant Communications Shanghai Editorial Office in association with Cell Press, an imprint of Elsevier Inc., on behalf of CSPB and CEMPS, CAS.

Supplemental information is available at Plant Communications Online.
==== Refs
References

Han X. Min Z. Wei M. Li Y. Wang D. Zhang Z. Hu X. Kong Q. QTL mapping for pumpkin fruit traits using a GBS-based high-density genetic map Euphytica 218 2022 106
Jain M. Khurana P. Tyagi A.K. Khurana J.P. Genome-wide analysis of intronless genes in rice and Arabidopsis Funct. Integr. Genomics 8 2008 69 78 17578610
Kong Q. Chen J. Liu Y. Ma Y. Liu P. Wu S. Huang Y. Bie Z. Genetic diversity of Cucurbita rootstock germplasm as assessed using simple sequence repeat markers Sci. Hortic. 175 2014 150 155
Liu Y. Yi C. Fan C. Liu Q. Liu S. Shen L. Zhang K. Huang Y. Liu C. Wang Y. Pan-centromere reveals widespread centromere repositioning of soybean genomes Proc. Natl. Acad. Sci. USA 120 2023 e2310177120
Melters D.P. Bradnam K.R. Young H.A. Telis N. May M.R. Ruby J.G. Sebra R. Peluso P. Eid J. Rank D. Comparative analysis of tandem repeats from hundreds of species reveals unique insights into centromere evolution Genome Biol. 14 2013 R10 23363705
Mo C. Wang H. Wei M. Zeng Q. Zhang X. Fei Z. Zhang Y. Kong Q. Complete genome assembly provides a high-quality skeleton for pan-NLRome construction in melon Plant J. 2024 10.1111/tpj.16705 Access published March 2, 2024
Naish M. Alonge M. Wlodzimierz P. Tock A.J. Abramson B.W. Schmücker A. Mandáková T. Jamge B. Lambing C. Kuo P. The genetic and epigenetic landscape of the Arabidopsis centromeres Science 374 2021 eabi7489
Song J.-M. Xie W.-Z. Wang S. Guo Y.-X. Koo D.-H. Kudrna D. Gong C. Huang Y. Feng J.-W. Zhang W. Two gap-free reference genomes and a global view of the centromere architecture in rice Mol. Plant 14 2021 1757 1767 34171480
Sun H. Wu S. Zhang G. Jiao C. Guo S. Ren Y. Zhang J. Zhang H. Gong G. Jia Z. Karyotype Stability and Unbiased Fractionation in the Paleo-Allotetraploid Cucurbita Genomes Mol. Plant 10 2017 1293 1306 28917590
Wang J. Yuan M. Feng Y. Zhang Y. Bao S. Hao Y. Ding Y. Gao X. Yu Z. Xu Q. A common whole-genome paleotetraploidization in Cucurbitales Plant Physiol. 190 2022 2430 2448 36053177
Wei M. Huang Y. Mo C. Wang H. Zeng Q. Yang W. Chen J. Zhang X. Kong Q. Telomere-to-telomere genome assembly of melon ( Cucumis melo L. var. inodorus ) provides a high-quality reference for meta-QTL analysis of important traits Hortic. Res. 10 2023 uhad189
