
==== Front
Microbiol Resour Announc
Microbiol Resour Announc
mra
Microbiology Resource Announcements
2576-098X
American Society for Microbiology 1752 N St., N.W., Washington, DC

39083682
mra00347-24
10.1128/mra.00347-24
mra.00347-24
Genome Sequences
genomics-and-proteomicsGenomics and ProteomicsChromosome-level genome assembly of an auxotrophic strain of the pathogenic yeast Candida parapsilosis
https://orcid.org/0009-0004-2969-3143
Mutalová Sofia 1
https://orcid.org/0000-0003-4152-8744
Hodorová Viktória 1
https://orcid.org/0000-0002-1566-5967
Brázdovič Filip 1 2
https://orcid.org/0000-0003-2652-3662
Cillingová Andrea 1
https://orcid.org/0000-0003-4886-1910
Tomáška Ľubomír 3
https://orcid.org/0000-0002-9483-1766
Brejová Broňa 4 brejova@dcs.fmph.uniba.sk

https://orcid.org/0000-0002-1020-5451
Nosek Jozef 1 jozef.nosek@uniba.sk

1 Department of Biochemistry, Faculty of Natural Sciences, Comenius University Bratislava , Bratislava, Slovak Republic
2 Laboratory of Regulation of Gene Expression, Institute of Microbiology of the Czech Academy of Sciences , Prague, Czech Republic
3 Department of Genetics, Faculty of Natural Sciences, Comenius University Bratislava , Bratislava, Slovak Republic
4 Department of Computer Science, Faculty of Mathematics, Physics and Informatics, Comenius University Bratislava , Bratislava, Slovak Republic
Editor Stajich Jason E. University of California Riverside , Riverside, California, USA

Address correspondence to Broňa Brejová, brejova@dcs.fmph.uniba.sk
Address correspondence to Jozef Nosek, jozef.nosek@uniba.sk
The authors declare no conflict of interest

9 2024
31 7 2024
31 7 2024
13 9 e00347-2403 4 2024
18 6 2024
Copyright © 2024 Mutalová et al.
2024
Mutalová et al.
https://creativecommons.org/licenses/by/4.0/ This is an open-access article distributed under the terms of the Creative Commons Attribution 4.0 International license.

ABSTRACT

We report the genome sequence of the pathogenic yeast Candida parapsilosis strain SR23 (CBS 7157) used in a number of experimental studies. The nuclear genome assembly consists of eight chromosome-sized contigs with a total size of 13.04 Mbp (N50 2.09 Mbp) and a G+C content of 38.7%.

KEYWORDS

yeasts
Candida
genome analysis
Agentúra na podporu výskumu a vývoja (APVV) 22-0144 Nosek Jozef Agentúra na podporu výskumu a vývoja (APVV) 19-0068 Tomáška Ľubomír Vedecká grantová agentúra MŠVVaŠ SR a SAV (VEGA) 1/0538/22 Brejová Broňa Vedecká grantová agentúra MŠVVaŠ SR a SAV (VEGA) 1/0234/23 Nosek Jozef Vedecká grantová agentúra MŠVVaŠ SR a SAV (VEGA) 1/0031/24 Tomáška Ľubomír EC | European Regional Development Fund(ERDF) ITMS2014+: 313021X329 cover-dateSeptember 2024
==== Body
pmcANNOUNCEMENT

Candida parapsilosis SR23 is an adenine and lysine auxotroph isolated about 40 years ago from contaminated culture of Saccharomyces cerevisiae as a yeast with peculiar physiological features and carrying a linear mitochondrial DNA (mtDNA) becoming one of the pillars for a systematic analysis of mtDNA architecture in yeasts (1, 2). Later studies showed that the linear mtDNA and its telomeric structures represent a typical feature of the “psilosis” group species (3–8). Importantly, SR23 became a model for investigations of telomere protection and telomerase-independent replication (9–16), and it was also instrumental in the development of tools for genetic manipulation of C. parapsilosis (17–19). To facilitate further studies, we determined the genome sequence of this strain originating from the laboratory of P. P. Slonimski (Gif-Sur-Yvette, France) and provided to us by L. Kováč in 1987.

High-molecular-weight DNA was isolated as described previously (20) from an overnight culture grown in yeast extract–peptone–dextrose (YPD) medium (Table 1) at 28°C and, without shearing or size selection, used for the preparation of a ligation library and sequenced in two runs on a MinION device with an R9.4.1 flow cell [Oxford Nanopore Technologies (ONT)]. The resulting data (Table 1) were basecalled using Guppy v.6.4.2 (ONT) in the HAC mode and quality-checked by NanoPlot v.1.33.1 (21). DNA was also sequenced on NovaSeq 6000 (Table 1), and the quality of the reads was checked by FastQC v.0.12.1 (22).

TABLE 1 Sequencing data used in the assembly and annotation of the C. parapsilosis SR23 genome

Sample	Sequencing platforma	Library kit	Total amount (Gbp)	Number of reads	ENA accession number	Cultivation medium	
DNA	MinION Mk1B	SQK-LSK109	4.96	659,024
(N50 = 13,492 nt)	ERR12736264, ERR12736265	YPD [1% (wt/vol) yeast extract, 2% (wt/vol) peptone, 2% (wt/vol) glucose]	
DNA	NovaSeq 6000	TruSeq DNA PCR-free (350 bp), paired ends (2 × 151 nt)	5.62	37,243,512	ERR12736259	YPD [1% (wt/vol) yeast extract, 2% (wt/vol) peptone, 2% (wt/vol) glucose]	
RNA	NovaSeq 6000	TruSeq Stranded mRNA LT Sample Prep, paired ends (2 × 151 nt)	4.32	28,597,422	ERR12736360	Synthetic with 4-hydroxybenzoate [0.17% (wt/vol) yeast nitrogen base without amino acids and ammonium sulfate (Difco), 0.5% (wt/vol) ammonium sulfate, 10 mM 4-hydroxybenzoate]	
RNA	NovaSeq 6000	TruSeq Stranded mRNA LT Sample Prep, paired ends (2 × 151 nt)	6.54	43,330,142	ERR12736361	Synthetic with galactose [0.17% (wt/vol) yeast nitrogen base without amino acids and ammonium sulfate (Difco), 0.5% (wt/vol) ammonium sulfate, 2% (wt/vol) galactose]	
RNA	NovaSeq 6000	TruSeq Stranded mRNA LT Sample Prep, paired ends (2 × 151 nt)	6.27	41,547,686	ERR12736362	Synthetic with hydroquinone [0.17% (wt/vol) yeast nitrogen base without amino acids and ammonium sulfate (Difco), 0.5% (wt/vol) ammonium sulfate, 10 mM hydroquinone]	
a Sequencing on Illumina NovaSeq 6000 platform was performed by Macrogen Europe (Amsterdam, the Netherlands).

Flye v.2.9.1-b1780 (23) was run on 40% of nanopore reads longer than 5 kbp. Nanopore reads were aligned to the assembly by minimap2 v.2.24-r1122 (24), and the assembly was manually adjusted based on the inspection of read alignments. Namely, low-quality sequences were truncated from two contigs, and one contig was split at a telomere. Two chromosomes were subsequently created by joining parts of two or three contigs, respectively. Mitochondrial DNA was replaced by the published version NC_005253.2 (4). Subsequently, the whole assembly was polished by three iterations of Pilon v.1.21 (25) using Illumina reads aligned by bwa mem v.0.7.17-r1188 (26). Copies of the rDNA repeat were replaced by a separately polished full-length copy. The nuclear genome assembly contains eight contigs with telomeres (5′-CCGGATGTTGATTATACTGAGGT-3′)n (27) on both ends corresponding to the electrophoretic karyotype and, thus, likely represent full-length chromosomes. Moreover, flow cytometry indicates that SR23 cells are diploid (Fig. 1).

Fig 1 (A) Nuclear chromosomes of C. parapsilosis SR23 visualized using Matplotlib library in Python (28). Genome regions with Illumina coverage higher than the overall median are shown in blue; these include rDNA (8.0×) and two segments with four genes (SCS7/CPARSR23_p51490, CPARSR23_p51500, CPARSR23_p51510, and BOR1/CPARSR23_p51520; 8.7×) and the gene ARR3/CPARSR23_p56030 (5.4×) on chromosomes 6, 7, and 8, respectively. (B) Electrophoretic karyotype of C. parapsilosis SR23. Chromosomal DNA was separated by pulsed-field gel electrophoresis in a 0.8% (wt/vol) agarose gel in 0.5× Tris-borate-Ethylenediaminetetraacetic acid (TBE) buffer (45 mM Tris-borate, 1 mM EDTA) using a Pulsaphor system (LKB) in contour-clamped homogeneous electric field configuration with pulse switching from 60 to 600 s for 72 h at 100 V and 9°C throughout (5). Note that chromosome 6 (labeled as 0.98R) contains multiple rDNA repeats and, therefore, appears longer than in the assembly. (C) Flow cytometry confirms the diploid nature of C. parapsilosis SR23 cells. DNA content was analyzed on a CytoFLEX S flow cytometer (Beckman Coulter) essentially as described in references (29, 30). Histograms show DNA content [propidium iodide (area)] vs cell counts. The peaks represent the cells in G1 and G2 phases of the cell cycle at the time of fixation. Diploid cells of the reference strain C. parapsilosis CDC317 (31) were used as a control.

For RNA-seq, SR23 cultures grew in three synthetic media differing by the carbon source (Table 1) at 28°C till optical density (OD)600 ~1. Total RNA was extracted using hot phenol (32), purified by an RNeasy Mini kit (Qiagen), and sequenced on NovaSeq 6000 (Table 1). The reads were trimmed by Trimmomatic v.0.36 (33) and assembled to transcripts by Trinity v.2.15.1 (34). These were mapped to the genome by blat v.36×7 (35) and used as evidence to predict protein-coding genes by Augustus v.3.2.3 (36) with species-specific parameters previously trained on C. parapsilosis CLIB214 (37). We manually adjusted 342 genes based on a comparison with RNA-seq data and genes from other strains. In total, 5,879 protein-coding genes were annotated in the nuclear genome. Full details of bioinformatics analyses can be found on Zenodo (38).

ACKNOWLEDGMENTS

We would like to thank Ladislav Kováč (Comenius University in Bratislava, Slovakia) and Geraldine Butler (Conway Institute, University College Dublin, Ireland) for providing us with the strains SR23 and CDC317, respectively. We also thank L. Kováč for long-term support and inspiring discussions. The research was supported by grants from the Slovak Research and Development Agency (APVV 22-0144, 19-0068), the Slovak Grant Agency (VEGA 1/0538/22, 1/0234/23, and 1/0031/24), and the Operation Program of Integrated Infrastructure for the project Advancing University Capacity and Competence in Research, Development and Innovation, ITMS2014+: 313021X329, co-financed by the European Regional Development Fund. The funders had no role in the study design, data collection and interpretation, or the decision to submit the work for publication.

DATA AVAILABILITY

The genome assembly has been deposited in the European Nucleotide Archive (ENA)/GenBank databases with accession number GCA_963989715. The genome sequence version described in this paper is the first version, GCA_963989715.1. The ENA accession numbers of Illumina and nanopore reads are shown in Table 1. The mitochondrial genome sequence was published previously (4) and is also available in GenBank and ENA databases under accession numbers NC_005253.2 and X74411.6, respectively. The results are also presented in a genome browser at http://genome.compbio.fmph.uniba.sk/.
==== Refs
REFERENCES

1 Kovác L, Lazowska J, Slonimski PP. 1984. A yeast with linear molecules of mitochondrial DNA. Mol Gen Genet 197 :420–424. doi:10.1007/BF00329938 6098800
2 Camougrand N, Mila B, Velours G, Lazowska J, Guérin M. 1988. Discrimination between different groups of Candida parapsilosis by mitochondrial DNA restriction analysis. Curr Genet 13 :445–449. doi:10.1007/BF00365667 2841034
3 Nosek J, Dinouël N, Kovac L, Fukuhara H. 1995. Linear mitochondrial DNAs from yeasts: telomeres with large tandem repetitions. Mol Gen Genet 247 :61–72. doi:10.1007/BF00425822 7715605
4 Nosek J, Novotna M, Hlavatovicova Z, Ussery DW, Fajkus J, Tomaska L. 2004. Complete DNA sequence of the linear mitochondrial genome of the pathogenic yeast Candida parapsilosis. Mol Genet Genomics 272 :173–180. doi:10.1007/s00438-004-1046-0 15449175
5 Rycovska A, Valach M, Tomaska L, Bolotin-Fukuhara M, Nosek J. 2004. Linear versus circular mitochondrial genomes: intraspecies variability of mitochondrial genome architecture in Candida parapsilosis. Microbiology (Reading) 150 :1571–1580. doi:10.1099/mic.0.26988-0 15133118
6 Kosa P, Valach M, Tomaska L, Wolfe KH, Nosek J. 2006. Complete DNA sequences of the mitochondrial genomes of the pathogenic yeasts Candida orthopsilosis and Candida metapsilosis: insight into the evolution of linear DNA genomes from mitochondrial telomere mutants. Nucleic Acids Res 34 :2472–2481. doi:10.1093/nar/gkl327 16684995
7 Valach M, Pryszcz LP, Tomaska L, Gacser A, Gabaldón T, Nosek J. 2012. Mitochondrial genome variability within the Candida parapsilosis species complex. Mitochondrion 12 :514–519. doi:10.1016/j.mito.2012.07.109 22824459
8 Mixão V, Del Olmo V, Hegedűsová E, Saus E, Pryszcz L, Cillingová A, Nosek J, Gabaldón T. 2022. Genome analysis of five recently described species of the CUG-Ser clade uncovers Candida theae as a new hybrid lineage with pathogenic potential in the Candida parapsilosis species complex. DNA Res 29 :dsac010. doi:10.1093/dnares/dsac010 35438177
9 Tomáska L, Nosek J, Fukuhara H. 1997. Identification of a putative mitochondrial telomere-binding protein of the yeast Candida parapsilosis. J Biol Chem 272 :3049–3056. doi:10.1074/jbc.272.5.3049 9006955
10 Nosek J, Tomáska L, Pagácová B, Fukuhara H. 1999. Mitochondrial telomere-binding protein from Candida parapsilosis suggests an evolutionary adaptation of a nonspecific single-stranded DNA-binding protein. J Biol Chem 274 :8850–8857. doi:10.1074/jbc.274.13.8850 10085128
11 Tomaska L, Nosek J, Makhov AM, Pastorakova A, Griffith JD. 2000. Extragenomic double-stranded DNA circles in yeast with linear mitochondrial genomes: potential involvement in telomere maintenance. Nucleic Acids Res 28 :4479–4487. doi:10.1093/nar/28.22.4479 11071936
12 Tomaska L, Makhov AM, Nosek J, Kucejova B, Griffith JD. 2001. Electron microscopic analysis supports a dual role for the mitochondrial telomere-binding protein of Candida parapsilosis. J Mol Biol 305 :61–69. doi:10.1006/jmbi.2000.4254 11114247
13 Tomaska L, Makhov AM, Griffith JD, Nosek J. 2002. t-Loops in yeast mitochondria. Mitochondrion 1 :455–459. doi:10.1016/s1567-7249(02)00009-0 16120298
14 Nosek J, Rycovska A, Makhov AM, Griffith JD, Tomaska L. 2005. Amplification of telomeric arrays via rolling-circle mechanism. J Biol Chem 280 :10840–10845. doi:10.1074/jbc.M409295200 15657051
15 Tomaska L, Nosek J, Kramara J, Griffith JD. 2009. Telomeric circles: universal players in telomere maintenance? Nat Struct Mol Biol 16 :1010–1015. doi:10.1038/nsmb.1660 19809492
16 Gerhold JM, Sedman T, Visacka K, Slezakova J, Tomaska L, Nosek J, Sedman J. 2014. Replication intermediates of the linear mitochondrial DNA of Candida parapsilosis suggest a common recombination based mechanism for yeast mitochondria. J Biol Chem 289 :22659–22670. doi:10.1074/jbc.M114.552828 24951592
17 Nosek J, Adamíková L, Zemanová J, Tomáska L, Zufferey R, Mamoun CB. 2002. Genetic manipulation of the pathogenic yeast Candida parapsilosis. Curr Genet 42 :27–35. doi:10.1007/s00294-002-0326-7 12420143
18 Zemanova J, Nosek J, Tomaska L. 2004. High-efficiency transformation of the pathogenic yeast Candida parapsilosis. Curr Genet 45 :183–186. doi:10.1007/s00294-003-0472-6 14648114
19 Kosa P, Gavenciakova B, Nosek J. 2007. Development of a set of plasmid vectors for genetic manipulations of the pathogenic yeast Candida parapsilosis. Gene 396 :338–345. doi:10.1016/j.gene.2007.04.008 17512139
20 Hodorová V, Lichancová H, Bujna D, Neboháčová M, Tomáška Ľ, Brejová B, Vinař T, Nosek J. 2018. De novo sequencing and high-quality assembly of yeast genomes using a MinION device. London Calling. London, UK. Available from: https://nanoporetech.com/resource-centre/de-novo-sequencing-and-high-quality-assembly-yeast-genomes-using-minion-device
21 De Coster W, Rademakers R. 2023. NanoPack2: population-scale evaluation of long-read sequencing data. Bioinformatics 39 :btad311. doi:10.1093/bioinformatics/btad311 37171891
22 Andrews S. 2010. FastQC: a quality control tool for high throughput sequence data. http://www.bioinformatics.babraham.ac.uk/projects/fastqc.
23 Kolmogorov M, Yuan J, Lin Y, Pevzner PA. 2019. Assembly of long, error-prone reads using repeat graphs. Nat Biotechnol 37 :540–546. doi:10.1038/s41587-019-0072-8 30936562
24 Li H. 2018. Minimap2: pairwise alignment for nucleotide sequences. Bioinformatics 34 :3094–3100. doi:10.1093/bioinformatics/bty191 29750242
25 Walker BJ, Abeel T, Shea T, Priest M, Abouelliel A, Sakthikumar S, Cuomo CA, Zeng Q, Wortman J, Young SK, Earl AM. 2014. Pilon: an integrated tool for comprehensive microbial variant detection and genome assembly improvement. PLoS One 9 :e112963. doi:10.1371/journal.pone.0112963 25409509
26 Li H, Durbin R. 2010. Fast and accurate long-read alignment with Burrows-Wheeler transform. Bioinformatics 26 :589–595. doi:10.1093/bioinformatics/btp698 20080505
27 Gunisova S, Elboher E, Nosek J, Gorkovoy V, Brown Y, Lucier J-F, Laterreur N, Wellinger RJ, Tzfati Y, Tomaska L. 2009. Identification and comparative analysis of telomerase RNAs from Candida species reveal conservation of functional elements. RNA 15 :546–559. doi:10.1261/rna.1194009 19223441
28 Hunter JD. 2007. Matplotlib: a 2D graphics environment. Comput Sci Eng 9 :90–95. doi:10.1109/MCSE.2007.55
29 Todd RT, Braverman AL, Selmecki A. 2018. Flow cytometry analysis of fungal ploidy. Curr Protoc Microbiol 50 :e58. doi:10.1002/cpmc.58 30028911
30 Gabaldón T, Martin T, Marcet-Houben M, Durrens P, Bolotin-Fukuhara M, Lespinet O, Arnaise S, Boisnard S, Aguileta G, Atanasova R, et al. . 2013. Comparative genomics of emerging pathogens in the Candida glabrata clade. BMC Genomics 14 :623. doi:10.1186/1471-2164-14-623 24034898
31 Butler G, Rasmussen MD, Lin MF, Santos MAS, Sakthikumar S, Munro CA, Rheinbay E, Grabherr M, Forche A, Reedy JL, et al. . 2009. Evolution of pathogenicity and sexual reproduction in eight Candida genomes. Nature 459 :657–662. doi:10.1038/nature08064 19465905
32 Collart MA, Oliviero S. 1993. Preparation of yeast RNA. Curr Protoc Mol Biol 23 :13. doi:10.1002/0471142727.mb1312s23
33 Bolger AM, Lohse M, Usadel B. 2014. Trimmomatic: a flexible trimmer for Illumina sequence data. Bioinformatics 30 :2114–2120. doi:10.1093/bioinformatics/btu170 24695404
34 Grabherr MG, Haas BJ, Yassour M, Levin JZ, Thompson DA, Amit I, Adiconis X, Fan L, Raychowdhury R, Zeng Q, Chen Z, Mauceli E, Hacohen N, Gnirke A, Rhind N, Palma F, Birren BW, Nusbaum C, Lindblad-Toh K, Friedman N, Regev A. 2011. Full-length transcriptome assembly from RNA-Seq data without a reference genome. Nat Biotechnol 29 :644. doi:10.1038/nbt.1883 21572440
35 Kent WJ. 2002. BLAT- the BLAST-like alignment tool. Genome Res 12 :656–664. doi:10.1101/gr.229202 11932250
36 Stanke M, Schöffmann O, Morgenstern B, Waack S. 2006. Gene prediction in eukaryotes with a generalized hidden Markov model that uses hints from external sources. BMC Bioinformatics 7 :62. doi:10.1186/1471-2105-7-62 16469098
37 Cillingová A, Tóth R, Mojáková A, Zeman I, Vrzoňová R, Siváková B, Baráth P, Neboháčová M, Klepcová Z, Brázdovič F, Lichancová H, Hodorová V, Brejová B, Vinař T, Mutalová S, Vozáriková V, Mutti G, Tomáška Ľ, Gácser A, Gabaldón T, Nosek J. 2022. Transcriptome and proteome profiling reveals complex adaptations of Candida parapsilosis cells assimilating hydroxyaromatic carbon sources. PLoS Genet 18 :e1009815. doi:10.1371/journal.pgen.1009815 35255079
38 Brejová B. 2024. Assembly of yeast Candida parapsilosis strain SR23 (fmfi-compbio/canParSR23: Version v1). Zenodo. doi:10.5281/zenodo.11214122.
