
==== Front
Microbiol Resour Announc
Microbiol Resour Announc
mra
Microbiology Resource Announcements
2576-098X
American Society for Microbiology 1752 N St., N.W., Washington, DC

39162455
mra00434-24
10.1128/mra.00434-24
mra.00434-24
Genome Sequences
applied-and-industrial-microbiologyApplied and Industrial MicrobiologyDraft genome sequence of rhamnolipid-producing bacteria Thermoanaerobacter thermocopriae strain CM-CNRG TB177 isolated from an oil reservoir in Mexico
https://orcid.org/0000-0002-3390-9283
Segovia Veronica 1 Methodology Writing – original draft
Martínez Pedro 2 Data curation Formal analysis Methodology Software Validation Visualization
https://orcid.org/0000-0002-8585-2469
Hernández-Gama Regina 2 Conceptualization Funding acquisition Project administration Resources Supervision Writing – review and editing rehernandez@ipn.mx

1 Unidad Profesional Interdisciplinaria de Ingeniería Campus Zacatecas, Instituto Politécnico Nacional , Zacatecas, Zacatecas, Mexico
2 Centro de Investigación en Ciencia Aplicada y Tecnología Avanzada Campus Querétaro, Instituto Politécnico Nacional , Querétaro, Querétaro, Mexico
Editor Stedman Kenneth M. Portland State University , Portland, Oregon, USA

Address correspondence to Regina Hernández-Gama, rehernandez@ipn.mx
Pedro Martínez and Regina Hernández-Gama contributed equally to this article.

The authors declare no conflict of interest.

9 2024
20 8 2024
20 8 2024
13 9 e00434-2430 4 2024
24 7 2024
Copyright © 2024 Segovia et al.
2024
Segovia et al.
https://creativecommons.org/licenses/by/4.0/ This is an open-access article distributed under the terms of the Creative Commons Attribution 4.0 International license.

ABSTRACT

The draft genome of Thermoanaerobacter thermocopriae CM-CNRG TB177 isolated from an oil reservoir in Mexico was determined and annotated. The organism is a thermophilic and strict anaerobe bacterium that produces rhamnolipids, using glucose as a carbon source. The predicted genome size is 2,496,169 bp and 2,550 genes.

KEYWORDS

Thermoanaerobacter
genomes
rhamnolipid
biosurfactants
oil reservoir
CONAHCYT 257164 Hernandez-Gama Regina cover-dateSeptember 2024
==== Body
pmcANNOUNCEMENT

The genus Thermoanaerobacter includes Gram-positive, thermophilic bacteria (1). Some species belonging to this genus have been studied for the production of biotechnologically important metabolites such as ethanol and enzymes, among others (2–4). The draft genome sequence presented here was particularly studied for rhamnolipid production (5). In addition to its biosurfactant production, the Thermoanaerobacter thermocopriae CM-CNRG TB177 strain is an extremophile bacterium, capable of thriving in temperatures ranging from 50°C to 80°C and in salt concentrations of up to 30 g/L of NaCl. These characteristics make this microorganism particularly valuable for industrial applications such as microbial enhanced oil recovery and mining.

The T. thermocopriae isolate was previously obtained from formation water at the Chicontepec oil field in Mexico (6). It was subsequently cultured in basal medium with glucose as the carbon source, as previously reported (7). Genomic DNA was extracted using the commercial ZymoBIOMICS DNA Miniprep Kit (Zymo Research, USA), and the DNA quality was verified using a NanoDrop One (Thermo Scientific, USA).

The DNA was sequenced twice. First, sequencing is performed using Illumina NextSeq 500 with the TruSeq DNA PCR-Free Sample Prep Kit in a 2 × 75 cycle configuration (Illumina, USA) by the Massive Sequencing Unit of the UNAM Institute of Biotechnology, constructing genomic DNA libraries with 350-bp fragments. The second sequencing was performed using Illumina MiSeq device in a 2 × 300 run format by the Genomic Services Laboratory of the National Laboratory of Genomics for Biodiversity (LANGEBIO), using the Illumina TruSeq DNA Nano Kit to construct the genomic DNA libraries using 550-bp fragments.

From the first sequencing, 28,471,116 paired raw sequences of 75 bp were obtained, and from the second sequencing, 1,341,482 paired raw sequences of 300 bp were obtained. The initial quality of the sequences was analyzed using the FastQC v0.11.9 program (8). Subsequently, the sequences were subjected to a cleaning process with the TrimmomaticPE v0.39 program (9). For the first sequencing data, the parameters used were a sliding window of 4:20, a head crop of one, trailing of two, and minimum length of 60. For the second sequencing data, the parameters used were a sliding window of 5:20 and the Ilumina clip to eliminate adaptors.

After cleaning 12,737,254 paired sequences of approximately 75 bp and 11,255,044 paired sequences of length 200–299 bp were obtained, with these reads, a coverage of 1,000 was estimated. From these sequences, genome hybrid assembly was performed using SPAdes version 3.15.3 (10), using only the assembly without sequence editing. The assembly generated 294 contigs, with a sequence length of 2,496,169, an N50 value of 94,187 bp; the longest contig was 252,232 bp, with a guanine and cytosine content of 34.68%. All tools were run with default parameters unless otherwise specified.

Genome annotation was performed using the Prokaryotic Genome Annotation Pipeline from the National Center for Biotechnology Information (11) version 6.7; as a result, 2,550 genes were obtained in the genome, including four complete rRNAs, 57 tRNAs, and four clustered regularly interspaced short palindromic repeats arrays. A genomic analysis was conducted using the GTDB-Tk toolkit (12), resulting in the assignment of T. thermocopriae.

ACKNOWLEDGMENTS

Financial support for this work was provided by Consejo Nacional de Humanidades Ciencias y Tecnologías (257164).

DATA AVAILABILITY

The sequence reads have been sumited to the Sequence Read Archive (SRA) under the accession number PRJNA1064994 (https://trace.ncbi.nlm.nih.gov/Traces/sra-reads-be/fastq?acc=SRR28816973 and https://trace.ncbi.nlm.nih.gov/Traces/sra-reads-be/fastq?acc=SRR28816972).

This Whole Genome Shotgun project has been deposited at DDBJ/ENA/GenBank under the accession JBBLMQ000000000 and the Bioproject accesion number PRJNA1064994. The version described in this announcement is version JBBLMQ010000000.
==== Refs
REFERENCES

1 Schleifer KH. 2009. Phylum XIII. Firmicutes Gibbons and Murray 1978, p 387–68489. In De Vos P (ed), Bergey’s manual of systematic bacteriology. Springer, New York.
2 Musa MM, Vieille C, Phillips RS. 2021. Secondary alcohol dehydrogenases from Thermoanaerobacter pseudoethanolicus and Thermoanaerobacter brockii as robust catalysts. Chembiochem 22 :1884–1893. doi:10.1002/cbic.202100043 33594812
3 Tian LF, Gao H, Yang S, Liu YP, Li M, Xu W, Yan XX. 2023. Structure and function of extreme TLS DNA polymerase TTEDbh from Thermoanaerobacter tengcongensis. Int J Biol Macromol 253 :126770. doi:10.1016/j.ijbiomac.2023.126770 37683741
4 Burger Y, Schwarz FM, Müller V. 2022. Formate-driven H2 production by whole cells of Thermoanaerobacter kivui. Biotechnol Biofuels Bioprod 15 :48. doi:10.1186/s13068-022-02147-5 35545791
5 Segovia V, Reyes A, Rivera G, Vázquez P, Velazquez G, Paz-González A, Hernández-Gama R. 2021. Production of rhamnolipids by the Thermoanaerobacter sp. CM-CNRG TB177 strain isolated from an oil well in Mexico. Appl Microbiol Biotechnol 105 :5833–5844. doi:10.1007/s00253-021-11468-8 34396489
6 Hernández R. 2014. Analysis of the microbial diversity of a Chicontepec well and the conditions in which it favors the emulsification of crude oil, PhD thesis. Instituto Politecnico Nacional. https://tesis.ipn.mx/bitstream/handle/123456789/19572/REGINA.pdf?sequence=1&isAllowed=y.
7 Feng X, Mouttaki H, Lin L, Huang R, Wu B, Hemme CL, He Z, Zhang B, Hicks LM, Xu J, Zhou J, Tang YJ. 2009. Characterization of the central metabolic pathways in Thermoanaerobacter sp. strain X514 via isotopomer-assisted metabolite analysis. Appl Environ Microbiol 75 :5001–5008. doi:10.1128/AEM.00715-09 19525270
8 Andrews S. 2010. FastQC: a quality control tool for high throughput sequence data. Available from: http://www.bioinformatics.babraham.ac.uk/projects/fastqc
9 Bolger AM, Lohse M, Usadel B. 2014. Trimmomatic: a flexible trimmer for Illumina sequence data. Bioinformatics 30 :2114–2120. doi:10.1093/bioinformatics/btu170 24695404
10 Bankevich A, Nurk S, Antipov D, Gurevich AA, Dvorkin M, Kulikov AS, Lesin VM, Nikolenko SI, Pham S, Prjibelski AD, Pyshkin AV, Sirotkin AV, Vyahhi N, Tesler G, Alekseyev MA, Pevzner PA. 2012. SPAdes: a new genome assembly algorithm and its applications to single-cell sequencing. J Comput Biol 19 :455–477. doi:10.1089/cmb.2012.0021 22506599
11 Tatusova T, DiCuccio M, Badretdin A, Chetvernin V, Nawrocki EP, Zaslavsky L, Lomsadze A, Pruitt KD, Borodovsky M, Ostell J. 2016. NCBI prokaryotic genome annotation pipeline. Nucleic Acids Res 44 :6614–6624. doi:10.1093/nar/gkw569 27342282
12 Chaumeil PA, Mussig AJ, Hugenholtz P, Parks DH. 2020. GTDB-Tk: a toolkit to classify genomes with the genome taxonomy database. Bioinformatics 36 :1925–1927. doi:10.1093/bioinformatics/btz848
