
==== Front
Chin Med J (Engl)
Chin Med J (Engl)
CM9
Chinese Medical Journal
0366-6999
2542-5641
Lippincott Williams & Wilkins Hagerstown, MD

38934052
CMJ-2024-517
10.1097/CM9.0000000000003182
00008
3
Original Article
Enhancing antimicrobial resistance detection with MetaGeneMiner: Targeted gene extraction from metagenomes
Liu Chang 1
Tang Zizhen 2
Li Linzhu 2
Kang Yan 1
Teng Yue 3
Yu Yan 2
Zhou Sihan
Hao Xiuyuan
1 Department of Critical Care Medicine, West China Hospital, Sichuan University, Chengdu, Sichuan 610041, China
2 Key Laboratory of Bio-Resources and Eco-Environment of Ministry of Education, College of Life Sciences, Sichuan University, Chengdu, Sichuan 610065, China
3 State Key Laboratory of Pathogen and Biosecurity, Beijing Institute of Microbiology and Epidemiology, Beijing 100071, China
Correspondence to: Yan Yu, Key Laboratory of Bio-Resources and Eco-Environment of Ministry of Education, College of Life Sciences, Sichuan University, Chengdu, Sichuan 610065, China E-Mail: yyu@scu.edu.cn;
Yue Teng, State Key Laboratory of Pathogen and Biosecurity, Beijing Institute of Microbiology and Epidemiology, Beijing 100071, China E-Mail: yueteng@me.com
27 6 2024
05 9 2024
137 17 20922098
18 2 2024
Copyright © 2024 The Chinese Medical Association, produced by Wolters Kluwer, Inc. under the CC-BY-NC-ND license.
2024
https://creativecommons.org/licenses/by-nc-nd/4.0/ This is an open access article distributed under the terms of the Creative Commons Attribution-Non Commercial-No Derivatives License 4.0 (CCBY-NC-ND), where it is permissible to download and share the work provided it is properly cited. The work cannot be changed in any way or used commercially without permission from the journal. http://creativecommons.org/licenses/by-nc-nd/4.0

Abstract

Background:

Accurately and efficiently extracting microbial genomic sequences from complex metagenomic data is crucial for advancing our understanding in fields such as clinical diagnostics, environmental microbiology, and biodiversity. As sequencing technologies evolve, this task becomes increasingly challenging due to the intricate nature of microbial communities and the vast amount of data generated. Especially in intensive care units (ICUs), infections caused by antibiotic-resistant bacteria are increasingly prevalent among critically ill patients, significantly impacting the effectiveness of treatments and patient prognoses. Therefore, obtaining timely and accurate information about infectious pathogens is of paramount importance for the treatment of patients with severe infections, which enables precisely targeted anti-infection therapies, and a tool that can extract microbial genomic sequences from metagenomic dataset would be of help.

Methods:

We developed MetaGeneMiner to help with retrieving specific microbial genomic sequences from metagenomes using a k-mer-based approach. It facilitates the rapid and accurate identification and analysis of pathogens. The tool is designed to be user-friendly and efficient on standard personal computers, allowing its use across a wide variety of settings. We validated MetaGeneMiner using eight metagenomic samples from ICU patients, which demonstrated its efficiency and accuracy.

Results:

The software extensively retrieved coding sequences of pathogens Acinetobacter baumannii and herpes simplex virus type 1 and detected a variety of resistance genes. All documentation and source codes for MetaGeneMiner are freely available at https://gitee.com/sculab/MetaGeneMiner.

Conclusions:

It is foreseeable that MetaGeneMiner possesses the potential for applications across multiple domains, including clinical diagnostics, environmental microbiology, gut microbiome research, as well as biodiversity and conservation biology. Particularly in ICU settings, MetaGeneMiner introduces a novel, rapid, and precise method for diagnosing and treating infections in critically ill patients. This tool is capable of efficiently identifying infectious pathogens, guiding personalized and precise treatment strategies, and monitoring the development of antibiotic resistance, significantly impacting the diagnosis and treatment of severe infections.

Keywords:

Metagenomics
Genomic sequencing
Clinical diagnostics
Reference-based assembly
Intensive care unit infections
Antibiotic resistance
OPEN-ACCESSTRUE
SDCT
==== Body
pmcIntroduction

Antimicrobial resistance (AMR) has emerged as a critical issue in contemporary medicine. It enables microbes to withstand the concentrations of antibiotics used in clinical treatments,[1] leading to the loss of efficacy of antimicrobial drugs and an increase in morbidity and mortality rates.[2,3] AMR represents a significant challenge in the field of global public health. In the United States alone, 2 million people annually contract infections resistant to first-line antimicrobials, resulting in healthcare costs of 20 billion dollars.[4,5] Tracking the emergence and prevalence of AMR is crucial for minimizing its threat to human health.[4] Antimicrobial susceptibility testing (AST) serves as the traditional method for determining the resistance of bacteria to antimicrobial drugs. These culture-based tests assess bacterial growth in the presence of antimicrobial agents. Although culture-based resistance assays can provide essential information for patient management and the epidemiology of resistant genes, they have limitations in implementation and the breadth of information they offer.[6] AST requires specialized facilities and personnel, and is limited to cultivable bacteria, excluding many unculturable species in diverse microbial communities.[7]

AMR is encoded by genes and arises through mechanisms like gene overexpression, mutations, or horizontal gene transfer. Advances in sequencing now make metagenomic sequencing a vital method for quickly identifying resistance genes. Techniques and methodologies based on metagenomic next-generation sequencing (mNGS) complement traditional culture-based approaches for clinical and surveillance applications,[8] offering opportunities for rapidly and sensitively determining the resistance of both cultivable and uncultivable bacteria.[91011] Several metagenomic analysis techniques have been established, including both assembly-based and read-based approaches.[12] Short reads, produced by technologies like Illumina, may be processed through assembly-based methods.[1314151617] This involves initially assembling sequencing reads into contiguous fragments (contigs), which are then annotated through comparison with either custom or public reference databases. Alternatively, reads can be directly analyzed using read-based methods, where resistance determinants are identified by mapping the reads directly to a reference database.[12]De novo metagenomic assembly is adept at discovering novel organisms within communities, although its sensitivity is frequently constrained by the complexity of the environment.[13] In contrast, read-based computational approaches, which accurately identify and quantify known taxonomic groups and genes in the microbiome through homology, serve as a valuable complement to assembly-based methods.[14,15] These methods have enabled a deeper exploration of the microbiome, uncovering correlations between microbes and health status,[16,17] dietary habits,[18,19] as well as the characterization of the evolution of microbial species and strains.[20,21]

Read-based methods are constrained by the availability of reference genomes, and the requirement for known microbial species limits the interpretation of metagenomes to some extent.[22] Fortunately, with the rapid development of read-based metagenomic analysis methods and sorting technologies in recent years, extensive metagenome-assembled genomes (MAG) catalogs have been compiled.[17] Tools built on these catalogs can accurately analyze the presence and abundance of various microbes in metagenomes, allowing for more in-depth and precise quantitative taxonomic analyses of human, host-associated, and environmental microbiomes.[10,15,23] In specific application scenarios, such as clinical diagnostics or indicator microbial testing, researchers often focus on particular organism groups.[24] Beyond the quantity of microbes, their genomic sequences can offer insights into evolutionary and dissemination characteristics. However, due to the typically low sequencing depth of metagenomes and the complexity of microbial community compositions, a substantial amount of metagenomic data remains underutilized.[18]

To fully leverage the sequence information in metagenomes, we have developed MetaGeneMiner, a read-based tool capable of mining specific microbial genomic sequences from metagenomes. MetaGeneMiner incorporates a k-mer-based, multi-reference filtering, and optimal reference sequence selection process specifically for metagenomes[25] and integrates it with a reference-based assembly pipeline constructed around minimap2.[26]

Methods

MetaGeneMiner is designed with the aim of recovering genomic data from shotgun metagenomic sequencing. We developed three modules to achieve an efficient metagenomic pipeline: read filtering, optimal reference selection, and reference-based assembly. We further validated MetaGene Miner using eight clinical samples with bacterial or viral metagenomes.

Read filtering

To utilize MetaGeneMiner, obtaining reference sequences for the target taxa is a prerequisite. Researchers can initially use tools like MetaPhlAn (Blanco-Míguez et al[15], https://huttenhower.sph.harvard.edu/metaphlan/) or Kraken (Derrick et al[27], https://ccb.jhu.edu/software/kraken2/) to determine the microbial composition of the metagenome. From this, the target taxa of interest are selected, and corresponding genes or genomic sequences are downloaded from the GenBank database (https://www.ncbi.nlm.nih.gov/genbank/) to serve as references. MetaGeneMiner has the capability to autonomously segment each coding sequence (CDS) within gb files, and, additionally, it offers support for importing standalone fasta files. To extract CDS from gb files, MetaGeneMiner iterates over the gene records in the file. From the records, MetaGeneMiner parses the coding range of the gene. Then the CDS, or a fragment of the CDS, can be retrieved from the genome with a simple lookup. When the parsing finishes, MetaGeneMiner joins multiple fragments as necessary and saves the full-length CDS in an intermediate file. If a gene is annotated more than once in the genome, MetaGeneMiner assumes there are multiple variations for that gene and saves all of them as possible candidates. Similarly, MetaGeneMiner will also consider sequences with the same name in different genomes as multiple possible versions of a gene. The sequence filtering method in MetaGeneMiner aligns with that of our previously developed GeneMiner [Figure 1A].[25] It starts by dividing the reference sequences of multiple genes into a k-mers dictionary, essentially a collection of length-k strings of nucleotides. The k-mer size, kf , is adjustable by the user. Similarly, each raw read is split into k-mers using the same kf value as set by the user. If a read’s k-mer matches one in the reference hash table, the entire read is retained and allocated to the filtered dataset of the target gene. This method allows MetaGeneMiner to process multiple genes and samples simultaneously, significantly enhancing analysis speed. In MetaGeneMiner, the default kf value is 25, and with a degree of variation (v) between the reference and target sequences at 0.1, and read length (r) at 150, the ideal conditions for the acquisition rate (p) of reads targeting a specific gene is p = 1−(1−vkf)r−kf+1= 99.99%. Generally, the lower the kf , the more unrelated reads are filtered out, resulting in slower assembly speed, lower accuracy of the target sequence, and longer assembly lengths. Conversely, excessively high kf values can lead to the loss of relevant reads, preventing the assembly of complete sequences. We recommend users start with the default kf value in their initial analysis and adjust it based on the results of the assembly, if necessary. Compared to direct mapping, splitting the CDS and filtering can significantly reduce the time and disk space consumed in subsequent assembly steps, thereby enhancing overall performance.

Figure 1 The workflow of MetaGeneMiner. (A) Read filtering. (B) Optimal reference selection. (C) Reference-based assembly. (D) The GUI of MetaGeneMiner. GUI: Graphical user interface.

Optimal reference selection

When a sample’s genome contains multiple reference sequences, the program automatically selects the most optimal reference sequence for the subsequent assembly step [Figure 1B]. The method involves re-establishing a k-mers hash table for the sample’s genome and comparing it with the filtered fastq file. During optimal reference selection, a k-mers hash table is created from filtered reads of each gene. All reference sequences that belong to a specific gene will be tested against the respective hash table. MetaGeneMiner then calculates the number of reads that each reference sequence hits. The reference sequence with the highest number of matching reads is designated as the optimal reference. This procedure is optimized for NGS data and intends to leverage as many raw reads as possible. In the process of reference-based assembly of microbial genomes, pre-selecting the optimal reference sequence offers several benefits. Firstly, it enhances the accuracy of the assembly by providing a closely related template, which helps in accurately aligning short sequencing reads. This is crucial for reconstructing regions with high similarity or repetitive sequences. Secondly, an optimal reference can improve the efficiency of the assembly process, as it reduces computational resources and time by directing the assembly algorithm toward more relevant genomic regions. Thirdly, using a well-chosen reference sequence can aid in the identification of novel genomic features, such as unique genes or variations, by providing a clear contrast between the reference and the newly assembled genome. In optimal reference selection, the k-mer-based approach offers significant advantages in speed and efficiency compared to sequence alignment tools like Basic Local Alignment Search Tool (BLAST, National Center for Biotechnology Information [NCBI], https://blast.ncbi.nlm.nih.gov).[28] This is because it analyzes and compares genomes using short sequence fragments rather than relying on local or global alignments between sequences. As such, it is particularly suited for processing metagenomic data.[27] Additionally, this method may be more effective when dealing with highly variable sequences, such as those found in microbial communities, as it does not depend on long-segment consistency between sequences.[29] However, for highly repetitive sequences, the k-mer method might not be as effective as alignment approaches and does not provide detailed alignment information, such as specific sequence differences, insertions, deletions, or rearrangements.

Reference-based assembly

In the assembly step, we initially utilize minimap2 to map the filtered fastq file onto the optimal reference sequence, generating a sam file. Subsequently, this sam file is used to construct a consensus sequence [Figure 1C]. The process of achieving consensus borrows from the sam2consensus code and employs the same methodology as Geneious (GraphPad Software, LLC, https://www.geneious.com).[30] In the consensus step, the threshold (-c) parameter is critical for determining which base is called in the consensus and can be set to a specific percentage. International Union of Pure and Applied Chemistry (IUPAC) ambiguity codes, such as R for an A or G nucleotide, provide fractional support for each nucleotide in the ambiguity set, and regions lacking coverage are filled with gaps (–). For instance, consider a column containing 6 As, 3 Gs, and 1 T. If the consensus threshold is set to 60% or lower, the consensus will be A. If set between 60% and 90%, the consensus will become R. If over 90%, it will be D. Due to the typically low sequencing depth for specific species in metagenomic sequencing, a higher threshold may lead to an abundance of degenerate bases. Therefore, the software defaults to a threshold of 25%. MetaGeneMiner supports the retrieval of results for multiple thresholds simultaneously and can export the nucleotide composition ratios at each site. This provides researchers with a wealth of information and facilitates integration with a broader range of downstream analytical processes.

User interface

MetaGeneMiner is a user-friendly Python program that, after substantial optimization, offers performance comparable to similar programs written in C. It is distributed with two distinct user interfaces: a graphical user interface (GUI) as shown in Figure 1D and a command-line interface (CLI). In the Windows GUI version, extensive enhancements have been made in terms of file management and usability, allowing users to conveniently save and reuse results with all operations achievable via mouse clicks. Leveraging the high performance of the k-mer method, users can efficiently process multiple metagenomes in parallel to mine genomic data of multiple target species, with the entire analysis feasible on a standard personal computer with 8 GB of memory or more. For those looking to integrate MetaGeneMiner into other workflows, the CLI offers a more detailed set of parameters for advanced users.

Clinical validation

To evaluate the effectiveness and efficiency of MetaGene Miner, we conducted a validation study using eight metagenomic samples from the Department of Critical Care Medicine at Sichuan University, which was approved by the Sichuan University Medical Ethics Committee (No. 2024-924). The samples were collected from patients in intensive care units (ICUs) who were infected by Acinetobacter baumannii (A. baumannii) or herpes simplex virus type 1 (HSV-1). Metagenomic sequencing was carried out on the Illumina NextSeq 550 platform, employing a 75-cycle single-end sequencing strategy. Post-sequencing data were processed using fastp (Chen et al[31], https://github.com/OpenGene/fastp) to remove low-quality sequences, adapter contamination, duplicate sequences, and low complexity sequences. The identification and exclusion of human-origin sequences were conducted by mapping to the human reference genome (hg38) using the Burrows-Wheeler Aligner (Li et al[32], https://bio-bwa.sourceforge.net). The sequencing data can be downloaded from https://sourceforge.net/projects/metageneminer/files/TEST/.

Results

In our study, MetaGeneMiner was employed to extract gene sequences of A. baumannii and HSV-1 based on one reference genome from Hamidian et al[33] (GenBank ID CP045110) and 11 reference genomes from Johnston et al[34], respectively. For A. baumannii, MetaGeneMiner successfully retrieved an average of 1764 CDS out of a total of 1974 identified, with an average of 1154 CDS having more than 50% coverage, 703 with over 80% coverage, and 77 with more than 99% coverage. For HSV-1, each reference genome contained 75 gene annotations, all of which were utilized as references in MetaGeneMiner. The software managed to recover all of the 75 genes, with all but one having more than 50% of bases recovered, an average of 68 genes with over 80% recovery, and 14 genes with more than 99% recovery. The details of samples and their respective gene recovery statistics are presented in Table 1.

Table 1 Information of the tested samples.

Sample ID	Reads count	Sample type*	Filtered reads	Recovered	Recovered bases	
50%	80%	99%	
HQ0011057811	10,273,503	A-1	4,581,012	1829	1285	870	123	
HQ0011069862	8,033,450	A-1	2,848,786	1645	1042	567	52	
HQ1011088925	8,005,360	A-2	3,031,534	1791	1142	678	63	
HQ1011091052	10,690,404	A-2	2,483,018	1792	1148	695	71	
HQ0011054748	1,369,073	B-1	726,297	75	75	73	19	
HQ0011055111	889,131	B-1	344,147	75	75	65	11	
HQ0011054847	759,185	B-2	326,091	75	74	63	4	
HQ0011074099	925,976	B-2	517,026	75	75	73	24	
*Sample Type: A: Acinetobacter baumannii; B: HSV-1; 1: Sputum; 2: Bronchoalveolar lavage fluid. HSV-1: Herpes simplex virus type 1.

In the test environment (i9-12900KF, 4.25 GHz, single-threaded), MetaGeneMiner analyzed A. baumannii and HSV-1 in 119.2 min and 7.4 min, respectively. In comparison, mapping the raw data to all best references sequentially using minimap2 took approximately 423.8 min for A. baumannii and 8.5 min for HSV-1. The results obtained through consensus were entirely consistent with those from MetaGeneMiner, highlighting its significant speed advantage, particularly in scenarios involving a large number of reference genes and extensive sequencing data.

Our resistance testing is carried out using the same sequencing data from the genome recovery test. We utilized the MEGARes V3.0 database,[35] which includes sequence data for nearly 9000 hand-curated resistance genes related to antimicrobial drugs, biocides, and metals, to test MetaGeneMiner’s competency to recover functional genes, especially drug resistance genes. We considered genes with a recovered sequence length greater than 50% and coverage greater than 2 as potential resistance genes, with their names and corresponding resistances listed in Table 2. Two of the HSV-1 samples (HQ0011055111 and HQ0011074099) do not harbor any genes related to AMR. The other two HSV-1 samples contain a few resistance genes, likely from bacterial coinfection. Detailed information about the recovered genes, including coverage and recovery rates, is provided in Supplementary Material, http://links.lww.com/CM9/C51.

Table 2 Results of resistance testing.

Sample ID	Drugs	Genes	
HQ0011057811	Elfamycins	TUFAB	
Sulfonamides	FOLP; SULI; SULII	
Lipopeptides	LPXC	
Beta-lactams	ADC; OXA; TEM	
Fluoroquinolones	PARC	
Aminocoumarins	PARE	
Rifampin	RPOB	
HQ0011069862	Aminoglycosides	A16S; RRSA; RRS	
Beta-lactams	CTX; OXA; TEM	
Macrolide, lincosamide, and streptogramin	ERMB; ERMC	
Sulfonamides	FOLP; SULI; SULII	
Lipopeptides	LPXC	
Fluoroquinolones	PARC	
HQ1011088925	Aminoglycosides	AAC6-PRIME; RRSC; RRSH	
Glycopeptides	BRP	
Cationic antimicrobial peptides	CAP16S	
Beta-lactams	CTX; NDM; OHIO; OXA; SHV; TEM; AAK	
Sulfonamides	FOLP; SULI; SULII	
Lipopeptides	LPXC; ARNT; EPTB	
Macrolide, lincosamide and streptogramin	MLS23S	
Multi-drug resistance	IBCR; OMPF; LPTD; OMPK	
Fluoroquinolones	PARC; QNRB	
Aminocoumarins	PARE	
Tetracyclines	TETA	
Elfamycins	TUFAB	
HQ1011091052	Aminoglycosides	APH3-DPRIME; APH3-PRIME; A16S; RRSC; RRSH	
Cationic antimicrobial peptides	CAP16S; UGD	
Beta-lactams	CTX; KPC; OHIO; OXA; SHV; TEM; AMPH; AAK	
Sulfonamides	FOLP; SULI; SULII	
Fosfomycin	FOSA; MURA	
Fluoroquinolones	GYRA; PARC; QNRS	
Lipopeptides	LPXC; ARNT; EPTB	
Macrolide, lincosamide, and streptogramin	MLS23S	
Multi-drug resistance	MSBA; OMPF; LPTD; OMPK	
Aminocoumarins	PARE	
Tetracyclines	TETA; TETR	
Elfamycins	TUFAB	
HQ0011054748	Macrolide, lincosamide, and streptogramin	ERMB	
Sulfonamides	FOLP; SULI; SULII	
Beta-lactams	OXA	
HQ0011054847	Beta-lactams	OXA; TEM	

In comparison to the AMRplusplus[35] workflow, the range of antibiotic resistance genes identified by MetaGeneMiner fully encompasses those found in its results. The higher the values of recovered sequence length and coverage used in the screening process, the fewer the types of resistance genes are identified. In clinical practice, the outcomes of metagenomic analysis are influenced by various factors such as sample quality, type, sequencing depth, and read length, implying that the parameters used in this study may not be suitable for all scenarios. Moreover, the presence of corresponding resistance genes in microbes does not necessarily translate to a resistant phenotype; confirmation of resistance requires consideration of patient symptoms and other auxiliary diagnostic methods and tests. Within the eight samples examined in this study, it was observed that patients infected with A. baumannii harbored a rich diversity of resistance genes, which explained the frequent recurrence of hospital-acquired A. baumannii infection in ICUs.

Discussion

In China, data from China Antimicrobial Surveillance Network (CHINET, www.chinets.com) indicate that the resistance rate of A. baumannii to most antimicrobial drugs has remained above 50%. From 2005 to 2019, the resistance rate of A. baumannii to various antimicrobials significantly increased, with resistance to imipenem and meropenem rising from 31.0% and 39%, respectively, to 73.6% and 75.1%, reaching 78.6% and 79.5% by 2023.[36] Looking forward, we hope to employ MetaGene Miner for microbial resistance molecular epidemiology, resistance mechanisms, and transmission studies. This will enable timely understanding of the epidemiological characteristics and trends of microbial resistance across different regions, populations, healthcare facilities, animals, and environmental settings. It may contribute to elucidating the mechanisms of microbial pathogenicity, resistance, and their transmission, providing scientific data to formulate strategies for resistance prevention and control, and to support the research and development of new drugs and technologies. Overall, MetaGeneMiner showcases proficient performance with metagenomic data from various sources, and the sequences retrieved are well suited for further analyses, including resistance locus analysis, selection pressure examination, and phylogenetic studies.

In addition to clinical testing, MetaGeneMiner has the potential to help with environmental biology, gut microbiome research, or biodiversity. The use of MetaGeneMiner can be expanded to pathogen gene extraction from environmental DNA samples, facilitating low-cost monitoring of the spread of zoonotic diseases. A promising application would be monitoring how climate change affects the adaptation and virulence of pathogens; for example, the hypermutation of virulence genes in Vibrio parahaemolyticus.[37] MetaGeneMiner can also be used to determine the presence of specific functional genes in the gut microbiota, empowering researchers to corroborate the metabolism invoked by gut microbes. However, MetaGeneMiner, in its current form, is optimized for pathogen genome extraction. Further validation is required for its use in other fields.

In summary, MetaGeneMiner represents a software tool with state-of-the-art performance, which is designed to efficiently and accurately extract specific microbial genomic sequences from metagenomes. Users simply need to provide reference sequences for their target species, enabling them to retrieve the desired genetic sequences from metagenomes with high precision. This tool exhibits tremendous potential for application across various fields, particularly in critical care medicine, where its significance lies in rapidly identifying pathogens, guiding personalized treatment strategies, and monitoring the development of antibiotic resistance. Moreover, MetaGeneMiner demonstrates its extensive applicability in areas such as clinical diagnostics, environmental microbiology, gut microbiome research, as well as in biodiversity and conservation biology. Designed for accessibility, MetaGeneMiner features both graphical and CLIs, making it a versatile tool for diverse genomic research applications. In response to the specific demands of critical care medicine, we will continue to refine MetaGeneMiner to enhance its precision, speed, and memory efficiency. MetaGeneMiner is especially valuable for handling urgent and complex clinical situations, offering rapid results while ensuring accuracy, which is crucial for the treatment of critically ill patients. MetaGeneMiner is available freely from https://gitee.com/sculab/MetaGeneMiner and is licensed under the terms of the MIT license.

Acknowledgments

We acknowledge Dr. Han Kang from the Life Science Core Facilities, College of Life Science, Sichuan University, for technical support in analysis.

Funding

This work was supported by grants from the National Natural Science Foundation of China (Nos. 32071666 and 32271552) and the Science & Technology Fundamental Resources Investigation Program (No. 2022FY101000).

Supplementary Material

SUPPLEMENTARY MATERIAL

Chang Liu and Zizhen Tang contributed equally to this work.

How to cite this article: Liu C, Tang ZZ, Li LZ, Kang Y, Teng Y, Yu Y. Enhancing antimicrobial resistance detection with MetaGeneMiner: Targeted gene extraction from metagenomes. Chin Med J 2024;137:2092–2098. doi: 10.1097/CM9.0000000000003182
==== Refs
References

1. Lee JH . Perspectives towards antibiotic resistance: From molecules to population. J Microbiol 2019;57 :181–184. doi: 10.1007/s12275-019-0718-8.30806975
2. Tillotson GS Zinner SH . Burden of antimicrobial resistance in an era of decreasing susceptibility. Expert Rev Anti Infect Ther 2017;15 :663–676. doi: 10.1080/14787210.2017.1337508.28580804
3. Hawkey P . The growing burden of antimicrobial resistance. J Antimicrob Chemother 2008;62 (Suppl_1 ):i1–i9. doi: 10.1093/jac/dkn241.18684701
4. Cassini A Högberg LD Plachouras D Quattrocchi A Hoxha A Simonsen GS , . Attributable deaths and disability-adjusted life-years caused by infections with antibiotic-resistant bacteria in the EU and the European Economic Area in 2015: A population-level modelling analysis. Lancet Infect Dis 2019;19 :56–66. doi: 10.1016/S1473-3099(18)30605-4.30409683
5. O’Neill J . Tackling a crisis for the health and wealth of nations. Review on Antimicrobial Resistance 2014. Available from: http://amr-review.org/ [Last accessed on 31 January 2024]
6. Didelot X Bowden R Wilson DJ Peto TE Crook DW . Transforming clinical microbiology with bacterial genome sequencing. Nat Rev Genet 2012;13 :601–612. doi: 10.1038/nrg3226.22868263
7. D’Costa VM McGrann KM Hughes DW Wright GD . Sampling the antibiotic resistome. Science 2006;311 :374–377. doi: 10.1126/science.1120800.16424339
8. Quince C Walker AW Simpson JT Loman NJ Segata N . Shotgun metagenomics, from sampling to analysis. Nat Biotechnol 2017;35 :833–844. doi: 10.1038/nbt.3935.28898207
9. De Filippis F Pasolli E Ercolini D . Newly explored Faecalibacterium diversity is connected to age, lifestyle, geography, and disease. Curr Biol 2020;30 :4932–4943.e4. doi: 10.1016/j.cub.2020.09.063.33065016
10. Levin D Raab N Pinto Y Rothschild D Zanir G Godneva A , . Diversity and functional landscapes in the microbiota of animals in the wild. Science 2021;372 :eabb5352. doi: 10.1126/science.abb5352.33766942
11. Tully BJ Graham ED Heidelberg JF . The reconstruction of 2,631 draft metagenome-assembled genomes from the global oceans. Sci Data 2018;5 :1–8. doi: 10.1038/sdata.2017.203.30482902
12. Boolchandani M D’Souza AW Dantas G . Sequencing-based methods and resources to study antimicrobial resistance. Nat Rev Genet 2019;20 :356–370. doi: 10.1038/s41576-019-0108-4.30886350
13. Ayling M Clark MD Leggett RM . New approaches for metagenome assembly with short reads. Brief Bioinform 2020;21 :584–594. doi: 10.1093/bib/bbz020.30815668
14. Milanese A Mende DR Paoli L Salazar G Ruscheweyh HJ Cuenca M , . Microbial abundance, activity and population genomic profiling with mOTUs2. Nat Commun 2019;10 :1014. doi: 10.1038/s41467-019-08844-4.30833550
15. Blanco-Míguez A Beghini F Cumbo F McIver LJ Thompson KN Zolfo M , . Extending and improving metagenomic taxonomic profiling with uncharacterized species using MetaPhlAn 4. Nat Biotechnol 2023;41 :1–12. doi: 10.1038/s41587-023-01688-w.36653493
16. Qin N Yang F Li A Prifti E Chen Y Shao L , . Alterations of the human gut microbiome in liver cirrhosis. Nature 2014;513 :59–64. doi: 10.1038/nature13568.25079328
17. Zhou W Sailani MR Contrepois K Zhou Y Ahadi S Leopold SR , . Longitudinal multi-omics of host–microbe dynamics in prediabetes. Nature 2019;569 :663–671. doi: 10.1038/s41586-019-1236-x.31142858
18. David LA Maurice CF Carmody RN Gootenberg DB Button JE Wolfe BE , . Diet rapidly and reproducibly alters the human gut microbiome. Nature 2014;505 :559–563. doi: 10.1038/nature12820.24336217
19. Wang DD Nguyen LH Li Y Yan Y Ma W Rinott E , . The gut microbiome modulates the protective association between a Mediterranean diet and cardiometabolic disease risk. Nat Med 2021;27 :333–343. doi: 10.1038/s41591-020-01223-3.33574608
20. Chen L Wang D Garmaeva S Kurilshikov A Vich Vila A Gacesa R , . The long-term genetic stability and individual specificity of the human gut microbiome. Cell 2021;184 :2302–2315.e12. doi: 10.1016/j.cell.2021.03.024.33838112
21. Brito IL Gurry T Zhao S Huang K Young SK Shea TP , . Transmission of human-associated microbiota along family and social networks. Nat Microbiol 2019;4 :964–971. doi: 10.1038/s41564-019-0409-6.30911128
22. Thomas AM Segata N . Multiple levels of the unknown in microbiome research. BMC Biol 2019;17 :1–4. doi: 10.1186/s12915-019-0667-z.30616566
23. Almeida A Nayfach S Boland M Strozzi F Beracochea M Shi ZJ , . A unified catalog of 204,938 reference genomes from the human gut microbiome. Nat Biotechnol 2021;39 :105–114. doi: 10.1038/s41587-020-0603-3.32690973
24. Chiu CY Miller SA . Clinical metagenomics. Nat Rev Genet 2019;20 :341–355. doi: 10.1038/s41576-019-0113-7.30918369
25. Xie P Guo Y Teng Y Zhou W Yu Y . GeneMiner: A tool for extracting phylogenetic markers from next-generation sequencing data. Mol Ecol Resour 2024;24 :e13924. doi: 10.1111/1755-0998.13924.38197287
26. Li H . Minimap2: Pairwise alignment for nucleotide sequences. Bioinformatics 2018;34 :3094–3100. doi: 10.1093/bioinformatics/bty191.29750242
27. Wood DE Lu J Langmead B . Improved metagenomic analysis with Kraken 2. Genome Biol 2019;20 :1–13. doi: 10.1186/s13059-019-1891-0.30606230
28. Ye J McGinnis S Madden TL . BLAST: Improvements for better sequence analysis. Nucleic Acids Res 2006;34 (Suppl_2 ):W6–9. doi: 10.1093/nar/gkl164.16845079
29. Garcia BJ Simha R Garvin M Furches A Jones P Gazolla JGFM , . A k-mer based approach for classifying viruses without taxonomy identifies viral associations in human autism and plant microbiomes. Comput Struct Biotechnol J 2021;19 :5911–5919. doi: 10.1016/j.csbj.2021.10.029.34849195
30. Kearse M Moir R Wilson A Stones-Havas S Cheung M Sturrock S , . Geneious Basic: An integrated and extendable desktop software platform for the organization and analysis of sequence data. Bioinformatics 2012;28 :1647–1649. doi: 10.1093/bioinformatics/bts199.22543367
31. Chen S Zhou Y Chen Y Gu J . fastp: An ultra-fast all-in-one FASTQ preprocessor. Bioinformatics 2018;34 :i884–i890. doi: 10.1093/bioinformatics/bty560.30423086
32. Abuín JM Pichel JC Pena TF Amigo J . BigBWA: Approaching the Burrows–Wheeler aligner to Big Data technologies. Bioinformatics 2015;31 :4003–4005. doi: 10.1093/bioinformatics/btv506.26323715
33. Hamidian M Blasco L Tillman LN To J Tomas M Myers GS . Analysis of complete genome sequence of Acinetobacter baumannii strain ATCC 19606 reveals novel mobile genetic elements and novel prophage. Microorganisms 2020;8 :1851. doi: 10.3390/microorganisms8121851.33255319
34. Johnston C Magaret A Son H Stern M Rathbun M Renner D , . Viral shedding 1 year following first-episode genital HSV-1 infection. JAMA 2022;328 :1730–1739. doi: 10.1001/jama.2022.19061.36272098
35. Bonin N Doster E Worley H Pinnell LJ Bravo JE Ferm P , . MEGARes and AMR++, v3. 0: An updated comprehensive database of antimicrobial resistance determinants and an improved software pipeline for classification using high-throughput sequencing. Nucleic Acids Res 2023;51 :D744–D752. doi: 10.1093/nar/gkac1047.36382407
36. Hu F Guo Y Yang Y Zheng Y Wu S Jiang X , . Resistance reported from China antimicrobial surveillance network (CHINET) in 2018. Eur J Clin Microbiol Infect Dis 2019;38 :2275–2281. doi: 10.1007/s10096-019-03673-1.31478103
37. Zhang W Chen K Zhang L Zhang X Zhu B Lv N , . The impact of global warming on the signature virulence gene, thermolabile hemolysin, of Vibrio parahaemolyticus. Microbiol Spectr 2023;11 :e0150223. doi: 10.1128/spectrum.01502-23.37843303
