
==== Front
Brief Bioinform
Brief Bioinform
bib
Briefings in Bioinformatics
1467-5463
1477-4054
Oxford University Press

10.1093/bib/bbae435
bbae435
Problem Solving Protocol
AcademicSubjects/SCI01060
ReadCurrent: a VDCNN-based tool for fast and accurate nanopore selective sequencing
https://orcid.org/0009-0009-3262-6909
Fan Kechen Advanced & Interdisciplinary Biotechnology, Academy of Military Medical Sciences, No. 27 Taiping Road, Haidian District, Beijing 100850, China
College of Information Science and Technology, Beijing University of Chemical Technology, No. 15 North Third Ring East Road, Chaoyang District, Beijing 100029, China

Li Mengfan Advanced & Interdisciplinary Biotechnology, Academy of Military Medical Sciences, No. 27 Taiping Road, Haidian District, Beijing 100850, China
Information Center, Academy of Military Medical Sciences, No. 27 Taiping Road, Haidian District, Beijing 100850, China

Zhang Jiarong Advanced & Interdisciplinary Biotechnology, Academy of Military Medical Sciences, No. 27 Taiping Road, Haidian District, Beijing 100850, China
School of Forensic Medicine, Shanxi Medical University, No. 55 Wenhua Street, Yuci District, Jinzhong 030600, China

Xie Zihan Advanced & Interdisciplinary Biotechnology, Academy of Military Medical Sciences, No. 27 Taiping Road, Haidian District, Beijing 100850, China
College of Life Science and Technology, Beijing University of Chemical Technology, No. 15 North Third Ring East Road, Chaoyang District, Beijing 100029, China

Jiang Daguang College of Information Science and Technology, Beijing University of Chemical Technology, No. 15 North Third Ring East Road, Chaoyang District, Beijing 100029, China

Bo Xiaochen Advanced & Interdisciplinary Biotechnology, Academy of Military Medical Sciences, No. 27 Taiping Road, Haidian District, Beijing 100850, China

https://orcid.org/0000-0003-2616-8891
Zhao Dongsheng Information Center, Academy of Military Medical Sciences, No. 27 Taiping Road, Haidian District, Beijing 100850, China

Shi Shenghui College of Information Science and Technology, Beijing University of Chemical Technology, No. 15 North Third Ring East Road, Chaoyang District, Beijing 100029, China

https://orcid.org/0000-0001-9465-2787
Ni Ming Advanced & Interdisciplinary Biotechnology, Academy of Military Medical Sciences, No. 27 Taiping Road, Haidian District, Beijing 100850, China

Corresponding authors. Advanced & Interdisciplinary Biotechnology, Academy of Military Medical Sciences, No. 27 Taiping Road, Haidian District, Beijing 100850, China. E-mail: niming@bmi.ac.cn; College of Information Science and Technology, Beijing University of Chemical Technology, No. 15 North Third Ring East Road, Chaoyang District, Beijing 100029, China. E-mail: shish@mail.buct.edu.cn
Kechen Fan and Mengfan Li contributed equally to this work.

9 2024
03 9 2024
03 9 2024
25 5 bbae43526 3 2024
20 7 2024
© The Author(s) 2024. Published by Oxford University Press.
2024
https://creativecommons.org/licenses/by-nc/4.0/ This is an Open Access article distributed under the terms of the Creative Commons Attribution Non-Commercial License (https://creativecommons.org/licenses/by-nc/4.0/), which permits non-commercial re-use, distribution, and reproduction in any medium, provided the original work is properly cited. For commercial re-use, please contact journals.permissions@oup.com

Abstract

Nanopore selective sequencing allows the targeted sequencing of DNA of interest using computational approaches rather than experimental methods such as targeted multiplex polymerase chain reaction or hybridization capture. Compared to sequence-alignment strategies, deep learning (DL) models for classifying target and nontarget DNA provide large speed advantages. However, the relatively low accuracy of these DL-based tools hinders their application in nanopore selective sequencing. Here, we present a DL-based tool named ReadCurrent for nanopore selective sequencing, which takes electric currents as inputs. ReadCurrent employs a modified very deep convolutional neural network (VDCNN) architecture, enabling significantly lower computational costs for training and quicker inference compared to conventional VDCNN. We evaluated the performance of ReadCurrent across 10 nanopore sequencing datasets spanning human, yeasts, bacteria, and viruses. We observed that ReadCurrent achieved a mean accuracy of 98.57% for classification, outperforming four other DL-based selective sequencing methods. In experimental validation that selectively sequenced microbial DNA from human DNA, ReadCurrent achieved an enrichment ratio of 2.85, which was higher than the 2.7 ratio achieved by MinKNOW using the sequence-alignment strategy. In summary, ReadCurrent can rapidly classify target and nontarget DNA with high accuracy, providing an alternative in the toolbox for nanopore selective sequencing. ReadCurrent is available at https://github.com/Ming-Ni-Group/ReadCurrent.

nanopore sequencing
selective sequencing
deep learning
VDCNN
classification
Ministry of Science and Technology of the People’s Republic of China 2021YFC0863400
==== Body
pmcIntroduction

Targeted sequencing is a method for sequencing selected DNA regions instead of the entire genomic DNA of samples. It enhances the signal-to-noise ratio of sequencing and is widely employed for research and diagnostic purposes. The two commonly used targeted sequencing approaches are hybridization capture with a set of probes targeting specific DNA regions, such as diagnostic gene sequencing panels [1], and multiplex polymerase chain reaction (PCR) with a primer mix, such as those used to recover pathogen genomes from clinical samples [2]. CRISPR-Cas9-targeted fragmentation can also facilitate target enrichment in sequencing [3–5]. All these approaches rely on wet-lab experiments.

Nanopore selective sequencing (also referred to as adaptive sequencing and adaptive sampling), developed by Oxford Nanopore Technologies (ONT, Oxford, UK), provides an alternative approach for targeted sequencing that primarily depends on computational methods. Electric current signals are generated in real time when DNA molecules pass through nanopores, and these signals are read and rapidly analyzed to determine whether the DNA molecules are the targets. If not, the DNA molecules are ejected from the nanopores by reversing the voltage, making the nanopores available for sequencing other molecules. ONT provides an application programming interface (API) named “Read Until” for researchers to implement selective sequencing.

The speed and accuracy of classifying DNA into target and nontarget sequences largely influence the enrichment ratio of selective sequencing. The enrichment ratio is defined as the ratio of target bases generated by selective sequencing to those by standard sequencing [6, 7]. Two strategies are employed by the previously reported tools or models for nanopore selective sequencing. The first strategy is to align base-called reads or raw current signals to a reference dataset. Three representative tools using this strategy include Readfish [6], UNCALLED [7], and adaptive sampling integrated into MinKNOW software [8]. Other tools that also employ this strategy include MoSS [9], RUBRIC [10], ReadBouncer [11], and BOSS-RUNS [12]. Classifying DNA based on alignment offers high accuracy, but the computational expense escalates with the expanding target regions. Consequently, for large references, using the alignment-based strategy may lead to a slow classification or become unfeasible. For example, the maximum reference size for UNCALLED is limited to ~1 Gbp [7]. Alignment-based tools have been used to facilitate pathogen detection [13, 14] and genome enrichment of rare and unknown species from complex microbiomes [15].

The second strategy is to use deep learning (DL) models for an alignment-free classification of DNA. DL models are pretrained before nanopore selective sequencing, and the inference conducted by these DL models is more rapid than achieved through alignment calculations. Recently, several studies have proposed DL-based tools or methods for nanopore selective sequencing. For example, SquiggleNet employs a convolutional architecture by applying 1D convolution to raw signals through bottleneck residual blocks adapted from ResNet [16]; with a training dataset of 2.4 million reads, SquiggleNet achieved a 90% accuracy in distinguishing between human and bacterial DNA [17]. Danilevsky et al. [18] conducted a collaborative study involving seven different DL models to assess their performance in distinguishing human mitochondrial DNA from genomic DNA. Trained by using two public and two home-made human datasets, the best-performing model by Danilevsky et al. [18] achieved an average accuracy of 92% to classify mitochondrial sequences from genomic sequences, resulting in a 2.3-fold enrichment ratio in a selective sequencing experiment. DeepSelectNet [19] is another approach based on DL to distinguish raw signals, which has a similar model structure to SquiggleNet, yet it mitigates model overfitting through the incorporation of dropout layers between convolutional layers. Additionally, a novel data preprocessing method has been introduced, enabling the extraction of more usable samples from a limited number of reads, thus allowing the model to capture a broader range of features. Based on a training dataset of 20 000 reads, DeepSelectNet achieved an average accuracy of 95% for the pair-wise classification involving SARS-CoV-2, Saccharomyces cerevisiae (S. cerevisiae), Chlamydomonas reinhardtii (C. reinhardtii), and microbiome. Recently, Lin et al. [20] reported a DL framework named NanoDeep for nanopore selective sequencing of microbes from mammals. NanoDeep utilizes the finding that the signals of 6-mers in bacterial genomes are distinct from those in the human genome, and it achieved an enrichment ratio of 1.8-fold. NanoDeep can be applied to microbial species that do not exist in training sets, albeit with a performance decrease of ~5%–15%. Moreover, the DL techniques are employed to detect and classify seven multilocus sequence typing (MLST) loci in the Klebsiella pneumoniae genome [21] and to detect DNA methylation states from nanopore sequencing reads [22].

The drawbacks of DL models are also evident. The DL models highly depend on the training datasets [23]. There is a lack of databases or resources for nanopore sequencing raw data, and therefore, conventional nanopore sequencing of the target DNA is usually required to obtain the training dataset prior to the application of selective sequencing. Therefore, DL strategies are particularly suited for applications in well-defined scenarios, such as determining a panel of viral or bacterial pathogens. Secondly, there is a need to improve the accuracy of DL models in identifying target molecules. Accuracy in existing DL-based methods exhibits great variability depending on the dataset [17–20]. Achieving a high target enrichment ratio critically depends on accuracy, necessitating continuous optimization of DL models. Consequently, it is essential to develop a new method that can perform well with any dataset.

In this study, we introduced a method named ReadCurrent, comprising an optimized current signal preprocessing approach and a DL model adapted from VDCNN [24]. Across 10 datasets, ReadCurrent exhibited higher mean accuracy than other DL-based methods. We employed ReadCurrent to customize the Read Until script, successfully enriching microbial DNA in real time from mixtures of the human cell line 293T and the microbial community standard.

Materials and methods

The design and architecture of ReadCurrent

To determine the architecture for DL-based selective sequencing, we examined the effects of three mainstream models for natural language processing (NLP), including VDCNN, long short-term memory network (LSTM) [25], and Transformer [26]. We also constructed hybrid models of convolutional neural network (CNN) with LSTM (CNN-LSTM) and CNN with Transformer (CNN-Transformer) for testing. Among these models, VDCNN achieved the highest accuracy (provided in the Results section). Therefore, ReadCurrent was developed based on VDCNN. VDCNN is designed for text processing and operates at the character level. VDCNN uses small convolutional kernel sizes, usually three or one, and thus has a very deep network. VDCNN performs well in text classification tasks with 1D data and contextual relationships [27], which are similar to the classification of nanopore sequencing signals. Typically, VDCNN requires the construction of a dictionary, representing a unique mapping of tokens to encoded forms, and subsequently, a word embedding layer is built based on the dictionary. However, for electrical current signals from nanopore sequencing, constructing a dictionary is not feasible. Thus, in the first layer of the model in ReadCurrent, instead of word embedding, a convolutional block with a large kernel size of 19 was used for primary feature extraction (Fig. 1). Compared to small kernels, the large-kernel convolution is expected to cover a larger receptive field, thus capturing a broader context [28]. Besides, one nucleotide roughly corresponds to 10 current signal points, and to ensure that at least one nucleotide was included in the receptive field of the convolution for feature extraction, we selected a kernel size of 19, like Guppy [29] and SquiggleNet. We also tested a larger kernel size of 61, and the performance of models in classification did not improve (Table S1). Additionally, to mitigate the computational expense of VDCNN, the convolution block substituting word embedding in ReadCurrent adopted a stride of three for convolution and a stride of two for max-pooling (Fig. 1). Although this approach may result in information loss [30, 31], inference speed is critical for the efficacy of nanopore selective sequencing.

Figure 1 The architecture of ReadCurrent. The convolutional block in the dashed box is used to replace the word embedding layer of the VDCNN. “Conv1D-32” represents a 1D convolution with an output dimension of 32. Except for the first convolution, which has a stride of three and a kernel size of 19, the rest have a stride of one and a kernel size of three or one. “ResBlock-64” denotes a 1D residual block with an output dimension of 64, while “FC-2048” indicates a fully connected layer with an output dimension of 2048. The output dimension of the adaptive max-pooling in the fifth layer is eight.

The architecture of ReadCurrent comprised six layers, as illustrated in Fig. 1. The second to fifth layers were composed of basic residual blocks [16] and max-pooling. When the input and output dimensions of the basic residual block matched, the residual connections performed identity mapping by directly adding the input and output. Otherwise, to align dimensions, the input was mapped to the output through a convolution with a kernel size and a stride of one before addition. The sixth layer of ReadCurrent functioned as a classifier, comprising three fully connected layers, ultimately resulting in the classification of input current signals as either on target or not.

Public nanopore sequencing datasets

Enrichment of specific microorganisms is a common application scenario for nanopore selective sequencing. Thus, as in previous studies [19], we obtained four public datasets for nanopore sequencing of microorganisms, which included both the “fast5” formatted raw data and the base-called “fastq” formatted data. Two datasets were downloaded from the National Center for Biotechnology Information (NCBI) Sequence Read Archive (SRA) database, including a dataset of a S. cerevisiae BY4741 isolate (2.6 Gbp, SRA BioSample SAMN10621887, Run identifier SRR8648517) [32] and a dataset of a C. reinhardtii isolate (1.4 Gbp, SRA BioSample SAMEA5522908, Run identifier ERR3237140) [33]. One dataset (0.6 Gbp) of genome sequencing of a SARS-CoV-2 variant (GISAID accession EPI_ISL_412964) [34] was obtained via the link https://cadde.s3.climb.ac.uk/SP1-raw.tgz. A dataset of the microbial community standard of 10 microorganism species (ZYMO RESEARCH, Irvine, USA) was downloaded from the European Bioinformatics Institute (EMBL-EBI) European Nucleotide Archive (ENA) database (16.5 Gbp, ENA BioSample SAMEA5065622, Run identifiers ERR3152366 and ERR2887850) [35].

Generating home-made nanopore sequencing datasets

Besides using public datasets, we generated home-made datasets through nanopore sequencing of DNA or complementary DNA (cDNA) samples from yeast, microbiome, SARS-CoV-2, and human. Two samples included the DNA standard of S. cerevisiae S288C (TSTO, Ningbo, China) and the microbial community standard of eight species (ZYMO RESEARCH, as detailed in Table S2). Ten synthetic SARS-CoV-2 RNA standards (Twist Bioscience, South San Francisco, USA, as detailed in Table S2) initially underwent reverse transcription before being amplified using the QuantiTect Whole Transcriptome Kit (Qiagen, Duesseldorf, Germany) with random primers. We extracted total DNA from the human cell line 293T (ATCC, Manassas, USA) using the DNeasy Blood & Tissue Kit (Qiagen), following the manufacturer’s instructions.

Nanopore sequencing of these DNA and cDNA samples was conducted using the ligation sequencing kit LSK109 (ONT, Oxford, UK) and the MinION Mk1B with R9.4 flow cells (ONT). Four flow cells were used to generate a total of 21.7 Gbp of base-called data, as detailed in Table S3.

Training, validation, and testing datasets

The training, validation, and testing datasets for ReadCurrent were constructed based on four public datasets and four home-made datasets of nanopore sequencing. The base-called reads were aligned to the corresponding reference genomes (listed in Table S4) by Minimap2 [36] (v2.17), and the alignments were analyzed by SAMtools [37] (v1.16.1). For each of the eight datasets, 40 000 reads assigned to the expected species were randomly selected, and then, we divided these reads into three subdatasets in a 2:1:1 ratio, designated for training, validation, and testing, respectively. The details of parameter settings of the bioinformatics tools are provided in Supplementary Methods.

Next, we in silico constructed the mixture datasets to include multiple species. The mixing was conducted separately within the four public datasets and the four home-made datasets. The training, validation, and testing subdatasets were mixed, respectively. There are six possible pair-wise mixtures among the four datasets; since the microbiome dataset included S. cerevisiae, the mixture of microbiome and yeast was disregarded.

Finally, regarding the home-made datasets, five mixture datasets were created, including combinations of microbiome and human, yeast and human, SARS-CoV-2 and human, microbiome and SARS-CoV-2, and yeast and SARS-CoV-2. For the public datasets, five mixture datasets were assembled, including combinations of microbiome and SARS-CoV-2, yeast and SARS-CoV-2, C. reinhardtii and SARS-CoV-2, microbiome and C. reinhardtii, and yeast and C. reinhardtii. Ultimately, 10 mixture datasets were prepared for training, validation, and testing.

Preprocessing of current signals

The raw electrical current signals from nanopore sequencing could not be directly utilized as input for the DL model and required preprocessing. The initial part of a read contains invalid signals, including white noise of nanopore before sequencing, a sequencing adapter, and a multiplex barcode sequence of the sample. The length of adapter and barcodes are fixed [38], but the lengths of white noise signals varied among reads. We randomly sampled 50 reads from each of the eight datasets and based on the 400 reads, we found that over 90.3% of the invalid parts of the initial reads were <1500 signal points (Table S5). Consequently, we trimmed the initial 1500 signal points from the reads, which was similar to that in SquiggleNet and DeepSelectNet.

Subsequently, after trimming the first 1500 signal points, the remaining signals from the training datasets were sampled using a sliding window to obtain training data. The size and stride of the window were optional. For example, employing strides of one-third and one-half of the window size resulted in 3-fold and 2-fold tiling of the reads, respectively. For the validation and testing datasets, after trimming the first 1500 signal points, we sampled the initial part of the remaining signals (e.g. 3000 signal points) to obtain validation and testing data.

Since the raw datasets from different experiments might exhibit a systematic bias in electrical current intensities, a modified Z-score was employed to normalize signal segments, using the median value rather than the mean value for Z-score calculation [39]. We applied the modified Z-score to each signal segment from the training, validation, and testing datasets, respectively. The modified Z-score is calculated as follows:

(1) \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} \begin{equation*} MAD= median\left|{x}_i- median(X)\right| \end{equation*}\end{document}

(2) \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} \begin{equation*} {Z}_i=\frac{x_i- median(X)}{MAD} \end{equation*}\end{document}

where MAD denotes the median absolute deviation (MAD), X denotes the signal segment, xi denotes the ith signal point, and Z denotes the modified Z-score. Furthermore, outliers exceeding 3.5 times the median Z-score were substituted with the mean value of their adjacent points. Finally, the normalized and processed Z-score signal segments of the raw currents were used as the input for ReadCurrent.

Training and evaluation of ReadCurrent and other tools for selective sequencing

In the training of ReadCurrent, we employed the Adam optimizer with a learning rate of 0.001 and a batch size of 1024. If the accuracy of the model did not improve after 2000 batches, the learning rate was to be halved. The evaluation metrics for assessing model performance primarily included accuracy, precision, recall, and F1 score.

For comparison, four previously proposed DL-based methods were included: SquiggleNet, Danilevsky et al.’s [18] method, DeepSelectNet, and NanoDeep. Notably, Danilevsky et al. [18] assessed several models, and the long short-term memory networks with recurrent batch normalization (BNLSTM) [40] exhibited the highest average classification accuracy. Thus, the BNLSTM model developed by Danilevsky et al. [18] was used in this study. The detailed configurations of the four comparison methods are listed in Supplementary Methods.

Experimental validation of ReadCurrent

Two mixture samples were constructed using human DNA (293T cell line) and microbial community standard samples (ZYMO RESEARCH), with ratios of 4:1 and 9:1, respectively. Subsequently, the mixtures were prepared for library construction using the ligation sequencing kit LSK109 (ONT).

Nanopore selective sequencing was performed using a MinION Mk1B device (R9.4.1 flow cell, ONT) with the MinKNOW software (version 23.07.12, ONT). The channels 1–256 on the flow cell were set for selective sequencing as experimental groups, while channels 257–512 were allocated for standard sequencing as controls. We compared ReadCurrent with DeepSelectNet and the adaptive sampling function integrated into MinKNOW (details provided in Supplementary Methods), with each method performing 2 h of sequencing. The MinKNOW software employs a sequence-alignment strategy for target classification. Several studies [41–44] highlighted the high performance of MinKNOW for selective sequencing, which might be partly due to superior software integration by the manufacturer (ONT).

In the selective sequencing experiments, the real-time current signals were first processed as we did in training, including removing the first 1500 signal points, normalization and using the next 3000 signal points as inputs to the model for classification.

Computational hardware

Model training and testing were conducted on a computation workstation (TRY, China) equipped with a 13th Gen Intel Core i9-13900 CPU, 128 GB RAM, and an NVIDIA GeForce RTX 4090 GPU. Nanopore selective sequencing was performed on a laptop (HP, USA), powered by a 12th Gen Intel Core i7-12700H CPU, 16 GB RAM, and an NVIDIA GeForce RTX 3080 Ti GPU.

Results

Sliding window sampling of current signals increases accuracy

The sampling methods used by previous DL-based selective sequencing tools do not fully utilize the raw current signal data. Thus, we employed a sliding window method to sample the current signals for the training of ReadCurrent (Fig. 2A). Compared to training using initial parts of the reads after removing the invalid signals, the sliding window method offered a more comprehensive characterization of the features of target and nontarget reads in comparable training sets. Various window sizes and strides were evaluated to identify the optimal setting. Seven in silico mixture datasets with long read lengths (mean >1300 bp) were included, and training of the ReadCurrent was conducted separately for each dataset.

Figure 2 The sliding window sampling with 3-fold tiling and a window size of 3000 achieved better performance. (A) A schematic diagram illustrates sliding window sampling, with examples depicted for strides of 1 and 0.5 times the window size, aimed at sampling nanopore sequencing current signals. Larger tiling-folds were achieved through the reduction of strides. (B, C) The median and mean accuracies of ReadCurrent, derived from training inputs with varying tiling-folds (B) and sliding window sizes (C), were assessed across seven datasets, excluding the public SARS-CoV-2 dataset due to its insufficiently short read lengths.

To determine the optimal stride length, five strides, corresponding to 1-fold through 5-fold tiling, were evaluated with a fixed window size of 3000 current signal points. The accuracy for target and nontarget inferences across the seven testing sets is depicted in Fig. 2B and Fig. S1A–G. By employing a 3-fold tiling for training, ReadCurrent demonstrated optimal performance, with mean and median accuracies being remarkably higher than those of 1-fold and 2-fold tiling and slightly superior to those of 4-fold and 5-fold tiling (Table S6). The training computational expense of ReadCurrent increased with the number of tiling-folds for sampling (Fig. S2).

Next, with a 3-fold tiling sampling, we evaluated the performance of ReadCurrent with window sizes ranging from 1000 to 6000 current signal points (Fig. 2C and Fig. S1H–N) to determine the optimal window size. The accuracy of ReadCurrent improved with increasing window sizes. However, beyond 3000 points, the accuracy improvement became marginal (Table S7). In previous studies [17, 19], 3000 signal points, which correspond to 300 nucleotides, were used for inference. Larger sizes of signals for inference required more time to generate the data, thereby decreasing the efficiency of selective sequencing.

Finally, a window size of 3000 and a 3-fold tiling were chosen for sampling current signals for training, resulting in ReadCurrent achieving a mean accuracy of 98.57 ± 1.1% in classifying target and nontarget reads. To assess the contribution of the sliding window sampling to accuracy, we compared it with the sampling methods used by SquiggleNet and DeepSelectNet. After removing the first 1500 signal points, SquiggleNet used the next 3000 signal points as input for the model, and DeepSelectNet randomly selected four segments of reads, each with a length of 3000 signal points. The accuracies were 93.15 ± 3.99% and 95.6 ± 3.07% for the sampling methods used by SquiggleNet and DeepSelectNet, respectively (Table S8), both of which were lower than that of the sliding window method.

ReadCurrent has a low computational expense for training and inference

Although the sliding window sampling method improved the accuracy of ReadCurrent, it also increased the training cost. Consequently, we used a modified VDCNN, incorporating an additional convolution block for primary feature extraction, to mitigate the computational expenses of the VDCNN. ReadCurrent exhibited a computational efficiency of 0.47 giga floating-point operations per second (GFLOPs), approximately one-sixth of the conventional VDCNN’s 2.74 GFLOPs. Utilizing a batch size of 400 for training, ReadCurrent required 6.7 GB of memory and took 21 s to iterate through 200 batches, in contrast to the conventional VDCNN, which consumed 22.9 GB of memory and 99 s. For inference with a 512-batch size, ReadCurrent required 2.53 GB of memory and 2.4 ms, significantly lower than those of the conventional VDCNN, which required 4.77 GB and 3 ms.

The modified VDCNN used in ReadCurrent greatly decreased the costs of training and inference compared to the conventional VDCNN, albeit with a slight reduction in accuracy. For the 10 mixed datasets, the mean accuracy of the conventional VDCNN was 98.66 ± 0.85% versus 98.57 ± 1.1% for ReadCurrent (Fig. 3A and B, Table S9). Especially in the yeast–human and microbiome–human mixture datasets, ReadCurrent had >0.8% drops in accuracy compared with the conventional VDCNN. This might be associated with the more complex genomes involved in the two mixture datasets compared to those in the other datasets. The architecture of the VDCNN in ReadCurrent was designed to consider the trade-off between inference speed and information loss, and it seems that ReadCurrent was more affected by the complex genomes than the conventional VDCNN.

Figure 3 The modified VDCNN of ReadCurrent improved in speed without a significant decrease in accuracy compared to the conventional VDCNN. (A, B) The accuracy of ReadCurrent, VDCNN, LSTM, CNNLSTM, Transformer, and CNN-Transformer across five public datasets (A) and five home-made datasets (B).

Besides VDCNN, we also evaluated other NLP models, including LSTM, Transformer, CNN-LSTM, and CNN-Transformer. With the identical preprocessing of sequencing signals, the modified VDCNN surpassed all the other models in accuracy (Fig. 3A and B, Table S9). Especially, the Transformer exhibited the lowest mean accuracy at 91.23 ± 6.08%, followed by the LSTM at 95.83 ± 3.06%. Both the Transformer and LSTM performed worse than hybrid models integrated with CNN (CNN-LSTM, 97.66 ± 0.97%; CNN-Transformer, 95.46 ± 2.71%). The Transformer and LSTM are designed to capture long-term dependencies; this result implies that in the scenario of classifying nanopore sequencing current signals, long-term dependencies within the signals are not essential and leveraging local features is sufficient for the classification.

Performance evaluation of ReadCurrent

Subsequently, to evaluate the performance of ReadCurrent and compare it with existing DL-based selective sequencing methods, we tested ReadCurrent and the other methods on the 10 mixed datasets. The accuracy and other classification metrics in testing datasets are shown in Fig. 4 and Table S10, and we found that the accuracy of ReadCurrent outperformed that of other tools. ReadCurrent had a mean accuracy of 98.57 ± 1.1% (median 99.23%) compared to 96.71 ± 2.45% for NanoDeep (median 95.99%), 95.07 ± 3.03% for DeepSelectNet (median 94.56%), 91.62 ± 2.93% for BNLSTM (median 91.72%), and 90.18 ± 6.07% for SquiggleNet (median 89.79%). ReadCurrent demonstrated relatively consistent performance, in contrast to the other tools, which exhibited larger variations in accuracy across the 10 datasets (Fig. 4A and B). The accuracy of ReadCurrent fell below 98% solely in two datasets containing data from the 293T human cell line set, potentially attributable to the large size of the human genome. While on the dataset of SARS-CoV-2 and human, ReadCurrent achieved high accuracy, likely due to the small size of the SARS-CoV-2 genome and its disparity from the human genome (Fig. 4B). Similarly, the F1 score of ReadCurrent exceeded those of other tools, with a mean F1 score of 98.57 ± 1.1% (median 99.23%, Fig. 4C). The precision and recall rate were also examined. BNLSTM had a higher mean precision (99.26 ± 0.9%) than ReadCurrent (98.66 ± 1.05%), but its mean recall rate (83.84 ± 5.36%) was the lowest among the five tools. Compared to the other tools, ReadCurrent achieved both high precision and recall rates (98.48 ± 1.24%) for determining target and nontarget reads (Fig. 4D and E).

Figure 4 ReadCurrent had high accuracy in classifying target and nontarget sequences. (A, B) The accuracy of ReadCurrent and other DL-based methods on five public datasets (A) and five home-made datasets (B). (C–E) The boxplots of the F1 score (C), precision (D), and recall (E) for ReadCurrent and other DL-based methods on 10 mixed datasets. The boxes represent the interquartile range (IQR) between the first and third quartiles. Horizontal lines and dots inside the boxes indicate the median and mean, respectively, and the lines outside represent values within 1.5 times the IQR. (F, G) The accuracy of ReadCurrent and other DL-based methods tested on the datasets different from the training datasets (F) and the differences in accuracy compared to training and testing with the same datasets (G).

Next, to evaluate the generalizability of ReadCurrent, we tested it on datasets different from the training datasets and compared it with other tools. The home-made and public datasets of SARS-CoV-2, microbiome, and yeast were generated by different experiments with different sequencing devices (MinION versus GridION). When the training was based on home-made datasets and the testing was on public datasets, ReadCurrent achieved ~95% accuracy compared to an average of 80% (60%–89%) of the other models (Fig. 4F and Table S11). Compared to training and testing both with the home-made datasets, the accuracy of ReadCurrent had a 3%–5% decrease (Fig. 4G). In contrast, the decreases for other models were much larger than for ReadCurrent. For instance, NanoDeep, which performed best except ReadCurrent, had a 10%–11% decrease. We also trained the models with public datasets and tested them on home-made datasets. Although ReadCurrent still surpassed other models (mean 73%, ranged from 65% to 80%), we found a large drop in the accuracy of ReadCurrent to 80%–83% (Fig. 4F). It is obvious that ReadCurrent was affected by the training datasets. The base quality of public datasets was lower than that of home-made datasets (Table S3), which might lead to a decrease in performance, and it indicated the importance of the quality of training datasets for ReadCurrent.

We also assessed the inference time and computer memory usage of ReadCurrent and other methods on 10 mixed datasets. Great differences were observed in the inference times of the five tools. ReadCurrent took an average of 1.98 ms to process 50 reads, marginally slower than NanoDeep (1.24 ms) and SquiggleNet (1.23 ms), yet greatly quicker than DeepSelectNet (12.17 ms) and BNLSTM (72.25 ms, Table S12). NanoDeep required ~5 GB of memory to process 50 reads, while ReadCurrent and the other methods required 1–2 GB of memory (Table S13).

Experimental validation of ReadCurrent on nanopore selective sequencing

To further evaluate the capability to enrich target sequences, we designed a nanopore selective sequencing experiment to enrich microbial reads from mixtures of human DNA and microbial DNA using ReadCurrent, DeepSelectNet, and the adaptive sampling of MinKNOW (Fig. 5A). Selective sequencing (channels 1–256) yielded 211.5, 197, and 179.7 Mbp of base-called data from MinKNOW, ReadCurrent, and DeepSelectNet, respectively. Concurrently, the within-run control sequencing (channels 257–512) generated 259, 249.4, and 243.5 Mbp, respectively. The details of the sequencing data are listed in Table S14. Initially, we examined the lengths of ejected reads, reflecting the rapidity of distinguishing between target and nontarget DNA. ReadCurrent had a median length of 439 bp, which was shorter than the 452 bp of DeepSelectNet, consistent with its quicker inference time based on the in silico evaluation (Fig. 5B). MinKNOW exhibited the largest median read length (575 bp), attributable to its sequence-alignment strategy for determining target and nontarget DNA. Furthermore, in all control groups, the median length of nontargeted reads was ~7300 bp, significantly exceeding that of the ejected reads.

Figure 5 ReadCurrent had high enrichment efficiency in selective sequencing. (A) The scheme for nanopore selective sequencing on two mixture samples using a MinION Mk1B with an R9.4 flow cell. (B) Comparison of the length distribution of reads ejected by ReadCurrent, MinKNOW, and DeepSelectNet. ***P < .001, two-sided Wilcoxon rank-sum test (from left to right, 0, 2.81 × 10–151, 0). (C, D) The accuracy, precision, recall, and F1 score of ReadCurrent, MinKNOW, and DeepSelectNet when the ratio of human to microbiome is 4:1 (C) and 9:1 (D). (E–G) The cumulative number of microbial bases increased with sequencing time using ReadCurrent (E), MinKNOW (F), and DeepSelectNet (G) for selective sequencing, respectively, when the ratio of microbiome to human is 1:4. (H–J) The cumulative number of microbial bases increased with sequencing time using ReadCurrent (H), MinKNOW (I), and DeepSelectNet (J) for selective sequencing, respectively, when the ratio of microbiome to human is 1:9.

Then, utilizing selective sequencing data, we obtained metrics that include accuracy, precision, recall, and F1 score for classifying target and nontarget reads in ReadCurrent, DeepSelectNet, and MinKNOW (Fig. 5C and D). Based on a DL method for classification, these metrics of ReadCurrent were comparable but marginally lower than those of MinKNOW, which employs the sequence-alignment strategy. For example, ReadCurrent achieved accuracies of 93.81% and 93.51% for the 1:4 and 1:9 mixture samples, respectively, compared to 94.61% and 94.26% of MinKNOW. Conversely, classification via DeepSelectNet, based on a DL approach, exhibited relative inaccuracy (1:4 mixture at 84.48%, 1:9 mixture at 82.01%).

Lastly, the enrichment ratio was evaluated. We found that ReadCurrent achieved the highest enrichment ratios (1:4 mixture at 2.86-fold, 1:9 mixture at 2.85-fold, Fig. 5E and H), surpassing MinKNOW (1:4 mixture at 2.7-fold, 1:9 mixture at 2.66-fold, Fig. 5F and I) and DeepSelectNet (1:4 mixture at 1.86-fold, 1:9 mixture at 1.84-fold, Fig. 5G and J). Throughout the 2-h sequencing period for each method, microbial bases were generated linearly, thereby maintaining the enrichment ratio consistently.

Discussion

We developed a VDCNN-based computational tool named ReadCurrent for nanopore selective sequencing, which can rapidly identify target DNA in real time through electrical signals. Adopting a large-kernel convolution block for primary feature extraction enabled ReadCurrent to significantly reduce computational expenses in both training and inference phases, compared to conventional VDCNN while maintaining accuracy. Compared with previous DL-based tools, ReadCurrent achieved a notable improvement in accuracy and showed superior performance in both in silico and experimental validations.

ReadCurrent was also compared with the adaptive selective function of the official MinKNOW software, which classifies target and nontarget DNA through sequence alignment. Alignment with references is considered the gold standard for classification, but it becomes slower with larger references. Hence, leveraging the speed of DL models for selective sequencing necessitates achieving accuracy comparable to alignment methods. Experimental validation for enriching microbial DNA from human genomic DNA showed that the accuracy of ReadCurrent in classifying targets was marginally lower than MinKNOW but largely exceeded DeepSelectNet. Consequently, with both high accuracy and speed in classification, ReadCurrent achieved a 6%–7% higher enrichment ratio than MinKNOW. Although DeepSelectNet had a comparable speed in classification with ReadCurrent (as indicated by similar ejected read lengths), its experimental performance was greatly impeded by low accuracy in classification. To the best of my knowledge, this is the first report of a DL-based tool outperforming the adaptive sampling function of MinKNOW in selective sequencing.

We also assessed the generalizability of ReadCurrent by using datasets generated from different experiments and sequencing devices. The results showed that compared to other methods, ReadCurrent had an advantage in maintaining performance when the training was conducted based on our home-made datasets. However, when we used public datasets for training, which had lower base quality scores than those of home-made datasets, the performance of ReadCurrent, as well as other tools, was largely affected. Therefore, DL models were more prone to overfitting when trained on low-quality datasets. It indicates the importance of the quality of training datasets for the DL models.

Different from the DL frameworks of the other models, NanoDeep includes a pretrained model to classify microbes from mammals based on the finding that the signals of 6-mers in bacterial genomes are distinct from those in the human genome [20]. Theoretically, the pretrained model of NanoDeep can be directly applied to enrich microbial species that do not exist in the pretraining datasets. In our validation, when the pretrained model of NanoDeep was directly used to classify the microbiome (home-made datasets) and human, the accuracy was 83.44%. After retraining, the accuracy of NanoDeep increased to 92.55%, still lower than the 96.38% of ReadCurrent. Therefore, further research may be needed on how pretrained models can promote DL-based selective sequencing.

This study has several limitations. Firstly, only human and a few microorganisms were included in the in silico and experimental validations, necessitating further assessment of the performance of ReadCurrent in selective sequencing across a broader range of species. Secondly, in the experimental validations, each tool was sequenced for 2 h, which was much shorter than the flow cells’ capacity. The impact of gradually decreasing active nanopores on the enrichment ratios could not be analyzed. Thirdly, although our data preprocessing method has improved classification accuracy, further optimization remains possible. For example, similar to previous tools, we directly removed the first 1500 signal points to eliminate the influence of invalid regions such as adapters and barcodes. Thus, developing an algorithm to precisely identify invalid regions in current signals may further enhance the performance of ReadCurrent.

In summary, ReadCurrent is a DL-based tool that can rapidly classify target and nontarget DNA with high accuracy and low computational cost. It provides an alternative in the toolbox for nanopore selective sequencing.

Key Points

The optimal tiling sampling (3-fold) for model training and input current signal length (3000 signal points, ~300 bp) were determined, resulting in a ~5% increase in accuracy.

For the VDCNN structure, a large-kernel convolutional block was employed instead of the word embedding layer for primary feature extraction, reducing computational complexity by approximately six times.

In ten mixed datasets, ReadCurrent achieved a mean accuracy of 98.57%, outperforming the 90.18%–96.71% accuracy range of SquiggleNet, BNLSTM, DeepSelectNet, and NanoDeep.

In experimental validation for microbial DNA enrichment, ReadCurrent obtained a 2.85-fold enrichment, compared to 2.7-fold of the MinKNOW adaptive sampling.

Supplementary Material

Supplementary_figures_bbae435

Supplementary_tables_bbae435

Supplementary_methods_bbae435

Funding

This work was supported by the National Natural Science Foundation of China (No. 31870079), the Young Scientists Fund of the National Natural Science Foundation of China (No. 82100130) and the Ministry of Science and Technology of the People’s Republic of China (No. 2021YFC0863400).

Conflict of interest: None declared.

Data availability

The nanopore sequencing data generated in this study have been deposited in the NCBI SRA database with BioProject accession number PRJNA1083903.

Author contributions

Development of ReadCurrent: F.K. and L.M. Conceived and designed the experiments: N.M., Z.J., and F.K. Performed the experiments: Z.J. and F.K. Analyzed the data: F.K., L.M., and X.Z. Wrote the paper: N.M. and F.K. Reviewed and edited the manuscript: N.M. and F.K. Project discussions: N.M., B.X., S.S., Z.D., and J.D. All authors contributed to the article and approved the submitted version.

Code availability

The source code for ReadCurrent is available on GitHub: https://github.com/Ming-Ni-Group/ReadCurrent.
==== Refs
References

1. Bean LJH , FunkeB, CarlstonCM. et al. Diagnostic gene sequencing panels: from design to report-a technical standard of the American College of Medical Genetics and Genomics (ACMG). Genet Med 2020;22 :453–61. 10.1038/s41436-019-0666-z.31732716
2. Liu H , LiJ, LinY. et al. Assessment of two-pool multiplex long-amplicon nanopore sequencing of SARS-CoV-2. J Med Virol 2022;94 :327–34. 10.1002/jmv.27336.34524690
3. Jiang W , ZhaoX, GabrieliT. et al. Cas9-assisted targeting of CHromosome segments CATCH enables one-step targeted cloning of large gene clusters. Nat Commun 2015;6 :8101. 10.1038/ncomms9101.26323354
4. Loose M . Finding the needle: targeted Nanopore sequencing and CRISPR-Cas9. CRISPR J 2018;1 :265–7. 10.1089/crispr.2018.29028.mlo.31021218
5. Gabrieli T , SharimH, FridmanD. et al. Selective nanopore sequencing of human BRCA1 by Cas9-assisted targeting of chromosome segments (CATCH). Nucleic Acids Res 2018;46 :e87–e87. 10.1093/nar/gky411.29788371
6. Payne A , HolmesN, ClarkeT. et al. Readfish enables targeted nanopore sequencing of gigabase-sized genomes. Nat Biotechnol 2021;39 :442–50. 10.1038/s41587-020-00746-x.33257864
7. Kovaka S , FanY, NiB. et al. Targeted nanopore sequencing by real-time mapping of raw electrical signal with UNCALLED. Nat Biotechnol 2021;39 :431–41. 10.1038/s41587-020-0731-9.33257863
8. Oxford Nanopore Technologies, Oxford, UK. Nanopore Documentation: MinKNOW. https://community.nanoporetech.com/docs/prepare/library_prep_protocols/experiment-companion-minknow/v/mke_1013_v1_revdc_11apr2016.
9. Masutani B , MorishitaS. A framework and an algorithm to detect low-abundance DNA by a handy sequencer and a palm-sized computer. Bioinformatics 2019;35 :1443. 10.1093/bioinformatics/bty771.30252019
10. Edwards HS , KrishnakumarR, SinhaA. et al. Real-time selective sequencing with RUBRIC: read until with Basecall and reference-informed criteria. Sci Rep 2019;9 :11475. 10.1038/s41598-019-47857-3.31391493
11. Ulrich JU , LutfiA, RutzenK. et al. ReadBouncer: precise and scalable adaptive sampling for nanopore sequencing. Bioinformatics 2022;38 :i153–60. 10.1093/bioinformatics/btac223.35758774
12. Weilguny L , De MaioN, MunroR. et al. Dynamic, adaptive sampling during nanopore sequencing using Bayesian experimental design. Nat Biotechnol 2023;41 :1018–25. 10.1038/s41587-022-01580-z.36593407
13. Lin Y , DaiY, ZhangS. et al. Application of nanopore adaptive sequencing in pathogen detection of a patient with chlamydia psittaci infection. Front Cell Infect Microbiol 2023;13 :1064317. 10.3389/fcimb.2023.1064317.36756615
14. Lin Y , DaiY, LiuY. et al. Rapid PCR-based Nanopore adaptive sequencing improves sensitivity and timeliness of viral clinical detection and genome surveillance. Front Microbiol 2022;13 :929241. 10.3389/fmicb.2022.929241.35783376
15. Sun Y , ChengZ, LiX. et al. Genome enrichment of rare and unknown species from complicated microbiomes by nanopore selective sequencing. Genome Res 2023;33 :612–21. 10.1101/gr.277266.122.37041035
16. He K , ZhangX, RenS. et al. Deep Residual Learning for Image Recognition. In: 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, Las Vegas, NV, USA, 2016, 770–8.
17. Bao Y , WaddenJ, Erb-DownwardJR. et al. SquiggleNet: real-time, direct classification of nanopore signals. Genome Biol 2021;22 :298. 10.1186/s13059-021-02511-y.34706748
18. Danilevsky A , PolskyAL, ShomronN. Adaptive sequencing using nanopores and deep learning of mitochondrial DNA. Brief Bioinform 2022;23 :bbac251. 10.1093/bib/bbac251.
19. Senanayake A , GamaarachchiH, HerathD. et al. DeepSelectNet: deep neural network based selective sequencing for oxford nanopore sequencing. BMC Bioinformatics 2023;24 :31. 10.1186/s12859-023-05151-0.36709261
20. Lin Y , ZhangY, SunH. et al. NanoDeep: a deep learning framework for nanopore adaptive sampling on microbial sequencing. Brief Bioinform 2023;25 :bbad499. 10.1093/bib/bbad499.
21. Nykrynova M , JakubicekR, BartonV. et al. Using deep learning for gene detection and classification in raw nanopore signals. Front Microbiol 2022;13 :942179. 10.3389/fmicb.2022.942179.36187947
22. Ni P , HuangN, ZhangZ. et al. DeepSignal: detecting DNA methylation state from Nanopore sequencing reads using deep-learning. Bioinformatics 2019;35 :4586–95. 10.1093/bioinformatics/btz276.30994904
23. Liu Z , JinL, ChenJ. et al. A survey on applications of deep learning in microscopy image analysis. Comput Biol Med 2021;134 :104523. 10.1016/j.compbiomed.2021.104523.34091383
24. Conneau A , SchwenkH, BarraultL. et al. Very Deep Convolutional Networks for Text Classification. In: Lapata M, Blunsom P, Koller A (eds). 2017 Conference of the European Chapter of the Association for Computational Linguistics (EACL). Association for Computational Linguistics, Valencia, Spain, 2017, 1107–16.
25. Hochreiter S , SchmidhuberJ. Long short-term memory. Neural Comput 1997;9 :1735–80. 10.1162/neco.1997.9.8.1735.9377276
26. Vaswani A , ShazeerN, ParmarN. et al. Attention Is all you Need. In: Luxburg U, Guyon I, Bengio S, et al. (ed). Proceedings of the 31st International Conference on Neural Information Processing Systems (NIPS'17). NY, USA: Curran Associates Inc., Red Hook, 6000–10.
27. Minaee S , KalchbrennerN, CambriaE. et al. Deep learning--based text classification: a comprehensive review. ACM computing surveys (CSUR) 2021;54 :1–40. 10.1145/3439726.
28. Ding X , ZhangX, HanJ. et al. Scaling up your Kernels to 31×31: Revisiting Large Kernel Design in CNNs. In: 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, New Orleans, LA, USA, 2022, 11953–65.
29. Oxford Nanopore Technologies, Oxford, UK. Nanopore Documentation: Guppy protocol. https://community.nanoporetech.com/protocols/Guppy-protocol/v/GPB_2003_v1_revT_14Dec2018.
30. Coates A , NgA, LeeH. An Analysis of Single-Layer Networks in Unsupervised Feature Learning. In: Gordon G, Dunson D, Dudík M (eds). 2011 International Conference on Artificial Intelligence and Statistics. PMLR, Fort Lauderdale, FL, USA, 2011, 215–23.
31. Liu Z , GaoJ, YangG. et al. Localization and classification of Paddy field pests using a saliency map and deep convolutional neural network. Sci Rep 2016;6 :20410. 10.1038/srep20410.26864172
32. Wang Y , WangA, LiuZ. et al. Single-molecule long-read sequencing reveals the chromatin basis of gene expression. Genome Res 2019;29 :1329–42. 10.1101/gr.251116.119.31201211
33. Liu Q , FangL, YuG. et al. Detection of DNA base modifications by deep recurrent neural network on Oxford Nanopore sequencing data. Nat Commun 2019;10 :2449. 10.1038/s41467-019-10168-2.31164644
34. Jesus JG , SacchiC, CandidoDDS. et al. Importation and early local transmission of COVID-19 in Brazil, 2020. Rev Inst Med Trop Sao Paulo 2020;62 :e30. 10.1590/s1678-9946202062030.32401959
35. Nicholls SM , QuickJC, TangS. et al. Ultra-deep, long-read nanopore sequencing of mock microbial community standards. Gigascience 2019;8 (5):giz043. 10.1093/gigascience/giz043.
36. Li H . Minimap2: Pairwise alignment for nucleotide sequences. Bioinformatics 2018;34 :3094–100. 10.1093/bioinformatics/bty191.29750242
37. Li H , HandsakerB, WysokerA. et al. The sequence alignment/map format and SAMtools. Bioinformatics 2009;25 :2078–9. 10.1093/bioinformatics/btp352.19505943
38. Oxford Nanopore Technologies, Oxford, UK. Nanopore Documentation: Chemistry technical document. https://community.nanoporetech.com/docs/sequence/sequencing_software/chemistry-technical-document/v/chtd_500_v1_revaq_07jul2016.
39. Al Shalabi L , ShaabanZ, KasasbehB. Data mining: a preprocessing engine[J]. J Comput Sci 2006;2 :735–9. 10.3844/jcssp.2006.735.739.
40. Cooijmans T , BallasN, LaurentC. et al. Recurrent batch normalization. arXiv preprint arXiv:1603.09025. 2016.
41. Ulrich JU , EppingL, PilzT. et al. Nanopore adaptive sampling effectively enriches bacterial plasmids. mSystems Published online February 9 2024;9 :e0094523. 10.1128/msystems.00945-23.38376263
42. Deserranno K , TillemanL, RubbenK. et al. Targeted haplotyping in pharmacogenomics using Oxford Nanopore Technologies' adaptive sampling. Front Pharmacol 2023;14 :1286764. 10.3389/fphar.2023.1286764.38026945
43. Martin S , HeavensD, LanY. et al. Nanopore adaptive sampling: a tool for enrichment of low abundance species in metagenomic samples. Genome Biol 2022;23 :11. 10.1186/s13059-021-02582-x.35067223
44. Wrenn DC , DrownDM. Nanopore adaptive sampling enriches for antimicrobial resistance genes in microbial communities. GigaByte 2023;2023 :gigabyte103.38111521
