
==== Front
iScience
iScience
iScience
2589-0042
Elsevier

S2589-0042(24)01943-6
10.1016/j.isci.2024.110718
110718
Article
An ensemble deep learning model for predicting minimum inhibitory concentrations of antimicrobial peptides against pathogenic bacteria
Chung Chia-Ru 1
Chien Chung-Yu 1
Tang Yun 2
Wu Li-Ching 3
Hsu Justin Bo-Kai 4
Lu Jang-Jih 567
Lee Tzong-Yi 28
Bai Chen baichen@cuhk.edu.cn
9∗
Horng Jorng-Tzong horng@db.csie.ncu.edu.tw
1510∗∗
1 Department of Computer Science and Information Engineering, National Central University, Taoyuan, Taiwan
2 Institute of Bioinformatics and Systems Biology, National Yang Ming Chiao Tung University, Hsinchu, Taiwan
3 Department of Biomedical Sciences and Engineering, National Central University, Taoyuan, Taiwan
4 Department of Computer Science and Engineering, Yuan Ze University, Taoyuan, Taiwan
5 Department of Laboratory Medicine, Chang Gung Memorial Hospital at Linkou, Taoyuan City, Taiwan
6 School of Medicine, Chang Gung University, Taoyuan City, Taiwan
7 Department of Medical Biotechnology and Laboratory Science, Chang Gung University, Taoyuan City, Taiwan
8 Center for Intelligent Drug Systems and Smart Biodevices (IDS2B), National Yang Ming Chiao Tung University, Hsinchu City, Taiwan
9 Warshel Institute for Computational Biology, School of Medicine, The Chinese University of Hong Kong (Shenzhen), Shenzhen 518172, China
∗ Corresponding author baichen@cuhk.edu.cn
∗∗ Corresponding author horng@db.csie.ncu.edu.tw
10 Lead contact

13 8 2024
20 9 2024
13 8 2024
27 9 11071829 12 2023
9 7 2024
8 8 2024
© 2024 The Authors
2024
https://creativecommons.org/licenses/by-nc-nd/4.0/ This is an open access article under the CC BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/).
Summary

The rise of antibiotic resistance necessitates effective alternative therapies. Antimicrobial peptides (AMPs) are promising due to their broad inhibitory effects. This study focuses on predicting the minimum inhibitory concentration (MIC) of AMPs against whom-priority pathogens: Staphylococcus aureus ATCC 25923, Escherichia coli ATCC 25922, and Pseudomonas aeruginosa ATCC 27853. We developed a comprehensive regression model integrating AMP sequence-based and genomic features. Using eight AI-based architectures, including deep learning with protein language model embeddings, we created an ensemble model combining bi-directional long short-term memory (BiLSTM), convolutional neural network (CNN), and multi-branch model (MBM). The ensemble model showed superior performance with Pearson correlation coefficients of 0.756, 0.781, and 0.802 for the bacterial strains, demonstrating its accuracy in predicting MIC values. This work sets a foundation for future studies to enhance model performance and advance AMP applications in combating antibiotic resistance.

Graphical abstract

Highlights

• Accurate MIC prediction of AMPs against key WHO-priority pathogens

• AI models, including BiLSTM and CNN, are used for superior MIC prediction

• Integration of sequence-based and genomic features enhances MIC prediction accuracy

• Ensemble model outperforms existing benchmarks in predicting MIC values

Chemistry; Biological sciences; Computer science

Subject areas

Chemistry
Biological sciences
Computer science
Published: August 13, 2024
==== Body
pmcIntroduction

The overuse of antibiotics has led to the emergence of microbial pathogens that are resistant to treatment, creating an urgent need to develop alternative treatments for infectious diseases. Antibiotic-resistant bacteria are now a global crisis, and the benefits of antibiotics are at risk.1 The ESKAPE pathogens, including Enterococcus faecium, Staphylococcus aureus, Klebsiella pneumoniae, Acinetobacter baumannii, Pseudomonas aeruginosa, and Enterobacter species, have been identified by the World Health Organization (WHO) as critical organisms in need of urgent research and development of new antibiotics.2 Antimicrobial peptides (AMPs) are a promising approach for the development of new therapeutic strategies. AMPs are part of the body’s immune defense system and possess several beneficial properties, such as antimicrobial, immunomodulating, wound healing, and anti-tumor activities.3,4 In contrast to traditional antibiotics, AMPs exhibit rapid bactericidal activity, lower concentrations required, and efficacy against resistant strains, making them a highly viable therapeutic option.5

AMPs are found in organisms from microorganisms to humans and are an essential part of the innate immune system.6 They have a high amount of cationic and hydrophobic amino acids, giving them a positive charge and amphiphilic properties that are essential to their effectiveness against bacteria by disrupting bacterial membranes.7,8 Additionally, some AMPs modulate the host defense systems.9 The minimum inhibitory concentration (MIC) is a crucial measure in antimicrobial research, as it determines the potency of antimicrobial agents. MIC is defined as the lowest concentration of an antimicrobial agent that prevents the visible growth of a microorganism after overnight incubation.10,11 This metric is important for evaluating new antimicrobial compounds, guiding optimal drug prescriptions, and developing effective treatment strategies, particularly in the increasing challenge of antibiotic resistance. In clinical practice, determining MIC values precisely informs therapeutic decision-making by aiding in selecting the most effective antimicrobial agent and dosage. In addition to improving patient outcomes, this helps antibiotic stewardship programs minimize the risk of developing resistance. For discovering and optimizing new antimicrobials, the predictive accuracy of MIC values for AMPs is particularly important. Improving our understanding and ability to predict MICs is essential, given the critical importance of the MIC in guiding antimicrobial therapy and the potential of AMPs in combating antibiotic resistance. This effort represents a significant step forward in antimicrobial research, promising to improve antimicrobial therapy and pave the way for novel approaches to combat resistant infections.

Several studies have investigated the activity of AMPs against different bacterial strains using various computational methods.12,13,14,15,16,17,18,19,20 Xiao et al. used pseudo-amino acid composition (PAAC) and Gaussian kernel regression to predict the MICs of AMPs against E. coli,12,21 while Chou’s PAAC was used to analyze the effects of sequence information22 from the Antimicrobial Peptide Database (APD).23 Dean et al. developed the PepVAE framework, which used variational autoencoder (VAE) for peptide sequence generation and gradient boost regression for MIC prediction,13 and compared various models, including convolutional neural network (CNN),24 elastic net,25 gradient boosting,26 kernel ridge,27 Lasso,28 random forest (RF),29 light gradient boosting machine (LGBM),30 and extreme gradient boosting (XGBoost).31 Yan et al.14 used a multi-branch CNN model with an attention mechanism that encodes peptide sequences via multiple methods for MIC prediction using data from DBAASP.32 Chen et al. developed the xDeep-AcPEP model using CNN,15 focusing on the anticancer activity of AMPs. Vishnepolsky et al.16 and Gull et al.17 combined genomic features and sequence-based properties to enhance predictions. Meanwhile, Sharma et al.18 created ESKAPEE-MICpred, which employs language models and the entire genome sequences of ESKAPEE pathogens to build their regressor model.

The recent advancements in artificial intelligence (AI) have significantly advanced the discovery of AMPs, including the generative adversarial networks (GANs) developed by Zervou et al.33 The designed AMPs exhibit enhanced antimicrobial properties and demonstrate high efficacy against drug-resistant bacteria. The feedback generative adversarial network (FBGAN) model incorporates classifiers to enhance the quality of the generated peptides.33 Transformer-based architecture, ProtTrans, has also demonstrated impressive performance in generating and predicting AMP activity. This is due to their ability to capture long-range dependencies and sequence motifs crucial for antimicrobial activity.34 Another promising approach is reinforcement learning (RL) for the design of AMPs. Wang et al. developed an RL framework that optimizes AMP sequences by iteratively refining peptide properties to improve antimicrobial potency and minimize toxicity.35 This method aimed to identify AMP candidates with optimal therapeutic profiles by exploring the extensive sequence space of AMPs. Furthermore, hybrid models that integrate empirical data with predictive algorithms have been developed. Lin et al. developed a hybrid deep learning model integrating CNNs with recurrent neural networks (RNNs) to predict AMP activity from high-dimensional screening data, achieving state-of-the-art performance.36

Computational models have the potential to impact the prediction of AMPs significantly. However, the practical application of these models is often hindered by the high synthesis costs and the inherent complexity of their molecular structures. Traditional models have relied on sequence or structural features to predict the effectiveness of AMPs or their MICs. However, this approach can overlook the broader genomic context significantly influencing bacterial behavior and treatment outcomes. Genomic features encompass variations in nucleotide compositions and patterns and offer a more comprehensive understanding of the microbial landscape. Integrating genomic data into predictive modeling allows for identifying the genetic mechanisms that govern bacterial resistance and susceptibility to AMPs. This integration enables the refinement of antimicrobial susceptibility test predictions, thus improving the efficacy of potential treatments by allowing for more precise tailoring of AMPs to specific bacterial strains. Therefore, the aim of our investigation is to develop regression models that can predict the MIC of AMPs against three ESKAPEE strains (S. aureus ATCC 25923, E. coli ATCC 25922, and P. aeruginosa ATCC 27853) by using both sequence-based and genomic features. The bacterial strains were chosen due to their clinical significance and prevalence in healthcare-associated infections. These strains represent a broad spectrum of Gram-positive and Gram-negative bacteria, offering a comprehensive challenge to the predictive capabilities of our model. Their inclusion in the study allows for a rigorous assessment of the MIC predictions across diverse bacterial types known for their varying resistance mechanisms. This selection aligns with the WHO’s priority list of pathogens that urgently require new therapeutic interventions, emphasizing the practical relevance of our research in addressing global health challenges posed by antibiotic-resistant bacteria.

Rather than settling for a binary output, the intent to capture the subtleties of AMP activity across a continuum drives the adoption of a regression framework over binary classification. This detailed quantification allows for more informed selection and optimization of AMPs, considering the specific nuances of each bacterial strain. Our approach aims to streamline the discovery of new AMPs and increase the precision of therapeutic interventions against pathogens by accurately predicting MIC values. This integrated methodological approach is expected to improve prediction accuracy and be essential in developing potent antimicrobial therapies. To achieve our research goals, we followed a systematic workflow, as illustrated in Figure 1. First, we applied rigorous preprocessing steps to ensure that only high-quality data were included by obtaining peptide sequence data from the DBAASP,32 dbAMP,37 and DRAMP,38 specifically targeting S. aureus ATCC 25923, E. coli ATCC 25922, and P. aeruginosa ATCC 27853. We then extracted relevant features for each peptide sequence and developed deep learning-based models. To improve model performance and interpretability, we attempted to visualize the adopted features. This comprehensive approach allowed us to accurately analyze the MICs of AMPs and identify the critical features essential for this analysis.Figure 1 Workflow of the proposed approach for predicting MICs against pathogenic bacteria of AMPs

The process begins with (1) data collection and preprocessing, sourcing AMP data from three primary databases—DBAASP, dbAMP, and DRAMP—and performing basic processing for target strains S. aureus, E. coli, and P. aeruginosa, along with the distribution analysis of peptide lengths and MICs. (2) Feature extraction transforms protein and genome sequences into numerical vectors using methods such as iFeature, sequence encoding, and pre-trained embeddings to generate both protein and genomic features. (3) Model construction involves developing machine learning models, such as Bi-LSTM, CNN, and a multi-branch model, that incorporate both peptide and genomic features, with an ensemble approach for prediction refinement. Finally, (4) performance evaluation and analysis are performed using techniques such as principal component analysis and comparative studies for the validation and benchmarking of model predictive performance.

Results

Analysis of antimicrobial peptides sequence characteristics and minimum inhibitory concentration distribution across targeted bacterial strains

The training and independent test sets of the study, as detailed in Table 1, form the fundamentals of our analysis of the MICs of the AMPs. Upon examination, we found that the sequence lengths and MIC values had unique distributions for each of the three bacterial strains. These findings are visually represented in Figure 2, which presents a spectrum of MIC values against varying peptide lengths. For S. aureus ATCC 25923, the peptides had an average log MIC of 1.08, with a wide range of antibacterial activity as shown in Figure 2A. Peptides targeting E. coli ATCC 25922 had a slightly lower average log MIC of 1.035 but exhibited greater variability, as indicated by a larger standard deviation (Figure 2B). Finally, peptides that target P. aeruginosa ATCC 27853 have an average log MIC of 1.179 with a smaller standard deviation, indicating a more consistent effectiveness among these peptides (Figure 2C). These distributions emphasize the importance of considering both peptide length and sequence intricacies when developing AMPs that are tailored to specific bacterial pathogens.Table 1 Number of peptides for the training and independent testing sets for each bacterium

Bacterial strains	Training set	Validation set	Independent testing set	Total	
S. aureus ATCC 25923	1692	423	529	2644	
E. coli ATCC 25922	2443	611	764	3818	
P. aeruginosa ATCC 27853	1572	394	492	2458	
Total number of peptides	5707	1428	1785	8920	

Figure 2 Correlation of AMP sequence length and MIC distribution across three bacterial strains

Scatterplots illustrating the correlation between sequence length and MIC values expressed in log(μM) for AMPs targeting three bacterial strains: (A) Staphylococcus aureus ATCC 25923, (B) Escherichia coli ATCC 25922, and (C) Pseudomonas aeruginosa ATCC 27853. The horizontal axis represents the MIC in log(μM), while the vertical axis shows the sequence length of the AMPs. Each point corresponds to a single AMP. Histograms on the left y axis show the distribution of sequence lengths. The green solid line indicates the mean MIC value, while the purple and orange lines indicate one standard deviation above and below the mean, respectively.

In this study, we categorized AMPs based on their MIC levels to investigate the relationship between MIC values and AMP sequence characteristics. The MIC categories were defined as follows: lower-MIC (below 0.5 logμM), medium-MIC (between 0.5 logμM and 1.8 logμM), and higher-MIC (above 1.8 logμM). These MIC thresholds were chosen to ensure a balanced representation between the lower and higher MIC groups for our analysis. Figure 3 shows the results, while Table 2 provides the details. The majority of AMPs exhibit medium MIC values against the targeted strains, with sequence lengths ranging from 10 to 25 amino acids. AMPs with higher MIC values are generally observed within a sequence length bracket of 10–20 amino acids. On the other hand, the sequence length distribution of lower-MIC AMPs appears more uniform across the spectrum. Lysine (K) and arginine (R) are consistently prevalent across all MIC categories, underscoring their potential role in antimicrobial efficacy. In the lower-MIC category, cysteine (C) occurs markedly more frequently, approximately 1.5–2 times, than in the medium-MIC and higher-MIC groups, suggesting its significant contribution to the potency of AMPs.Figure 3 Comprehensive analysis of AMPs categorized by MIC against three bacterial strains

(A) shows the number and percentage of AMPs classified into low, medium, and high MIC categories for S. aureus ATCC 25923 (SA), E. coli ATCC 25922 (EC), and P. aeruginosa ATCC 27853 (PA).

(B, D, and F) show the distribution of AMP sequence lengths relative to their MIC values for S. aureus (B), E. coli (D), and P. aeruginosa (F), respectively.

(C, E, and G) show the distribution of amino acid composition in relation to MIC levels in the respective strains for S. aureus (C), E. coli (E), and P. aeruginosa (G), respectively.

Table 2 Number of AMPs against specific strains after classification

Bacterial strains	Higher MIC	Medium MIC	Lower MIC	
Staphylococcus aureus ATCC 25923	569	1467	608	
Escherichia coli ATCC 25922	728	2220	870	
Pseudomonas aeruginosa ATCC 27853	617	1415	426	

To visualize and interpret the behavior of the three sequence-based features in our study, we used principal component analysis (PCA).39 PCA is a technique that reduces dimensionality in complex datasets, making it an invaluable tool for elucidating patterns in high-dimensional data by projecting it onto a lower-dimensional subspace. This simplifies the complexity while retaining critical structural information. Figure 4 presents a graphical representation of the dimensionality reduction applied to the three distinct features of our bacterial strains. The PCA plots aim to differentiate AMPs with higher and lower MIC values and illustrate the underlying structure of the data. However, the plots also highlight the challenge of distinguishing between higher-MIC and lower-MIC samples through these reduced dimensions alone. This suggests that although PCA can reveal broad trends and clustering within the data, more sophisticated or alternative analytical strategies are required to efficiently distinguish different MIC values.Figure 4 Two-dimensional PCA visualization of AMPs based on three different feature extraction methods for S. aureus ATCC 25923, E. coli ATCC 25922, and P. aeruginosa ATCC 27853

(A, D, and G) display the PCA plots for AMPs against S. aureus, generated using iFeature, sequence encoding, and pre-trained embeddings, respectively. Similarly, (B, E, and H) illustrate the PCA results for E. coli, and (C, F, and I) for P. aeruginosa. Each point represents an AMP, color-coded by MIC category: red for higher MIC, green for medium MIC, and blue for lower MIC. These visualizations aid in assessing the feature space differentiation and clustering tendencies of AMPs based on the applied feature extraction technique.

Performance of different models with sequence-based features

Our analysis used four traditional machine learning models—RF, XGBoost, categorical boosting (CatBoost), and LGBM— to predict MICs for different strains. The models were trained using a set of features derived from the iFeature, including amino acid composition (AAC), PAAC, composition, transition and distribution descriptor (CTDD), and grouped amino acid composition (GAAC). These features were selected for their relevance and compatibility with the traditional machine learning approaches employed. Table 3 presents the performance metrics of each model for the bacterial strains S. aureus, E. coli, and P. aeruginosa. The metrics include mean squared error (MSE), root mean squared error (RMSE), coefficient of determination (R2), and Pearson correlation coefficient (PCC), providing a comprehensive view of each model’s predictive capabilities. The results indicate that all models perform competently, but some models have a slight edge over others in specific contexts. For example, the RF model shows consistent performance across all bacteria in terms of MSE and RMSE, suggesting its robustness. In contrast, the LGBM exhibits slightly higher error metrics, particularly with S. aureus, which may indicate potential overfitting or a need for parameter tuning. The R2 values indicate that a significant portion of the variance in MIC values can be explained by the models, particularly for E. coli, which has the highest R2 values overall. All models maintain a strong PCC with actual MIC values, supporting the validity of the features and models used. These findings not only confirm the effectiveness of the selected features and models but also guide future optimization and selection of traditional machine learning strategies for MIC prediction in antimicrobial research.Table 3 Evaluation of machine learning models for MIC prediction using iFeature-derived features on the independent testing set

Metrics	Bacteria	Models	
RF	XGBoost	CatBoost	LGBM	
MSE	S. aureus	0.369	0.385	0.369	0.397	
E. coli	0.294	0.302	0.306	0.299	
P. aeruginosa	0.312	0.318	0.315	0.310	
RMSE	S. aureus	0.607	0.620	0.607	0.630	
E. coli	0.542	0.549	0.553	0.546	
P. aeruginosa	0.558	0.563	0.561	0.556	
R2	S. aureus	0.421	0.396	0.420	0.377	
E. coli	0.480	0.466	0.459	0.472	
P. aeruginosa	0.451	0.440	0.445	0.453	
PCC	S. aureus	0.670	0.647	0.659	0.620	
E. coli	0.697	0.688	0.682	0.692	
P. aeruginosa	0.681	0.672	0.671	0.680	
The table presents the predictive accuracy of four machine learning algorithms—random forest (RF), extreme gradient boosting (XGBoost), categorical boosting (CatBoost), and light gradient boosting machine (LGBM)—for estimating the MIC across three bacterial strains. The models were trained on a comprehensive set of features extracted via the iFeature. These features include amino acid composition (AAC), pseudo-amino acid composition (PAAC), composition, transition, and distribution descriptor (CTDD), and grouped amino acid composition (GAAC). The evaluation metrics include mean squared error (MSE), root mean squared error (RMSE), coefficient of determination (R2), and Pearson correlation coefficient (PCC) for S. aureus, E. coli, and P. aeruginosa.

Table 4 summarizes our investigation into the predictive ability of deep learning models for estimating the MIC of AMPs against different bacterial species. We used various sequence encoding techniques to capture the essential evolutionary and physicochemical properties of amino acid sequences through four distinct matrices: AAIndex, BLOSUM62, Z-Scale, and one-hot encoding. These methodologies, which are tailored for deep learning applications, provide a multidimensional perspective on the characteristics of peptides. The evaluation of the models, including bi-directional long short-term memory (BiLSTM), CNN, and multi-branch model (MBM), is presented through a series of metrics: MSE, RMSE, R2, and PCC. The analysis shows that CNNs are particularly proficient, especially with E. coli, as indicated by the lowest MSE and RMSE values. The BiLSTM and MBM also demonstrate commendable performance. The transformer model was not included in this comparative study because its complex architecture was not suitable for the selected features, which could lead to overfitting even with parameter adjustments. Table 4 provides valuable insights into the effectiveness of different deep learning architectures in antimicrobial research and highlights the importance of feature selection in model training. The results indicate that the CNN performs better in capturing the nuances of peptide sequences, suggesting potential avenues for further optimization of MIC prediction models.Table 4 Evaluation of deep learning models for MIC prediction using sequence encoding features on the independent testing set

Metrics	Bacteria	Models	
BiLSTM	CNN	Transformer	MBM	
MSE	S. aureus	0.507	0.410	0.637	0.481	
E. coli	0.394	0.291	0.573	0.406	
P. aeruginosa	0.445	0.358	0.571	0.401	
RMSE	S. aureus	0.712	0.640	0.798	0.693	
E. coli	0.627	0.539	0.757	0.637	
P. aeruginosa	0.667	0.598	0.756	0.633	
R2	S. aureus	0.204	0.356	<0.001	0.245	
E. coli	0.303	0.486	<0.001	0.283	
P. aeruginosa	0.216	0.369	<0.001	0.293	
PCC	S. aureus	0.461	0.613	NaN	0.499	
E. coli	0.578	0.698	NaN	0.547	
P. aeruginosa	0.524	0.667	NaN	0.552	
The table presents the predictive accuracy of four deep learning algorithms—Bi-directional long short-term memory (BiLSTM), convolutional neural network (CNN), transformer, and multi-branch model (MBM)—for estimating the MIC across three bacterial strains. These models utilized a comprehensive set of sequence encoding features, including AAIndex, BLOSUM62, Z-Scale, and One-Hot Encoding, to capture the nuanced characteristics of peptide sequences. The evaluation metrics include mean squared error (MSE), root mean squared error (RMSE), coefficient of determination (R2), and Pearson correlation coefficient (PCC) for S. aureus, E. coli, and P. aeruginosa.

The models utilize pre-trained embeddings, generated by protein language models, to analyze complex biological data and extract essential information from protein sequences. This information is critical for predicting the activity of AMPs. Our analysis utilizes transfer learning to enhance the models with pre-existing biological knowledge. In this study, the same deep learning models and evaluation metrics were considered. Table 5 shows that CNNs have high accuracy in predicting MIC, with the lowest MSE and RMSE for E. coli, indicating their robust modeling capabilities. The Transformer model was generally effective, but it presented higher errors in the context of S. aureus. This finding suggests that further investigation is needed to determine model fit and to adjust the features. The R2 values suggest that the models explain a moderate to high proportion of variance, with the CNN model being particularly effective for E. coli. The PCC results confirm strong positive correlations across all models, especially with the CNN model for E. coli, indicating a significant alignment with empirical MIC values. These observations demonstrate the potential of pre-trained embeddings to improve the predictive modeling of the MIC values of AMPs. These findings have implications for designing AMP therapeutics against specific bacterial infections.Table 5 Evaluation of deep learning models for MIC prediction using pre-trained embeddings features on the independent testing set

Metrics	Bacteria	Models	
BiLSTM	CNN	Transformer	MBM	
MSE	S. aureus	0.359	0.387	0.508	0.369	
E. coli	0.283	0.256	0.293	0.288	
P. aeruginosa	0.363	0.350	0.322	0.349	
RMSE	S. aureus	0.599	0.622	0.712	0.607	
E. coli	0.531	0.505	0.541	0.536	
P. aeruginosa	0.602	0.591	0.576	0.590	
R2	S. aureus	0.437	0.393	0.203	0.421	
E. coli	0.501	0.548	0.482	0.491	
P. aeruginosa	0.360	0.383	0.415	0.386	
PCC	S. aureus	0.665	0.628	0.470	0.661	
E. coli	0.709	0.740	0.706	0.701	
P. aeruginosa	0.641	0.634	0.649	0.649	
This table displays the performance of four deep learning models—bi-directional long short-term memory (BiLSTM), convolutional neural network (CNN), transformer, and multi-branch model (MBM)—in predicting MIC for three bacterial strains using pre-trained embedding features. The evaluation metrics include mean squared error (MSE), root mean squared error (RMSE), coefficient of determination (R2), and Pearson correlation coefficient (PCC) for S. aureus, E. coli, and P. aeruginosa.

Performance of models integrating genomic and antimicrobial peptide sequence features

Table 6 illustrates a detailed comparative analysis of four traditional machine learning models—RF, XGBoost, CatBoos, and LGBM—using both sequence-based and genomic features to predict MICs for three bacterial strains. The models were enhanced with genomic features obtained from encoding microbial strain genomes, in addition to the sequence-based features from iFeature. This integration aimed to provide a more comprehensive biological context for each AMP, treating every unique AMP-strain combination as an individual data sample. In this extensive analysis, we combined the individual training datasets of the three strains, resulting in a final training dataset of 5707 samples. This approach facilitated a robust model training phase by utilizing both the diversity and specificity of the data. The results indicate that the inclusion of genomic features in conjunction with sequence-based features significantly enhances the models' predictive accuracies. Specifically, the RF and XGBoost models perform particularly well, as evidenced by lower MSE and RMSE and higher R2 and PCC, indicating more accurate and correlated prediction of MIC values. The improvements in all areas indicate that the combined features offer a more detailed comprehension of microbial resistance, resulting in more precise MIC predictions. These findings highlight the potential of an integrated approach that combines genomic and sequence-based features to enhance the predictive capability of traditional machine learning models in antimicrobial research. By utilizing all available biological data, we can make more accurate and comprehensive predictions of MIC values, which can aid in the development of more effective antimicrobial treatments.Table 6 Evaluation of machine learning models for MIC prediction using iFeature and genomic features on the independent testing set

Metrics	Bacteria	Models	
RF	XGBoost	CatBoost	LGBM	
MSE	S. aureus	0.336	0.334	0.396	0.385	
E. coli	0.281	0.289	0.317	0.311	
P. aeruginosa	0.251	0.274	0.300	0.285	
RMSE	S. aureus	0.579	0.577	0.629	0.620	
E. coli	0.530	0.537	0.563	0.557	
P. aeruginosa	0.500	0.523	0.547	0.533	
R2	S. aureus	0.473	0.475	0.379	0.396	
E. coli	0.503	0.489	0.440	0.451	
P. aeruginosa	0.558	0.517	0.472	0.497	
PCC	S. aureus	0.691	0.702	0.637	0.656	
E. coli	0.716	0.706	0.673	0.689	
P. aeruginosa	0.761	0.722	0.691	0.718	
The table presents the predictive accuracy of four machine learning algorithms—random forest (RF), extreme gradient boosting (XGBoost), categorical boosting (CatBoost), and light gradient boosting machine (LGBM)—for estimating the MIC across three bacterial strains. The models were trained on a comprehensive set of features extracted via the iFeature and genomic features. These features include amino acid composition (AAC), pseudo-amino acid composition (PAAC), composition, transition, and distribution descriptor (CTDD), grouped amino acid composition (GAAC), and genomic features. The evaluation metrics include mean squared error (MSE), root mean squared error (RMSE), coefficient of determination (R2), and Pearson correlation coefficient (PCC) for S. aureus, E. coli, and P. aeruginosa.

Table 7 demonstrates the culmination of our efforts to enhance MIC predictions by exploiting the synergy of deep learning models, pre-trained embeddings, and genomic features. We focused exclusively on deep learning architectures integrated with potent pre-trained embeddings and comprehensive genomic features. For each bacterial strain, we implemented a blend of BiLSTM, CNN, Transformer, and MBM. The genomic features underwent processing via a basic artificial neural network (ANN) and were then combined with sequence-based information. This guided the data through the respective deep learning architectures before culminating at the regressor. Table 7 outlines the metrics used to evaluate model performance. The MBM for S. aureus and BiLSTM for P. aeruginosa show promising results, aligning well with actual MIC values. This is evidenced by their lower MSE and RMSE and higher R2 and PCC values. The integration of genomic and pre-trained embedding features could effectively capture the complex biological interactions underlying antimicrobial resistance. These results demonstrate the potential of deep learning to advance MIC prediction. Through the use of advanced architectural strategies and rich biological data, these models offer a path to more accurate and nuanced predictions of bacterial resistance, which would facilitate more effective antimicrobial strategies.Table 7 Evaluation of deep learning models for MIC prediction using pre-trained embeddings and genomic features on the independent testing set

Metrics	Bacteria	Models	
BiLSTM	CNN	Transformer	MBM	
MSE	S. aureus	0.332	0.331	0.394	0.310	
E. coli	0.235	0.271	0.313	0.274	
P. aeruginosa	0.230	0.248	0.330	0.266	
RMSE	S. aureus	0.576	0.575	0.627	0.556	
E. coli	0.484	0.520	0.559	0.523	
P. aeruginosa	0.479	0.497	0.574	0.515	
R2	S. aureus	0.495	0.481	0.382	0.514	
E. coli	0.584	0.521	0.447	0.516	
P. aeruginosa	0.594	0.564	0.419	0.532	
PCC	S. aureus	0.720	0.697	0.622	0.738	
E. coli	0.772	0.732	0.680	0.747	
P. aeruginosa	0.785	0.756	0.698	0.757	
This table displays the performance of four deep learning models—bi-directional long short-term memory (BiLSTM), convolutional neural network (CNN), transformer, and multi-branch model (MBM)—in predicting MIC for three bacterial strains using pre-trained embedding features and genomic features. The evaluation metrics include mean squared error (MSE), root mean squared error (RMSE), coefficient of determination (R2), and Pearson correlation coefficient (PCC) for S. aureus, E. coli, and P. aeruginosa.

Comparison of individual and ensemble models

Table 8 illustrates the predictive performance of our ensemble model versus individual models for MIC prediction. Our ensemble strategy was carefully designed to combine AMP sequence-based and genomic features, leveraging the strengths of selected models to achieve improved prediction accuracy. After a thorough evaluation, we identified the BiLSTM, CNN, and MBM models as the most effective for predicting performance across different feature sets for each bacterial strain. The BiLSTM model, known for its ability to unravel sequential dependencies, stood out significantly when augmented with pre-trained embeddings.Table 8 Comparison of performance of best-performing individual and ensemble models among bacterial strains

Models	S. aureus	E. coli	P. aeruginosa	
BiLSTM	MBM	Ensemble Model	CNN	BiLSTM	Ensemble Model	LGBM	BiLSTM	Ensemble Model	
Feature	Pre-trained embeddings	Pre-trained embeddingsa	Pre-trained embeddingsa	Pre-trained embeddings	Pre-trained embeddingsa	Pre-trained embeddingsa	iFeature	Pre-trained embeddingsa	Pre-trained embeddingsa	
MSE	0.359	0.310	0.274	0.256	0.235	0.225	0.310	0.230	0.205	
RMSE	0.599	0.556	0.523	0.505	0.484	0.474	0.556	0.479	0.452	
R2	0.437	0.514	0.570	0.548	0.584	0.603	0.453	0.594	0.638	
PCC	0.665	0.738	0.756	0.740	0.772	0.781	0.680	0.785	0.802	
The bold type indicates better performance.

This table contrasts the predictive performance of the best-performing individual models for each bacterial strain—BiLSTM and MBM for S. aureus, CNN and BiLSTM for E. coli, and LGBM and BiLSTM for P. aeruginosa—against the ensemble model. The ensemble model integrated these individual models using a combination of pre-trained embeddings and genomic features, highlighting its superior predictive accuracy through metrics such as mean squared error (MSE), root mean squared error (RMSE), coefficient of determination (R2), and Pearson correlation coefficient (PCC). The selection of individual models for comparison was based on their demonstrated excellence in specific feature sets, underlining the ensemble model’s robustness across varied analytical contexts.

a Means with genomic features.

The ensemble model was created using a weighted averaging scheme, primarily influenced by the performance metrics of the individual models. The BiLSTM model was given significant weight due to its consistent superiority in prediction, while the CNN and MBM models also made notable contributions to the ensemble’s performance. This approach ensured a balanced integration of the models, effectively utilizing their unique strengths while compensating for their weaknesses. The ensemble model outperformed the others, achieving the lowest MSE and RMSE, as well as the highest R2 and PCC for most of the bacterial strains assessed. These results demonstrate the robust predictive ability of the ensemble model, highlighting the effectiveness of combining the strengths of individual models through ensemble techniques.

To clarify the selection process for the individual models compared to the ensemble model in Table 8, it is essential to note that the models were chosen based on their exemplary performance within their respective feature sets. This was evidenced by their PCC across prior analyses. For example, the BiLSTM model, which employed pre-trained embeddings, demonstrated unparalleled performance for S. aureus, as highlighted in previous tables. Similarly, the CNN model performed well in predicting MIC values for E. coli. This focused selection allowed for a meaningful comparison, further supporting the ensemble model’s superior predictive capacity. The ensemble model strategically combines the strengths of individual models tailored to each bacterial strain.

Comparative analysis with existing models

After demonstrating the superior performance of our ensemble model compared to our previous models, we recognized the importance of benchmarking it against other established models in the AMP prediction. In the benchmarking section of our study, we aimed to objectively evaluate the performance of our ensemble model against existing models in the AMP prediction domain, particularly its ability to predict MICs against specific bacterial strains. Although related studies often use heterogeneous datasets and source codes and testing platforms are often limited, we identified the ESKAPEE-MICpred18 model as a significant benchmark due to its focus on predicting MIC values and providing a public web server. To ensure a fair and robust comparison, we curated an independent testing dataset specifically designed for benchmarking. This dataset was carefully curated to prevent any overlap with the training data of the ESKAPEE-MICpred model, specifically excluding data from the DBAASP32 database. This was done to minimize bias and ensure the integrity of our comparison. Our independent dataset consisted of peptides that target three bacterial strains: S. aureus ATCC 25923, E. coli ATCC 25922, and P. aeruginosa ATCC 27853. The number of AMPs for each strain is detailed in Table S1. The composition of our benchmarking dataset and the deliberate exclusion of data derived from DBAASP highlight our commitment to achieving a balanced and representative sample for model evaluation.

Table 9 compares our ensemble model and ESKAPEE-MICpred for three bacterial strains. We chose PCC and MSE as our primary metrics for evaluating model performance. This choice was driven by the need, in the context of the biological variability and complex patterns inherent in MIC data, to accurately measure the models' predictive accuracy (PCC) and precision (MSE). Although the R2 metric is frequently used in regression analyses, we carefully reconsidered its suitability for the MIC prediction task. Due to the unique distribution of MIC values and the specific objectives of our study, we prioritized metrics that directly align with our goal of developing highly accurate and precise models for predicting the MIC of AMPs against pathogenic bacteria. Our ensemble model consistently demonstrated lower MSE and RMSE values, along with higher PCC values than ESKAPEE-MICpred. This indicates that our model has a more precise and reliable prediction capability on the independent dataset. Figure 5 further visualizes these findings by comparing the prediction scatterplots of both models.Table 9 Performance of ESKAPEE-MICpred and our ensemble model for the prediction of MICs across three bacterial strains: Staphylococcus aureus, Escherichia coli, and Pseudomonas aeruginosa

	S. aureus	E. coli	P. aeruginosa	
ESKAPEE-MICpred	Ensemble Model	ESKAPEE-MICpred	Ensemble Model	ESKAPEE-MICpred	Ensemble Model	
MSE	0.443	0.310	0.439	0.190	0.491	0.354	
RMSE	0.665	0.556	0.662	0.435	0.700	0.594	
PCC	0.202	0.500	0.265	0.645	0.146	0.429	
The metrics include mean squared error (MSE), root mean squared error (RMSE), and Pearson correlation coefficient (PCC), reflecting the accuracy and correlation of the predicted MIC values with the experimental data. Lower values of MSE and RMSE indicate better model performance, while higher values of PCC indicate a stronger positive correlation between predicted and experimental MIC values.

Figure 5 Comparative analysis of MIC predictions for three bacterial strains using our ensemble model and ESKAPEE-MICpred

Scatterplots illustrate the relationship between experimentally determined MIC values (log μM) and the predicted MIC values by both methods.

(A) S. aureus ATCC 25923, (B) E. coli ATCC 25922, and (C) P. aeruginosa ATCC 27853. Each point represents an individual antibiotic compound. The blue dots denote predictions made by the ESKAPEE-MICpred model, and the red dots denote predictions made by our ensemble model. The solid lines indicate the regression line for each prediction model with corresponding colors, and the gray line represents the line of perfect positive correlation, indicating an ideal prediction match with the experimental results.

Discussion

To address the growing challenge of bacterial resistance to conventional antibiotics, our study explores AMPs as a promising alternative. AMPs have demonstrated potential in effectively inhibiting various bacterial strains, positioning themselves as a key solution in this ongoing battle. Understanding and accurately predicting the activity of AMPs is crucial to exploiting their potential, and this has been the primary goal of our research.

An ensemble model was developed to regress the experimental MIC of AMPs against three specific bacterial strains: S. aureus ATCC 25923, E. coli ATCC 25922, and P. aeruginosa ATCC 27853. The approach was comprehensive, integrating both AMP sequence-based features and genomic features to enrich the predictive power of the model. Eight different artificial intelligence-based architectures were used for the rigorous evaluation and training of our models. In our study, the incorporation of genomic features was a significant improvement in model performance. The BiLSTM, CNN, and MBM consistently demonstrated high performance. We constructed an ensemble model that leveraged the strengths of these top models and exhibited improved predictive accuracy in estimating MIC values. This study found that the model outperformed others, demonstrating superior efficacy across all three bacterial strains with PCC exceeding 0.75.

Our analysis of various bacterial strains has led to notable observations. Specifically, for S. aureus, the BiLSTM model using pre-trained embedding features was the top performer. For E. coli, the CNN utilizing pre-trained embedding features demonstrated superior performance. In contrast, the LGBM integrated with iFeature features outperformed other models for P. aeruginosa. These findings highlight the effectiveness of deep learning models paired with pre-trained embeddings, especially in scenarios with abundant data. However, for datasets with fewer samples, traditional machine learning models combined with iFeature provide a viable and robust alternative. Consequently, we identified the best-performing model for each bacterial strain, which sets the stage for a focused and comparative analysis of these optimized approaches.

Analyzing the performance of our model showed a significant improvement in MIC predicting accuracy when genomic features are integrated with AMP sequence-based features. The MBM was found to be the most effective for S. aureus, demonstrating its robustness to this strain. The BiLSTM model demonstrates superior performance for both E. coli and P. aeruginosa, indicating its versatility across different bacterial types. In our comparative assessment, we compared the top performing sequence-only models with those using the combined genomic and sequence-based approach. The top three models for MIC prediction across different strains were consistently BiLSTM, CNN, and MBM. These models are characterized by their adaptability to pre-trained embeddings and genomic features, demonstrating the potential of a multifaceted approach to improve prediction accuracy. These results confirm the importance of integrating diverse biological data sources and set a precedent for future studies to explore the synergy of sequence and genomic features in predicting MIC. The BiLSTM, CNN, and MBM collectively perform well, adapting to advanced features and paving the way for the development of more accurate and reliable antimicrobial treatment strategies.

The results of our study are promising. Several factors contribute to the enhanced performance of our ensemble model. First, our model focuses on three bacterial strains. This provides a focused and tailored approach compared to the broader target set of ESKAPEE-MICpred. Furthermore, our dataset, derived from three distinct AMP databases, offers a wider spectrum of AMP variations, potentially enriching our model’s learning and prediction capacity. Finally, it is possible that the differences in how ESKAPEE-MICpred and our model collect and integrate MIC values could have influenced the comparative performance outcomes. Overall, these results confirm the effectiveness of our approach, which combines AMP sequence-based and genomic features with an ensemble of top-performing models, in achieving improved predictive accuracy. This methodology establishes a new benchmark in predicting AMP activity and highlights the potential of detailed, feature-rich models in advancing the field.

Although these results are promising, there is still room for improvement. Increasing the volume of data in our training dataset, particularly with the continuous verification of new AMPs, could further enhance the accuracy of our models. Experimenting with alternative methods for encoding both AMP sequences and genomic data may unlock new avenues in AMP prediction. Our study provides valuable contributions to the field of AMP research. Our findings could assist researchers in several ways, such as developing new AMPs, categorizing their efficacy, and exploring other aspects of AMP-related research. This work is a stepping stone toward more effective strategies in the fight against antibiotic resistance.

STAR★Methods

Key resources table

REAGENT or RESOURCE	SOURCE	IDENTIFIER	
Deposited data	
	
DBAASP	Pirtskhalava et al.32	https://dbaasp.org/home	
dbAMP	Jhong et al.37	https://awi.cuhk.edu.cn/dbAMP/	
DRAMP	Shi et al.38	http://dramp.cpu-bioinfor.org/	
	
Software and algorithms	
	
Python version 3.8.12	Python Software Foundation	https://www.python.org	
iFeature	Chen et al.40	https://ifeature.erc.monash.edu/	
MathFeature	Bonidia et al.41	https://bio.tools/mathfeature	
scikit-learn version 1.0.1	Pedregosa et al.42	https://scikit-learn.org/stable/	
Tensorflow version 2.3.0	Abadi et al.43	https://pypi.org/project/tensorflow/	

Resource availability

Lead contact

Further information and requests for resources should be directed to and will be fulfilled by the lead contact, Prof. Jorng-Tzong Horng (horng@db.csie.ncu.edu.tw).

Materials availability

No new unique reagents were generated in this study.

Data and code availability

• The datasets used in this study are available on Github: https://github.com/chungcr/esAMPMIC. We have also shared the key codes. Additional information to reanalyze the data reported in this paper is available from the lead contact.

Method details

Data collection and preprocessing

In recent years, the availability of datasets related to AMPs has increased significantly. These datasets provide easy access to detailed information. These datasets are essential for AMP research, providing valuable insights into various aspects such as chemical modifications, amino acid sequences, bioactivities, 3D structures, and toxicities. In this study, three databases (DBAASP,32 dbAMP,37 and DRAMP38) were used to collect the data for AMPs specifically targeting S. aureus ATCC 25923, E. coli ATCC 25922, and P. aeruginosa ATCC 27853.

Our first step was to refine the peptide sequences by selecting only those peptides containing the 20 basic amino acids commonly found in biological systems, including Alanine (A), Arginine (R), Asparagine (N), Aspartic Acid (D), Cysteine (C), Glutamine (Q), Glutamic Acid (E), Glycine (G), Histidine (H), Isoleucine (I), Leucine (L), Lysine (K), Methionine (M), Phenylalanine (F), Proline (P), Serine (S), Threonine (T), Tryptophan (W), Tyrosine (Y), and Valine (V). To preserve the accuracy and relevance of our analysis, peptides containing amino acids B, J, O, U, X and Z were excluded from our investigation, in line with the typical short peptide nature of AMPs as noted in various studies.12,13,14,15,16,18 We focused on peptides with a length ranging from 6 to 40 amino acids, with MIC as our primary activity measure. In particular, there were several duplicates in the dataset, so for equivalent sequences we selected the sequence with the lowest MIC. We furthermore observed disparity in MIC units, frequently denoted in μM (micromolar) and μg/mL (micrograms per milliliter). For normalization of the data, we transformed all MIC values to μM. This process involved calculating the molecular weight (MW) of each sequence, changing the MIC values from μg/mL to μM, and finally converting all values to logμM. A logarithmic scale was used for MIC values due to the diverse and broad range of MIC values encountered across different AMPs and bacterial strains. This choice aligns with common practices in microbiological research where MIC values span several orders of magnitude. Utilizing a logarithmic scale allows for a more balanced representation of data, reducing the skewness associated with the wide range of MIC concentrations. This approach enables a more precise and significant comparison of MIC values among various AMPs and target bacteria. It ensures that variations in high and low concentrations are accurately captured and reflected in our analysis. Moreover, the logarithmic transformation of MIC values helps to stabilize variance, improve the homoscedasticity of the data, and make statistical analyses more robust and interpretable.

After completing the preprocessing steps, we identified a total of 2644, 3818, and 2458 AMPs targeting the bacterial strains of S. aureus ATCC 25923, E. coli ATCC 25922, and P. aeruginosa ATCC 27853, respectively. The dataset was then partitioned into training and independent testing sets in a 4:1 ratio to prepare for robust model training and evaluation. This division strategy ensures that the model is exposed to a diverse range of data while reserving a significant portion for unbiased evaluation. We then refined our training set by dividing it into a smaller training set and a validation dataset, using a 4:1 ratio. This additional split facilitates a more nuanced model tuning process, allowing for iterative improvements while monitoring and preventing overfitting.

Utilizing iFeature for feature extraction

In our study, the transformation of peptide and genome sequences into numerical vectors is crucial for feeding them into traditional machine learning and deep learning models. Several feature extraction approaches were used for this task, with a focus on using the iFeature platform.40,44 This platform provides a variety of techniques for extracting significant features from sequences.

We analyzed four main feature categories, computed using the iFeature: AAC, PAAC, CTD, and GAAC. To facilitate ML model implementation, we have summarized these categories in Table S2.

AAC represents the proportion of each amino acid in a peptide sequence. It is calculated using the formulaF(t)=N(t)N,

where t is one of the 20 natural amino acids (A, C, D, …, Y), N(t) is the number of amino acid t, and N is the total sequence length. For instance, in the peptide sequence “PPIVGSIPLGCG”, the AAC value for “C” is 0.0833, computed based on its sole occurrence in a 12-residue sequence.

The PAAC22 is a widely used feature in AMPs research. PAAC incorporates various physicochemical properties such as hydrophobicity values, hydrophilicity, side chain masses, and sequence-order information. The feature set of PAAC comprises over 20 attributes, with the first 20 features sharing similarities with AAC but encompassing supplementary information. The remaining features primarily relate to sequence order and are referred to as sequence order-correlated factors.

The CTD features show the patterns of amino acid distribution associated with specific physicochemical or structural properties within a peptide sequence. These features consist of 13 different physicochemical properties, including hydrophobicity, normalized van der Waals volume, polarity, polarizability, charge, secondary structures, and solvent accessibility. These properties have been previously utilized to calculate these features.40 In this study, we specifically focus on the distribution descriptor as part of our feature set. For each physicochemical property, the 20 natural amino acids are divided into three groups, and their distributions in the sequence are calculated at different positions: the first residue, 25%, 50%, 75%, and 100% of the sequence.

The GAAC divides the 20 amino acids into five groups according to their physicochemical properties. These groups consist of the aliphatic group (g1: GAVLMI), aromatic group (g2: FYW), positive charge group (g3: KRH), negative charge group (g4: DE), and uncharged group (g5: STCPNQ). The GAAC descriptor measures the occurrence of each group, with the formulaF(g)=N(g)N,

where N(g) is the count of amino acids in group g and N is the sequence length.

Protein sequence-based feature encoding

To represent the different AMPs as fixed-size numerical vectors, four different feature matrices were used to capture evolutionary, physicochemical, and nominal properties of the sequences, based on the work of Chen et al.15 The descriptors were integrated while preserving the sequence’s positional data, standardizing by adding zeros at the C-terminal. The summary of these features is presented in the Table S3.

AAindex is a comprehensive database that houses numerical indices representing diverse biochemical and physicochemical properties associated with amino acids.45 The database encompasses 531 physicochemical properties. Each of the 20 standard amino acids is assigned a specific numerical value corresponding to each property within the AAindex database.

We employed a sequence encoding method that retains information about amino acids. This approach is commonly used to extract features from peptide sequences. The BLOSUM62 feature matrix represents the probability of amino acid mutations occurring during the course of evolution.46 Each row in the BLOSUM62 matrix represents one of the 20 amino acids and is used for encoding purposes. Originally designed for protein sequence alignment, BLOSUM62 has found extensive application as a feature set in predicting peptide properties.

The Z-Scale descriptor, developed by Sandberg et al.,47 assigns five physicochemical descriptor variables to each amino acid. These variables, denoted as Z1 to Z5, capture various physicochemical properties. Z1 represents hydrophobicity and hydrophilicity, Z2 corresponds to steric bulk properties and polarizability, Z3 represents polarity, while Z4 and Z5 capture electronic effects.

One-hot encoding is a straightforward and direct approach employed to encode the individual residues of a peptide sequence. It involves creating new columns equal to the number of unique amino acids in the property column, and these new columns are filled with 0s and 1s. This encoding technique is widely employed in various traditional machine learning tasks. In our study, since we have 20 unique amino acids, each amino acid is represented by a binary vector consisting of 20 bits.

Language models for protein sequence encoding

In this study, we employed transfer learning, a method that uses pre-trained embeddings of sequences to encode peptides into vectors. This method exploits recent advances in natural language technology (NLP), which has been greatly enhanced by the development of contextualized language models (LMs) and advances in high-performance computing (HPC). This technology, typically used for language processing, is currently being effectively applied to the analysis of biological data, in particular to the understanding and prediction of AMPs.34

Protein language models (pLMs) consider protein sequences as if they were sentences in a language, where each amino acid represents a unique word. These models are built using the Transformer architecture,48 a framework designed primarily for language-related applications. The pLMs are pre-trained using large datasets, such as UniRef100 or UniRef50,49 which contain millions of protein sequences, providing access to a wealth of information for training purposes.

In this study, we used the results of ProtTrans,34 which provides sophisticated pre-trained models that are customized for protein sequences. The models were created using a variety of GPUs from the Summit supercomputer and several Google TPUs. They generate embedding vectors that provide context-rich representations of protein sequences using various Transformer models. Our selection among the many ProtTrans models was guided by the research of Dee et al.50 They evaluated various models for multiple tasks, including AMP identification and determined that the “T5-UniRef50” architecture was especially efficient. Therefore, we incorporated the “T5-UniRef50” model in our study to generate contextualized embeddings for peptide sequences. Peptide sequences were coded to no more than 40 amino acids. To achieve this fixed length, sequences shorter than 40 amino acids were padded with zeros. Using the “T5-UniRef50” architecture, we generated feature vectors with dimensions of 40 × 1024, which successfully ensured a uniform and comprehensive depiction of peptide sequences, critical to our analysis.

Feature extraction from genome sequences

In this study, encoding genome sequences constituted a pivotal step for integrating genomic features into our predictive models. We utilized genome sequences from GenBank51 for three critical bacterial strains relevant to our study: Escherichia coli str. K-12, Staphylococcus aureus subsp. NCTC 8325, and Pseudomonas aeruginosa PAO1. These strains were selected for their clinical relevance and well-documented genetic profiles, which are essential for comprehending the responses of bacteria to AMPs. We employed the MathFeature software41 to process this genomic data, a tool designed for extracting and quantifying genomic and proteomic features. A detailed summary of the encoded features and their specific attributes is provided in Table S4, ensuring transparency and reproducibility of our data processing methods.

The nucleic acid composition (NAC) was determined via the formulaF(i)=N(i)N,

where i represents a nucleic acid (A, C, G, T), and N(i) is its count. Di-nucleic acid composition (DAC) and tri-nucleic acid composition (TAC) were also computed to determine nucleotide proportions within genome sequences, critical for our study. The calculations followed a similar procedure, considering pairs and triplets of nucleotides, respectively. The NAC provides a basic profile of the genomic content by illustrating the proportion of each nucleotide within the genome. Extending this, the DAC and TAC offer deeper insights into the nucleotide pairing and triplet frequencies, which are critical for understanding the regulatory mechanisms and potential mutation sites within the bacterial genome. These genomic features are crucial for our models as they allow us to capture a comprehensive genetic snapshot that influences bacterial behavior and susceptibility to AMPs. By integrating these detailed genomic insights with traditional sequence data, our models can predict how different bacterial strains might react to specific AMPs more accurately. This enhanced predictive capability is essential for developing effective antimicrobial strategies tailored to the genetic profiles of target bacteria, thereby improving the efficacy and specificity of treatment interventions.

Development of prediction models

To achieve our research goals, we used a strategic combination of traditional machine learning and deep learning models in this study. These models, highly effective in regression tasks, accurately predicted MICs against pathogenic bacteria. Four ML models were developed: RF,29 XGBoost,31 CatBoost,52 and LGBM.30 In addition, four DL models were included in our study. BiLSTM,53 CNN,24 Transformer,48 and a Multi-Branch Model were all used in this study. All models were implemented in Python 3.9.16, with DL models developed using TensorFlow2.5.054 and Keras 2.5.0. The study was supported by the use of Pandas 1.3.4 and NumPy 1.19.5. Computations were performed using an RTX 1080 Ti.

We built a 400-tree RF regression model using the scikit-learn library.55 This model increases accuracy and reduces the possibility of overfitting by unifying decision tree classifiers through bagging and random feature sampling. Our XGBoost regression model was built using the xgboost package version 1.7.3.31 It consists of 400 trees with a learning rate of 0.01. The ensemble approach of XGBoost improves the ability of the model to effectively deal with structured data by using gradient boosting. We created a CatBoost regression model with 400 trees, a learning rate of 0.05, and a maximum depth of 5 using version 1.1.1 of the catboost package. CatBoost is particularly adept at processing categorical features and uses ordered boosting to improve performance. Our regression model using LGBM was developed with the lightgbm package version 3.3.5,30 consisting of 400 trees and a learning rate of 0.01. The efficiency and speed of LGBM can be attributed to its unique leaf-wise tree growth and its gradient-based one-sided sampling (GOSS) approach.

The foundation of our sequence modeling is the BiLSTM architecture, an extension of long short-term memory (LSTM).53 BiLSTM skillfully incorporates information from both forward and backward contexts, making it not only advantageous but essential to the success of our research study. The architecture of the LSTM is a precisely tailored form of a recurrent neural network (RNN) that is designed to maintain long-term dependencies in sequential data. The system functions through a series of memory cells and gates that govern the flow of information, constantly updating an internal concealed state with current and prior data inputs at every time step. The operation within each LSTM cell at each time step t involves multiple computations, as shown in Figure S1 and represented in the following equations:Forgetgate:ft=σ(Wf[ht−1,Xt]+bf)

Inputgate:it=σ(Wi[ht−1,Xt]+bi)

Cellcandidate:Ctˇ=tanh(WC[ht−1,Xt]+bC)

Outputgate:Ot=σ(Wf[ht−1,Xt]+bf)

Cellstateupdate:Ct=ft⊗Ct−1⊕it⊗Ctˇ

Hiddenstateupdate:ht=Ot⊗tanh(Ct)

Here, Xt signifies the embedding vector of the amino acid at time time step t, with ft, it, Ot, Ct, and ht representing the various gates and states within the cell. The weights and biases associated with it, ot, ft, and Ct are denoted by (Wi, bi), (Wo, bo), (Wf, bf), and (Wc, bc), respectively. The element-wise multiplication and addition operations are represented by ⊗ and ⊕, respectively. In the BiLSTM framework, one LSTM unit processes the sequence in its original order, while another LSTM unit processes it in reverse. This dual approach allows the model to capture information from both past and future contexts, thereby enhancing its ability to understand and interpret sequential data in a comprehensive manner. After passing through the BiLSTM architecture, we merged the genomic features and applied batch normalization to improve our training. We then fed this matrix into a fully connected regressor to predict the MIC value. The Figure S2 below shows our model architecture.

Our study includes a CNN, a powerful architecture that is commonly used in image classification and that is also effective in our domain. The structure of the CNN consists of convolutional layers, pooling layers, and fully connected layers. The convolutional layers use filters to analyze the input data, extracting features through element-wise operations. This is followed by pooling layers, specifically max pooling, which reduces computational complexity by downsizing feature maps and decreasing spatial dimensions. The fully connected layers then interpret these features to generate final predictions. The CNN learning process is further refined through the integration of genomic features and batch normalization, as shown in Figure S3.

The Transformer model,48 an innovative architecture in the field of natural language processing (NLP), is also used in our research. The Transformer model is known for its self-attention mechanism. This mechanism has proven to be effective in capturing long-range dependencies in sequences. To evaluate the importance of different parts of the input sequence with respect to each other, the model’s self-attention mechanism computes attention scores. The computations are based on key, query, and value elements obtained from input embeddings. The process is described as follows:Keyprojection:Key(k)=XWk

Queryprojection:Query(q)=XWq

Valueprojection:Value(v)=XWv

AttentionScores:AttentionScores(q,k,v)=softmax(qkTdk)v

Here, X dentes input embeddings, and Wk, Wq, and Wv are the learnable weight matrices for key, query and value projections, respectively. qkT represents the dot product between the query q and the transpose of the keys k. The division by dk is an optional scaling factor, where dk represents the dimensionality of the key and query representations. The scaling helps to prevent extremely small or large values in the attention scores. The multi-head attention mechanism of the Transformer enhances its ability to process information at various levels of detail. The architecture integrates genomic features through an encoder and decoder, applying batch normalization before proceeding to fully connected layers for final prediction, as shown in Figure S4.

A crucial aspect of this study is the multi-branch model, which combines different architectures of neural networks in order to handle different perspectives of the data. To achieve this, we merged the CNN and BiLSTM architectures, taking advantage of their individual capabilities. Each branch has a different interpretation of the data, and we used a cross-view attention mechanism56 to evaluate the importance of information across branches. This integration of data results in a comprehensive illustration shown in Figure S5.

Ensemble learning approach

To further improve predictive performance, we implemented an ensemble learning approach that integrates multiple deep learning models. This ensemble model utilizes three selected architectures, BiLSTM, CNN, and MBM, to leverage their combined strengths. These models were selected based on their individual performance metrics, particularly their PCC against the validation set, which indicated their ability to accurately predict MIC values of AMPs against the target bacterial strains.

Our ensemble method relied on weighted averaging to combine predictions from selected models, where we assigned weights based on how each model performs on a separate validation dataset. This dataset is distinct from those used for training and testing and is critical for performance evaluation. The weights are assigned based on the PCC values achieved by each model, which reflect their accuracy in predicting MIC values. Models with higher PCC values, indicating superior predictive accuracy, are given greater weight in the ensemble calculation. This approach ensures that the most reliable predictions have the most significant influence on the ensemble outcome. Specifically, the weights were normalized to ensure they summed up to one, maintaining a balanced contribution from each model. The final MIC prediction for each AMP was calculated as a weighted average of the predictions from the BiLSTM, CNN, and MBM models.

Using an ensemble learning approach, our model takes advantage of the distinct predictive capabilities of BiLSTM, CNN, and MBM while mitigating potential weaknesses that may arise from reliance on a single predictive model. This collaborative synergy between the different models will enhance the robustness and accuracy of our predictions, particularly in the complex domain of the analysis of the activity of AMPs.

Evaluation metrics

In our study, the effectiveness and accuracy of the regression models were rigorously evaluated using four well-established statistical metrics: MSE, RMSE, R2, and PCC. These metrics are useful in understanding the performance of the models in terms of prediction of outcomes relative to actual values.

MSE is a fundamental measure used to assess the average magnitude of errors in a set of predictions. It calculates the average squared difference between the predicted values (yiˆ) and the actual values (yi). The formula for MSE is expressed as:MSE=1N∑i=1N(yiˆ−yi)2,

where N represents the number of samples. The MSE is sensitive to outliers because it squares the errors before averaging. Therefore, larger deviations are penalized more severely.

RMSE measures the standard deviation of the prediction errors. It is obtained by taking the square root of the MSE, resulting in an error metric in the same unit as the output variable. The RMSE formula is as follows:RMSE=1N∑i=1N(yiˆ−yi)2.

R2, the coefficient of determination, is a statistical measure that indicates the percentage of variability in the dependent variable that can be explained by the independent variables present in a regression model. It is computed as follows:R2=1−∑i=1N(yiˆ−yi)2∑i=1N(μy−yi)2

where μy is the mean of the actual values. A higher R2 value is an indication of a better fit of the model to the observed data.

The magnitude and direction of the linear relationship between two variables is quantified by the PCC. It ranges from −1 to +1, with +1 indicating a perfect positive linear association, −1 indicating a perfect negative linear association, and 0 indicating no linear association. The PCC formula is:PCC=∑i=1N(yi−μy)(yiˆ−μyˆ)∑i=1N(yi−μy)2∑i=1N(yiˆ−μyˆ)2

where μyˆ is the mean of the predicted value.

In this study, MSE was used as the primary loss function during model training, providing a reliable measure of model accuracy, and PCC was used as the primary performance metric, providing information on the linear correlation between predicted and actual values. These metrics jointly provided a comprehensive evaluation of our models, ensuring their effectiveness and applicability in predictive tasks.

Quantification and statistical analysis

All computations were performed in the Python programming language. The graphic abstract and Figure 1 were generated by Microsoft PowerPoint, other plots appearing in this study were generated by the Python package.

Supplemental information

Document S1. Figures S1–S5 and Tables S1–S4

Acknowledgments

This work was supported by the 10.13039/100020595 National Science and Technology Council (NSTC112-2321-B-A49-016 , 112-2740-B-400-005 , 112-2221-E-008-045 , 113-2321-B-A49-025 , 113-2634-F-039-001 , 113-2221-E-A49-160-MY3 , 113-2221-E-008 -005 , and 113-2222-E008-001-MY2 ) and The National Health Research Institutes (NHRI-EX113-11320BI ), Taiwan. This work was also financially supported by the Center for Intelligent Drug Systems and Smart Bio-devices (IDS2B) from The Featured Areas Research Center Program within the framework of the Higher Education Sprout Project and Yushan Young Fellow Program (112C1N084C ) by the 10.13039/100010002 Ministry of Education in Taiwan.

Author contributions

CYC carried out the data collection and curation. CRC and CYC participated in the data analyses, model construction, and drafted the article. CRC, CYC, YT, LCW, JBKH, and TYL participated in the design of the study and performed the draft revision. JJL, CB and JTH conceived of the study, and participated in its design and coordination and helped to revise the article. All authors read and approved the final article.

Declaration of interests

The authors have no affiliations with or involvement in any organization or entity with any financial interest, or non-financial interest in the subject matter or materials discussed in this article.

Supplemental information can be found online at https://doi.org/10.1016/j.isci.2024.110718.
==== Refs
References

1 Ventola C.L. The antibiotic resistance crisis: part 1: causes and threats P T. 40 2015 277 283 25859123
2 Kritsotakis E.I. Lagoutari D. Michailellis E. Georgakakis I. Gikas A. Burden of multidrug and extensively drug-resistant ESKAPEE pathogens in a secondary hospital care setting in Greece Epidemiol. Infect. 150 2022 e170 10.1017/S0950268822001492
3 Luong H.X. Thanh T.T. Tran T.H. Antimicrobial peptides - Advances in development of therapeutic applications Life Sci. 260 2020 118407 10.1016/j.lfs.2020.118407
4 Lei J. Sun L. Huang S. Zhu C. Li P. He J. Mackey V. Coy D.H. He Q. The antimicrobial peptides and their potential clinical applications Am. J. Transl. Res. 11 2019 3919 3931 31396309
5 Huang K.Y. Chang T.H. Jhong J.H. Chi Y.H. Li W.C. Chan C.L. Robert Lai K. Lee T.Y. Identification of natural antimicrobial peptides from bacteria through metagenomic and metatranscriptomic analysis of high-throughput transcriptome data of Taiwanese oolong teas BMC Syst. Biol. 11 2017 131 10.1186/s12918-017-0503-4
6 Pasupuleti M. Schmidtchen A. Malmsten M. Antimicrobial peptides: key components of the innate immune system Crit. Rev. Biotechnol. 32 2012 143 171 10.3109/07388551.2011.594423 22074402
7 Jenssen H. Hamill P. Hancock R.E.W. Peptide antimicrobial agents Clin. Microbiol. Rev. 19 2006 491 511 10.1128/CMR.00056-05 16847082
8 Toke O. Antimicrobial peptides: New candidates in the fight against bacterial infections Biopolymers 80 2005 717 735 10.1002/bip.20286 15880793
9 Mahlapuu M. Håkansson J. Ringstad L. Bjorn C. Antimicrobial Peptides : An Emerging Category of Therapeutic Agents Front Cell Infect. Mi. 6 2016 235805 10.3389/fcimb.2016.00194
10 Andrews J.M. Determination of minimum inhibitory concentrations J Antimicrob Chemoth 48 2002 1049 10.1093/jac/dkf083
11 Yasir M. Karim A.M. Malik S.K. Bajaffer A.A. Azhar E.I. Prediction of antimicrobial minimal inhibitory concentrations for Neisseria gonorrhoeae using machine learning models Saudi J. Biol. Sci. 29 2022 3687 3693 10.1016/j.sjbs.2022.02.047 35844400
12 Xiao X. You Z.-B. Predicting Minimum Inhibitory Concentration of Antimicrobial Peptides by the Pseudo-amino Acid Composition and Gaussian Kernel Regression 2015 IEEE 301 305
13 Dean S.N. Alvarez J.A.E. Zabetakis D. Walper S.A. Malanoski A.P. PepVAE: Variational Autoencoder Framework for Antimicrobial Peptide Generation and Activity Prediction Front. Microbiol. 12 2021 725727 10.3389/fmicb.2021.725727
14 Yan J. Zhang B. Zhou M. Campbell-Valois F.-X. Siu S.W. A Deep Learning Method for Predicting the Minimum Inhibitory Concentration of Antimicrobial Peptides against Escherichia coli Using Multi-Branch-CNN and Attention 2023
15 Chen J. Cheong H.H. Siu S.W.I. xDeep-AcPEP: Deep Learning Method for Anticancer Peptide Activity Prediction Based on Convolutional Neural Network and Multitask Learning J. Chem. Inf. Model. 61 2021 3789 3803 10.1021/acs.jcim.1c00181 34327990
16 Vishnepolsky B. Grigolava M. Managadze G. Gabrielian A. Rosenthal A. Hurt D.E. Tartakovsky M. Pirtskhalava M. Comparative analysis of machine learning algorithms on the microbial strain-specific AMP prediction Brief. Bioinform. 23 2022 bbac233
17 Gull S. Minhas F. AMP(0): Species-Specific Prediction of Anti-microbial Peptides Using Zero and Few Shot Learning IEEE/ACM Trans. Comput. Biol. Bioinform. 19 2022 275 283 10.1109/TCBB.2020.2999399 32750857
18 Sharma R. Shrivastava S. Singh S.K. Kumar A. Singh A.K. Saxena S. Artificial intelligence-based model for predicting the minimum inhibitory concentration of antibacterial peptides against ESKAPEE pathogens IEEE J. Biomed. Health Inform. 28 2023 1949 1958 10.1109/JBHI.2023.3271611
19 Chung C.-R. Kuo T.-R. Wu L.-C. Lee T.-Y. Horng J.-T. Characterization and identification of antimicrobial peptides with different functional activities Brief. Bioinform. 21 2019 bbz043 10.1093/bib/bbz043
20 Chung C.-R. Liou J.-T. Wu L.-C. Horng J.-T. Lee T.-Y. Multi-label classification and features investigation of antimicrobial peptides with various functional classes iScience 26 2023 108250 10.1016/j.isci.2023.108250
21 Wang X.X. Chen S. Brown D.J. An approach for constructing parsimonious generalized Gaussian kernel regression models Neurocomputing 62 2004 441 457 10.1016/j.neucom.2004.06.003
22 Chou K.C. Prediction of signal peptides using scaled window Peptides 22 2001 1973 1979 10.1016/s0196-9781(01)00540-x 11786179
23 Wang Z. Wang G. APD: the Antimicrobial Peptide Database Nucleic Acids Res. 32 2004 D590 D592 10.1093/nar/gkh025 14681488
24 O'Shea K. Nash R. An Introduction to Convolutional Neural Networks Preprint at arXiv 2015 10.48550/arXiv.1511.08458
25 Zou H. Hastie T. Regularization and variable selection via the elastic net J. R. Stat. Soc. B 67 2005 768 10.1111/j.1467-9868.2005.00527.x
26 Natekin A. Knoll A. Gradient boosting machines, a tutorial Front Neurorobotics 7 2013 21 10.3389/fnbot.2013.00021
27 Vovk V. Kernel ridge regression. Empirical Inference: Festschrift in Honor of Vladimir N. Vapnik 2013 105 116
28 Ranstam J. Cook J.A. LASSO regression J. British Surg. 105 2018 1348
29 Breiman L. Random forests Mach. Learn. 45 2001 5 32
30 Fan J. Ma X. Wu L. Zhang F. Yu X. Zeng W. Light Gradient Boosting Machine: An efficient soft computing model for estimating daily reference evapotranspiration with local and external meteorological data Agric. Water Manag. 225 2019 105758
31 Chen T. He T. Benesty M. Khotilovich V. Tang Y. Cho H. Chen K. Mitchell R. Cano I. Zhou T. Xgboost: Extreme Gradient Boosting 2015 1 4 R package version 0.4-2 1
32 Pirtskhalava M. Amstrong A.A. Grigolava M. Chubinidze M. Alimbarashvili E. Vishnepolsky B. Gabrielian A. Rosenthal A. Hurt D.E. Tartakovsky M. DBAASP v3: database of antimicrobial/cytotoxic activity and structure of peptides as a resource for development of new therapeutics Nucleic Acids Res. 49 2021 D288 D297 33151284
33 Zervou M.A. Doutsi E. Pantazis Y. Tsakalides P. De Novo Antimicrobial Peptide Design with Feedback Generative Adversarial Networks Int. J. Mol. Sci. 25 2024 5506 38791544
34 Elnaggar A. Heinzinger M. Dallago C. Rehawi G. Wang Y. Jones L. Gibbs T. Feher T. Angerer C. Steinegger M. Prottrans: Toward understanding the language of life through self-supervised learning IEEE Trans. Pattern Anal. Mach. Intell. 44 2022 7112 7127 34232869
35 Wang R. Wang T. Zhuo L. Wei J. Fu X. Zou Q. Yao X. Diff-AMP: tailored designed antimicrobial peptide framework with all-in-one generation, identification, prediction and optimization Brief. Bioinform. 25 2024 bbae078 10.1093/bib/bbae078
36 Lin T.-T. Yang L.-Y. Lin C.-Y. Wang C.-T. Lai C.-W. Ko C.-F. Shih Y.-H. Chen S.-H. Intelligent De Novo Design of Novel Antimicrobial Peptides against Antibiotic-Resistant Bacteria Strains Int. J. Mol. Sci. 24 2023 6788 37047760
37 Jhong J.H. Yao L. Pang Y. Li Z. Chung C.R. Wang R. Li S. Li W. Luo M. Ma R. dbAMP 2.0: updated resource for antimicrobial peptides with an enhanced scanning method for genomic and proteomic data Nucleic Acids Res. 50 2022 D460 D470 10.1093/nar/gkab1080 34850155
38 Shi G. Kang X. Dong F. Liu Y. Zhu N. Hu Y. Xu H. Lao X. Zheng H. DRAMP 3.0: an enhanced comprehensive data repository of antimicrobial peptides Nucleic Acids Res. 50 2022 D488 D496 10.1093/nar/gkab651 34390348
39 Abdi H. Williams L.J. Principal component analysis WIREs Comput. Stats. 2 2010 433 459
40 Chen Z. Zhao P. Li F. Leier A. Marquez-Lago T.T. Wang Y. Webb G.I. Smith A.I. Daly R.J. Chou K.C. Song J. iFeature: a Python package and web server for features extraction and selection from protein and peptide sequences Bioinformatics 34 2018 2499 2502 10.1093/bioinformatics/bty140 29528364
41 Bonidia R.P. Domingues D.S. Sanches D.S. de Carvalho A.C.P.L.F. MathFeature: feature extraction package for DNA, RNA and protein sequences based on mathematical descriptors Brief. Bioinform. 23 2022 bbab434 10.1093/bib/bbab434
42 Pedregosa F. Varoquaux G. Gramfort A. Michel V. Thirion B. Grisel O. Blondel M. Prettenhofer P. Weiss R. Dubourg V. Scikit-learn: Machine Learning in Python J. Mach. Learn. Res. 12 2011 2825 2830
43 Abadi M. Barham P. Chen J. Chen Z. Davis A. Dean J. Devin M. Ghemawat S. Irving G. Isard M. {TensorFlow}: A System for {Large-Scale} Machine Learning 2016 265 283
44 Chen Z. Liu X. Zhao P. Li C. Wang Y. Li F. Akutsu T. Bain C. Gasser R.B. Li J. iFeatureOmega: an integrative platform for engineering, visualization and analysis of features from molecular sequences, structural and ligand data sets Nucleic Acids Res. 50 2022 W434 W447 10.1093/nar/gkac351 35524557
45 Kawashima S. Kanehisa M. AAindex: amino acid index database Nucleic Acids Res. 28 2000 374 10592278
46 Lee T.Y. Chen S.A. Hung H.Y. Ou Y.Y. Incorporating distant sequence features and radial basis function networks to identify ubiquitin conjugation sites PLoS One 6 2011 e17331 10.1371/journal.pone.0017331
47 Sandberg M. Eriksson L. Jonsson J. Sjöström M. Wold S. New chemical descriptors relevant for the design of biologically active peptides. A multivariate characterization of 87 amino acids J. Med. Chem. 41 1998 2481 2491 9651153
48 Vaswani A. Shazeer N. Parmar N. Uszkoreit J. Jones L. Gomez A.N. Kaiser Ł. Polosukhin I. Attention is all you need Adv. Neural Inf. Process. Syst. 30 2017
49 Suzek B.E. Huang H. McGarvey P. Mazumder R. Wu C.H. UniRef: comprehensive and non-redundant UniProt reference clusters Bioinformatics 23 2007 1282 1288 10.1093/bioinformatics/btm098 17379688
50 Dee W. LMPred: predicting antimicrobial peptides using pre-trained language models and deep learning Bioinform. Adv. 2 2022 vbac021 10.1093/bioadv/vbac021
51 O'Leary N.A. Wright M.W. Brister J.R. Ciufo S. Haddad D. McVeigh R. Rajput B. Robbertse B. Smith-White B. Ako-Adjei D. Reference sequence (RefSeq) database at NCBI: current status, taxonomic expansion, and functional annotation Nucleic Acids Res. 44 2016 D733 D745 26553804
52 Dorogush A.V. Ershov V. Gulin A. CatBoost: Gradient Boosting with Categorical Features Support Preprint at arXiv 2018 10.48550/arXiv.1810.11363
53 Hochreiter S. Schmidhuber J. Long short-term memory Neural Comput. 9 1997 1735 1780 9377276
54 Abadi M. Barham P. Chen J. Chen Z. Davis A. Dean J. Devin M. Ghemawat S. Irving G. Isard M. Tensorflow: A System for Large-Scale Machine Learning 2016 265 283
55 Pedregosa F. Varoquaux G. Gramfort A. Michel V. Thirion B. Grisel O. Blondel M. Prettenhofer P. Weiss R. Dubourg V. Scikit-learn: Machine learning in Python J. Mach. Learn. Res. 12 2011 2825 2830
56 Zhao W. Luo S. Wu H. Jiang X. He T. Hu X. A multi-label learning framework for predicting antibiotic resistance genes via dual-view modeling Brief. Bioinform. 23 2022 bbac052
