
==== Front
RSC Adv
RSC Adv
RA
RSCACL
RSC Advances
2046-2069
The Royal Society of Chemistry

d4ra04502e
10.1039/d4ra04502e
Chemistry
Identification of lead inhibitors for 3CLpro of SARS-CoV-2 target using machine learning based virtual screening, ADMET analysis, molecular docking and molecular dynamics simulations†
† Electronic supplementary information (ESI) available. See DOI: https://doi.org/10.1039/d4ra04502e

Chhetri Sandeep Poudel a
Bhandari Vishal Singh b
Maharjan Rajesh a
https://orcid.org/0000-0002-3422-0808
Lamichhane Tika Ram a
a Central Department of Physics, Tribhuvan University Kathmandu 44600 Nepal tika.lamichhane@cdp.tu.edu.np

b Central Department of Chemistry, Tribhuvan University Kathmandu 44600 Nepal
18 9 2024
12 9 2024
18 9 2024
14 40 2968329692
20 6 2024
4 9 2024
This journal is © The Royal Society of Chemistry
2024
The Royal Society of Chemistry
https://creativecommons.org/licenses/by-nc/3.0/ This article is licensed under a Creative Commons Attribution-Non Commercial 3.0 Unported Licence. You can use material from this article in other publications without requesting further permissions from the RSC, provided that the correct acknowledgement is given and it is not used for commercial purposes.
The SARS-CoV-2 3CLpro is a critical target for COVID-19 therapeutics due to its role in viral replication. We employed a screening pipeline to identify novel inhibitors by combining machine learning classification with similarity checks of approved medications. A voting classifier, integrating three machine learning classifiers, was used to filter a large database (∼10 million compounds) for potential inhibitors. This ensemble-based machine learning technique enhances overall performance and robustness compared to individual classifiers. From the screening, three compounds M1, M2 and M3 were selected for further analysis. Absorption, distribution, metabolism, excretion, and toxicity (ADMET) analysis compared these candidates to nirmatrelvir and azvudine. Molecular docking followed by 200 ns MD simulations showed that only M1 (6-[2,4-bis(dimethylamino)-6,8-dihydro-5H-pyrido[3,4-d]pyrimidine-7-carbonyl]-1H-pyrimidine-2,4-dione) remained stable. For azvudine and M1, the estimated median lethal doses are 1000 and 550 mg kg−1, respectively, with maximum tolerated doses of 0.289 and 0.614 log mg per kg per day. The predicted inhibitory activity of M1 is 7.35, similar to that of nirmatrelvir. The binding free energy based on Molecular Mechanics Poisson-Boltzmann Surface Area (MM-PBSA) of M1 is −18.86 ± 4.38 kcal mol−1, indicating strong binding interactions. These findings suggest that M1 merits further investigation as a potential SARS-CoV-2 treatment.

Identification of novel drug candidate with appropriate pharmacokinetic properties and drug-likeness for SARS-CoV-2.

pubstatusPaginated Article
==== Body
pmc1 Introduction

The SARS coronavirus (SARS-CoV) causes severe acute respiratory syndrome (SARS).1 As of January 21, 2024, there were 774 395 593 confirmed cases of SARS-CoV-2 infection, resulting in 7 023 271 deaths.2 SARS-CoV-2, an enveloped positive-sense single-stranded RNA virus,3 belongs to the genus Betacoronavirus. Viral proteases, crucial for replication, are well-validated targets for treating hepatitis C and HIV.4 The primary protease, 3CLpro (also known as Mpro or Nsp5),5 cleaves polyproteins at 11 sites, essential for viral protein maturation.6 Inhibiting 3CLpro halts viral replication by preventing the production of necessary enzymes like RNA-dependent RNA polymerase.7 Human proteases lack 3CLpro's cleavage specificity, making these inhibitors safe for human use.8 Known oral 3CLpro inhibitors9 are shown in ESI Fig. S1.†

The COVID-19 pandemic has increased the demand for new antiviral drugs. Traditional high-throughput screening (HTS) of 1 to 2 million compounds is expensive and operationally challenging.10,11 Artificial Intelligence (AI) can accelerate drug discovery by evaluating vast data, predicting drug efficacy, and reducing the time and resources needed for clinical trials, enhancing the chances of developing effective treatments. Drug discovery has been revolutionized over the last ten years by AI models.12–14

As discussed in ref. 15, we used machine learning combined with similarity analysis, ADMET analysis, molecular docking and MD simulation in our study. We used a voting classifier to screen a large database (∼10 million compounds) for potential inhibitors. Selected compounds were compared to known 3CLpro inhibitors and analyzed for ADMET properties. Stability was assessed using molecular docking and molecular dynamics simulations.16Fig. 1 illustrates our study's workflow.

Fig. 1 Schematic workflow of the study.

2 Materials and methods

2.1 Data collection and curation

The OpenCADD platform, an open-source tool for cheminformatics, was employed to obtain compound data and develop machine learning models.17 Simplified Molecular Input Line Entry System (SMILES) for 903 inhibitors of 3CLpro, along with their respective Half-maximal inhibitory concentration (IC50) values, were retrieved from the Chemical European Molecular Biology Laboratory (CHEMBL) database.18 After downloading the data, we filtered out SMILES entries lacking IC50 values, retained only bioactivity entries measured in nanomolar (nM), and removed duplicate molecules, resulting in 744 data points. Due to the varied scales of IC50 values, they were converted into corresponding negative logarithms, known as pIC50 values. Pfizer's rule, also known as Lipinski's Rule of Five (RO5), was utilized at this stage to filter the data according to drug-likeness.19,20 Meeting most of the Ro5 parameters does not ensure that a compound will become a drug; it merely indicates drug-likeness and assists in eliminating weaker compounds during the preclinical phase. Our models were trained using the 659 data points that remained after the RO5 filter was applied. The spider plots of the compounds in the dataset that are either inside or outside RO5 domain are displayed in Fig. 2.

Fig. 2 Physio-chemical radar plots of the compounds in the dataset (a) inside RO5 domain or (b) outside RO5 domain.

2.2 Model building and evaluation

Molecular fingerprints21 encode structural data into numerical vectors or fixed-length bit-strings, which enable fast similarity comparisons crucial for virtual screening,22 structure–activity relationship studies, and chemical space maps creation.23 In our work, molecular fingerprints derived from SMILES were computed using RDKit24 and used as inputs for machine learning models. The dataset was split into 332 active and 327 inactive compounds based on a pIC50 cut-off value of 6.2. We built twenty machine learning classifiers using Morgan3 fingerprints for quantitative structure–activity relationship (QSAR) classification,25 selecting the top three classifiers based on various learning methods and evaluation metrics. The classifiers were built using Scikit-learn and Lightgbm.26,27 The hyperparameters of the top three classifiers were fine-tuned and combined to form a voting classifier, enhancing overall performance and robustness compared to individual classifiers.28 Similar approach was used for QSAR regression.

To assess classifiers, metrics including accuracy, precision, sensitivity, specificity, and AUC (Area Under Curve) were computed based on the confusion matrix.29 Regressors were evaluated based on mean absolute error (MAE), root-mean-squared error (RMSE), and R2 – score.30

2.3 Ligand and similarity based virtual screening

We employed the voting classifier for ligand-based virtual screening31,32 of the eMolecules databases33 to screen for active compounds, selecting molecules with a predicted probability exceeding 90% as potential active inhibitors. The database was filtered before screening to remove entries with invalid SMILES, Pan Assay Interference Molecules (PAINS), and those not meeting Lipinski's Rule of Five (RO5) criteria, using RDKit.

Similarity-based virtual screening measures the similarity between database structures and reference structures, based on the principle that similar structures likely have similar bioactivities.34–36 For 3CLpro inhibitors, the chemical similarity between potential active inhibitors and known inhibitors was calculated using Molecular ACCess System (MACCS) and Morgan2 fingerprints, with Tanimoto and Dice similarity indices ensuring consistent comparisons.37 Three potential compounds, consistently ranking in the top five during analysis, were selected for further assessment. Similarity maps of these candidates, created using Morgan2 fingerprints, visualized their similarity to known inhibitors.38

2.4 ADMET analysis of potential inhibitors

Assessing the absorption, distribution, metabolism, excretion, and toxicity (ADMET) properties is a crucial yet complex part of the drug discovery process, as these factors contribute to a significant portion of clinical failures.39,40 In this study, we conducted a preliminary ADMET analysis using the SwissADME platform41 for ADMET profiling and the ProTox-II tool42 for toxicity predictions of candidate compounds relative to known inhibitors. Additionally, the maximum tolerated dose (MTD) for humans was estimated using the pkCSM tool.43 Although these tools offer useful preliminary insights, their results are speculative and should be interpreted carefully.

2.5 Molecular docking

Molecular docking was used to determine how drugs attach to and interact with a protein. The crystal structure of 3CLpro (PDB ID: 5R82) complexed with an Z219104216 inhibitor was retrieved from the Protein Data Bank (PDB)44,45 and refined using I-TASSER.46 AutoDock4 was used to perform molecular docking.47 To validate the docking parameters, the native ligand was redocked in the same binding pocket and the root mean square deviation (RMSD) between the initial pose and the docked pose was calculated using PyMOL.48 Proteins and ligands were prepared using AutoDockTools by removing water molecules, adding Kollmann's charges, integrating polar hydrogens, and converting to protein data bank with partial charge and atom type (PDBQT) format. A cubic grid box (50 Å sides) centered at coordinates 10.364, 1.549, and 20.182 was used for site-specific docking. Docking parameters included a grid spacing of 0.375 Å, a population size of 300; 2 500 000 energy evaluations; and 100 docking runs using the Lamarckian Genetic Algorithm.49–51 Protein–ligand interactions for the best-scored poses were analyzed with Protein–Ligand Interaction Profiler (PLIP).52

2.6 Molecular dynamics simulation

Molecular dynamics (MD) simulations were conducted for three candidate compounds to refine binding affinities, stability, and interactions. Using GROMACS53 with the CHARMM36 forcefield,54 the highest-scoring protein–ligand complex from docking was simulated. Ligands were parameterized via SwissParam,55 and the system was neutralized with 0.15 mol L−1 concentration of Cl− and Na+ ions,56 solvated in a dodecahedron box of SPC water.57 Energy minimization used 50 000 steps of the steepest descent method, followed by equilibration for 100 ps at 300 K in an NVT ensemble with a V-rescale thermostat.58 Further equilibration for 100 ps at 1 bar and 300 K used an NPT ensemble with isotropic Berendsen pressure coupling. An unrestrained 200 ns MD simulation was then run with a 2 fs timestep, using a Parrinello-Rahman barostat and V-rescale thermostat.59

Stability was assessed by analyzing the root mean square deviation (RMSD), root mean square fluctuation (RMSF), protein solvent accessible surface area (SASA), radius of gyration (Rg), number of hydrogen bonds (H-bonds), and Dictionary of Secondary Structure in Proteins (DSSP).60 Ligand–protein binding free energies were calculated using gmx_MMPBSA and gmx_MMPBSA_ana, following the MM-PBSA approach, over the final 20 ns of equilibrated trajectories.61

3 Results and discussions

3.1 Model building and database screening

The performance of 20 classifiers (CLF) is summarized in ESI Table S1.† We selected Nu-Support Vector Classifier (NuSVC), ExtraTreesClassifier (ET), and Light Gradient Boosting Machine (LGBM) Classifier to construct a Voting Classifier (VC). For NuSVC, parameters were set to nu = ‘0.2’, kernel = ‘rbf’, and gamma = ‘scale’; for ET, n_estimators = ‘1000’, criterion = ‘gini’, and max_features = ‘sqrt’; for LGBM, n_estimators = ‘200’, learning_rate = ‘0.2’, max_depth = ‘4’, and num_leaves = ‘50’; all other parameters were left at their default values. The VC employed a ‘soft’ voting mechanism. The confusion matrices of the individual classifiers and the VC are presented in ESI Fig. S2,† with their evaluations detailed in Table 1.

Evaluation of three individual classifiers and voting classifier

Classifier	Accuracy	Precision	Sensitivity	Specificity	AUC	
NuSVC	0.89	0.94	0.86	0.93	0.96	
ET	0.90	0.96	0.86	0.95	0.95	
LGBM	0.88	0.91	0.86	0.90	0.95	
VC	0.88	0.91	0.86	0.90	0.96	

Table 2 presents the five-fold cross-validation results for individual classifiers and the voting classifier using a 20% random data selection.

Five-fold cross validation of individual classifiers and voting classifier

Classifier	Accuracy	Precision	Sensitivity	Specificity	AUC	
NuSVC	0.87 (±0.04)	0.87 (±0.06)	0.88 (±0.03)	0.87 (±0.06)	0.94 (±0.01)	
ET	0.88 (±0.04)	0.89 (±0.07)	0.87 (±0.04)	0.89 (±0.06)	0.94 (±0.03)	
LGBM	0.85 (±0.04)	0.84 (±0.06)	0.87 (±0.04)	0.83 (±0.06)	0.92 (±0.03)	
VC	0.87 (±0.04)	0.86 (±0.06)	0.87 (±0.05)	0.86 (±0.05)	0.94 (±0.03)	

Fig. 3 illustrates the ROC curves for these classifiers. With AUC scores of 0.96, 0.95, 0.95, and 0.96, all classifiers demonstrated strong classification performance. The voting classifier was chosen for screening the eMolecules database due to its superior robustness, identifying 39 molecules with prediction probabilities above 90% as potential active inhibitors.

Fig. 3 ROC curve of the individual classifiers and voting classifier.

3.2 Similarity measures analysis

We assessed the chemical similarity between 39 potential active inhibitors and known inhibitors of 3CLpro-azvudine, ensitrelvir, nirmatrelvir, and simnotrelvir using MACCS and Morgan2 fingerprints. Tanimoto and Dice similarity metrics were computed for both MACCS and Morgan2 fingerprints. Table 3 displays the top five compounds with the highest similarity to each reference, considering Tanimoto and Dice similarities for both MACCS and Morgan fingerprints.

Similarity checking between azvudine, ensitrelvir, nirmatrelvir and simnotrelvir with top-ranked molecules using Morgan2 and MACCS fingerprints

Known inhibitors	Top 5 molecules	Tanimoto_morgan	Dice_morgan	Top 5 molecules	Tanimoto_maccs	Dice_maccs	
Azvudine	5	0.147287	0.256757	37	0.578313	0.732824	
26	0.144330	0.252252	34	0.556818	0.715328	
15	0.136364	0.240000	35	0.547619	0.707692	
19	0.135922	0.239316	31	0.534884	0.696970	
14	0.134615	0.237288	38	0.530120	0.692913	
Ensitrelvir	20	0.186667	0.314607	30	0.653846	0.790698	
16	0.174497	0.297143	37	0.636364	0.777778	
11	0.169935	0.290503	35	0.623377	0.768000	
27	0.168919	0.289017	13	0.623377	0.768000	
8	0.166667	0.285714	24	0.618421	0.764228	
Simnotrelvir	34	0.154930	0.268293	4	0.535211	0.697248	
14	0.133803	0.236025	17	0.532468	0.694915	
5	0.130178	0.230366	14	0.531646	0.694215	
17	0.120805	0.215569	34	0.524390	0.688000	
38	0.118881	0.212500	3	0.520548	0.684685	
Nirmatrelvir	5	0.148148	0.258065	34	0.587500	0.740157	
34	0.135714	0.238994	30	0.550000	0.709677	
38	0.123188	0.219355	31	0.544304	0.704918	
31	0.117241	0.209877	33	0.525641	0.689076	
30	0.116438	0.208589	38	0.519481	0.683761	

For further analysis, we selected three structures-M1 (PubChem CID 56879830), M2 (PubChem CID 70722105), and M3 (PubChem CID 72893585)-based on their higher frequency of occurrence and higher similarity indexes among the similar compounds.

Fig. 4 illustrates the similarity map generated for these three compounds with four known inhibitors using the Morgan2 fingerprint.

Fig. 4 Similarity maps between azvudine, ensitrelvir, nirmatrelvir, and simnotrelvir as references and candidates M1, M2, and M3 using Morgan2 fingerprint. Coloring method: green: positive difference, gray: no change in similarity, and pink: negative difference.

3.3 ADMET analysis and drug-likeness

ESI Table S2† presents the physicochemical properties, pharmacokinetics, and drug-likeness of the molecules. ADMET analysis with reference to oral 3CLpro inhibitors azvudine and nirmatrelvir in phase 4 trials,62–64 shows all candidate compounds meet Lipinski's rule of five, suggesting favorable drug-likeness with good absorption and permeability.65 Solubility analysis indicates that M2 and nirmatrelvir are soluble, while azvudine, M1, and M3 are highly soluble.

ESI Fig. S3† presents the bioavailability radar diagram comparing candidate compounds with reference molecules across various physicochemical properties: lipophilicity, size, polarity, solubility, flexibility, and saturation. The pink region indicates the ideal drug-likeness zone, while the red hexagon represents drug-likeness profile of molecules. A bioavailability score of 0.55 suggests favorable pharmacokinetic characteristics. The log Kp values of the candidate compounds suggest good skin permeability, falling within the range of −9.7 to −3.5.39

The Brain Or IntestinaL EstimateD permeation method (BOILED-Egg) was used to predict molecular permeability, estimating the potential for passive human gastrointestinal absorption (HIA) and blood–brain barrier (BBB) penetration.66 ESI Fig. S4† presents a boiled-egg graph comparing known inhibitors with potential inhibitors.

The yolk portion represents the physicochemical space indicating molecules most likely to penetrate the brain, while the white part denotes molecules with a high probability of gastrointestinal (GI) absorption. Molecules predicted to have low human gastrointestinal absorption (HIA) and blood–brain barrier (BBB) penetration are depicted in the gray zone. Blue points indicate molecules that are substrates of P-glycoprotein (P-gp) and actively effluxed, while red points represent non-substrates. Fig. S4† shows that azvudine, M2 and M3 are not P-gp substrates, whereas M1 and nirmatrelvir are P-gp substrates.67 All candidate compounds and known inhibitors, except for azvudine are predicted to exhibit favorable absorption characteristics and are not expected to penetrate the BBB.

Table 4 displays the oral toxicity assessment of the candidate compounds using azvudine as the reference drug.

Oral toxicity assessment of the candidate compounds with azvudine as reference drug

Chemical compound	Predicted LD50 (mg kg−1)	Predicted toxicity class	Prediction accuracy (%)	Average similarity (%)	
Azvudine	1000	4	67.38	59.91	
M1	550	4	54.26	46.19	
M2	500	4	54.26	48.36	
M3	200	3	67.38	55.68	

Compared to azvudine, all candidates showed lower LD50 values,68 suggesting potentially higher toxicity. M1, M2, and azvudine are predicted to be in class IV, while M3 may fall into class III based on toxicity classification criteria. Additionally, The MTDs of azvudine, M1, M2, and M3 are predicted as 0.289, 0.614, 0.615, and 0.542 (log mg per kg per day), respectively.

3.4 Prediction of pIC50 values

The performance of 20 regressors (RG) is summarized in ESI Table S3.† To construct the Voting Regressor (VR), we chose the Random Forest (RF), Hist Gradient Boosting (HGB), and Light Gradient Boosting Machine (LGBM) regressors. The RF regressor had n_estimators set to ‘200’, criterion set to ‘squared_error’, max_features set to ‘sqrt’ and min_samples_split to ‘2’; the HGB regressor had max_iter set to ‘200’ and learning_rate to ‘0.1’; the LGBM regressor had n_estimators set to ‘200’ and learning_rate to ‘0.1’; all other parameters were left at their default values. These three regressors were combined to build the voting regressor. Evaluation metrics for the voting regressor in training, testing, and 5-fold CV are presented in Table 5.

Evaluation metrics of voting regressor in training, testing and 5-fold CV

Statistical metrics	Training	Testing	5-fold CV	
R2	0.97	0.71	0.73	
MAE	0.13	0.45	0.41	
RMSE	0.18	0.62	0.57	

Experimental and predicted pIC50 values are compared in ESI Fig. S5.† Predicted pIC50 values for M1, M2, and M3 were 7.35, 7.59, and 7.71 which are comparable to activity of nirmatrelvir (7.70).

3.5 Molecular docking analysis

Molecular docking was used to generate the 3CLpro protein–ligand complexes of candidate compounds. The RMSD between the initial pose and the re-docked pose of the native ligand was found to be 0.426 Å (Fig. S6†). These validated parameters were used for the docking of 3CLpro and the candidate compounds. The protein–ligand interactions of candidate compounds are shown in Fig. 5.

Fig. 5 Protein–ligand interactions of 3CLpro and candidate compounds.

The binding energy (with each contributing factor) of candidate compounds with 3CLpro for best docking pose is shown in ESI Table S4.† The binding energies of M1 (6-[2,4-bis(dimethylamino)-6,8-dihydro-5H-pyrido[3,4-d]pyrimidine-7-carbonyl]-1H pyrimidine-2,4-dione), M2 (6-[3-(2,5-dimethoxyphenyl)pyrrolidine-1-carbonyl]-1H-pyrimidine-2,4-dione) and M3 ([(3R,4R)-4-hydroxy-3-methyl-4-(oxan-4-yl)piperidine-1-carbonyl]-1H-pyrimidine-2,4-dione) are −8.64 kcal mol−1, −8.22 kcal mol−1 and −8.00 kcal mol−1 respectively which suggests a good binding affinity with target protein.

The interactions between the active residues of 3CLpro and the best docked pose of candidate compounds are shown in ESI Table S5.†

3.6 Molecular dynamics simulation analysis

We performed MD simulations for the complexes of the three candidate compounds and target protein to verify the outcomes of our virtual screening using machine learning and docking. Through trajectory analysis, only M1 was found to be stable during MD simulation among the three candidate compounds. From RMSD data we found that all our systems reached stability after 180 ns (Fig. 6a), so we defined the productive phase of our simulations as the time between 180 and 200 ns for all the runs. The RMSD, Rg and, RMSF plots of MD simulation for apo and M1-complex are shown in Fig. 6.

Fig. 6 (a) RMSD and (b) Rg plots for apo and M1 binding forms of 3CLpro during 200 ns MD simulation.

The stability of the ligand and protein in a complex was studied using RMSD analysis. The average RMSD of protein backbone in apo and M1 binding forms is 2.02 ± 0.21 Å and 1.91 ± 0.31 Å, respectively. The RMSD value of the protein backbone was less than 3 Å, indicating a minor change for globular proteins. These results demonstrate the stability of apo and ligand binding forms.

Next, we examined the Rg, which is a reliable indicator of protein folding. The average value of Rg throughout the simulation for apo and M1 binding forms is 22.26 ± 0.17 Å and 21.29 ± 0.12 Å respectively, which shows the overall stable protein folding in the complex without any significant expansion or condensation.

The average RMSF values for apo and ligand binding forms are 1.21 ± 0.59 Å and 1.09 ± 0.61 Å respectively, with the majority of residues showing similar RMSF values, while some regions – like SER1 (6.51 Å), GLY2 (4.38 Å), SER301 (3.03 Å), THR304 (3.27 Å), and GLN306 (3.07 Å) – showed larger fluctuations (Fig. S7†). These residues are not critical because they are found in the inactive regions of protein. On the other hand, key residues in the active site, like HIS41, SER144, CYS145, GLU166, and HIS172, showed reduced fluctuations with RMSF values below 1.1 Å, indicating that the formed hydrogen bonds stabilize the ligand complexation with protein 3CLpro.

Furthermore, we used the DSSP module installed in GROMACS to examine the stability of their secondary structure.69,70 During our simulation, the M1 and apo binding forms both kept a stable secondary structure on a global scale (Fig. S8†).

The GROMACS Hbond module71 with default parameters and the HbMap2Grace program72 were utilized to assess the hydrogen bond pattern, while the SurfinMD program73 was employed to evaluate the molecular surface area. The hydrogen bond data indicates that there were notable interactions between the M1 and the active residues (Fig. 7). M1 displayed hydrogen bonding with the SER144 complex for nearly the whole simulation period.

Fig. 7 Hydrogen bond stability in 3CLpro-M1 complex for the productive phase.

Additionally, we calculated the atomic contacts between M1 and SARS-CoV-2 Mpro (Fig. 8). The contact surface area disclosed interactions with key residues in the active site.

Fig. 8 Surface molecular area of 3CLpro-M1 complex for the productive phase.

By using MM-PBSA calculations, the post-MD free energy of M1 in complex with 3CLpro has been examined. Van der Waals energy (VDWAALS), electrostatic energy (EEL), polar solvation energy (EPB), and nonpolar solvation energy (ENPOLAR) are the main contributors to the total binding free energy. Fig. 9a shows the overall binding free energy contributors of M1 in complex with 3CLpro over the last 200 frames. The major contributors to the total MM-PBSA free energy of −18.86 ± 4.38 kcal mol−1, expressed as average ± SD, are electrostatic energy (−25.86 ± 8.88 kcal mol−1) and vdW energy (−37.85 ± 3.24 kcal mol−1), as shown in ESI Table S6.†

Fig. 9 MM-PBSA results of M1 in complex with 3CLpro during last 20 ns MD simulations: (a) binding free energy contribution by different interactions, (b) binding free energy contributions by active residues and ligand.

Fig. 9b displays the binding free energies that are contributed by the active residues of 3CLpro and M1. The decomposition analysis indicated that M1 has a strong binding affinity. The ligand engages with critical residues in 3CLpro, notably forming a significant interaction with the CYS145–HIS41 catalytic dyad, which is essential for the enzyme's functionality.74 Of the total MM-PBSA free energy, M1 contributes −8.13 ± 2.28 kcal mol−1. The lowest binding free energies of −1.70 ± 0.54 kcal mol−1 and −1.19 ± 0.52 kcal mol−1 are displayed by CYS145 and PHE140 respectively, out of all the residues (ESI Table S7†). Additionally, ESI Fig. S9† displays the heatmap of the binding free energy contribution by active residues and ligand.

4 Conclusion

Drug development is costly and time-consuming. We utilized a workflow integrating ligand-based virtual screening with similarity assessments of approved drugs to identify potential 3CLpro inhibitors. Using three machine learning classifiers, we created a voting classifier to predict activity probabilities, analyzing approximately 10 million molecules. We selected three compounds M1, M2 and M3 for further investigation. ADMET analysis, with azvudine and nirmatrelvir as references, and 200 ns MD simulations identified M1 (6-[2,4-bis(dimethylamino)-6,8-dihydro-5H-pyrido[3,4-d]pyrimidine-7-carbonyl]-1H-pyrimidine-2,4-dione) as stable. Predicted LD50 values for M1 and azvudine were 550 and 1000 mg kg−1, respectively. The pIC50 value for M1 was approximately 7.35, similar to nirmatrelvir. MM-PBSA calculations showed a binding energy of −18.86 ± 4.38 kcal mol−1 for the M1-3CLpro complex. Our study suggests that M1 warrants further investigation as a potential SARS-CoV-2 therapeutic, potentially improving drug discovery efficiency and conserving resources.

Data availability

The data supporting the findings of this study are available within the article and its ESI.†

Author contributions

Sandeep Poudel Chhetri: experiment design, data generation, analyzed data, and drafted the manuscript. Vishal Singh Bhandari: technical support and revised the manuscript. Rajesh Maharjan: technical support, data generation and revised the manuscript. Tika Ram Lamichhane: critical feedback, graphical and statistical analysis, and revised the manuscript.

Conflicts of interest

The authors declare that there are no conflicts of interest.

Supplementary Material

RA-014-D4RA04502E-s001
==== Refs
References

Corman V. M. Landt O. Kaiser M. Molenkamp R. Meijer A. Chu D. K. Bleicker T. Brünink S. Schneider J. Schmidt M. L. Mulders D. G. Haagmans B. L. Van Der Veer B. Van Den Brink S. Wijsman L. Goderski G. Romette J.-L. Ellis J. Zambon M. Peiris M. Goossens H. Reusken C. Koopmans M. P. Drosten C. Eurosurveillance 2020 25 10.2807/1560-7917.ES.2020.25.3.2000045
COVID-19 cases | WHO COVID-19 dashboard, https://data.who.int/dashboards/covid19/cases, (accessed January 21, 2024)
Hu B. Guo H. Zhou P. Shi Z.-L. Nat. Rev. Microbiol. 2021 19 141 154 33024307
Agbowuro A. A. Huston W. M. Gamble A. B. Tyndall J. D. A. Med. Res. Rev. 2018 38 1295 1331 29149530
Anand K. Ziebuhr J. Wadhwani P. Mesters J. R. Hilgenfeld R. Science 2003 300 1763 1767 12746549
Unoh Y. Uehara S. Nakahara K. Nobori H. Yamatsu Y. Yamamoto S. Maruyama Y. Taoda Y. Kasamatsu K. Suto T. Kouki K. Nakahashi A. Kawashima S. Sanaki T. Toba S. Uemura K. Mizutare T. Ando S. Sasaki M. Orba Y. Sawa H. Sato A. Sato T. Kato T. Tachibana Y. J. Med. Chem. 2022 65 6499 6512 35352927
Ullrich S. Nitsche C. Bioorg. Med. Chem. Lett. 2020 30 127377 32738988
Zhang L. Lin D. Sun X. Curth U. Drosten C. Sauerhering L. Becker S. Rox K. Hilgenfeld R. Science 2020 368 409 412 32198291
Li G. Hilgenfeld R. Whitley R. De Clercq E. Nat. Rev. Drug Discovery 2023 22 449 475 37076602
Lavecchia A. Giovanni C. CMC 2013 20 2839 2860
Gloriam D. E. Nature 2019 566 193 194
Zhong F. Xing J. Li X. Liu X. Fu Z. Xiong Z. Lu D. Wu X. Zhao J. Tan X. Li F. Luo X. Li Z. Chen K. Zheng M. Jiang H. Sci. China: Life Sci. 2018 61 1191 1204 30054833
Duan Y. Edwards J. S. Dwivedi Y. K. J. Inf. Manag. 2019 48 63 71
Lavecchia A. Drug Discovery Today 2019 24 2017 2032 31377227
Salimi A. Lim J. H. Jang J. H. Lee J. Y. Sci. Rep. 2022 12 18825 36335233
Maharjan R. Gyawali K. Acharya A. Khanal M. Ghimire M. P. Lamichhane T. R. Mol. Simul. 2024 50 717 728
Sydow D. Morger A. Driller M. Volkamer A. J. Cheminf. 2019 11 29
Mendez D. Gaulton A. Bento A. P. Chambers J. De Veij M. Félix E. Magariños M. P. Mosquera J. F. Mutowo P. Nowotka M. Gordillo-Marañón M. Hunter F. Junco L. Mugumbate G. Rodriguez-Lopez M. Atkinson F. Bosc N. Radoux C. J. Segura-Cabrera A. Hersey A. Leach A. R. Nucleic Acids Res. 2019 47 D930 D940 30398643
Lipinski C. A. Lombardo F. Dominy B. W. Feeney P. J. Adv. Drug Delivery Rev. 2012 64 4 17
Doak B. C. Over B. Giordanetto F. Kihlberg J. Chem. Biol. 2014 21 1115 1142 25237858
Bajusz D. , Rácz A. and Héberger K. , in Comprehensive Medicinal Chemistry III, Elsevier, 2017, pp. 329–378
Willett P. Drug Discovery Today 2006 11 1046 1053 17129822
Awale M. Visini R. Probst D. Arús-Pous J. Reymond J.-L. CHIMIA 2017 71 661 29070411
Landrum G. , Rdkit: Open-source cheminformatics software. (version 2023.9.4), 2016
Kwon S. Bae H. Jo J. Yoon S. BMC Bioinf. 2019 20 521
Pedregosa F. Varoquaux G. Gramfort A. Michel V. Thirion B. Grisel O. Blondel M. Prettenhofer P. Weiss R. Dubourg V. Vanderplas J. Passos A. Cournapeau D. Brucher M. Perrot M. Duchesnay É. J. Mach. Learn. Res. 2011 12 2825 2830
Ke G. , Meng Q. , Finley T. , Wang T. , Chen W. , Ma W. , Ye Q. and Liu T.-Y. , in Advances in Neural Information Processing Systems, Curran Associates, Inc., 2017, vol. 30
Dietterich T. G. , in Multiple Classifier Systems, Springer, Berlin, Heidelberg, 2000, pp. 1–15
Luque A. Carrasco A. Martín A. de las Heras A. Pattern Recognit. 2019 91 216 231
Chicco D. Warrens M. J. Jurman G. PeerJ Comput. Sci. 2021 7 e623
Quimque M. T. J. Notarte K. I. R. Fernandez R. A. T. Mendoza M. A. O. Liman R. A. D. Lim J. A. K. Pilapil L. A. E. Ong J. K. H. Pastrana A. M. Khan A. Wei D.-Q. Macabeo A. P. G. J. Biomol. Struct. Dyn. 2021 39 4316 4333 32476574
Lima A. N. Philot E. A. Trossini G. H. G. Scott L. P. B. Maltarollo V. G. Honorio K. M. Expert Opin. Drug Discovery 2016 11 225 239
eMolecules, https://search.emolecules.com/, (accessed January 4, 2024)
Sheridan R. P. Kearsley S. K. Drug Discovery Today 2002 7 903 911 12546933
Maldonado A. G. Doucet J. P. Petitjean M. Fan B.-T. Mol. Divers. 2006 10 39 79 16404528
Cheng G. Lajiness M. Johnson M. A. J. Chem. Inf. Comput. Sci. 1996 36 909 915
Bajusz D. Rácz A. Héberger K. J. Cheminf. 2015 7 20
Riniker S. Landrum G. A. J. Cheminf. 2013 5 43
Bojarska J. Remko M. Breza M. Madura I. D. Kaczmarek K. Zabrocki J. Wolf W. M. Molecules 2020 25 1135 32138329
Kola I. Landis J. Nat. Rev. Drug Discovery 2004 3 711 716 15286737
Daina A. Michielin O. Zoete V. Sci. Rep. 2017 7 42717 28256516
Banerjee P. Eckert A. O. Schrey A. K. Preissner R. Nucleic Acids Res. 2018 46 W257 W263 29718510
Pires D. E. V. Blundell T. L. Ascher D. B. J. Med. Chem. 2015 58 4066 4072 25860834
Rose P. W. Prlić A. Altunkaya A. Bi C. Bradley A. R. Christie C. H. Costanzo L. D. Duarte J. M. Dutta S. Feng Z. Green R. K. Goodsell D. S. Hudson B. Kalro T. Lowe R. Peisach E. Randle C. Rose A. S. Shao C. Tao Y.-P. Valasatava Y. Voigt M. Westbrook J. D. Woo J. Yang H. Young J. Y. Zardecki C. Berman H. M. Burley S. K. Nucleic Acids Res. 2017 45 D271 D281 27794042
Douangamath A. Fearon D. Gehrtz P. Krojer T. Lukacik P. Owen C. D. Resnick E. Strain-Damerell C. Aimon A. Ábrányi-Balogh P. Brandão-Neto J. Carbery A. Davison G. Dias A. Downes T. D. Dunnett L. Fairhead M. Firth J. D. Jones S. P. Keeley A. Keserü G. M. Klein H. F. Martin M. P. Noble M. E. M. O'Brien P. Powell A. Reddi R. N. Skyner R. Snee M. Waring M. J. Wild C. London N. von Delft F. Walsh M. A. Nat. Commun. 2020 11 5047 33028810
Yang J. Zhang Y. Nucleic Acids Res. 2015 43 W174 W181 25883148
Morris G. M. Huey R. Lindstrom W. Sanner M. F. Belew R. K. Goodsell D. S. Olson A. J. J. Comput. Chem. 2009 30 2785 2791 19399780
DeLano W. L. CCP4 Newslett. Protein Cryst. 2002 40 82
Morris G. M. Goodsell D. S. Halliday R. S. Huey R. Hart W. E. Belew R. K. Olson A. J. J. Comput. Chem. 1998 19 1639 1662
Khanal M. Acharya A. Maharjan R. Gyawali K. Adhikari R. Mulmi D. D. Lamichhane T. R. Lamichhane H. P. PLoS One 2024 19 e0307501 39037973
Acharya A. Khanal M. Maharjan R. Gyawali K. Luitel B. R. Adhikari R. Mulmi D. D. Lamichhane T. R. Lamichhane H. P. Acharya A. Khanal M. Maharjan R. Gyawali K. Luitel B. R. Adhikari R. Mulmi D. D. Lamichhane T. R. Lamichhane H. P. AIMSBPOA 2024 11 142 165
Adasme M. F. Linnemann K. L. Bolz S. N. Kaiser F. Salentin S. Haupt V. J. Schroeder M. Nucleic Acids Res. 2021 49 W530 W534 33950214
Abraham M. J. Murtola T. Schulz R. Páll S. Smith J. C. Hess B. Lindahl E. SoftwareX 2015 1–2 19 25
Huang J. MacKerell Jr A. D. J. Comput. Chem. 2013 34 2135 2145 23832629
Zoete V. Cuendet M. A. Grosdidier A. Michielin O. J. Comput. Chem. 2011 32 2359 2368 21541964
Lamichhane T. R. Ghimire M. P. Heliyon 2021 7 e08220 34693066
Mark P. Nilsson L. J. Phys. Chem. A 2001 105 9954 9960
Bussi G. Donadio D. Parrinello M. J. Chem. Phys. 2007 126 014101 17212484
Parrinello M. Rahman A. J. Appl. Phys. 1981 52 7182 7190
Silva R. C. Freitas H. F. Campos J. M. Kimani N. M. Silva C. H. T. P. Borges R. S. Pita S. S. R. Santos C. B. R. Int. J. Mol. Sci. 2021 22 11739 34769170
Valdés-Tresanco M. S. Valdés-Tresanco M. E. Valiente P. A. Moreno E. J. Chem. Theory Comput. 2021 17 6281 6291 34586825
Yu B. Chang J. Sig. Transduct. Target Ther. 2020 5 1 2
Owen D. R. Allerton C. M. N. Anderson A. S. Aschenbrenner L. Avery M. Berritt S. Boras B. Cardin R. D. Carlo A. Coffman K. J. Dantonio A. Di L. Eng H. Ferre R. Gajiwala K. S. Gibson S. A. Greasley S. E. Hurst B. L. Kadar E. P. Kalgutkar A. S. Lee J. C. Lee J. Liu W. Mason S. W. Noell S. Novak J. J. Obach R. S. Ogilvie K. Patel N. C. Pettersson M. Rai D. K. Reese M. R. Sammons M. F. Sathish J. G. Singh R. S. P. Steppan C. M. Stewart A. E. Tuttle J. B. Updyke L. Verhoest P. R. Wei L. Yang Q. Zhu Y. Science 2021 374 1586 1593 34726479
Study Details | A Study of Efficacy and Safety of Azvudine vs. Nirmatrelvir-Ritonavir in the Treatment of COVID-19 Infection | ClinicalTrials.gov, https://clinicaltrials.gov/study/NCT05697055, (accessed March 3, 2024)
Delaney J. S. J. Chem. Inf. Comput. Sci. 2004 44 1000 1005 15154768
Daina A. Zoete V. ChemMedChem 2016 11 1117 1121 27218427
Chen C. Lee M.-H. Weng C.-F. Leong M. K. Molecules 2018 23 1820 30037151
Lohohola P. O. Mbala B. M. Bambi S.-M. N. Mawete D. T. Matondo A. Mvondo J. G. M. Int. J. Trop. Dis. Health 2021 42 1 12
Touw W. G. Baakman C. Black J. te Beek T. A. H. Krieger E. Joosten R. P. Vriend G. Nucleic Acids Res. 2015 43 D364 D368 25352545
Kabsch W. Sander C. Biopolymers 1983 22 2577 2637 6667333
van der Spoel D. van Maaren P. J. Larsson P. Tîmneanu N. J. Phys. Chem. B 2006 110 4393 4398 16509740
Gomes D. E. B. , Silva A. W. , Linis R. D. , Pascutti P. G. and Soares T. A. , HbMap2Grace, https://lmdm.biof.ufrj.br/software/hbmap2grace/index.html-2002
Gomes D. E. B. , Sousa G. L. S. C. , Silva A. W. S. D. and Pascutti P. G. , SurfinMD, https://lmdm.biof.ufrj.br/software/surfinmd/index.html-2012
Ferreira J. C. Fadl S. Villanueva A. J. Rabeh W. M. Front. Chem. 2021 9 692168 34249864
